Method, device and equipment for generating video

By receiving videos to be edited and using target atomic capability programs and preset parameters for editing, the problem of low video generation efficiency is solved, and the effect of quickly generating promotional videos is achieved.

CN121509743APending Publication Date: 2026-02-10ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511676471.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies for generating promotional videos are inefficient, and the manual production process is cumbersome, failing to meet rapidly changing market demands and the need to produce large quantities of promotional videos in the context of fierce competition for traffic.

Method used

By receiving the video to be edited, the system obtains the target atomic capability program and preset parameters selected by the user, and then uses the target atomic capability program to edit the video to generate the target video.

Benefits of technology

It improves the efficiency of video generation, enabling the rapid processing of multiple videos to meet the needs of mass production, and reducing manpower input and production cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509743A_ABST
    Figure CN121509743A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device and equipment for generating a video. The method for generating the video comprises the following steps: receiving a to-be-edited video for generating a target video; obtaining a target atomic power program selected by a user and used for editing the to-be-edited video; wherein one atomic power program has an atomic power for editing the video to be edited; different atomic energy programs have different types of atomic energy; acquiring preset parameters for the target atomic power program; and editing the to-be-edited video by using the target atomic power program based on the preset parameters so as to generate a target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video generation technology, and in particular to a method, apparatus, and device for generating video. Background Technology

[0002] As the industry's digital transformation accelerates, promotional videos have become a core medium connecting products and users. Currently, promotional videos are typically produced manually, a process that is cumbersome and inefficient. This high-manpower, long-cycle model struggles to meet rapidly changing market demands and cannot cope with the massive production needs arising from fierce competition for online traffic.

[0003] Therefore, how to provide a method for generating videos to improve the efficiency of video generation has become an urgent technical problem to be solved. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus, and device for generating video to solve the problem of low efficiency in existing methods for generating video.

[0005] According to a first aspect of the embodiments of this application, a method for generating video is provided, comprising: Receive a video to be edited for generating a target video; the target video is a video used to promote or introduce a preset product; Obtain the target atomic capability program selected by the user for editing the video to be edited; an atomic capability program has one atomic capability for editing the video to be edited; different atomic capability programs have different types of atomic capabilities; Obtain preset parameters for the target atom capability program; Based on the preset parameters, the target atomic capability program is used to edit the video to be edited, thereby generating the target video.

[0006] According to a second aspect of the embodiments of this application, an apparatus for generating video is provided, comprising: The receiving module is used to receive the video to be edited for generating the target video; the target video is a video used to promote or introduce a preset product. The first acquisition module is used to acquire the target atomic capability program selected by the user for editing the video to be edited; an atomic capability program has one atomic capability for editing the video to be edited; different atomic capability programs have different types of atomic capabilities; The second acquisition module is used to acquire preset parameters for the target atom capability program; The target video generation module is used to edit the video to be edited based on the preset parameters and using the target atomic capability program to generate the target video.

[0007] According to a third aspect of the present application, a computing device is provided, including a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor executes the computer instructions to implement the steps of the method for generating video.

[0008] One embodiment of this specification can achieve at least the following beneficial effects: by acquiring a video to be edited for generating a target video, acquiring a target atomic capability program selected by the user for editing the video to be edited, wherein an atomic capability program has one atomic capability for editing the video to be edited; different atomic capability programs have different types of atomic capabilities; acquiring preset parameters for the target atomic capability program, and using the target atomic capability program to edit the video to be edited based on the preset parameters to generate the target video. By modularizing the various processing steps for the video to be edited, various atomic capability programs are obtained. This allows for the acquisition of target atomic capability programs for processing the video to be edited according to requirements, enabling rapid processing of the video to be edited using the target atomic capabilities, thereby improving the efficiency of generating the target video. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram illustrating an application scenario of a video generation method provided in the embodiments of this specification; Figure 2 This is a flowchart illustrating a method for generating video provided in an embodiment of this specification; Figure 3 This is a schematic diagram of the structure of a video generation device provided in the embodiments of this specification; Figure 4 This is a schematic diagram of the structure of a video generation device provided in the embodiments of this specification. Detailed Implementation

[0011] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0012] The terminology used in one or more embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this application. The singular forms “a,” “the,” and “the” used in one or more embodiments of this application and in the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” used in one or more embodiments of this application refers to and includes any or all possible combinations of one or more associated listed items.

[0013] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this application, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0014] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0015] In existing technologies, video editing software is typically required for manual editing of video content, making the production process cumbersome and inefficient. Furthermore, since existing video editing software can only process one file at a time, the manual editing method cannot be reused, failing to meet the production needs of batch videos. This high-manpower-input and long-production-cycle model is ill-suited to rapidly changing market demands and cannot cope with the large-scale production needs of promotional videos in the context of fierce competition for traffic.

[0016] To address the shortcomings of the existing technology, the following embodiments are provided in this solution.

[0017] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0018] Figure 1 This is a schematic diagram illustrating an application scenario of a video generation method provided in an embodiment of this specification.

[0019] like Figure 1As shown, server 101 can be a server for generating a target video; a user can provide a video to be edited, an identifier of a target atomic capability program for processing the video to be edited, and preset parameters required for processing the video to be edited using the target atomic capability program through terminal device 102; terminal device 102 can send the video to be edited, the identifier of the target atomic capability program, and the preset parameters to server 101. After receiving the above data, server 101 can retrieve the corresponding target atomic capability program based on the identifier of the target atomic capability program, and edit the video to be edited based on the preset parameters using the identifier of the target atomic capability program, thereby generating the target video.

[0020] although Figure 1 The diagram shows that after receiving data such as the video to be edited, the identifier of the target atomic capability program, and preset parameters, the terminal device 102 can send the aforementioned data to the server 101. In practical applications, if the computing resources of the terminal device 102 meet the conditions for running the method for generating the target video, the scheme for generating the target video can also be executed on the terminal device 102.

[0021] In such Figure 1 In the application scenario shown, server 101 can connect to one or more terminal devices via a local area network (LAN), a wide area network (WAN), an internet connection, or other types of data networks. Figure 1 The server 101 in the text may include, but is not limited to, any device, equipment, platform, or equipment cluster with computing and processing capabilities. Figure 1 The terminal device 102 may include, but is not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices.

[0022] Next, a method for generating video provided in the embodiments of the specification will be described in detail with reference to the accompanying drawings.

[0023] Figure 2 This is a flowchart illustrating a method for generating video provided in an embodiment of this specification. From a programming perspective, the entity executing the process can be a program hosted on an application server or an application client. It is understood that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.

[0024] like Figure 2 As shown, the process may include the following steps: Step 202: Receive the video to be edited for generating the target video.

[0025] Target videos can be used to promote or introduce pre-defined products. In practical applications, target videos can be used to promote or introduce the features and target audience of pre-defined products. These pre-defined products can include insurance products, smart devices, online education courses, beauty and skincare sets, customized travel services, and compliant financial products, among others.

[0026] If the preset product is an insurance product, the target video can be a video introducing or promoting the insurance product. Specifically, the target video can be a commercial advertisement promoting or introducing commercial insurance, or a public service advertisement promoting or introducing public welfare insurance. Typically, a single target video can focus on introducing one preset product. However, to help users select the right preset product from multiple similar products, a single target video can also introduce multiple preset products of the same category. For example, it could introduce multiple different mobile phone models, or introduce multiple insurance products with similar coverage, sum insured, payment periods, and claims conditions, thus helping users intuitively compare core differences and quickly select a product that meets their needs. The specific content of the target video is not specifically limited here.

[0027] The video to be edited can be a source file used to generate the target video. Specifically, the video to be edited can be a video, an image, or other file that can generate the target video. In practice, there can be one or more videos to be edited; no specific limitation is made here.

[0028] In practical applications, the executing entity (server) can receive the video to be edited sent by the terminal device; the video to be edited provided by the user through the terminal device can be selected from an existing video database, created by the user, or downloaded from other platforms or websites.

[0029] Step 204: Obtain the target atomic capability program selected by the user for editing the video to be edited; an atomic capability program has one atomic capability for editing the video to be edited; different atomic capability programs have different types of atomic capabilities.

[0030] In the embodiments of this specification, the atomic capability program can be a basic functional unit built for the indivisible basic processing sub-processes in the target video generation process.

[0031] In practical applications, the various processing steps in the video generation process can be abstracted into several indivisible basic processing sub-processes in advance. For each sub-process, independent basic functional units can be constructed to obtain atomic capability programs. These atomic capability programs can include video analysis atomic capability programs, video cropping atomic capability programs, video stitching atomic capability programs, image overlay atomic capability programs, text filling atomic capability programs, format conversion atomic capability programs, digital human video generation atomic capability programs, resolution adjustment atomic capability programs, audio extraction atomic capability programs, etc.

[0032] Understandably, an atomic capability program is a single modular program; the video analysis atomic capability program is a modular program that can be used to analyze highlight video segments; the video cropping atomic capability program is a modular program that can be used to crop videos; the video splicing atomic capability program is a modular program that can be used to splice videos together, and also to splice videos with images; the image overlay atomic capability program is a modular program that can be used to overlay videos together, and also to overlay videos with images; the text fill atomic capability program is a modular program that can be used to fill text content into image templates; the format conversion atomic capability program is a modular program that can be used to convert non-video format files into video format files; the digital human video generation atomic capability program is a modular program that can be used to generate digital human videos; and the audio extraction atomic capability program is a modular program that can be used to extract audio information from video files.

[0033] A target atomic capability program is an atomic capability program selected by the user for editing the video to be edited. The target atomic capability program can be one or more atomic capability programs. In practical applications, the video requester can input the identifier of the target atomic capability program through the data input interface of the terminal device, and the executing entity can obtain the target atomic capability program based on the identifier. The identifier of the target atomic capability program can be the name or ID of the target atomic capability program.

[0034] In practical applications, atomic capability programs can be stored on a local server or terminal device, or on other devices or servers that are connected to the local server. Understandably, if the atomic capability program is stored on a local server or terminal device, the executing entity can directly obtain the target atomic capability program from the local server or terminal device; if the atomic capability program is stored on other devices or terminals that are connected to the local server, it can be retrieved through an interface call.

[0035] In the embodiments of this specification, the various processing steps in the video generation process are abstracted into several indivisible sub-processes, and atomic capability programs are formed for each sub-process. During the generation of the target video, the target atomic capability program used for editing the video to be edited is obtained. This allows for direct use of the target atomic capability program to replace manual editing of the video, thereby improving the efficiency of target video generation.

[0036] Step 206: Obtain preset parameters for the target atom capability program.

[0037] In the embodiments described in this specification, preset parameters can refer to the conditions required by the target atomic capability program during execution. Preset parameters can be used to control the execution logic, output effects, processing rules, etc. of the target atomic capability program.

[0038] In practical applications, if there is only one target atomic capability program, the preset parameters can usually be input by the user through the data input interface of the terminal device. For example, if the target atomic capability program is a video cropping atomic capability program, the preset parameters for the video cropping atomic capability program can include the start time and end time of the cropping, which can be input by the user through the data input interface of the terminal device. If there are multiple target atomic capability programs, the preset parameters can include user input, and can also include intermediate parameters generated after the previous target atomic capability program has edited the video to be edited. For example, the target atomic capability programs include a video analysis atomic capability program and a video cropping atomic capability program; the video analysis atomic capability program can analyze the highlight video in the video to be edited, and can also determine the time information of the highlight video in the video to be edited; the video cropping atomic capability program performs cropping on the video to be edited based on the above time information; therefore, the preset parameters of the video cropping atomic capability program, the start time and end time of the cropping, are obtained based on the analysis of the video analysis atomic capability program.

[0039] In practical applications, some atomic capability programs may not require preset parameters during execution. Understandably, for atomic capability programs that do not require preset parameters, obtaining those parameters is also unnecessary. For example, a video analysis atomic capability program may not need parameters when analyzing highlight video; in this case, obtaining preset parameters is also not required.

[0040] Step 208: Based on the preset parameters, use the target atom capability program to edit the video to be edited to generate the target video.

[0041] In the embodiments of this specification, during the target video generation process, by utilizing the target atomic capability program to directly replace the manual editing of the video to be edited, the generation efficiency of the target video can be significantly improved.

[0042] Furthermore, when performing the same processing procedure on different videos to be edited, the target atomic capability program can be reused to process each video. For example, the video cropping atomic capability program can be used to crop different videos to be edited. Reusing basic functional units improves the efficiency of target video generation.

[0043] In practical applications, target atomic capability programs can also be used to process multiple videos to be edited in parallel, enabling the generation of a large number of target videos in a short time to meet the needs of mass production. For example, resolution adjustment atomic capability programs can be used to adjust the resolution of multiple videos to be edited in parallel. As one implementation, the target resolution of each video to be encoded undergoing resolution adjustment can be the same, i.e., the preset parameter values ​​are the same. As another implementation, the target resolution of each video to be encoded can also be different, i.e., the target resolution of each video to be edited can be set separately.

[0044] It should be understood that in the methods described in one or more embodiments of this specification, the order of some steps may be adjusted according to actual needs, or some steps may be omitted.

[0045] Figure 2 The method described herein involves acquiring a video to be edited for generating a target video, obtaining a user-selected target atomic capability program for editing the video, where each atomic capability program possesses a specific atomic capability for editing the video; different atomic capability programs have different types of atomic capabilities; acquiring preset parameters for the target atomic capability program, and then using the target atomic capability program to edit the video based on the preset parameters to generate the target video. Basic functional units are constructed for the indivisible basic processing sub-processes in the target video generation process to obtain various atomic capability programs. Therefore, during the target video generation process, the target atomic capability program for processing the video to be edited is acquired as needed, allowing direct use of the target atomic capability program to replace manual editing of the video, thus improving the efficiency of target video generation.

[0046] based on Figure 2 In addition to the method described herein, this specification also provides some improved implementation methods, which will be described below.

[0047] In practical applications, if there are multiple target atomic capability programs, and the processing order of each program on the video to be edited is different, the final target video generated may be different. For example, a target video obtained by first using a video trimming atomic capability program to trim the video to be edited from the [t1-t2] time period, and then using a video splicing atomic capability program to splice the trimmed video with another video, is different from a target video obtained by first using a video splicing atomic capability program to splice the video to be edited with another video, and then using a video trimming atomic capability program to trim the spliced ​​video from the [t1-t2] time period.

[0048] Therefore, in order to ensure that the generated target video meets the user's needs, the execution order of each target atomic capability program can also be set.

[0049] In one or more embodiments of this specification, optionally, if there are multiple target atomic capability programs, the process before generating the target video may further include: Obtain the execution order of the capability programs for each of the target atoms.

[0050] The step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video may specifically include: Based on the preset parameters, the target atomic capability programs are used sequentially according to the execution order to edit the video to be edited, thereby generating the target video.

[0051] In the embodiments of this specification, if there are multiple target atomic capability programs, the execution order of each target atomic capability program can be set based on the processing flow for the video to be edited.

[0052] As one implementation method, the execution order of each target atomic capability program can be obtained based on its arrangement. In practical applications, users can select the target atomic capability programs sequentially according to the processing flow of the video to be edited; alternatively, they can select the target atomic capability programs and then adjust their arrangement order. As one implementation method, the target atomic capability program listed first is executed first.

[0053] As another implementation, the execution order of each of the target atomic capability programs can be obtained based on the execution serial numbers of the target atomic capability programs. In practical applications, after the user selects a target atomic capability program, an execution serial number can be set for each target atomic capability program. The execution serial number can be a number in any form, such as ①②③④⑤⑥⑦⑧⑨⑩, etc. or 12345678910, etc.; it can also be a letter in any form, such as abcdefg, etc., or ABCDEFG, etc.; when the execution serial number is a number or a letter, the execution serial numbers can be continuous or discontinuous, and no specific limitation is made here.

[0054] Specifically, when the execution serial number is an Arabic numeral, it can be executed in ascending order of the numbers; if the execution serial number is the Chinese numerals "one, two, three", the Chinese serial numbers can be converted into digital serial numbers through a preset mapping rule (such as one → 1, two → 2, three → 3) and then executed. If the execution serial number is a letter, the execution order of each target atomic capability program can be determined according to the alphabetical order (a → b → c or A → B → C).

[0055] In the embodiments of this specification, by flexibly arranging multiple target atomic capability programs, a production pipeline for generating a target video can be quickly built, effectively improving the production efficiency of the target video. On the other hand, based on the processing flow of the video to be edited, the execution order of the target atomic capability programs is set, and each target atomic capability program is called in sequence according to this execution order to edit the video to be edited, which can effectively ensure that the generated target video meets the user's requirements.

[0056] Next, the embodiments of this specification will be described with two target atomic capability programs.

[0057] Optionally, if the target atomic capability program includes a first atomic capability program and a second atomic capability program; based on the preset parameters, editing the video to be edited by using each of the target atomic capability programs in sequence according to the execution order to generate the target video may specifically include: Based on the preset parameters, using the first atomic capability program and the second atomic capability program in sequence to edit the video to be edited to generate the target video.

[0058] In the embodiments of this specification, it is assumed that among the first atomic capability program and the second atomic capability program selected by the user, the first atomic capability program is before and the second atomic capability program is after; or the execution serial number of the first atomic capability program is 1 and the execution serial number of the second atomic capability program is 2; it means that the first atomic capability program is first used to perform a first processing on the video to be edited, and after obtaining a first processing result, the second atomic capability program is used to perform a second processing on the first processing result to obtain the target video.

[0059] In practical applications, the number of target atomic capability programs can be three, four, or more. The target atomic capability programs described above, including the first atomic capability program and the second atomic capability program, are merely illustrative examples for the purpose of explaining the technical principles and are not intended to limit the scope of application of this solution.

[0060] In practical applications, after obtaining the execution order of each of the target atomic capability programs, the target atomic capability programs can be arranged based on the execution order to obtain the target application program; then, based on the preset parameters, the target application program is used to edit the video to be edited, thereby generating the target video.

[0061] During the arrangement of target atomic capability programs based on execution order, intermediate linkers can be used to connect two target atomic capability programs. It should be understood that the intermediate linkers between different target atomic capability programs can differ. For example, the linker between the first and second atomic capability programs may be different from the linker between the first and third atomic capability programs; similarly, when the execution order of two target atomic capability programs is different, the intermediate linkers between them will also be different, such as the linker between the first and second atomic capability programs and the linker between the second and first atomic capability programs.

[0062] In practical applications, due to the highly fragmented nature of video user attention and the fact that platform algorithms prioritize initial interaction data when recommending videos to users (such as the number of likes and viewing time at the beginning of the video), highlight videos can be spliced ​​into the opening sequence to quickly capture attention through the primacy effect and increase the target video's visibility. Based on this, this specification provides an embodiment of a method for generating a target video with a highlight video segment at the beginning using a target atomic capability program.

[0063] Optionally, if the video to be edited includes a first video; the target atomic capability program further includes a third atomic capability program; the first atomic capability program is a video analysis atomic capability program, the second atomic capability program is a video trimming atomic capability program, and the third atomic capability program is a video splicing atomic capability program.

[0064] The step of editing the video to be edited based on the preset parameters by sequentially using the first atomic capability program and the second atomic capability program to generate the target video may specifically include: The video analysis atomic capability program is used to identify the highlight video in the first video and determine the position information of the highlight video in the time dimension; the highlight video is used to reflect the theme of the first video.

[0065] Based on the temporal location information of the highlight video, the first video is cropped using the video cropping atomic capability program to obtain a highlight video segment containing the highlight video.

[0066] The target video is obtained by using the video splicing atomic capability program to splice the highlight video clip with the first video.

[0067] In the embodiments of this specification, the video analysis atomic capability program is a modular program with video content analysis function; the video analysis atomic capability program can be used to identify key frames or segments (i.e., highlight videos) in a video and determine their temporal location information in the video.

[0068] The first video can refer to the original video to be processed, such as insurance promotional video material to be edited. The highlight video can be a captivating segment reflecting the core theme of the first video, such as key shots showcasing the product's protection effects in an insurance promotional video, or a climax in user reviews. The highlight video's positional information in the time dimension can be its time coordinates within the first video, represented as "start time - end time" (e.g., 00:00:10-00:00:15). In the embodiments of this specification, using the video analysis atomic capability program, the highlight video segments within the first video can be identified, and the specific time point of the highlight video within the first video can be determined, for example, locating the highlight video at the position "starting from the 10th second and lasting for 5 seconds."

[0069] In this embodiment, the video cropping atomic capability program is a modular tool for cropping videos. Specifically, the video cropping atomic capability program can crop videos according to a specified time range, resolution, or specific video frames at a specified position. In this embodiment, the video cropping atomic capability program can be used to accurately extract segments containing highlight video from a first video. Continuing with the above embodiment, assuming the highlight video's time coordinates in the first video are 00:00:10-00:00:15, the video cropping atomic capability program can cut out the video segment from seconds 10 to 15 from the first video.

[0070] A video splicing atomic capability program can be a modular program that splices multiple video clips sequentially. In the embodiments of this specification, the target video is obtained by splicing a highlight video clip with a first video using the video splicing atomic capability program. In the embodiments of this specification, the target video is a video with a highlight video clip preceding it, that is, a video whose opening sequence is a highlight video clip.

[0071] In the embodiments of this specification, during the process of generating a video clip with a highlight at the beginning, each target atomic capability program is determined. By flexibly arranging each target atomic capability program, a video production pipeline is quickly built, effectively improving the production efficiency of the target video.

[0072] In addition, if multiple highlight video clips need to be generated as the target video, the user can provide multiple videos to be edited when providing the video to be edited using the terminal device, so as to reuse the video production pipeline built above.

[0073] In practical applications, to help users understand the core information of the preset product more intuitively and clearly, objects used to introduce the preset product can be embedded in the video during the generation of the target video.

[0074] Based on this, if the video to be edited includes a second video and a first object; the target atomic capability program includes an image overlay atomic capability program; the preset parameters include a first time point in which the first object is displayed in the second video; and the step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video can specifically include: Using the image overlay atomic capability program, the first object and the second video are overlaid, so that the first object is overlaid in the second video at the first time point.

[0075] In the embodiments of this specification, the second video and the first object can refer to the source files to be processed. Specifically, the second video and the first object can both be video source files to be edited; alternatively, the second video can be the video source files to be edited, and the first object can be the image source files to be edited.

[0076] Image overlay atomic capabilities are modular programs used to overlay videos onto each other and videos onto images. Based on computer vision and image processing technologies, they can perform pixel-level compositing operations to ultimately output a target video with an overlay effect.

[0077] In the embodiments of this specification, the first time point can be the time point at which the first object and the second video are superimposed; during the playback of the target video obtained using image superposition atomic capabilities, only the second video is played before the first time point; when playback reaches the first time point, the first object and the second video are superimposed and played. In practical applications, preset parameters may also include display duration or the second time point. The display duration and the second time point can be used to reflect the duration of the superimposed playback of the second object and the second video.

[0078] After obtaining the second video and the first object, the superimposed object and the superimposed object in the second video and the first object can be further determined. In practical applications, users can set the first attribute of the second video and / or the second attribute of the first object through the device terminal. Based on the attribute values ​​set by the user, the executing entity can determine the superimposed object and the superimposed object according to the following rules: If a user sets both the first attribute of the second video and the second attribute of the first object, the executing entity can use a preset attribute matching algorithm (such as a priority weight model) to perform semantic parsing and logical comparison of the two attribute values, thereby determining the overlaid object and the overlay object. For example, if the first attribute of the second video is set to "background layer" and the second attribute of the first object is set to "foreground element", then the second video is determined to be the overlaid object, and the first object is determined to be the overlay object.

[0079] If the user configures the attribute value of only one of the second video and the first object, the execution entity can initiate the attribute inference mechanism; for example, if the first attribute of the second video is set to "background layer"; or if the second attribute of the first object is set to "foreground element", then the second video is determined to be the overlay object, and the first object is determined to be the overlay object.

[0080] Additionally, if there are logical conflicts in the attribute values ​​set by the user (such as both the second video and the first object being set to "foreground layer"), the execution entity will trigger the conflict and prompt the user to correct the conflict settings through the interactive interface; alternatively, the system default configuration can be used (such as setting the first input object in the time dimension as the overlaid object).

[0081] In practical applications, to help users understand the core information of the preset product more intuitively and clearly, images used to introduce the preset product can be embedded into the video during the generation of the target video.

[0082] Optionally, if the first object is an image template, the target atomic capability program further includes a text fill atomic capability program; Before performing the image overlay atomic capability procedure on the first object and the second video, the procedure may further include: The script information of a preset product involved in the target video is obtained using the text filling atomic capability program; the script information includes the product information of the preset product.

[0083] The product information of the preset product is obtained by parsing the script information.

[0084] The product information of the preset product is filled into the corresponding position in the image template to obtain the target image; The process of overlaying the first object and the second video using the image overlay atomic capability program can specifically include: Using the image overlay atomic capability program, the target image and the second video are overlaid, so that the target image is overlaid in the second video at the first time point.

[0085] In this embodiment, the image template is a standardized framework or template used to generate images. The database can store various types of image templates. These different templates differ in layout, element arrangement, color scheme, and size. In this embodiment, content can be quickly filled in based on the image template to generate a target image that meets the requirements.

[0086] A text-fill atomic capability program is a modular program that automatically generates images by filling structured information into specified areas of an image template using preset information extraction rules and position mapping relationships. In the embodiments of this specification, the text-fill atomic capability program can be used to extract product information of a preset product from script information, and can also fill the extracted product information into the corresponding positions in the image template.

[0087] Script information can be used to record and describe specific content. Script information can be a structured collection of text information; in this embodiment, the script information may include product information for a pre-defined product, such as the product name, performance parameters, core functions, advantages, price range, usage methods, and after-sales support. Specifically, for insurance products, the script information may include the product name, type, coverage, sum insured, policy period, and product features. For smart devices, the script information may include the smart device model, features, performance parameters, and price range. The above product information can be integrated in a structured format to ensure accurate and clear presentation of key product characteristics. In this embodiment, the script information provides standardized information support for target video production and product explanation.

[0088] In the embodiments of this specification, the first time point can be the time point at which the target image and the second video are superimposed. It is understood that during the playback of the target video, before the first time point, the playback content includes the second video but not the target image; when playback reaches the first time point, the target image and the second video are superimposed. In practical applications, preset parameters may also include display duration or the second time point. The display duration and the second time point represent the duration for which the target image and the second video are superimposed.

[0089] The first time point can be the point in time when the content in the image is introduced; for example, when discussing insurance products, the target image describing the coverage amount can be added at this time point when explaining the coverage discount; when discussing claims cases, the target image about user reviews can be added at this time point. The first time point, as well as the second time point and display duration mentioned above, can be set by the user based on the video content of the second video, or it can be determined by the image overlay atomic capability program or the connection program connected to the image overlay atomic capability. Specifically, the image overlay atomic capability program or the connection program connected to the image overlay atomic capability can determine the time point in the second video that introduces the core elements of the preset product, that is, the time point in the second video that introduces the content in the target image. Based on the start and end times of the time point, the first time point, the second time point, or the first time point or display duration can be determined.

[0090] In this embodiment of the specification, adding images or text descriptions to the target video to explain the preset product can avoid misunderstandings that can easily arise from relying solely on voice explanations. Images and text can transform abstract concepts into concrete visual information, simultaneously stimulating the brain's language and visual systems, thereby improving the retention rate of key information.

[0091] In practical applications, the text-filling atomic ability program can not only fill text information into image templates to generate images, but also fill text into video files. For example, the text-filling atomic ability program can be used to fill subtitle files into video files, thereby generating video files containing subtitles.

[0092] In addition, in the embodiments of this specification, during the process of generating the target video with overlaid images, a production pipeline for generating the target video with overlaid images is quickly built by flexibly arranging the atomic capability program for text filling and the atomic capability program for image overlay, which effectively improves the production efficiency of the target video.

[0093] In addition, if it is necessary to generate a target video with multiple overlaid images, the video production pipeline built above can be reused to quickly generate a target video with multiple overlaid images.

[0094] In practical applications, during the generation of the target video, other video materials can be embedded in the second video to illustrate the core information of the preset product based on the other video materials; or to enrich the visual layers of the target video and improve its presentation effect.

[0095] Optionally, if the first object is a Lottie file, the target atomic capability program further includes a format conversion atomic capability program.

[0096] Before performing the image overlay atomic capability procedure on the first object and the second video, the procedure may further include: The Lottie file is converted using the aforementioned format conversion atomic capability program to obtain a third video, which is a video file in the format obtained from the Lottie file.

[0097] The process of overlaying the first object and the second video using the image overlay atomic capability program can specifically include: Using the image overlay atomic capability program, the third video and the second video are overlaid, so that the third video is overlaid on the second video at the first time point.

[0098] Lottie files are a JSON-based vector animation format. In the embodiments described in this specification, Lottie files can be used to enrich the visual layers of a target video; for example, files with animation effects, such as fireworks. Additionally, Lottie files can also be used to supplement the introduction of a preset product; for example, for insurance products, Lottie files can be files containing information about the insurance product's claims process, product terms and conditions, user application records, etc.; furthermore, for smart devices, Lottie files can be files containing information about the smart device's appearance parameters.

[0099] In this embodiment of the specification, the format conversion atomic capability program is a modular program that can convert Lottie files into video formats. In this embodiment of the specification, the format conversion atomic capability program can be used to parse Lottie files and render the parsed animation data into standard video formats such as MP4 and MOV. The third video is a video file obtained by converting the Lottie file using the format conversion atomic capability program. The third video has universal video format compatibility.

[0100] In the embodiments of this specification, the first time point can be the time point at which the third video and the second video are superimposed; it is understood that during the playback of the target video, before the first time point, the playback content includes the second video but not the third video; the first time point is when the third video and the second video are played superimposed. In practical applications, the preset parameters may also include the display duration or the second time point. The display duration and the second time point represent the duration of the superimposed playback of the third video and the second video.

[0101] The methods for determining the first time point, the second time point, and the display duration in the embodiments described in this specification have been detailed above and will not be repeated here.

[0102] In practical applications, the third video can be integrated into the second video in the form of picture-in-picture, split-screen display, dynamic pop-ups, etc. For example, during a specific segment explaining an insurance product, an animation of the claims process can be embedded in a floating window, or the product terms and conditions analysis and user application records can be displayed in a split-screen format on both sides of the screen, transforming abstract insurance concepts into concrete visual scenes, thereby improving the retention rate of key information about insurance products.

[0103] In practical applications, the text content contained in the Lottie file can be modified to make the content in the third video obtained based on the Lottie file match the preset product.

[0104] Based on this, optionally, before using the format conversion atomic capability program to perform format conversion on the Lottie file to obtain the third video, it may further include: The script information of the preset product involved in the second video is obtained using the format conversion atomic capability program; the script information includes the product information of the preset product.

[0105] The product information of the preset product is obtained by parsing the script information.

[0106] Based on the product information of the preset product, the product information contained in the Lottie file is modified to obtain the target Lottie file.

[0107] The process of using the aforementioned format conversion atomic capability program to convert the Lottie file to obtain the third video may specifically include: The target Lottie file is converted using the aforementioned format conversion atomic capability program to obtain the third video.

[0108] In this embodiment of the specification, script information can be used to record and describe specific content. Script information can be a structured collection of text information; in this embodiment, the script information can include core elements of a pre-defined product, such as the product's name, performance parameters, core functions, advantages, price range, usage methods, and after-sales support. Specifically, for insurance products, script information can include the product's name, type, coverage, sum insured, policy period, and product features. For smart devices, script information can include the smart device's model, features, performance parameters, and price range. The above product information can be integrated in a structured form to ensure accurate and clear presentation of the product's key characteristics.

[0109] In the embodiments described in this specification, the format conversion atomic capability program can identify product information of a preset product contained in structured text information. Specifically, the format conversion atomic capability program can identify product information based on rule and pattern matching methods, or it can identify product information based on natural language processing (NLP) methods. It is understood that the method for identifying product information is related to the format conversion atomic capability program itself, and is not specifically limited here.

[0110] As one implementation method, the format conversion atomic capability program can identify only a portion of the product information contained in the script information; specifically, the format conversion atomic capability program can first determine the type of product information contained in the Lottie file, obtain the corresponding product information from the script information, and use the obtained preset product information to update the corresponding product information in the Lottie file.

[0111] As another implementation, the format conversion atomic capability program can identify all product information contained in the script information. Based on the identification results, the format conversion atomic capability program can filter out the product information types contained in the Lottie file from all product information types. Then, using these filtered preset product information, the corresponding product information in the Lottie file is updated. For example, if the product information types contained in the Lottie file are name, coverage, and sum insured, the format conversion atomic capability program can filter out the corresponding name, coverage, and sum insured information from all product information, and then update the corresponding product information content in the Lottie file.

[0112] The product information contained in the target Lottie file is information about a preset product that needs to be introduced or promoted. In the embodiments of this specification, the format conversion atomic capability program is used to update the product information contained in the script information, so that the Lottie file (third video) can dynamically adjust the displayed content according to the actual preset product information, without the need to manually modify the Lottie file, thereby improving efficiency and ensuring the accuracy and consistency of information display.

[0113] In practical applications, digital human videos can also be generated to promote and introduce pre-defined products.

[0114] Optionally, the video to be edited may include a digital human template; the target atomic capability program may include a digital human video generation atomic capability program; and the preset parameters may include digital human script information, which may include the digital human's dialogue.

[0115] The step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video may specifically include: Based on the digital human script, the audio of the digital human is generated using the digital human video generation atomic capability program.

[0116] The audio is then fused with the digital human to generate the target video of the digital human.

[0117] A digital human is a virtual avatar constructed using computer technology, artificial intelligence, graphics, and other technologies, possessing human appearance, behavioral characteristics, and interactive capabilities. Digital humans can simulate human facial expressions, body movements, and speech. A digital human template refers to a pre-designed digital human model or framework, containing preset parameters and styles for the digital human's appearance, movements, expressions, and speech; in the embodiments of this specification, the database stores various types of digital human templates. The appearances of different digital human templates vary. In the embodiments of this specification, content can be quickly filled into image templates to generate target images that meet the requirements.

[0118] In the embodiments described in this specification, the digital human's dialogue can be in text form.

[0119] The Digital Human Video Generation Atomic Capability Program is a modular program for generating digital human videos.

[0120] In the embodiments of this specification, the digital human video generation atomic capability program can analyze the digital human's dialogue, identify key information (such as product name, coverage terms, digital data, etc.), and mark content that needs to be emphasized (such as interest rate, compensation ratio). It can also analyze the text's emotional tone (such as professional, friendly, stable) using NLP technology to provide emotional parameters for speech synthesis. For example, product introductions typically require a neutral to slightly professional tone. Based on deep learning models (such as Transformer-TTS, VITS), the text-based dialogue can be converted into audio, adjusting the audio's speed, tone, and stress according to the text's semantics. For example, when emphasizing "1 million" in the phrase "This product provides 1 million in critical illness coverage," the tone is raised, and there is an appropriate pause. Simultaneously, corresponding lip-sync parameter sequences can be generated. After generating the audio, phoneme information (such as pronunciation and duration) can be extracted. The phonemes are mapped to key points on the digital human's facial skeleton using a deep learning model, driving lip movements. For example, when pronouncing "b / p / m," the lips close and then open; when pronouncing "a / o / e," the mouth presents different circular or oval shapes. The system selects appropriate gestures and postures from a predefined motion library based on the dialogue. For example, when explaining terms, an open palm gesture indicates "demonstration"; when emphasizing key points, a raised palm gesture indicates a pause. Facial expressions (such as smiling and nodding) are adjusted according to the audio's emotional tone (e.g., professionalism, concern). For example, when saying, "We are always attentive to your protection needs," the digital human simultaneously displays a gentle smile. Audio, lip-sync animation, body language, and background scenes can be integrated to generate the final target video through a rendering engine.

[0121] In practical applications, the audio files used to generate digital human videos can also be extracted from video files using an audio extraction atomic ability program.

[0122] Optionally, if the video to be edited includes a digital human template and a fourth video; the target atomic capability program includes a digital human video generation atomic capability program and an audio extraction atomic capability program; the preset parameters include the initial time and the deadline for audio extraction from the fourth video.

[0123] The step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video may specifically include: Based on the initial time and the cutoff time, the audio files between the initial time and the cutoff time in the fourth video are extracted using the audio extraction atomic capability program.

[0124] The audio file is then fused with the digital human to generate the target video of the digital human.

[0125] The audio extraction atomic capability program is a modular program that can be used to extract audio information from video files. In the embodiments of this specification, the audio of the digital human can also be extracted from existing videos based on the audio extraction atomic capability program.

[0126] In the embodiments described in this specification, the target video of the digital human is generated through an atomic capability program for digital human video generation, achieving low-cost, high-efficiency, and standardized video content production. Compared to traditional live-action filming, digital human videos can be generated at any time, offer a wide variety of video formats, and ensure consistency in the content presented each time.

[0127] In practical applications, the audio extraction atomic capability program can also be combined with other atomic capability programs to process the video to be edited.

[0128] Optionally, if the video to be edited includes a seventh video; the target atomic capability program includes an audio extraction atomic capability program and a text filling atomic capability program.

[0129] The step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video may specifically include: The audio information of the seventh video is extracted using the aforementioned audio extraction atomic capability program; The audio information is extracted using an automatic speech recognition method to obtain a text file of the audio information; based on the text file, subtitles are added to the seventh video using a text filling atomic capability program.

[0130] In the embodiments of this specification, an Automatic Speech Recognition (ASR) method can be used to extract text content from audio. The resulting text file can also contain timestamp information. For example, the text file obtained for the audio "When we insiders buy insurance" is: [{"I": 0.1}, {"we": 0.2}, {"insider": 0.5}, {"industry": 0.6}, {"person": 0.8}, {"buy": 1.0}, {"insurance": 1.1}, {"insurance": 1.3}]. The timestamp information in the text file ensures that the generated subtitles are consistent with the audio.

[0131] The step of "using an automatic speech recognition method to extract text from the audio information to obtain a text file of the audio information" can be executed by a connection program that connects the audio extraction atomic capability program and the text filling atomic capability program, or it can be executed by a separate atomic capability program.

[0132] Optionally, if the video to be edited includes a fifth video and a sixth video; the target atomic capability program includes a video splicing atomic capability program; and the preset parameters are used to indicate the splicing order of the fifth video and the sixth video.

[0133] The step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video may specifically include: Based on the splicing order, the fifth video and the sixth video are spliced ​​using the video splicing atomic capability program to obtain the target video.

[0134] In practical applications, two videos can be spliced ​​together to generate a new video. The fifth and sixth videos can be provided by the user through their terminal device. Specifically, the fifth and sixth videos can be obtained by the user from other terminal devices or servers, or selected by the user from a local database. For example, the fifth video might be obtained from another terminal device or server, while the sixth video might be selected from a local database. In practice, the database can store various types of intro and outro videos, and the sixth video can be either an intro or outro video selected by the user from the database.

[0135] In the embodiments of this specification, the video splicing atomic capability program is a modular program that splices multiple video segments in a splicing order.

[0136] As one implementation method, if the user pre-sets the splicing order of the fifth and sixth videos and specifies that the fifth video comes first and the sixth video comes last, then the video splicing atomic capability program will be used to perform splicing processing on the fifth and sixth videos. The final target video will strictly follow this order and present a playback effect with the fifth video first and the sixth video last.

[0137] As another implementation method, if the user does not set the video splicing order, but the selected sixth video is the intro video, the video splicing atomic capability program will automatically determine the splicing logic according to preset rules: by default, the sixth video with the intro attribute will be placed before the fifth video to conform to the conventional logic of "intro first, main content last" in video production, so as to realize automated sequential arrangement and content splicing.

[0138] Similarly, if the user does not set the video splicing order, but the selected sixth video is the end credits video, the video splicing atomic capability program will place the fifth video before the sixth video based on the industry convention of "main content first, end credits last". Specifically, the video splicing atomic capability program will first identify the "end credits" attribute tag of the sixth video, and then, according to the logical closed-loop requirements of the video content, splice the fifth video as the main content first, and then insert the end credits video.

[0139] This mechanism, through intelligent recognition and rule-based processing of video attributes, automatically completes the sequential arrangement in accordance with industry standards without explicit user intervention, thereby improving video generation efficiency and ensuring the standardization of content structure.

[0140] In practical applications, in order to increase the diversity of the generated target videos, there can be multiple fifth videos and multiple sixth videos.

[0141] Optionally, if there are multiple fifth videos and multiple sixth videos, the step of using the video splicing atomic capability program to splice the fifth and sixth videos based on the splicing order to obtain the target video may specifically include: For any of the fifth videos, based on the splicing order, the video splicing atomic capability program is used to splice each of the sixth videos with each of the fifth videos to obtain multiple target videos; each of the multiple target videos includes a fifth video and a sixth video.

[0142] In the embodiments of this specification, for a fifth video, each of the sixth videos can be spliced ​​together with the fifth video; the same splicing process can be performed on other fifth videos. It is understood that, assuming there are m fifth videos and n sixth videos, the final generated target videos can be m*n.

[0143] In the embodiments of this specification, the target atom capability program is reused to process multiple videos to be edited, which can generate a large number of target videos in a short time and also improve the diversity of the generated target videos.

[0144] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they have not been described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.

[0145] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods.

[0146] Figure 3 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a device for generating video.

[0147] like Figure 3 As shown, the device may include: The receiving module 302 is used to receive the video to be edited for generating the target video.

[0148] The first acquisition module 304 is used to acquire the target atomic capability program selected by the user for editing the video to be edited; an atomic capability program has one atomic capability for editing the video to be edited; different atomic capability programs have different types of atomic capabilities.

[0149] The second acquisition module 306 is used to acquire preset parameters for the atomic capability program.

[0150] The target video generation module 308 is used to edit the video to be edited based on the preset parameters and using the target atomic capability program to generate the target video.

[0151] based on Figure 3 The embodiments of this specification also provide some specific implementation schemes of the method, which are described below.

[0152] Optionally, if there are multiple target atom capability programs, Figure 3 The device may further include: The third acquisition module is used to acquire the execution order of each of the target atomic capability programs.

[0153] The target video generation module 308 may specifically include: The first target video generation unit is used to edit the video to be edited by sequentially using each of the target atomic capability programs according to the preset parameters and the execution order, based on the preset parameters, to generate the target video.

[0154] Optionally, if the target atomic capability program includes a first atomic capability program and a second atomic capability program; the first target video generation unit may specifically include: The first target video generation subunit is used to edit the video to be edited by sequentially using the first atomic capability program and the second atomic capability program based on the preset parameters, thereby generating the target video.

[0155] Optionally, if the video to be edited includes a first video; the target atomic capability program further includes a third atomic capability program; the first atomic capability program is a video analysis atomic capability program, the second atomic capability program is a video trimming atomic capability program, and the third atomic capability program is a video splicing atomic capability program.

[0156] The first target video generation subunit can be specifically used for: The video analysis atomic capability program is used to identify the highlight video in the first video and determine the position information of the highlight video in the time dimension; the highlight video is used to reflect the theme of the first video.

[0157] Based on the temporal location information of the highlight video, the first video is cropped using the video cropping atomic capability program to obtain a highlight video segment containing the highlight video.

[0158] The target video is obtained by using the video splicing atomic capability program to splice the highlight video clip with the first video.

[0159] Optionally, if the video to be edited includes a second video and a first object; the target atomic capability program includes an image overlay atomic capability program; the preset parameters include a first time point in which the first object is displayed in the second video; the target video generation module 308 may specifically include: The second target video generation unit is used to perform superposition processing on the first object and the second video using the image superposition atomic capability program, so that the first object is superimposed on the second video at the first time point.

[0160] Optionally, if the first object is an image template, the target atomic capability program further includes a text fill atomic capability program.

[0161] Figure 3 The device may further include: The fourth acquisition module is used to acquire script information of preset products involved in the target video using the text filling atomic capability program; the script information includes product information of the preset products.

[0162] The first parsing module is used to parse the script information to obtain the product information of the preset product.

[0163] The image generation module is used to fill the product information of the preset product into the corresponding position in the image template to obtain the target image.

[0164] The second target video generation unit can specifically be used for: Using the image overlay atomic capability program, the target image and the second video are overlaid, so that the target image is overlaid in the second video at the first time point.

[0165] Optionally, if the first object is a Lottie file, the target atomic capability program further includes a format conversion atomic capability program.

[0166] Figure 3 The device may further include: The format conversion module is used to perform format conversion on the Lottie file using the format conversion atomic capability program to obtain a third video, wherein the third video is a video file obtained based on the Lottie file.

[0167] The second target video generation unit can specifically be used for: Using the image overlay atomic capability program, the third video and the second video are overlaid, so that the third video is overlaid on the second video at the first time point.

[0168] Optionally, the Figure 3 The device may further include: The fifth acquisition module is used to acquire script information of a preset product involved in the second video using the format conversion atomic capability program; the script information includes product information of the preset product.

[0169] The second parsing module is used to parse the script information to obtain the product information of the preset product.

[0170] The update module is used to modify the product information contained in the Lottie file based on the product information of the preset product, so as to obtain the target Lottie file.

[0171] The format conversion module can be specifically used for: The target Lottie file is converted using the aforementioned format conversion atomic capability program to obtain the third video.

[0172] Optionally, the video to be edited may include a digital human template; the target atomic capability program may include a digital human video generation atomic capability program; and the preset parameters may include digital human script information, which may include the digital human's dialogue.

[0173] The target video generation module 308 can be specifically used for: Based on the digital human script, the audio of the digital human is generated using the digital human video generation atomic capability program.

[0174] The audio is then fused with the digital human to generate the target video of the digital human.

[0175] Optionally, if the video to be edited includes a digital human template and a fourth video; the target atomic capability program includes a digital human video generation atomic capability program and an audio extraction atomic capability program; the preset parameters include the initial time and the deadline for audio extraction from the fourth video.

[0176] The target video generation module 308 can be specifically used for: The audio file between the initial time and the cutoff time in the fourth video is extracted using the audio extraction atomic capability program.

[0177] The audio file is then fused with the digital human to generate the target video of the digital human.

[0178] Optionally, if the video to be edited includes a fifth video and a sixth video; the target atomic capability program includes a video splicing atomic capability program; and the preset parameters are used to indicate the splicing order of the fifth video and the sixth video.

[0179] The target video generation module 308 may specifically include: The third target video generation unit is used to perform splicing processing on the fifth video and the sixth video based on the splicing order and using the video splicing atomic capability program to obtain the target video.

[0180] Optionally, if there are multiple fifth videos, there are multiple sixth videos; the third target video generation unit can specifically be used for: For any of the fifth videos, based on the splicing order, the video splicing atomic capability program is used to splice each of the sixth videos with each of the fifth videos to obtain multiple target videos; each of the multiple target videos includes a fifth video and a sixth video.

[0181] It is understood that the modules mentioned above refer to computer programs or program segments used to perform one or more specific functions. Furthermore, the distinction between these modules does not imply that the actual program code must also be separate.

[0182] The above is an illustrative scheme of a video generation apparatus according to this embodiment. It should be noted that the technical solution of this video generation apparatus and the technical solution of the video generation method described above belong to the same concept. For details not described in detail in the technical solution of the video generation apparatus, please refer to the description of the technical solution of the video generation method described above.

[0183] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.

[0184] Figure 4 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this specification. Figure 4 As shown, device 400 may include: At least one processor 410; and, Memory 430 communicatively connected to the at least one processor; wherein, The memory 430 stores instructions 420 that can be executed by the at least one processor 410, the instructions being executed by the at least one processor 410 to enable the at least one processor 410 to: Receive the video to be edited for generating the target video.

[0185] Obtain the target atomic capability program selected by the user for editing the video to be edited; an atomic capability program has one atomic capability for editing the video to be edited; different atomic capability programs have different types of atomic capabilities.

[0186] Obtain preset parameters for the target atom capability program.

[0187] Based on the preset parameters, the target atomic capability program is used to edit the video to be edited, thereby generating the target video.

[0188] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatuses, devices, and embodiments, since they are basically similar to the method embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method embodiments. The apparatuses, devices, and methods provided in the embodiments of this specification are corresponding to each other, and therefore the apparatuses and devices also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatuses and devices will not be repeated here.

[0189] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0190] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0191] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0192] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0193] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0194] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0195] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0196] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0197] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0198] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0199] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0200] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0201] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0202] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0203] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for generating video, comprising: Receive the video to be edited for generating the target video; Obtain the target atomic capability program selected by the user for editing the video to be edited; An atomic capability program has one atomic capability for editing the video to be edited; different atomic capability programs have different kinds of atomic capabilities; Obtain preset parameters for the target atom capability program; Based on the preset parameters, the target atomic capability program is used to edit the video to be edited, thereby generating the target video.

2. The method for generating video as described in claim 1, wherein if there are multiple target atomic capability programs, the method further includes, before generating the target video: Obtain the execution order of the capability programs for each of the target atoms; The step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video specifically includes: Based on the preset parameters, the target atomic capability programs are used sequentially according to the execution order to edit the video to be edited, thereby generating the target video.

3. The method for generating a video as described in claim 2, wherein the target atomic capability program includes a first atomic capability program and a second atomic capability program; the step of editing the video to be edited sequentially using each of the target atomic capability programs according to the execution order based on the preset parameters to generate the target video specifically includes: Based on the preset parameters, the first atomic capability program and the second atomic capability program are used sequentially to edit the video to be edited, thereby generating the target video.

4. The method for generating video as described in claim 3, wherein the video to be edited includes a first video; the target atomic capability program further includes a third atomic capability program; the first atomic capability program is a video analysis atomic capability program, the second atomic capability program is a video cropping atomic capability program, and the third atomic capability program is a video splicing atomic capability program; The step of editing the video to be edited based on the preset parameters, using the first atomic capability program and the second atomic capability program sequentially, to generate the target video, specifically includes: The video analysis atomic capability program is used to identify the highlight video in the first video and determine the location information of the highlight video in the time dimension; The highlight video is used to reflect the theme of the first video; Based on the positional information of the highlight video in the time dimension, the first video is cropped using the video cropping atomic capability program to obtain a highlight video segment containing the highlight video. The target video is obtained by using the video splicing atomic capability program to splice the highlight video clip with the first video.

5. The method for generating a video as described in claim 1, wherein the video to be edited includes a second video and a first object; the target atomic capability program includes an image overlay atomic capability program; the preset parameters include a first time point in which the first object is displayed in the second video; and the step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video specifically includes: Using the image overlay atomic capability program, the first object and the second video are overlaid, so that the first object is overlaid in the second video at the first time point.

6. The method for generating video as described in claim 5, wherein if the first object is an image template; the target atomic capability program further includes a text filling atomic capability program; Before performing the image overlay atomic capability procedure on the first object and the second video, the method further includes: The text-filling atomic capability program is used to obtain script information of a preset product involved in the target video; the script information includes product information of the preset product. The product information of the preset product is obtained by parsing the script information; The product information of the preset product is filled into the corresponding position in the image template to obtain the target image; The process of using the image overlay atomic capability program to overlay the first object and the second video specifically includes: Using the image overlay atomic capability program, the target image and the second video are overlaid, so that the target image is overlaid in the second video at the first time point.

7. The method for generating video as described in claim 5, wherein if the first object is a Lottie file; the target atomic capability program further includes a format conversion atomic capability program; Before performing the image overlay atomic capability procedure on the first object and the second video, the method further includes: The Lottie file is converted using the aforementioned format conversion atomic capability program to obtain a third video, which is a video format file obtained based on the Lottie file. The process of using the image overlay atomic capability program to overlay the first object and the second video specifically includes: Using the image overlay atomic capability program, the third video and the second video are overlaid, so that the third video is overlaid on the second video at the first time point.

8. The method for generating video as described in claim 7, further comprising, before performing format conversion on the Lottie file using the format conversion atomic capability program to obtain the third video: The script information of the preset product involved in the second video is obtained using the format conversion atomic capability program; The script information includes the product information of the preset product; The product information of the preset product is obtained by parsing the script information; Based on the product information of the preset product, the product information contained in the Lottie file is modified to obtain the target Lottie file; The process of using the aforementioned format conversion atomic capability program to convert the Lottie file to obtain the third video specifically includes: The target Lottie file is converted using the aforementioned format conversion atomic capability program to obtain the third video.

9. The method for generating a video as described in claim 1, wherein the video to be edited includes a digital human template; and the target atomic capability program includes a digital human video generation atomic capability program; The preset parameters include digital human script information, which includes the digital human's lines. The step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video specifically includes: Based on the digital human script, the audio of the digital human is generated using the digital human video generation atomic capability program; The audio is then fused with the digital human to generate the target video of the digital human.

10. The method for generating a video as described in claim 1, wherein the video to be edited includes a digital human template and a fourth video; and the target atomic capability program includes a digital human video generation atomic capability program and an audio extraction atomic capability program; The preset parameters include the initial time and the end time for audio extraction from the fourth video; The step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video specifically includes: The audio extraction atomic capability program is used to extract the audio file between the initial time and the cutoff time in the fourth video; The audio file is then fused with the digital human to generate the target video of the digital human.

11. The method for generating video as described in claim 1, wherein the video to be edited includes a fifth video and a sixth video; and the target atomic capability program includes a video splicing atomic capability program; The preset parameters are used to indicate the splicing order of the fifth video and the sixth video; The step of editing the video to be edited using the target atomic capability program based on the preset parameters to generate the target video specifically includes: Based on the splicing order, the fifth video and the sixth video are spliced ​​using the video splicing atomic capability program to obtain the target video.

12. The method for generating a video as described in claim 11, wherein if there are multiple fifth videos and multiple sixth videos; the step of splicing the fifth videos and the sixth videos using the video splicing atomic capability program based on the splicing order to obtain the target video specifically includes: For any of the fifth videos, based on the splicing order, the video splicing atomic capability program is used to splice each of the sixth videos with each of the fifth videos to obtain multiple target videos; each of the multiple target videos includes a fifth video and a sixth video.

13. An apparatus for generating video, comprising: The receiving module is used to receive the video to be edited for generating the target video; The first acquisition module is used to acquire the target atomic capability program selected by the user for editing the video to be edited; An atomic capability program has one atomic capability for editing the video to be edited; different atomic capability programs have different kinds of atomic capabilities; The second acquisition module is used to acquire preset parameters for the atomic capability program; The target video generation module is used to edit the video to be edited based on the preset parameters and using the target atomic capability program to generate the target video.

14. An apparatus for generating video, comprising: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Receive the video to be edited for generating the target video; Obtain the target atomic capability program selected by the user for editing the video to be edited; an atomic capability program has one atomic capability for editing the video to be edited; different atomic capability programs have different types of atomic capabilities; Obtain preset parameters for the atomic capability program; Based on the preset parameters, the target atomic capability program is used to edit the video to be edited, thereby generating the target video.