Multi-track video editing method and device, electronic equipment and storage medium

By configuring intelligent agent controls within the video editing interface, the system can obtain editing requirement text and automatically adjust the target editing track, thus solving the problems of complex interaction and low efficiency in multi-track video editing scenarios in existing technologies and realizing an automated video editing process.

CN121462829APending Publication Date: 2026-02-03BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511633116.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing AI-based video editing functions cannot achieve integrated, fully automated editing in multi-track video editing scenarios, resulting in complex interaction processes and low video editing efficiency.

Method used

By configuring intelligent agent controls within the video editing interface, the system responds to user operations to obtain editing request text, automatically identifies and adjusts the target editing track, and generates a second video draft, enabling automatic editing in multi-track scenarios.

Benefits of technology

It requires no additional user intervention, reduces interaction complexity, improves video editing efficiency, and achieves automatic recognition and matching in multi-editing track scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121462829A_ABST
    Figure CN121462829A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-track video editing method and device, electronic equipment and a storage medium, and the method comprises the steps: configuring an agent control in a video editing interface, responding to a first operation for the agent control, obtaining a first demand text, determining a target editing track from at least two editing tracks based on the first demand text, and editing the target editing track according to the target editing track. And the configuration content of the target editing track is adjusted, editing of the first video draft is completed, and the second video draft is generated, so that automatic identification and matching of editing tracks based on editing requirements in a multi-editing-track scene are realized, additional intervention of a user is not needed, the complexity of an interaction process is reduced, and the video editing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a multi-track video editing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Currently, with the development of artificial intelligence (AI) technology, AI-based video editing functions are gradually being applied to various video editing software content. For example, in existing technologies, users can generate corresponding videos or images by inputting descriptive words, or process a specific material.

[0003] However, existing AI-based video editing functions can only perform fragmented and unit-based processing operations. When faced with more complex multi-track video editing scenarios, they cannot complete integrated, fully automated editing. Users still need to manually operate to edit videos, resulting in complex interaction processes and low video editing efficiency. Summary of the Invention

[0004] This disclosure provides a multi-track video editing method, apparatus, electronic device, and storage medium to overcome the problems of complex interactive processes and low video editing efficiency.

[0005] In a first aspect, embodiments of this disclosure provide a multi-track video editing method, including:

[0006] A video editing interface for displaying a first video draft is provided, the video editing interface being used to display at least two editing tracks and an intelligent agent control, wherein at least one of the editing tracks is configured with media material constituting the first video draft; in response to a first operation on the intelligent agent control, a first requirement text is obtained, the first requirement text being used to describe the editing requirements for the first video draft; based on the first requirement text, a target editing track among the at least two editing tracks is determined, and the configuration content of the target editing track is adjusted to generate a second video draft.

[0007] Secondly, embodiments of this disclosure provide a multi-track video editing apparatus, including:

[0008] The display unit is used to display the video editing interface of the first video draft. The video editing interface is used to display at least two editing tracks and a smart agent control, wherein at least one of the editing tracks is configured with media materials constituting the first video draft.

[0009] The acquisition unit is configured to, in response to a first operation on the intelligent agent control, acquire a first requirement text, wherein the first requirement text describes the editing requirements for the first video draft.

[0010] The processing unit is configured to determine the target editing track among the at least two editing tracks based on the first requirement text, adjust the configuration content of the target editing track, and generate a second video draft.

[0011] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;

[0012] The memory stores computer-executed instructions;

[0013] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the multitrack video editing method as described in the first aspect and various possible designs of the first aspect.

[0014] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the multi-track video editing method described in the first aspect and various possible designs of the first aspect.

[0015] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the multitrack video editing method as described in the first aspect and various possible designs of the first aspect.

[0016] The multi-track video editing method, apparatus, electronic device, and storage medium provided in this embodiment, through a video editing interface displaying a first video draft, wherein the video editing interface is used to display at least two editing tracks and an intelligent agent control, wherein at least one of the editing tracks is configured with media materials constituting the first video draft; in response to a first operation on the intelligent agent control, a first requirement text is obtained, the first requirement text being used to describe the editing requirements for the first video draft; based on the first requirement text, a target editing track among the at least two editing tracks is determined, and the configuration content of the target editing track is adjusted to generate a second video draft. By configuring an intelligent agent control in the video editing interface, and in response to a first operation on the intelligent agent control, obtaining the first requirement text, and then determining the target editing track from at least two editing tracks based on the first requirement text, and adjusting the configuration content of the target editing track to complete the editing of the first video draft and generate a second video draft, the automatic identification and matching of editing tracks based on editing requirements in a multi-editing track scenario is realized, without the need for additional user intervention, reducing the complexity of the interaction process and improving video editing efficiency. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an application scenario diagram of the multi-track video editing method provided in the embodiments of this disclosure;

[0019] Figure 2 Flowchart of the multi-track video editing method provided in the embodiments of this disclosure Figure 1 ;

[0020] Figure 3 A schematic diagram of a video editing interface provided in an embodiment of this disclosure;

[0021] Figure 4 for Figure 2 A flowchart illustrating the specific implementation of step S103 in the illustrated embodiment;

[0022] Figure 5 for Figure 4 A flowchart illustrating the specific implementation of step S1032 in the illustrated embodiment;

[0023] Figure 6A schematic diagram illustrating the process of determining a target editing track provided in this embodiment of the disclosure;

[0024] Figure 7 Flowchart of the multi-track video editing method provided in the embodiments of this disclosure Figure 2 ;

[0025] Figure 8 A schematic diagram illustrating the display state of an intelligent agent control provided in an embodiment of this disclosure;

[0026] Figure 9 A schematic diagram of a functional panel provided for a disclosed embodiment;

[0027] Figure 10 This is a schematic diagram illustrating a process for generating a second media material, provided in an embodiment of this disclosure.

[0028] Figure 11 A structural block diagram of the multi-track video editing apparatus provided in the embodiments of this disclosure;

[0029] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure;

[0030] Figure 13 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0033] The application scenarios of the embodiments of this disclosure are explained below:

[0034] The multi-track video editing method provided in this disclosure can be applied to applications (APPs) with video editing functions, such as video editing applications and AI assistant applications. More specifically, it can be applied to multi-track video editing application scenarios. The execution subject of this embodiment can be a terminal device running the aforementioned application with video editing functions, a server deploying the server corresponding to the aforementioned application, or other electronic devices that perform similar functions. Specifically, when the execution subject is a terminal device, the terminal device executes the method provided in this embodiment by running the aforementioned application; when the execution subject is a server, the server of the aforementioned application with video editing functions can run partially or entirely on the server, and the method provided in this embodiment is executed on the server side, while the terminal device runs the client of the application. Communication between the server and the terminal device is based on server-client communication, thereby enabling the terminal device to obtain the execution result of the method provided in this embodiment and display it as needed.

[0035] In some embodiments, the terminal device or server can implement the video editing method provided in this disclosure by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be program-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be local applications, i.e., programs that need to be installed in the operating system to run, or mini-programs embedded in any app, i.e., programs that run in a browser environment. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin; the specific implementation can be configured as needed. Furthermore, in implementing the multi-track video editing method provided in this disclosure, the terminal device can execute the method by running computer-executable instructions or computer programs set locally, or by calling computer-executable instructions or computer programs set in an external server. In some embodiments, the server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. Among these, cloud services may be interactive processing services that can be invoked by terminal devices.

[0036] Figure 1This is an application scenario diagram of the multi-track video editing method provided in this disclosure embodiment. Taking a terminal device as an example, the terminal device runs a target application with video editing functions. (Refer to...) Figure 1 As shown, the target application's video editing interface has at least two editing tracks for configuring media materials, such as tracks #1, #2, and #3 shown in the figure. Media materials include, for example, videos, images, music, and text. Based on user selection or automatic media material selection logic, media materials from the media library can be loaded onto one or more editing tracks to generate a video draft. Different editing tracks can be configured with the same or different types of media materials. When the same type of media material is configured on two or more editing tracks, the configured media materials typically correspond to different display layers. Alternatively, an existing video draft can be directly loaded, and corresponding media materials can be configured on the editing tracks based on the draft's content. Based on this, the user can preview the video content generated from the current video draft through a preview window, and adjust the video content by modifying the configuration content of the editing tracks; or export the video as a corresponding video file based on the configuration content of the current editing track.

[0037] In existing technologies, in the aforementioned multi-track video editing application scenarios, the AI-based video editing functions provided by video editing programs can typically only perform fragmented, unit-based processing operations. For example, by selecting a media clip already configured in the editing track and then triggering the corresponding function control, the processing of that media clip or the generation of new media clips can be completed. This process still requires the user to manually perform the operations of selecting media clips and triggering the target function, resulting in a complex interaction process and low video editing efficiency.

[0038] This disclosure provides a multi-track video editing method to solve the above-mentioned problems.

[0039] refer to Figure 2 , Figure 2 Flowchart of the multi-track video editing method provided in the embodiments of this disclosure Figure 1The method of this embodiment can be applied to a terminal device or a server. In one possible implementation, for a terminal device executing the method provided in this embodiment, the terminal device can execute program code deployed locally and / or externally to implement the multi-track video editing method provided in this embodiment. In another possible implementation, a server can be used to deploy functional services based on the multi-track video editing method provided in this embodiment, and the terminal device can access the server and call the corresponding functional services to implement the multi-track video editing method provided in this embodiment. For example, the multi-track video editing method provided in this embodiment includes:

[0040] Step S101: Display the video editing interface of the first video draft. The video editing interface is used to display at least two editing tracks and smart agent controls, wherein at least one editing track is configured with media materials that constitute the first video draft.

[0041] Step S102: In response to the first operation on the smart agent control, obtain the first requirement text, which describes the editing requirements for the first video draft.

[0042] Step S103: Based on the first requirement text, determine the target editing track from at least two editing tracks.

[0043] Step S104: Adjust the configuration of the target editing track to generate a second video draft.

[0044] refer to Figure 1 The illustrated application scenario diagram illustrates the multi-track video editing method provided in this embodiment, using a terminal device as the execution subject. For example, the terminal device runs a target application (e.g., a video editing program) and provides video editing functionality to the user through the target application. Specifically, the target application has a video editing interface that can display at least two editing tracks for configuring media materials. (Refer to...) Figure 1 As shown, for example, track #1 is an editing track for configuring video footage, and track #2 is an editing track for configuring audio footage. In addition, the video editing interface also includes an intelligent agent control. In one possible implementation, this intelligent agent control is in a floating state and can be dragged to any position on the video editing interface, thus avoiding obscuring the effective content within the interface.

[0045] Subsequently, the user can trigger the intelligent agent control by performing a first operation, and obtain the first request text through the intelligent agent control. In one possible implementation, the first operation is, for example, a click operation. When the user performs a click operation on the intelligent agent control, the intelligent agent control triggers the voice input function and collects the user's voice input. Then, the user's voice input is converted into the first request text describing the editing requirements for the first video draft, such as "generate video subtitles" or "remove filler words from the video." The above-mentioned voice input scheme can reduce the steps of display and interaction, improve the response speed of the intelligent agent, and thus improve the efficiency of video editing. In another possible implementation, the first operation consists of at least two steps, that is, the first operation includes a first trigger operation and a first input operation. Accordingly, the specific implementation of obtaining the first request text in step S102 includes:

[0046] Step S1021: In response to the first trigger operation on the smart agent control, display the input window in the video editing interface.

[0047] Step S1022: In response to the first input operation on the input window, obtain the first request text.

[0048] For example, in this embodiment, in response to a first trigger operation on the agent control, an input window is first displayed to receive text input by the user and to display the text content of the agent's reply. Then, in response to a first input operation on the input window, a first required text is obtained, which includes voice input and / or text input. Optionally, the input window is also configured with an input control. After the input control is triggered, the terminal device determines the content input in the input window as the first required text and executes subsequent processing steps through the agent.

[0049] Figure 3 This is a schematic diagram of a video editing interface provided in an embodiment of the present disclosure, such as... Figure 3 As shown, the video editing interface is configured with multiple editing tracks (such as track #1, track #2, and track #3 shown in the figure), as well as an intelligent agent control. Subsequently, in response to the first trigger operation on the intelligent agent control, an input window is displayed in the video editing interface. Then, the user enters the command "Remove filler words from the video" through this input window, enabling the intelligent agent to obtain the first request text. The intelligent agent then executes subsequent processing steps based on the editing requirements described in the first request text (displayed as "Processing..." in the figure).

[0050] Furthermore, the terminal device then processes the aforementioned first requirement text by invoking an intelligent agent (large model), and determines the target editing track from at least two editing tracks based on the first requirement text. One possible implementation is as follows: Figure 4 As shown, the specific implementation of step S103 includes:

[0051] Step S1031: Based on the first requirement text, determine at least one target clip item, which includes the clip target and the corresponding clip function.

[0052] Step S1032: Determine the target editing track based on the target clip item and the media material in the editing track.

[0053] For example, in one possible implementation, the first requirement text includes a first type of requirement text. This first requirement text describes a target editing item for the first video draft. The editing item is an information object composed of two sub-information items: "editing target" and "editing function," which can be expressed as an array, key-value pairs, or text. The target editing item is [video_1, func_1], where video_1 is the identifier of the video material, and func_1 is the identifier of the editing function, such as the "removal of filler words" function. The editing requirement represented by the target editing item is "removing filler words from video_1." Then, based on the target editing item, the editing track whose configuration needs to be changed is located to determine the target editing track. For example, as shown in the example above, if the editing requirement represented by the target editing item is "removing filler words from video_1," then the editing track containing video_1 is determined as the target editing track. If the editing requirement represented by the target editing item is "generating subtitles for video_1," then the subtitle track within the video editing interface used for configuring subtitles is determined as the target editing track. That is, the target editing track is determined based on the "editing function" corresponding to the first requirement text.

[0054] Furthermore, in another possible implementation, the first requirement text includes a second type of requirement text. This second type of requirement text describes at least two target clip items for the first video draft. One possible content of the second type of requirement text is "remove watermarks from video_1 and generate subtitles." Referring to the first type of requirement text described above, the second type of requirement text corresponds to at least two target clip items, such as target clip item A and target clip item B. Target clip item A represents the clip requirement of "removing watermarks from video_1," and target clip item B represents the clip requirement of "generating subtitles for video_1." Accordingly, the target editing tracks corresponding to the second type of requirement text include the editing track containing video_1 corresponding to target clip item A, and the subtitle track corresponding to target clip item A.

[0055] Furthermore, in another possible implementation, the first requirement text includes a third type of requirement text. This third type of requirement text describes the editing effect on the first video draft. One possible content of the third type of requirement text is "beautify video_1". ​​Specifically, compared to the first and second types of requirement texts, the third type of requirement text is a more ambiguous expression of the requirement. The agent understands and reasones about the third type of requirement text, breaks it down into the aforementioned first and second type of requirement texts, and then performs similar processing steps to determine the target editing track corresponding to the third type of requirement text. That is, based on the semantics of the third type of requirement text, at least one first or second type of requirement text is generated. For example, based on the third type of requirement text "beautify video_1", three target editing items are generated: "delete non-highlight scenes from video_1", "add transition effects to video_1", and "add music to video_1". ​​Then, based on the above three target editing items, the corresponding target editing track is determined. The specific implementation process will not be elaborated here; please refer to the introduction for the first and second type of requirement texts.

[0056] Furthermore, in one possible implementation, such as Figure 5 As shown, the specific implementation of step S1032 includes:

[0057] Step S1032-1: Perform semantic understanding on the media materials in each editing track corresponding to the first video draft to obtain the material semantics of the media materials configured in the editing track.

[0058] Step S1032-2: Perform semantic matching between the editing target and the material semantics in the target editing item to obtain the target media material.

[0059] Step S1032-3: Determine the editing track corresponding to the target editing item based on the track position of the editing result of the editing function corresponding to the editing target of the matching target media material.

[0060] For example, after identifying at least one target clip item based on the first requirement text, a mapping relationship between the target clip item and the edit track is established based on the content of the target clip item. Figure 6 This is a schematic diagram illustrating a process for determining a target editing track according to an embodiment of the present disclosure. The following is in conjunction with... Figure 6 For further details on the above process, please refer to [link / reference]. Figure 6As shown, for example, firstly, semantic understanding is performed on the configuration content in each editing track corresponding to the first video draft. For example, editing track #1 is configured with video_1 and video_2, and editing track #2 is configured with audio_1. The agent performs overall semantic understanding on the content in each editing track of the first video draft, determining that the content of video_1 is a "movie clip," the content of video_2 is an "advertisement clip," and the content of audio_1 is a "narration voice clip" of video_1, i.e., the semantic meaning of each media material. The target editing item generated based on the first requirement text is [narration, voice word removal], where "narration" is the editing target, and "voice word removal" is the editing function. The editing target ("narration") is semantically matched with the above material semantics, and the media material that matches it is audio_1. Next, based on the track position of the editing result of the editing function corresponding to the editing target of the target media material, that is, the track position of the editing result of the editing function (e.g., "voice removal") corresponding to the editing target "narration" mentioned above, the editing track corresponding to the target editing item is determined. Specifically, the track position of the editing result corresponding to the editing function "voice removal" in the above example, that is, the editing track #2 where audio_1 is located, is confirmed as the target editing track corresponding to the target editing item.

[0061] Furthermore, after (or before) determining the target editing track, based on the editing requirements described in the first requirement text, the editing target is processed to generate corresponding new or adjusted media materials, which are then placed onto the track to adjust the configuration content of the target editing track, thereby generating a second video draft. In one possible implementation, the editing requirements described in the first requirement text, i.e., the target clip item, are used to call the corresponding functional modules based on preset processing logic to process the corresponding editing target, generate new or adjusted media materials, and place them onto the target editing track.

[0062] Specifically, in one possible implementation, step S104 includes the following:

[0063] Step S1041: Generate an instruction sequence based on the first requirement text. The instruction sequence includes at least one ordered editing function instruction, which is used to call a corresponding editing function.

[0064] Step S1042: Based on the media material configured in the editing track, obtain the function parameters of at least one editing function instruction pair in the instruction sequence.

[0065] Step S1043: Adjust the configuration of the target editing track according to the instruction sequence and function parameters to generate a second video draft.

[0066] For example, firstly, the aforementioned first requirement text generates corresponding target clip items, and based on the implementation logic of the editing functions in the target clip items, the execution order of each target clip item is determined. For example, the target clip items include clip item A and clip item B, where the editing function corresponding to clip item A is "filtering exciting clips" and the editing function corresponding to clip item B is "adding transition effects". Based on the preset processing logic, clip item A should be executed first, followed by clip item B. Alternatively, the processing logic of the above clip items can be judged by calling an intelligent agent to determine the execution order of the above clip items A and B, thereby generating an instruction sequence. The instruction sequence includes at least one ordered editing function instruction, which is a trigger instruction based on the editing function. The editing function instruction is used to call a corresponding editing function.

[0067] Furthermore, after determining the execution order of the editing functions, based on the media materials configured within the editing track, the functional parameters of at least one pair of editing function instructions in the instruction sequence are obtained. Specifically, when the media materials configured within the editing track are different, the functional parameters used when executing the corresponding editing function instructions are dynamically adjusted accordingly. For example, for the "highlight clip selection" editing function, there are two corresponding functional parameters: "number of highlight clips" and "duration of highlight clips." When determining these functional parameters, the terminal device first dynamically determines the corresponding "number of highlight clips" and "duration of highlight clips" based on the media materials configured within the editing track, more specifically, the total duration of the video materials configured within the editing track. The longer the total duration of the video materials, the higher the ratio of "number of highlight clips" to "duration of highlight clips"; conversely, the shorter the total duration of the video materials, the lower the ratio of "number of highlight clips" to "duration of highlight clips," thereby increasing the content coverage of highlight clips while avoiding the problem of excessively long video length after trimming. It is understandable that the above example is just one instance of dynamically adjusting the function parameters of the editing function command pair based on the media materials configured in the editing track. Depending on the purpose and method, there may be other ways to dynamically adjust the function parameters of the editing function command pair based on the media materials configured in the editing track, which will not be listed here.

[0068] Finally, based on the instruction sequence and function parameters, the corresponding editing function instructions in the instruction sequence are executed according to the function parameters dynamically determined in the above steps, thereby improving the execution effect of the editing function. In this embodiment, the editing function instructions are sorted based on the special effects of the editing function, and then the function parameters matching the editing function instructions are dynamically determined based on the media materials configured in the editing track, so that each editing function has a better execution effect and improves the quality of the generated second video draft.

[0069] In this embodiment, a video editing interface for displaying a first video draft is used. This interface displays at least two editing tracks and an intelligent agent control. At least one editing track contains media materials constituting the first video draft. In response to a first operation on the intelligent agent control, a first requirement text is obtained, describing the editing requirements for the first video draft. Based on the first requirement text, a target editing track among the at least two editing tracks is determined, and the configuration of the target editing track is adjusted to generate a second video draft. By configuring an intelligent agent control within the video editing interface, responding to a first operation on the intelligent agent control, obtaining the first requirement text, determining the target editing track from the at least two editing tracks based on the first requirement text, and adjusting the configuration of the target editing track, the editing of the first video draft is completed, generating a second video draft. This achieves automatic identification and matching of editing tracks based on editing requirements in multi-editing track scenarios, eliminating the need for additional user intervention, reducing the complexity of the interaction process, and improving video editing efficiency.

[0070] refer to Figure 7 , Figure 7 Flowchart of the multi-track video editing method provided in the embodiments of this disclosure Figure 2 This embodiment is in Figure 2 Based on the illustrated embodiment, step S103 is further refined, and the multi-track video editing method includes:

[0071] Step S201: Display the video editing interface of the first video draft. The video editing interface is used to display at least two editing tracks and smart agent controls, wherein at least one editing track is configured with media materials that constitute the first video draft.

[0072] Step S202: In response to the second operation, select the first media material within the first editing track.

[0073] Step S203: In response to the first operation on the smart agent control, obtain the first requirement text, which describes the editing requirements for the first video draft.

[0074] Step S204: By calling the intelligent agent to process the first requirement text and the first media material, determine at least one editing function and the corresponding target editing track.

[0075] For example, in this embodiment, after displaying the video editing interface of the first video draft, the user can select at least one first media material within the first editing track by applying a second operation. Then, the aforementioned first media material is used as candidate media material to perform subsequent operations. Specifically, after obtaining the first requirement text, the intelligent agent processes the first requirement text and the first media material. In this case, the intelligent agent prioritizes the first media material as the editing target to construct the editing item and determines the corresponding landing track (target editing track). This results in the editing function determined based on the first media material and the corresponding target editing track.

[0076] In this embodiment, the above steps enable users to actively specify media materials, allowing them to first determine a subset of candidate media materials based on their needs, and then perform subsequent processing steps on these candidate media materials (see details in [link to relevant documentation]). Figure 2 The implementation method in the illustrated embodiment is equivalent to realizing media material selection based on specific needs. Especially when multiple types of media materials are configured within the editing track, it can effectively improve the accuracy of target media material identification, thereby improving the efficiency and accuracy of the intelligent agent video editing function and the target editing track. The specific implementation process of determining the target editing track and editing function (i.e., the target editing item) based on the first requirement text can be found in [reference needed]. Figure 2 The descriptions of the target editing track and editing functions (i.e., target editing items) determined based on all media materials in the illustrated embodiments will not be repeated here.

[0077] Step S205: The intelligent agent calls the corresponding function module of the editing function to edit the first media material and generate the corresponding second media material.

[0078] Step S206: Configure the second media material into the target editing track to generate the second video draft.

[0079] Furthermore, based on the editing functions determined in the above steps, the intelligent system then calls the corresponding functional modules to process the user-selected first media material. This includes processes such as removing filler words and adding transition effects, generating the corresponding processing result, i.e., the second media material. Subsequently, based on the drop track position determined in the previous steps, i.e., the target editing track, the second media material is placed within the target editing track, thus generating the edited second video draft.

[0080] Furthermore, optionally, the above process also includes:

[0081] Step S200A: After processing the first requirement text by calling the agent, configure the agent control to the first display state, and / or after generating the media materials corresponding to each editing function, configure the agent control to the second display state.

[0082] For example, Figure 8 This is a schematic diagram illustrating the display state of an intelligent agent control provided in an embodiment of this disclosure, such as... Figure 8 As shown, exemplarily, after entering the first requirement text in the input window of the intelligent agent control and clicking the "OK" control, the process of generating a second video draft begins, specifically, step S103 or step S204 and subsequent steps in the above embodiment. Afterwards, the intelligent agent control in the video editing interface is configured to a first display state, which is a thumbnail state, that is, the interface including the input window is shrunk to an icon, for example, as shown in the figure, the intelligent agent control is displayed as the icon "Process". During the process of calling the intelligent agent to execute the above steps, the terminal device can respond to other user operation commands in parallel to execute corresponding tasks, achieving the purpose of parallel editing (user manual editing + AI automatic editing). Further, optionally, after the intelligent agent generates the media materials corresponding to each editing function corresponding to the editing requirements expressed in the first requirement text, the intelligent agent control is configured to a second display state, where the second display state may still be a thumbnail state, but the icon style configured for the intelligent agent control changes, for example, as shown in the figure, the intelligent agent control is displayed as the icon "Finish". In another possible second display state, the second display state is an interface state where the smart agent control includes an input window, and the processing result of the above steps, such as the second media material, is displayed in the input window. In one possible implementation, the second display state can be implemented using both of the above methods. For example, after the smart agent control displays the icon "Finish", clicking the icon again switches the smart agent control to an interface state including an input window, and the second media material produced by the above steps is displayed in the input window.

[0083] The steps in this embodiment enable parallel processing for the video editing interface, thereby improving video editing efficiency.

[0084] Step S200B: In the video editing interface, the function panel corresponding to the editing function is displayed. The function panel is used to configure the function parameters corresponding to the editing function.

[0085] Figure 9 A schematic diagram of a functional panel provided for a disclosed embodiment, such as Figure 9As shown, exemplarily, after the terminal device analyzes the first request text by invoking the intelligent agent, it can determine the editing function indicated by the first request text. Then, in response to user operation or automatically, it can display the corresponding function panel within the video editing interface. The function panel contains the function parameters corresponding to the editing function. (Refer to...) Figure 9 As shown, after the agent generates the second media clip and places it on the target editing track, clicking on the second media clip displays the function panel for the corresponding editing function within the video editing interface. For example, if the editing function for the second media clip is "Add Transition Effects," then this function panel will have parameters such as "Number of Transition Shots," "Transition 1 Method," and "Transition 2 Method." The parameter value for "Number of Transition Shots" is 2, indicating that there are two transition effects; the parameter value for "Transition 1 Method" is 001, indicating that the name of the transition effect for Transition 1 is "003"; and the parameter value for "Transition 2 Method" is 002, indicating that the name of the transition effect for Transition 2 is "003." Afterwards, the user can adjust the above function parameters by applying modification operations and regenerate the previously generated second media clip by clicking the "Generate" control.

[0086] Step S200C: Display a preview image or access link of the media material within the smart agent control, configure the media material on the target editing track, and in response to a third operation on the preview image or access link, configure it on the target editing track.

[0087] Figure 10 This is a schematic diagram illustrating a process for generating a second media material according to an embodiment of the present disclosure, such as... Figure 10As shown, after the terminal device generates a second media material based on the first requirement text by invoking an intelligent agent, it determines the complexity of the second media material. More specifically, the complexity is related to the material type; for example, AI-generated video clips (one type of material) have higher complexity, while subtitles (another type of material) have lower complexity. When the complexity of the second media material is high, the terminal device first displays a preview image and access link of the generated media material (e.g., the second media material) within the intelligent agent control. The access link is used to display the preview image of the media material after it is accessed. Referring to the figure, when the material complexity is greater than the complexity threshold (shown as path Y in the figure), the preview image of the second media material is displayed in the dialog window within the intelligent agent control. In this way, the purpose of showing the preview effect to the user is achieved. Afterwards, the user can configure the second media material on the target editing track based on their subjective feeling about the preview image. When the material complexity is less than the complexity threshold (shown as path N in the figure), the second media material is directly configured on the corresponding target editing track. By using the steps described above in this embodiment, it is possible to preview complex editing content in advance, reduce the number of repeated modifications, and improve the efficiency of video editing.

[0088] In this embodiment, the implementation methods of steps S201 and S203 are the same as those of this disclosure. Figure 2 The implementation methods of steps S101 and S103 in the illustrated embodiment are the same, and will not be described in detail here.

[0089] Corresponding to the multi-track video editing method in the above embodiments, Figure 11 This is a structural block diagram of a multi-track video editing device provided in an embodiment of this disclosure. The method described in the above embodiments can be executed by this multi-track video editing device, which can be implemented by software and / or hardware, and can be integrated into an electronic device with certain data processing capabilities. The electronic device may include, but is not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities such as desktop computers and supercomputers.

[0090] For ease of explanation, only the parts relevant to embodiments of this disclosure are shown. (Refer to...) Figure 11 The multi-track video editing device 3 includes:

[0091] Display unit 31 is used to display the video editing interface of the first video draft. The video editing interface is used to display at least two editing tracks and smart agent controls, wherein at least one editing track is configured with media materials constituting the first video draft.

[0092] The acquisition unit 32 is used to acquire first request text in response to a first operation on the smart agent control. The first request text describes the editing requirements for the first video draft.

[0093] The processing unit 33 is used to determine the target editing track among at least two editing tracks based on the first requirement text, and adjust the configuration content of the target editing track to generate a second video draft.

[0094] According to one or more embodiments of this disclosure, the first operation includes a first trigger operation and a first input operation. The acquisition unit 32 is specifically used to: display an input window in the video editing interface in response to the first trigger operation on the smart agent control; and obtain the first demand text in response to the first input operation on the input window.

[0095] According to one or more embodiments of this disclosure, the first requirement text includes at least one of a first type of requirement text, a second type of requirement text, and a third type of requirement text; wherein, the first type of requirement text is used to describe a target clip item for a first video draft; the second type of requirement text is used to describe at least two target clip items for the first video draft; and the third type of requirement text is used to describe the editing effect for the first video draft.

[0096] According to one or more embodiments of this disclosure, when the processing unit 33 determines a target editing track among at least two editing tracks based on the first requirement text, it is specifically configured to: determine at least one target clip item based on the first requirement text, wherein the target clip item includes a clip target and a corresponding clipping function; and determine a target editing track based on the target clip item and media material in the editing track.

[0097] According to one or more embodiments of this disclosure, when the processing unit 33 determines the target editing track based on the target clip item and the media material in the editing track, it is specifically used to: perform semantic understanding on the media material in each editing track corresponding to the first video draft to obtain the material semantics of the media material configured in the editing track; perform semantic matching on the clip target and material semantics in the target clip item to obtain the target media material; and determine the editing track corresponding to the target clip item based on the track position of the clipping result of the clipping function corresponding to the clipping target of the matched target media material.

[0098] According to one or more embodiments of this disclosure, the acquisition unit 32 is further configured to: select a first media material within a first editing track in response to a second operation; when the processing unit 33 determines a target editing track among at least two editing tracks based on the first requirement text, it is specifically configured to: determine a target editing track among at least two editing tracks based on the first requirement text and the first media material; when the processing unit 33 adjusts the configuration content of the target editing track to generate a second video draft, it is specifically configured to: edit the first media material based on the editing requirements represented by the first requirement text to generate a corresponding second media material; configure the second media material into the target editing track to generate a second video draft.

[0099] According to one or more embodiments of this disclosure, when the processing unit 33 adjusts the configuration content of the target editing track to generate a second video draft, it is specifically configured to: generate an instruction sequence based on the first requirement text, the instruction sequence including at least one ordered editing function instruction, the editing function instruction being used to call a corresponding editing function; obtain the function parameters of at least one pair of editing function instructions in the instruction sequence based on the media materials configured in the editing track; and adjust the configuration content of the target editing track based on the instruction sequence and the function parameters to generate a second video draft.

[0100] According to one or more embodiments of this disclosure, when the processing unit 33 determines the target editing track among at least two editing tracks based on the first requirement text, adjusts the configuration content of the target editing track, and generates a second video draft, it is specifically used to: determine at least one editing function and the corresponding target editing track by calling an intelligent agent to process the first requirement text; generate corresponding media materials by calling the functional module corresponding to the editing function through the intelligent agent, and configure the media materials on the target editing track.

[0101] According to one or more embodiments of this disclosure, the processing unit 33 is further configured to: configure the intelligent agent control to a first display state after processing the first request text by invoking the intelligent agent; configure the intelligent agent control to a second display state after generating media materials corresponding to each editing function; display a function panel corresponding to the editing function in the video editing interface, the function panel being used to configure the function parameters corresponding to the editing function; display a preview image or access link of the media material in the intelligent agent control, configure the media material on the target editing track, and in response to a third operation on the preview image or access link, configure it on the target editing track.

[0102] The display unit 31, the acquisition unit 32, and the processing unit 33 are connected in sequence. The multi-track video editing device 3 provided in this embodiment can execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.

[0103] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, such as... Figure 12 As shown, the electronic device 4 includes:

[0104] Processor 41, and memory 42 communicatively connected to processor 41;

[0105] Memory 42 stores instructions executed by the computer;

[0106] The processor 41 executes computer execution instructions stored in the memory 42 to achieve, for example, Figures 2-10 The multi-track video editing method in the illustrated embodiment.

[0107] Optionally, the processor 41 and the memory 42 are connected via a bus 43.

[0108] For relevant instructions, please refer to the corresponding text. Figures 2-10 The relevant descriptions and effects of the steps in the corresponding embodiments are understood, and will not be elaborated on here.

[0109] This disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement this disclosure. Figures 2-10 The multi-track video editing method provided in any of the corresponding embodiments.

[0110] This disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements this disclosure. Figures 2-10 The multi-track video editing method provided in any of the corresponding embodiments.

[0111] To implement the above embodiments, this disclosure also provides an electronic device.

[0112] refer to Figure 13 The diagram illustrates a structural schematic of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 13The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0113] like Figure 13 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0114] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 13 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0115] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.

[0116] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0117] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0118] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0119] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0121] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the specific unit itself.

[0122] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0124] In a first aspect, according to one or more embodiments of this disclosure, a multi-track video editing method is provided, comprising:

[0125] A video editing interface for displaying a first video draft is provided. The video editing interface is used to display at least two editing tracks and an intelligent agent control. At least one editing track is configured with media materials constituting the first video draft. In response to a first operation on the intelligent agent control, a first requirement text is obtained, which describes the editing requirements for the first video draft. Based on the first requirement text, a target editing track among the at least two editing tracks is determined, and the configuration content of the target editing track is adjusted to generate a second video draft.

[0126] According to one or more embodiments of this disclosure, the first operation includes a first trigger operation and a first input operation. In response to the first operation on the smart agent control, obtaining the first demand text includes: in response to the first trigger operation on the smart agent control, displaying an input window on the video editing interface; and in response to the first input operation on the input window, obtaining the first demand text.

[0127] According to one or more embodiments of this disclosure, the first requirement text includes at least one of a first type of requirement text, a second type of requirement text, and a third type of requirement text; wherein, the first type of requirement text is used to describe a target clip item for a first video draft; the second type of requirement text is used to describe at least two target clip items for the first video draft; and the third type of requirement text is used to describe the editing effect for the first video draft.

[0128] According to one or more embodiments of this disclosure, determining a target editing track in at least two editing tracks based on a first requirement text includes: determining at least one target clip item based on the first requirement text, the target clip item including a clip target and a corresponding clip function; and determining a target editing track based on the target clip item and media material in the editing track.

[0129] According to one or more embodiments of this disclosure, determining a target editing track based on a target clip item and media materials in an editing track includes: performing semantic understanding on the media materials in each editing track corresponding to the first video draft to obtain the material semantics of the media materials configured in the editing track; performing semantic matching on the editing target and material semantics in the target clip item to obtain target media materials; and determining the editing track corresponding to the target clip item based on the track position of the editing result of the editing function corresponding to the editing target of the matched target media materials.

[0130] According to one or more embodiments of this disclosure, the method further includes: in response to a second operation, selecting a first media material within a first editing track; determining a target editing track among at least two editing tracks based on a first requirement text, including: determining a target editing track among at least two editing tracks based on the first requirement text and the first media material; adjusting the configuration content of the target editing track to generate a second video draft, including: editing the first media material based on the editing requirements represented by the first requirement text to generate a corresponding second media material; configuring the second media material into the target editing track to generate a second video draft.

[0131] According to one or more embodiments of this disclosure, adjusting the configuration of a target editing track to generate a second video draft includes: generating an instruction sequence based on a first requirement text, the instruction sequence including at least one ordered editing function instruction, the editing function instruction being used to call a corresponding editing function; obtaining function parameters of at least one pair of editing function instructions in the instruction sequence based on media materials configured within the editing track; and adjusting the configuration of the target editing track based on the instruction sequence and function parameters to generate a second video draft.

[0132] According to one or more embodiments of this disclosure, based on a first requirement text, a target editing track is determined from at least two editing tracks, and the configuration content of the target editing track is adjusted to generate a second video draft, including: determining at least one editing function and a corresponding target editing track by calling an intelligent agent to process the first requirement text; generating corresponding media materials by calling the functional module corresponding to the editing function through the intelligent agent, and configuring the media materials on the target editing track.

[0133] According to one or more embodiments of this disclosure, it further includes at least one of the following: after processing the first request text by invoking the agent, configuring the agent control to a first display state; after generating media materials corresponding to each editing function, configuring the agent control to a second display state; in the video editing interface, displaying a function panel corresponding to the editing function, the function panel being used to configure the function parameters corresponding to the editing function; displaying a preview image or access link of the media material within the agent control, configuring the media material on the target editing track, and in response to a third operation on the preview image or access link, configuring it on the target editing track.

[0134] Secondly, according to one or more embodiments of this disclosure, a multi-track video editing apparatus is provided, comprising:

[0135] The display unit is used to display the video editing interface of the first video draft. The video editing interface is used to display at least two editing tracks and smart agent controls, wherein at least one editing track is configured with media materials constituting the first video draft.

[0136] The acquisition unit is used to acquire first request text in response to a first operation on the smart agent control, the first request text describing the editing requirements for the first video draft.

[0137] The processing unit is used to determine the target editing track among at least two editing tracks based on the first requirement text, and to adjust the configuration content of the target editing track to generate a second video draft.

[0138] According to one or more embodiments of this disclosure, the first operation includes a first trigger operation and a first input operation. The acquisition unit is specifically used for: displaying an input window in the video editing interface in response to the first trigger operation on the smart agent control; and obtaining the first request text in response to the first input operation on the input window.

[0139] According to one or more embodiments of this disclosure, the first requirement text includes at least one of a first type of requirement text, a second type of requirement text, and a third type of requirement text; wherein, the first type of requirement text is used to describe a target clip item for a first video draft; the second type of requirement text is used to describe at least two target clip items for the first video draft; and the third type of requirement text is used to describe the editing effect for the first video draft.

[0140] According to one or more embodiments of this disclosure, when the processing unit determines a target editing track among at least two editing tracks based on a first requirement text, it is specifically configured to: determine at least one target clip item based on the first requirement text, wherein the target clip item includes a clip target and a corresponding clipping function; and determine a target editing track based on the target clip item and media material in the editing track.

[0141] According to one or more embodiments of this disclosure, when the processing unit determines the target editing track based on the target clip item and the media material in the editing track, it is specifically used to: perform semantic understanding on the media material in each editing track corresponding to the first video draft to obtain the material semantics of the media material configured in the editing track; perform semantic matching on the clip target and material semantics in the target clip item to obtain the target media material; and determine the editing track corresponding to the target clip item based on the track position of the clipping result of the clipping function corresponding to the clipping target of the matched target media material.

[0142] According to one or more embodiments of this disclosure, the acquisition unit is further configured to: select a first media material within a first editing track in response to a second operation; when the processing unit determines a target editing track among at least two editing tracks based on a first requirement text, it is specifically configured to: determine a target editing track among at least two editing tracks based on the first requirement text and the first media material; when the processing unit adjusts the configuration content of the target editing track to generate a second video draft, it is specifically configured to: edit the first media material based on the editing requirements represented by the first requirement text to generate a corresponding second media material; configure the second media material into the target editing track to generate a second video draft.

[0143] According to one or more embodiments of this disclosure, when the processing unit adjusts the configuration content of the target editing track to generate a second video draft, it is specifically configured to: generate an instruction sequence based on a first requirement text, the instruction sequence including at least one ordered editing function instruction, the editing function instruction being used to call a corresponding editing function; obtain the function parameters of at least one pair of editing function instructions in the instruction sequence based on the media materials configured in the editing track; and adjust the configuration content of the target editing track based on the instruction sequence and the function parameters to generate a second video draft.

[0144] According to one or more embodiments of this disclosure, when the processing unit determines the target editing track among at least two editing tracks based on the first requirement text, adjusts the configuration content of the target editing track, and generates a second video draft, it is specifically used to: determine at least one editing function and the corresponding target editing track by calling an intelligent agent to process the first requirement text; generate corresponding media materials by calling the functional module corresponding to the editing function through the intelligent agent, and configure the media materials on the target editing track.

[0145] According to one or more embodiments of this disclosure, the processing unit is further configured to: configure the agent control to a first display state after processing the first request text by invoking the agent; configure the agent control to a second display state after generating media materials corresponding to each editing function; display a function panel corresponding to the editing function in the video editing interface, the function panel being used to configure the function parameters corresponding to the editing function; display a preview image or access link of the media material in the agent control, configure the media material on the target editing track, and in response to a third operation on the preview image or access link, configure it on the target editing track.

[0146] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;

[0147] The memory stores instructions that the computer executes;

[0148] At least one processor executes computer execution instructions stored in memory, causing at least one processor to perform the multitrack video editing method as described in the first aspect above and various possible designs of the first aspect.

[0149] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions that, when executed by a processor, implement the multitrack video editing method described above as the first aspect and various possible designs of the first aspect.

[0150] Fifthly, according to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the multitrack video editing method as described in the first aspect above and various possible designs of the first aspect.

[0151] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0152] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0153] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A multi-track video editing method, characterized in that, include: A video editing interface for displaying a first video draft is provided, the video editing interface being used to display at least two editing tracks and an intelligent agent control, wherein at least one of the editing tracks is configured with media material constituting the first video draft; In response to a first operation on the smart agent control, a first request text is obtained, the first request text being used to describe the editing requirements for the first video draft; Based on the first requirement text, the target editing track among the at least two editing tracks is determined, and the configuration content of the target editing track is adjusted to generate a second video draft.

2. The method according to claim 1, characterized in that, The first operation includes a first trigger operation and a first input operation. The step of obtaining the first request text in response to the first operation on the smart agent control includes: In response to a first trigger operation on the smart agent control, an input window is displayed on the video editing interface; In response to a first input operation on the input window, a first request text is obtained.

3. The method according to claim 2, characterized in that, The first requirement text includes at least one of a first type of requirement text, a second type of requirement text, and a third type of requirement text; wherein, The first type of requirement text is used to describe a target clip item for the first video draft; The second type of requirement text is used to describe at least two target clip items for the first video draft; The third type of requirement text is used to describe the editing effect on the first video draft.

4. The method according to claim 1, characterized in that, The step of determining the target editing track among the at least two editing tracks based on the first requirement text includes: Based on the first requirement text, at least one target clip item is determined, wherein the target clip item includes a clip target and a corresponding clip function; The target editing track is determined based on the target clip item and the media material in the editing track.

5. The method according to claim 4, characterized in that, The step of determining the target editing track based on the target clip item and the media material in the editing track includes: Semantic understanding is performed on the media materials in each editing track corresponding to the first video draft to obtain the material semantics of the media materials configured in the editing track; Semantic matching is performed between the editing target in the target clip item and the semantics of the material to obtain the target media material; The editing track corresponding to the target editing item is determined based on the track position of the editing result of the editing function corresponding to the editing target of the matching target media material.

6. The method according to claim 1, characterized in that, The method further includes: In response to the second operation, select the first media material within the first editing track; The step of determining the target editing track among the at least two editing tracks based on the first requirement text includes: Based on the first requirement text and the first media material, determine the target editing track among the at least two editing tracks; The step of adjusting the configuration of the target editing track to generate a second video draft includes: Based on the editing requirements represented by the first requirement text, the first media material is edited to generate the corresponding second media material; The second media material is configured into the target editing track to generate a second video draft.

7. The method according to claim 1, characterized in that, The step of adjusting the configuration of the target editing track to generate a second video draft includes: Based on the first requirement text, an instruction sequence is generated, the instruction sequence including at least one ordered editing function instruction, the editing function instruction being used to call a corresponding editing function; Based on the media materials configured within the editing track, the functional parameters of at least one editing function instruction pair in the instruction sequence are obtained; Based on the instruction sequence and the function parameters, the configuration of the target editing track is adjusted to generate a second video draft.

8. The method according to claim 1, characterized in that, The step of determining the target editing track among the at least two editing tracks based on the first requirement text, adjusting the configuration content of the target editing track, and generating a second video draft includes: By invoking the intelligent agent to process the first required text, at least one editing function and the corresponding target editing track are determined. The intelligent agent invokes the corresponding functional modules of the editing function to generate corresponding media materials, and configures the media materials on the target editing track.

9. The method according to claim 8, characterized in that, It also includes at least one of the following: After the first request text is processed by invoking the intelligent agent, the intelligent agent control is configured to a first display state; After generating the media materials corresponding to each of the aforementioned editing functions, the intelligent agent control is configured to a second display state; The video editing interface displays a function panel corresponding to the editing function, which is used to configure the function parameters corresponding to the editing function. Displaying a preview image or access link of the media material within the smart agent control, and configuring the media material on the target editing track, includes: in response to a third operation on the preview image or access link, configuring the media material corresponding to the preview image or access link on the target editing track.

10. A multi-track video editing device, characterized in that, include: The display unit is used to display the video editing interface of the first video draft. The video editing interface is used to display at least two editing tracks and a smart agent control, wherein at least one of the editing tracks is configured with media materials constituting the first video draft. The acquisition unit is configured to, in response to a first operation on the intelligent agent control, acquire a first requirement text, wherein the first requirement text describes the editing requirements for the first video draft. The processing unit is configured to determine the target editing track among the at least two editing tracks based on the first requirement text, adjust the configuration content of the target editing track, and generate a second video draft.

11. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the multitrack video editing method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the multitrack video editing method as described in any one of claims 1 to 9.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the multitrack video editing method as described in any one of claims 1 to 9.