Method and apparatus for performing video editing on basis of dialog, and electronic device and storage medium

Through the dialogue-based video editing method, users can realize complex video editing needs through simple dialogue input, solving the problem of inefficient video editing in the prior art, and realizing a more efficient editing process.

WO2025128001A1PCT designated stage expired Publication Date: 2025-06-19LEMON INC(GB)

Patent Information

Application Number
PCT/SG2024/050786
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing video editing tools are inefficient, and users need to manually select and trigger editing props repeatedly, resulting in a cumbersome editing process.

Method used

The dialogue-based video editing method is adopted to obtain the editing requirements text input by the user through dialogue controls, and call the editing ability item to automatically execute the editing processing steps to achieve the rapid implementation of editing requirements.

Benefits of technology

Improve the efficiency of video editing, and users can achieve complex editing needs through simple dialogue input without paying attention to specific editing steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2024050786_19062025_PF_FP_ABST
    Figure SG2024050786_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present disclosure are a method and apparatus for performing video editing on the basis of a dialog, and an electronic device and a storage medium. The method comprises: after user material is loaded, acquiring, by means of a dialog control, first dialog text that is input by a user, wherein the first dialog text is used for representing an editing requirement for the user material; on the basis of the first dialog text, calling an editing capability item to edit the user material, so as to obtain a first editing result, wherein the editing capability item is used for executing a processing step for realizing the editing requirement; and displaying the first editing result. The first dialog text, which represents the editing requirement and is input by the user, is obtained in a dialog form, and after user intention analysis is performed on the basis of the first dialog text, the corresponding editing capability item is called to process the user material, the first editing result, which meets the editing requirement, is obtained and displayed, and the user does not need to pay attention to the implementation mode of the editing requirement in the process, so that the aim of quickly completing video editing on the basis of the editing requirement is achieved, and the efficiency of video editing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]This application claims priority to Chinese Patent Application No. 202311713285.9, filed on December 12, 2023. The disclosure of the aforementioned Chinese Patent Application is hereby incorporated by reference in its entirety as a part of this application. Technical Field: Embodiments of the present disclosure relate to a conversation-based video editing method, apparatus, electronic device, and storage medium. Background: Video editing typically involves a series of processing steps, such as video segmentation, merging, adding assets, and special effects, and is an essential step in the production of video works. Currently, applications (apps) with video editing capabilities typically provide users with a fixed set of editing tools to meet their video editing needs. However, video editing solutions in related technologies suffer from low video editing efficiency, impacting the user experience. SUMMARY: Embodiments of the present disclosure provide a conversation-based video editing method, apparatus, electronic device, and storage medium to overcome this issue. In a first aspect, embodiments of the present disclosure provide a conversation-based video editing method, comprising: after loading a user material, obtaining, through a conversation control, first conversation text input by a user, where the first conversation text is used to represent an editing requirement for the user material; based on the first conversation text, invoking an editing capability item to edit the user material to obtain a first editing result, where the editing capability item is used to execute processing steps to implement the editing requirement; and displaying the first editing result. In a second aspect, an embodiment of the present disclosure provides a dialogue-based video editing device, comprising: an interaction unit, configured to obtain, through a dialogue control, a first dialogue text input by a user after loading a user material, wherein the first dialogue text is used to represent an editing requirement for the user material; a processing unit, configured to call an editing capability item based on the first dialogue text to edit the user material to obtain a first editing result, wherein the editing capability item is used to execute a processing step for realizing the editing requirement; and a display unit, configured to display the first editing result. In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: a processor and a memory; the memory storing computer-executable instructions; the processor executing the computer-executable instructions stored in the memory, so that the at least one processor executes the dialogue-based video editing method as described in the first aspect and various possible designs of the first aspect.In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, the computer-readable storage medium implements the conversation-based video editing method described in the first aspect and various possible designs of the first aspect. In a fifth aspect, embodiments of the present disclosure provide a computer program product comprising a computer program. When executed by a processor, the computer program implements the conversation-based video editing method described in the first aspect and various possible designs of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below illustrate some embodiments of the present disclosure. Persons skilled in the art can derive other drawings based on these drawings without inventive effort. FIG1 is a diagram of an application scenario of a conversation-based video editing method according to an embodiment of the present disclosure; FIG2 is a flowchart of a first embodiment of a conversation-based video editing method according to an embodiment of the present disclosure; FIG3 is a schematic diagram of a conversation control according to an embodiment of the present disclosure; FIG4 is a flowchart of a specific implementation of step S102 in the embodiment shown in FIG2; FIG5 is a flowchart of a specific implementation of step S103 in the embodiment shown in FIG2; FIG6 is a schematic diagram of displaying a first editing result according to an embodiment of the present disclosure; FIG7 is a schematic diagram of updating a first editing result according to an embodiment of the present disclosure; FIG8 is a flowchart of a second embodiment of a conversation-based video editing method according to an embodiment of the present disclosure; FIG9 is a flowchart of a specific implementation of step S202 in the embodiment shown in FIG8; FIG10 is a schematic diagram of a process for generating a first editing link according to an embodiment of the present disclosure; FIG11 is a flowchart of a specific implementation of step S2022 in the embodiment shown in FIG9; FIG12 is a flowchart of a specific implementation of step S206 in the embodiment shown in FIG8; FIG13 is a flowchart of a specific implementation of step S2062 in the embodiment shown in FIG12; FIG14 is a flowchart of a specific implementation of step S2062C in the embodiment shown in FIG13; Figure 15 is a flowchart of another specific implementation method of step S2062 in the embodiment shown in Figure 12; Figure 16 is a flowchart diagram of the third flow chart of the dialogue-based video editing method provided in an embodiment of the present disclosure; Figure 17 is a structural block diagram of the dialogue-based video editing device provided in an embodiment of the present disclosure; Figure 18 is a structural diagram of an electronic device provided in an embodiment of the present disclosure; and Figure 19 is a hardware structure diagram of the electronic device provided in an embodiment of the present disclosure.DETAILED DESCRIPTION To further clarify the objectives, technical solutions, and advantages of the embodiments of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. It should be noted that the described embodiments represent only a portion of the embodiments of the present disclosure, and are not exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are within the scope of protection of the present disclosure. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in the present disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to grant or deny access. The following explains the application scenarios of the embodiments of the present disclosure. Figure 1 illustrates an application scenario of the conversation-based video editing method provided by the embodiments of the present disclosure. The conversation-based video editing method provided by the embodiments of the present disclosure can be applied to applications with video editing capabilities, more specifically, in video editing scenarios. The execution entity of this embodiment can be a terminal device running the aforementioned application with video editing capabilities, a server deploying the server corresponding to the aforementioned application, or other electronic devices performing similar functions. Referring to Figure 1 , taking a terminal device as an example, after running the aforementioned application with video editing capabilities (hereinafter referred to as the application), the terminal device first loads user material, such as a video, into the video editing interface. The video editing interface also includes several editing controls for video editing, such as Edit Control #1, Edit Control #2, and Edit Control #3 (shown as #1, #2, and #3). Afterwards, while playing the user material within the video editing interface, the terminal device can edit the user material by responding to user triggers on the aforementioned editing controls. For example, special effects, music, and text subtitles can be inserted into the video. For example, in the example shown, the video background of the user material is replaced with a "starry sky" image. This achieves the purpose of video editing. The edited target video can then be output and published to generate the edited video work, completing the video editing and publishing process.In related art applications targeting video editing, the video editing tools provided within applications are typically configured based on editing steps. Specifically, when a tool is triggered, a corresponding editing step is executed. However, user editing needs often require multiple editing steps to achieve. Therefore, users typically first need to use their experience to identify multiple editing tools that meet their editing needs, then trigger the tools and attempt an edit. If the result is poor, meaning the editing needs are not met, the process must be repeated until the desired effect is achieved. Because the quality of the translation from editing needs to editing steps is influenced by user experience, this process typically requires repeated manual selection and triggering of editing tools, as well as setting corresponding execution parameters for the editing tools, resulting in inefficient video editing. The present disclosure provides a conversation-based video editing method to address this issue. Refer to Figure 2, which is a first flow diagram of the conversation-based video editing method provided by the present disclosure. The method of this embodiment can be applied in a terminal device. The dialogue-based video editing method includes the following steps: Step S101: After loading user material, a first dialogue text input by the user is obtained through a dialogue control. The first dialogue text is used to represent the editing requirements for the user material. For example, referring to the application scenario diagram shown in FIG1 , the user material can be media data such as video, image, audio, or mixed media data containing two or more types of material. A loading control is provided within the video editing interface, and the terminal device can select and load the user material in response to user instructions. The specific implementation process is not further described here. Subsequently, a dialogue control for human-computer interaction can be displayed within the video editing interface. The user can enter the first dialogue text into the terminal device through the dialogue control to convey the editing requirements for the user material to the terminal device. Figure 3 is a schematic diagram of a dialog control provided in an embodiment of the present disclosure. As shown in Figure 3 , the dialog control may illustratively include a dialog interface and a text input box. The dialog control is overlaid on a video playback interface displaying user material (in another possible implementation, the video playback interface may not be displayed), facilitating quick user input of a first dialog text. Furthermore, the dialog control's dialog interface may have a certain degree of transparency to reduce the rigidity of the subsequent video playback interface, and this can be configured as needed. Furthermore, the user enters the first dialog text in the text input box. For example, as shown in the figure, the terminal device first displays a welcome or introductory message on the dialog interface, "Hi, how would you like to edit your video?" The user then enters the first dialog text: "Help me add a starry sky background to my video."After the user enters the first dialogue text, the terminal device displays the first dialogue text within the dialogue interface and executes subsequent processing steps based on the first dialogue text. Optionally, the terminal device may also display the current processing step within the dialogue interface, such as "Executing XX" as shown in the figure. In another possible implementation, the user may also input the first dialogue text into the terminal device via voice input. For example, the user may trigger a voice input control and input a spoken voice into the terminal device. The terminal device recognizes the spoken voice and obtains the first dialogue text, and then executes subsequent steps. Step S102: Based on the first dialogue text, the editing capability item is invoked to edit the user material, obtaining a first edit result. The editing capability item is used to execute the processing steps to implement the editing requirements. Exemplarily, after receiving the first dialogue text, the terminal device processes the first dialogue text, predicts the user's editing intent based on the first dialogue text, and then determines and invokes corresponding editing capability items based on the editing intent to process the user material. The editing capability items are used to execute processing steps to achieve the editing requirements, such as adding filters, cropping, merging, and adding background music. By invoking one or more editing capability items to execute corresponding processing steps, the user material is multimodally edited, ultimately achieving the editing requirements. After editing the user material, the resulting editing result that meets the editing requirements represented by the first dialogue text is referred to as the first editing result. Furthermore, in one possible implementation, the first dialogue text includes modality information that indicates the material type used in editing the user material. Exemplarily, as shown in FIG4 , the specific implementation of step S102 includes: Step S1021: Generating corresponding target prompt words based on the modality information. Step S1022: Obtain the editing capability item corresponding to the target material type through the language model and the target prompt word. Step S1023: Edit the user material by calling the editing capability item corresponding to the target material type to obtain the first editing result.Exemplarily, the first dialogue text input by the user includes modal information. The modal information is used to characterize the material type of the material used in the process of editing the user material. For example, the content of the first dialogue text is: "Help me add background music to the video" or "Help me add a voice-over to the video." "Background music" and "Voice-over" are modal information, correspondingly representing the material types of sound material and text material, respectively. Of course, modal information can also represent other material types such as image material and video material. The details are not repeated here. After parsing and extracting the modal information from the first dialogue text, the modal information is converted into corresponding prompt word parameters. Subsequently, the prompt word parameters are combined with a preset prompt word template to generate a target prompt word. The target prompt word can restrict the output content of the language model, thereby guiding the language model to determine the editing capability item corresponding to the target material type. For example, the modal information in the first dialogue text Text_1 characterizes the target material type as "music material." After the target prompt word prompt 1 generated by the modal information in 1 is input into the language model, the editing capability item M_1 is determined and invoked to load a music template for the user material, thereby inserting a piece of background music into the corresponding music track of the user material, thereby obtaining a first editing result. Furthermore, the specific implementation of step S1022 includes: Step S1022A: Generating a target material matching the target material type using the language model and the target prompt word. Step S1022B: Generating an editing capability item based on the target material. Exemplarily, after obtaining the target prompt word and inputting it into the language model, the language model searches for a target material matching the target material type from an existing material library based on the target material type indicated by the target prompt word, or generates a target material matching the target material type using content generation (AIGC) capabilities. The target material is then used to construct or configure an editing capability item, so that the editing capability item edits the user material based on the target material type, thereby obtaining an editing capability item corresponding to the target material type. In this embodiment, by parsing the first dialogue text, Modal information representing the material type is obtained, and corresponding target prompt words are generated and used to edit the user material. This enables content control during the editing process of the user material, ensuring that the resulting first editing result better meets the user's editing needs. Furthermore, optionally, before loading the user material, the following steps are further included: Step S100: Registering at least one candidate editing capability item in an editing capability library.Based on the introduction to editing capability items in the above embodiment, an editing capability item is the ability to edit material, corresponding to one or more processing steps for orderly processing the material. For example, editing capability item #1, which implements "blurring an image," corresponds to (and requires execution of) processing steps P1 and P2. Processing step P1 involves downsampling the image, while processing step P2 adds Gaussian noise to the downsampled image. In one possible implementation, editing capability item #1 must be pre-registered in an editing capability library, which can be located on a terminal device or server. When determining the corresponding editing capability item based on the first conversation text, feature extraction is performed on the first conversation text to obtain content features. Based on these content features, the editing capability item is searched in the editing capability library to obtain the corresponding target editing capability item and execute the item to obtain a first editing result. Furthermore, in one possible implementation, step S102 specifically includes: invoking the target editing capability item based on the first conversation text to edit the user material, obtaining the first editing result. The target editing capability item is obtained by searching the editing capability library. Exemplarily, the specific steps for searching the target editing capability item include: Step S1024: Searching the editing capability library based on keywords in the first conversation text to obtain at least two candidate editing capability items. Step S1025: Searching the at least two candidate editing capability items based on content features of the first conversation text to obtain the target editing capability item. For example, first, by parsing the first dialogue text, keywords in the first dialogue text are obtained, where keywords are text used to represent the main content and key content of the first dialogue text. For example, the content of the first dialogue text is "Help me add a relaxing and lively background music to the video." Among them, the recognition model can identify "relaxed", "lively", and "background music" as keywords. The specific implementation method of keyword recognition is not detailed here, and can be implemented using a pre-trained model. Then, based on the above keywords, the editing capability library is searched to obtain candidate editing capability items that can implement the content indicated by the above keywords. Then, the content features of the first dialogue text are extracted, for example, to obtain a content feature matrix. The sum of the above candidate editing capability items is further screened from the perspective of content features to obtain target editing capability items that meet the editing requirements.Regarding the solutions in the related art, keyword-based searches fail to identify negative phrases. For example, both "lively" and "inactive" can be found based on the keyword "lively," resulting in low accuracy. Content-feature-based searches also suffer from low hit rates and efficiency. Furthermore, dual-dimensional searches using both keywords and content features require the use of a threshold vector consisting of two different search thresholds, which is costly and difficult to achieve with guaranteed search accuracy. In this embodiment, by first searching based on keywords and then based on content features, the inability to identify negative phrases in keyword searches is effectively avoided. Furthermore, the keyword search reduces the amount of data required, thereby improving the efficiency and time consumption of content feature searches. This combined search approach enables accurate and efficient searches for the target editing capability item. Step S103: Displaying the first editing result. For example, after editing the user material by invoking the editing capability item, the first editing result is displayed. Exemplarily, the first editing result includes a target video generated after editing the user material and / or an execution result corresponding to the editing capability item. The target video here can be a draft or temporary file generated after editing the user material, which can be previewed on the video playback page. After further processing of the target video, a video file can be generated. Accordingly, as shown in FIG5 , the specific implementation of step S103 includes: step S1031: displaying the execution result via a dialog control; step S1032: playing the target video within the video playback interface for playing the user material. Steps S1031 and S1032 can be executed either or both; if both are executed, they can be executed synchronously or asynchronously. Furthermore, the target video and the execution result can be displayed simultaneously or separately. Figure 6 is a schematic diagram showing a first editing result according to an embodiment of the present disclosure. Referring to Figure 6 , after a user enters a first dialogue text (e.g., "Help me add a starry sky background to my video"), the terminal device, based on the first dialogue text, completes editing of the user's material and displays the resulting target video on the video playback interface. Furthermore, the terminal device displays the execution process and results on the dialogue interface. For example, as shown in the figure, the following are displayed in sequence: "Executing process step A," "Executing process step B," and "A starry sky background has been added to the video" (the execution result).Furthermore, as shown in the figure, the dialogue interface and the video player interface can be displayed overlapping. In this case, the dialogue interface for human-computer interaction is displayed on top of the video player interface, facilitating the user to enter multiple rounds of dialogue text, thereby improving interaction efficiency. In another implementation, the dialogue interface and the video player interface can also be displayed in split-screen mode, meaning they do not overlap each other. Furthermore, when the video player interface is reduced due to the split-screen display, the size of the target video played in the video player interface can remain unchanged or be proportionally reduced, depending on the needs. Subsequently, based on the current first editing result, the user can further enter new dialogue text (second dialogue text) through the dialogue interface to further edit the target video displayed in the video player interface. Accordingly, in response to the new dialogue text entered by the user, the execution result corresponding to the new dialogue text is displayed in the dialogue interface. Furthermore, in one possible implementation, the execution result includes at least two editing templates corresponding to the editing capability item. After step S1031, the process further includes: step S1033: in response to a user selection operation, setting one of the at least two editing templates as a target editing template. step S1034: editing the target video based on the target editing template to obtain an updated first editing result. In one possible implementation, the execution result includes at least two editing templates corresponding to the editing capability item. Both of the at least two editing templates can meet the user's editing requirements, but different editing modules can achieve different editing effects, thereby obtaining different editing results. In this case, the terminal device may display both of the at least two editing templates in a dialog interface, with one of the at least two editing templates including a currently applied editing board. Subsequently, in response to a user selection operation on the editing interface, one of the editing templates is determined as the target editing template, which is equivalent to specifying an editing effect. The target video is edited based on the target editing template to obtain an updated first editing result, thereby adjusting the first editing result. FIG7 is a schematic diagram of updating a first editing result provided by an embodiment of the present disclosure. Referring to FIG7 , exemplarily, based on the introduction in the previous steps, after editing the user material based on the first dialogue text input by the user (the content is “Help me add a natural theme background”), while the target video is played in the video playback interface, at least two editing templates are displayed in the dialogue interface. As shown in the reference figure, templates temp_1 and temp_2 are displayed in the dialogue interface. Furthermore, the execution result also includes the following text content: “Two starry sky theme background image templates have been found for you.”The current target editing template is the target video 1 currently being played in the video playback interface. This is the video after the background of the user material is replaced based on the template temp_1. Afterwards, the user selects template temp_2. In response to the user's selection of template temp_2, the terminal device determines template temp_2 as the target editing template. The terminal device then reprocesses the user material based on template temp_2 to generate a corresponding first editing result, for example, target video_2, after the user material's background is replaced based on template temp_2. In this embodiment, by displaying multiple editing boards within the dialogue interface and enabling rapid updates of the first editing result based on user selections, video editing efficiency is improved. This process also eliminates the need for the user to re-enter dialogue text or reuse the language model for processing, thereby reducing computing resource overhead. In this embodiment, after loading the user material, the dialogue control obtains first dialogue text entered by the user. The first dialogue text is used to represent the editing requirements for the user material. Based on the first dialogue text, the editing capability item is invoked to edit the user material to obtain a first editing result. The editing capability item is used to execute the processing steps to implement the editing requirements. The first editing result is then displayed. By obtaining the first dialogue text representing the editing requirements entered by the user in the form of a dialogue, After analyzing user intent based on the first conversation text, the corresponding editing capability item is invoked to process the user material, obtaining and displaying a first editing result that meets the editing requirements. During this process, the user does not need to be concerned with how the editing requirements are implemented. This achieves the goal of quickly completing video editing based on the editing requirements, thereby improving video editing efficiency. Referring to Figure 8, Figure 8 is a schematic flow diagram of a conversation-based video editing method provided by an embodiment of the present disclosure. Based on the embodiment shown in Figure 2, this embodiment further refines step S102 and adds a step for achieving continuous video editing through multiple rounds of conversation. This conversation-based video editing method includes the following: Step S201: After loading the user material, a first conversation text input by the user is obtained through a conversation control. The first conversation text is used to represent the editing requirements for the user material. Step S202: The first conversation text is processed using a language model to obtain a first editing chain. The first editing chain represents the editing process for implementing the editing requirements and includes at least two sequentially invoked editing capability items.Exemplarily, the language model includes, for example, a pre-trained Large Language Model (LLM). It has the capabilities of understanding, summarizing, and reasoning about natural language, as well as generating and executing corresponding instructions. The training and usage of the language model are not described in detail here. After the first dialogue text is processed by the language model, the editing requirements represented by the first dialogue text are extracted using the language model's capabilities and then mapped to a corresponding editing process, namely, a first editing link. The first editing link includes at least two sequentially invoked editing capability items. Specifically, for example, if the first dialogue text contains the content "Replace background image for video material," after being processed by the language model, it is mapped to a first editing link List_1. This first editing link List_1 includes several editing capability items, such as Editing Capability Item #1 and Editing Capability Item #2. More specifically, Editing Capability Item #1 is used to achieve target segmentation of video material, Editing Capability Item #2 is used to achieve background image search, and so on. The editing process that implements the editing requirement is executed through this first editing link List_1. Thus, a first editing result that meets the editing requirements is obtained. The first editing link can be an ordered set of editing capability items outputted at one time by the language model, wherein each editing capability item can be called in a serial or parallel manner. For example, as shown in FIG9 , the specific implementation of step S202 includes: Step S2021: Generate target prompt words for the language model based on the first dialogue text. Step S2022: Process the target prompt words using the language model to obtain an initial editing link and dynamic parameters, wherein the initial editing link includes a processing template corresponding to the processing steps constituting the editing process, and the dynamic parameters are template parameters for the processing template. Step S2023: Configure the initial editing link using the dynamic parameters to generate the first editing link. For example, first, the terminal device generates target prompt words for the language model based on the first dialogue text. Prompt words are parameters used to guide the output content of the language model. The specific implementation will not be repeated here. The first dialogue text can be parsed to obtain keywords therein, and then combined with a preset prompt word template to generate the corresponding target prompt words. Thereafter, The target prompt word is processed using a language model to generate an initial editing chain and dynamic parameters. The initial editing chain includes processing templates corresponding to the processing steps that make up the editing process. These templates can be understood as processing templates used to implement specific functions, such as object segmentation, material search, image color adjustment, and background music addition.A processing template may include one or more input parameters for adjusting the specific functions it implements. For example, if the processing template implements a material search function, the input parameters are used to determine the characteristics of the searched material, such as its content, type, and style, thereby enabling accurate material search. Furthermore, dynamic parameters are parameters determined by the target prompt word or the first dialogue text and capable of expressing editing requirements. For example, dynamic parameters include target segmentation margin, material content, and screen brightness. The initial editing chain is then configured based on the dynamic parameters. For example, the target segmentation margin, material content, and screen brightness are used to configure the target segmentation, material search, and screen color adjustment steps in the initial editing chain. This determines the specific implementation of each processing step in the initial editing chain, thereby generating the first editing chain. FIG10 is a schematic diagram of a process for generating a first editing link according to an embodiment of the present disclosure. As shown in FIG10 , illustratively, first, based on the target prompt word and language model, an initial editing link and dynamic parameters are generated. The initial editing link includes multiple processing templates, such as processing template M1, processing template M2, and processing template M3. The processing templates are arranged in an orderly manner, with each processing board corresponding to a processing step. Dynamic parameters include parameter P1 corresponding to processing template M1, parameter P2 corresponding to processing board M2, and parameter P3 corresponding to processing template M3. Parameters P1, P2, and P3 are then assigned to the corresponding processing templates to generate fully functional editing capability items #1, #2, and #3, thereby generating the first editing link. Furthermore, illustratively, as shown in FIG11 , the specific implementation of obtaining the initial editing link in step S2022 includes the following: Step S2022A: Processing the target prompt word using the language model to obtain search information, where the search information represents the search scope corresponding to the editing capability item. Step S2022B: Based on the search information, perform a search for editing capability items to obtain corresponding initial editing capability items. Step S2022C: Based on the initial editing capability items, obtain an initial editing chain. For example, in one possible implementation, the initial editing chain is a collection of processing steps directly generated by the language model based on the target prompt word. This is specifically implemented based on the generation capabilities of a pre-trained language model. This implementation requires the language model to analyze and deduce the target prompt word, possibly combined with user material description information (e.g., text describing the user material content or an image feature matrix), to generate the initial editing chain. This results in a high computational load.In another implementation, the initial editing link may be obtained by searching existing editing links, and the transport volume is relatively small. Specifically, the language model processes the target prompt word to obtain a search scope corresponding to the editing capability items. This search information can specify a search scope or search rule. For example, using the material search function as an example, the search information obtained after processing the target prompt word through the language model limits the types of "theme" to include "travel," "sports," "food," "pets," and "none." Subsequently, when searching for editing capability items related to "theme," only the aforementioned "travel," "sports," "food," "pets," and "none" editing capability item tags are searched. Editing capability item tags not within the search scope indicated by this search information are not searched. This is equivalent to further understanding and optimizing the target prompt word, thereby ensuring the effectiveness and efficiency of editing capability item searches. Subsequently, a search for editing capability items is performed based on the search information, and editing capability items that meet the requirements are determined as initial editing capability items. Furthermore, based on the set of initial editing capability items, they are sorted in an ordered manner (the sorting order of the initial editing capability items is also output by the language model), thereby obtaining an initial editing link. Step S203: The user material is edited by executing the first editing chain to obtain a first editing result. Exemplarily, based on the first editing chain obtained in the above steps, each editing capability item in the first editing chain is sequentially invoked to execute the corresponding processing steps, thereby editing the user material and obtaining the first editing result. Depending on the specific implementation of the first editing chain, in one possible implementation, the editing capability items in the first editing chain can be invoked in parallel by the terminal device, while each processing step is executed asynchronously (non-serially), and the processing results are merged to obtain the first editing result. This implementation allows for asynchronous execution of each editing capability item in the first editing chain, improving editing efficiency. In another possible implementation, the first editing chain is executed serially by the language model. In this case, the first editing chain is the initial editing chain referred to in the above embodiment. The first editing chain generated by the language model using the target prompt word does not include the input parameters of each editing capability item. Therefore, during the execution of this initial editing chain, based on the initial parameters, the first editing capability item is executed from the first editing capability item. The editing capability items in the first editing chain (initial editing chain) are called sequentially, and the input parameters (dynamic parameters) of the next editing capability item are determined based on the execution result of each editing capability item. This process is completed after all editing capability items in the first editing chain have been called.The above approach enables a more complex video editing process and more accurate fulfillment of editing requirements, improving the accuracy of the first editing result. Step S204: Display the first editing result. Step S205: If content input into the dialogue control is detected, the dialogue control obtains a second dialogue text input by the user. The second dialogue text is used to represent the editing requirements for the first editing result, which includes the target video generated after editing the user material. Step S206: Based on the second dialogue text, corresponding processing steps are executed to obtain and display a second editing result corresponding to the second dialogue text. For example, based on the description in the previous embodiment, after obtaining the first editing result, it can be displayed through the video playback interface and the dialogue control, thereby achieving the purpose of previewing the video editing results and interactive dialogue during the editing process. Afterward, if content input into the dialogue control is detected, i.e., the user enters a new dialogue via the dialogue control, the terminal device further obtains a second dialogue text representing the editing requirements for the first editing result. Based on the second dialogue text, the terminal device then continues to perform corresponding processing steps on the first editing result. These processing steps include at least one of the following: invoking at least one editing capability item to edit the target video; or revoking the editing effects corresponding to the at least one editing capability item. Specifically, the terminal device further edits the target video generated after editing the user material to obtain a second editing result, further satisfying the user's editing requirements. In one possible implementation, the terminal device performs corresponding processing steps based on the second dialogue text to obtain a second editing result corresponding to the second dialogue text. This can be similar to the process for obtaining the first editing result in the embodiment shown in FIG. 2 , namely, using the target video in the first editing result as the user material and the second dialogue text as the first dialogue text, and repeating the steps in the embodiment shown in FIG. 2 to obtain a new editing result, i.e., the second editing result. The specific implementation process will not be repeated here; please refer to the relevant descriptions of the embodiments corresponding to FIG. 2-11. In one possible implementation, the terminal device performs corresponding processing steps based on the second dialogue text to cancel the editing effect achieved based on the editing capability item in the first editing result. For example, if the content of the second dialogue text is "Undo the video background image just added," the terminal device uses the voice model and combines the previous context information (the first dialogue text) to cancel the background image added to the user material based on the first dialogue text and restore it to the original background image of the user material.This process can be implemented through a preset retraction process (retraction editing link), the implementation of which is not limited. Its specific execution process is similar to that of the first editing link and will not be further described. Furthermore, after step S206 is completed, the process can return to step S205 to continue detecting the dialogue control, thereby implementing a continuous video editing process based on multiple rounds of dialogue. During this continuous editing process, users can implement continuous video editing based on their editing needs solely through dialogue, without having to pay attention to or operate specific processing steps, thereby greatly improving video editing efficiency. Furthermore, in the aforementioned multi-round continuous editing scenario, each editing process and each edit involves caching and transmitting editing data (i.e., draft data). Since the models and algorithms used in this editing process are not necessarily all provided by the terminal device, data exchange between the terminal device and the external cloud server may cause delays, reduced real-time performance, and excessive server load. To address the above issues, in this embodiment, after obtaining the first editing result, the following further steps are included: Step S200: Storing the first editing result as a corresponding local universal draft. A draft refers to the editing data generated during the video editing process. Saving and loading drafts enables breakpoint editing during the video editing process. The drafts involved in the video editing process will not be discussed further here. The universal draft in this embodiment refers to a draft that is synchronized across multiple terminals (cloud or multiple terminals) based on a unified draft protocol, such as a non-linear editor (NLE) draft. The draft protocol for implementing universal drafts is typically developed by the developer of the application (with video editing functionality), and its specific implementation is not limited here. Furthermore, the local universal draft in this embodiment refers to a universal draft stored locally on the terminal device. Based on the introduction of the previous steps, after creating the editing project, the terminal device will create a blank local universal draft, then load the user material and edit it based on the steps of the previous embodiment. After obtaining the corresponding first editing result, or in the process of generating the first editing result, the terminal device saves the editing data generated during the editing process to the blank local universal draft, thereby generating a local universal draft containing the editing data.Accordingly, in the previously described embodiment, during the process of editing the user material based on the first dialogue text and obtaining the first edited result, each processing step can also be implemented based on the edit data stored in the local universal draft. For example, data describing the first edit link, edit capability items, edit templates, and dynamic parameters are all stored as edit data in the local universal draft, thereby utilizing the local universal draft to complete the video editing process. For specific implementations, please refer to the description of the previous embodiment. Accordingly, as shown in FIG12 , the specific implementation of step S206 includes: Step S2061: Obtaining a second edit link corresponding to the second dialogue text. Step S2062: Processing the local universal draft based on the second edit link to obtain an updated local universal draft. Step S2063: Generating the corresponding second edited result based on the local universal draft. For example, after obtaining the second dialogue text, and in the process of continuously editing the first edit result based on the editing requirements represented by the second dialogue text, a second editing link corresponding to the second dialogue text is first obtained. The specific implementation method is similar to the implementation method for obtaining the first editing link corresponding to the first dialogue text in the previous embodiment and will not be repeated here. The second editing link is then used to process the local universal draft to obtain an updated local universal draft. The second editing link includes multiple editing capability items. Since the editing data in the local universal draft includes at least the target video (data), the editing capability items are invoked to process the local universal draft. This is similar to the implementation method for invoking editing capability items to process user materials in the steps of the previous embodiment and will not be repeated here. Afterwards, based on the updated local universal draft, that is, the editing data in the local universal draft after the second editing link is executed, a second editing result is obtained, such as the target video after further editing (the updated target video). In this embodiment, by storing the universal draft generated during multiple rounds of editing locally on the terminal device to form a local universal draft, and completing the local universal draft to the second editing result locally on the terminal device, the resource overhead of storing the universal draft on the external server is minimized, thereby reducing the occupation of network resources and server resources, reducing latency, and improving data processing efficiency. Optionally, after step S2061, the step of determining the target execution end based on the complexity information is also included. Accordingly, in the subsequent step S2062, the second editing result is generated based on the target execution end. Specifically, in one possible implementation, as shown in Figure 13, the specific implementation method of step S2062 includes: Step S2062A: Obtaining complexity information corresponding to the second editing link.Step S2062B: Based on the complexity information, a target execution end is determined, which may include the cloud or a terminal. Step S2062C: The local general draft is placed on the target execution end and the second editing chain is executed to obtain a local general draft. Exemplarily, the complexity information refers to the complexity of implementing the second editing chain, which can be assessed using one or more indicators including the execution time of the second editing chain, the number of editing capability items in the second editing chain, and the complexity of the editing capability items. Subsequently, based on the complexity information, a target execution end is determined, i.e., whether the second processing chain is executed on the cloud or the terminal. More specifically, if the complexity represented by the complexity information is greater than or equal to a preset complexity threshold, indicating a more complex editing process, the corresponding target execution end is determined to be the cloud, i.e., processing is performed on the cloud server. On the other hand, if the complexity is less than or equal to the preset complexity threshold, indicating a simpler editing process, the corresponding target execution end is determined to be the terminal, i.e., processing is performed on the terminal device. Furthermore, since the local general draft typically resides locally on the terminal device, when the target execution end is the cloud, the local general draft is sent to the cloud server (cloud). After cloud processing, the result (i.e., the updated local general draft) is obtained and stored locally on the terminal device. When the target execution end is a terminal, the second editing chain is directly executed using the terminal device's capabilities to obtain the processing result. In this embodiment, the target execution end is determined based on the complexity information corresponding to the second editing chain to be processed, and the second editing chain is then processed at the target execution end. This fully utilizes the computing power of the terminal device, avoids excessive use of cloud data processing, and avoids straining the cloud server's network and computing resources, thereby improving processing efficiency. In one possible implementation, as shown in Figure 14, the specific implementation of step S2062C includes: Step S2062C-1: Obtaining the execution priority corresponding to each editing capability item in the second editing chain. Step S2062C-2: Based on the execution priority, the corresponding editing capability items are sequentially invoked to iteratively process the local general draft to obtain the processing result. Step S2062C-3: Based on the processing result, an updated local general draft is obtained. Furthermore, in one possible implementation, the processing steps corresponding to different editing capability items can be sorted based on execution priority, thereby improving the execution efficiency of the editing link.Specifically, editing capabilities can be categorized as first-category editing capabilities (hereinafter referred to as "first-category editing capabilities"), such as setting image saturation, super-resolution, and beautification, and non-pixel-level image processing capabilities (hereinafter referred to as "second-category editing capabilities"), such as highlighting, slicing, and aspect ratio adjustment. Second-category editing capabilities are assigned a higher execution priority, ensuring they are executed before first-category video editing capabilities (pixel-level AIGC video modification). Second-category editing capabilities typically involve selecting and cropping existing video / image material, which does not compromise the local integrity of the image and therefore does not affect AIGC applications. Furthermore, these temporal and spatial content cropping can reduce the amount of data AIGC must process, improving processing efficiency. First-category editing capabilities, on the other hand, are assigned a lower execution priority, ensuring they are executed after pixel-level AIGC video modification. This is because these processes require processing the entire image, necessitating global processing of the AIGC-modified content. Then, based on the execution priority corresponding to each editing capability item, the corresponding editing capability item is sequentially called to iteratively process the local general draft, obtaining corresponding processing results. Finally, the local general draft is updated based on the processing results to obtain an updated local general draft. Furthermore, in one possible implementation, the specific implementation of step S2062C-3 includes: obtaining processing results corresponding to at least two second editing links. The at least two processing results are merged on the terminal to obtain an updated local general draft. Exemplarily, when asynchronous processing is performed in the second editing link, that is, when the editing capability items can be processed asynchronously and in parallel, the processing results generated by each editing capability item are merged on the terminal device. For example, the second editing link includes asynchronously executed editing capability item #1 and editing capability item #2, where editing capability item #1 is used to add an editing element, such as a sticker, to a video; and editing capability item #2 is used to add background music to the video. After both are executed, the processing results corresponding to editing capability item #1 and editing capability item #2 are merged. The terminal device then performs a fusion process, resulting in a local, universal draft (video) containing both stickers and background music. In this embodiment, by merging the processing results corresponding to each editing capability item on the terminal device, cloud computing resources are conserved, reducing load and improving the overall processing efficiency and stability of video editing.In this embodiment, by obtaining the execution priorities corresponding to different editing capability items, the second-category editing capability items are prioritized and the first-category editing capability items are processed later. This improves processing efficiency across multiple processing steps and, based on the resulting editing chain, further enhances the efficiency of the editing chain. It is understood that the solution of executing the second editing chain based on execution priority in this embodiment can also be applied to the first editing chain corresponding to the first dialogue text in the embodiment shown in Figure 2 . The specific implementation steps are not further described. The following is a flowchart of another implementation of step S2062. As shown in Figure 15 , the specific implementation of step S2062 includes: Step S2062-1: Obtaining complexity information corresponding to the second editing chain. Step S2062-2: Determining a target execution terminal based on the complexity information, which may include the cloud or a terminal. Step S2062-3: Placing the local general draft on the target execution terminal. Step S2062-4: Obtaining the execution priority corresponding to each editing capability item in the second editing chain. Step S2062-5: Based on the execution priority, the corresponding editing capability items are sequentially invoked to iteratively process the local universal draft, obtaining the processing results of the second editing chain. Step S2062-6: The processing results of multiple second editing chains are merged to obtain an updated local universal draft. For example, in another possible implementation, in a scenario where a video is continuously edited based on multiple rounds of conversation, each conversation text generates a corresponding second editing chain, and a processing result is obtained based on the second editing chain. Subsequently, when the next conversation text (e.g., the second conversation text in this embodiment) is received, the processing result corresponding to the next conversation text is merged with the processing result corresponding to the previous conversation text, thereby updating the local universal draft and obtaining an updated local universal draft. This implements conversation-based continuous video editing. In this embodiment, the implementation of displaying the second editing results in steps S201 and S206 is the same as the implementation of steps S101 and S103 in the embodiment shown in FIG. 2 of this disclosure, and will not be further described here. Referring to Figure 16 , Figure 16 is a third flow diagram of a conversation-based video editing method provided in an embodiment of the present disclosure. This embodiment further refines step S102 based on the embodiment shown in Figure 2 . The solution provided in this embodiment can also be used in conjunction with the embodiment shown in Figure 8 . The conversation-based video editing method provided in this embodiment includes the following steps: Step S301: After loading user material, obtaining a first conversation text input by the user via a conversation control. The first conversation text is used to represent editing requirements for the user material.Step S302: Obtain descriptive information of the user material. The descriptive information is used to characterize the content characteristics of the user material. Step S303: Process the first conversation text and the descriptive information using a language model to obtain a first editing link. For example, in this embodiment, in addition to obtaining the first conversation text, the terminal device also obtains descriptive information characterizing the content characteristics of the user material. Specifically, the descriptive information may be a content description text of the user material, more specifically, a title or content summary of the user material. This descriptive information may be obtained from the user material file or generated by performing content recognition on the user material based on a model, and this is not limited here. Alternatively, the descriptive information may be an image characterizing the content characteristics of the user material, such as a pixel matrix or a feature matrix, and the specific configuration can be as needed. Furthermore, the first conversation text and the descriptive information are input as input parameters to a language model for processing to obtain the first editing link. The specific implementation process can refer to the implementation process of outputting the first editing link using a speech model in the embodiment shown in FIG8 , and will not be further described here. In this embodiment, based on the first conversation text, descriptive information of the user material is further obtained. Based on the content characteristics of the user material represented by the descriptive information, the obtained first editing link is further restricted. This allows the generation of the obtained first editing link to take into account the content characteristics of the user material, thereby obtaining an editing result that better matches the editing requirements. Exemplarily, the specific implementation of step S303 includes: Step S3031: Searching a cache module based on the first conversation text and the descriptive information to obtain a search result. The cache module stores at least one set of corresponding model input information and model output information. Step S3032: Based on the search result, obtaining the first editing link using the model output information stored in the cache module, or generating the first editing link using a language model. Furthermore, the cache module is a space for storing pairs of [model input information - model output information]. The model input information includes historical text input to a language model and, optionally, historical description information corresponding to the input language model. The model output information includes search information or historical editing links corresponding to the historical text. Historical text refers to the conversation text previously entered through the dialog control, and historical description information refers to the description information previously obtained based on historical user materials. Historical editing links refer to the search information or editing links generated by previously processing the conversation text based on the language model. The search information represents the search scope corresponding to the editing capability item. The specific meaning of the search information and the method for generating editing links based on the search information have been described in previous embodiments and will not be repeated here.Furthermore, after obtaining the first conversation text, the terminal device may first search the preset cache module. If a hit is found, subsequent steps are directly executed based on the historical hit results, thereby avoiding the need to regenerate the edit chain and improving processing efficiency. Specifically, in an exemplary embodiment, based on the search results, if the first conversation text matches the corresponding model input information, the first edit chain is determined based on the model output information corresponding to the model input information, i.e., the first edit chain is determined based on historical data. If the first conversation text does not match the corresponding model input information, the first conversation text is processed using a language model to obtain the first edit chain, i.e., a new first edit chain is regenerated. Furthermore, when the model input information includes the first conversation text and description information of the user material, the search results include a first sub-search result and a second sub-search result. The specific implementation of step S3031 includes: Step S3031A: Obtaining a first search threshold corresponding to the first conversation text and a second search threshold corresponding to the description information, wherein the first search threshold is greater than the second search threshold. Step S3031B: Based on the first search threshold and the cache module, the first conversation text is searched to obtain a first sub-search result. Step S3031C: Based on the second search threshold and the cache module, the description information is searched to obtain a second sub-search result. Exemplarily, when the model input information includes the first conversation text and the description information of the user material, searches are performed based on the first conversation text and the description information of the user material using different search thresholds. The first search threshold corresponding to the first conversation text is higher than the second search threshold corresponding to the description information. That is, the first conversation text describing the editing requirements corresponds to a higher search threshold, equivalent to stricter search conditions, while the description information corresponds to a lower search threshold, equivalent to looser search conditions, thereby improving the overall hit rate. Furthermore, after obtaining the first and second sub-search results, the first edit link is obtained based on the model output information, or the first edit link is generated using the language model, based on the first and second sub-search results, respectively. Specifically, if both the first and second sub-search results are hits, the corresponding model output information is obtained, and then a first edit link is generated based on the historical edit links or historical search results in the model output information. Step S304: The user material is edited by executing the first edit link to obtain a first edit result. Step S305: The first edit result is displayed. In this embodiment, the implementation of steps S301, S304, and S305 is the same as the corresponding steps in the embodiments shown in Figures 2 and 8 of this disclosure, and will not be further described here.Corresponding to the conversation-based video editing method described in the preceding embodiment, FIG17 is a block diagram of a conversation-based video editing apparatus provided in an embodiment of the present disclosure. For ease of illustration, only portions relevant to the present embodiment are shown. Referring to FIG17 , the conversation-based video editing apparatus 4 includes: an interaction unit 41, configured to, after loading a user material, obtain a first conversation text input by the user via a conversation control. The first conversation text represents an editing requirement for the user material; a processing unit 42, configured to, based on the first conversation text, invoke editing capability items to edit the user material, obtaining a first editing result. The editing capability items are configured to execute processing steps to implement the editing requirement; and a display unit 43, configured to display the first editing result. In one embodiment of the present disclosure, the first editing result includes a target video generated after editing the user material and / or an execution result corresponding to the editing capability items. The display unit 43 is specifically configured to perform at least one of the following: display the execution result via a conversation control; or play the target video within a video playback interface for playing the user material. In one embodiment of the present disclosure, the execution result includes at least two editing templates corresponding to the editing capability item. After displaying the editing result via the dialog control, the process further includes: an interaction unit 41 further configured to, in response to a user selection operation, set one of the at least two editing templates as a target editing template; and a processing unit 42 further configured to edit the target video based on the target editing template to obtain an updated first editing result. In one embodiment of the present disclosure, before loading the user material, the processing unit 42 further configured to register at least one candidate editing capability item with an editing capability library. When invoking the editing capability item to edit the user material based on the first dialog text to obtain the first editing result, the processing unit 42 specifically configures: searching the editing capability library based on keywords in the first dialog text to obtain at least two candidate editing capability items; searching the at least two candidate editing capability items based on content features of the first dialog text to obtain a target editing capability item; and invoking the target editing capability item to edit the user material to obtain the first editing result. In one embodiment of the present disclosure, the processing unit 42 is specifically configured to: process the first dialogue text using a language model to obtain a first editing link, where the first editing link is used to represent an editing process for implementing editing requirements and includes at least two editing capability items called in an ordered manner; and edit the user material by executing the first editing link to obtain a first editing result.In one embodiment of the present disclosure, when processing the first dialogue text using a language model to obtain a first editing link, the processing unit 42 is specifically configured to: generate a target prompt word for the language model based on the first dialogue text; process the target prompt word using the language model to obtain an initial editing link and dynamic parameters, wherein the initial editing link includes a processing template corresponding to the processing steps constituting the editing process, and the dynamic parameters are template parameters for the processing template; and configure the initial editing link using the dynamic parameters to generate the first editing link. In one embodiment of the present disclosure, when processing the target prompt word using the language model to obtain the initial editing link, the processing unit 42 is specifically configured to: process the target prompt word using the language model to obtain search information representing a search range corresponding to an editing capability item; perform an editing capability item search based on the search information to obtain a corresponding initial editing capability item; and obtain the initial editing link based on the initial editing capability item. In one embodiment of the present disclosure, processing unit 42 is further configured to: obtain descriptive information of the user material, where the descriptive information is used to characterize content features of the user material. When processing the first conversation text using a language model to obtain the first edit link, processing unit 42 is specifically configured to: process the first conversation text and the descriptive information using the language model to obtain the first edit link. In one embodiment of the present disclosure, processing unit 42 is specifically configured to: search a cache module based on the first conversation text to obtain a search result, where the cache module stores at least one set of corresponding model input information and model output information. Based on the search result, if the first conversation text matches the corresponding model input information, determining the first edit link as the first edit link based on the model output information corresponding to the model input information; and if the first conversation text does not match the corresponding model input information, processing the first conversation text using the language model to obtain the first edit link, where the model input information includes historical text input into the language model, and the model output information includes search information or historical edit links corresponding to the historical text. In one embodiment of the present disclosure, model input information includes a first conversation text and description information of a user material, and the retrieval results include a first sub-search result and a second sub-search result. Processing unit 42 searches a cache module based on the first conversation text to obtain the retrieval results, including: obtaining a first retrieval threshold corresponding to the first conversation text and a second retrieval threshold corresponding to the description information, wherein the first retrieval threshold is greater than the second retrieval threshold; searching the first conversation text based on the first retrieval threshold and the cache module to obtain a first sub-search result; and searching the description information based on the second retrieval threshold and the cache module to obtain a second sub-search result.In one embodiment of the present disclosure, the first dialog text includes modality information that indicates the material type of the material used in editing the user material. The processing unit 42 is specifically configured to: generate a corresponding target prompt word based on the modality information; obtain an editing capability item corresponding to the target material type using a language model and the target prompt word; and edit the user material by invoking the editing capability item corresponding to the target material type to obtain a first editing result. In one embodiment of the present disclosure, upon obtaining the editing capability item corresponding to the target material type using the language model and the target prompt word, the processing unit 42 is specifically configured to: generate a target material matching the target material type using the language model and the target prompt word; and generate the editing capability item based on the target material. In one embodiment of the present disclosure, after displaying the first editing result, the interaction unit 41 is further configured to: obtain a second dialog text input by the user via a dialog control, the second dialog text being used to represent editing requirements for the first editing result, the first editing result comprising a target video generated by editing the user material; and the processing unit 42 is further configured to: execute corresponding processing steps based on the second dialog text to obtain a second editing result corresponding to the second dialog text. The processing steps may include at least one of the following: invoking at least one editing capability item to edit the target video; and canceling the editing effect corresponding to the at least one editing capability item. In one embodiment of the present disclosure, the processing unit 42 is further configured to: store the first editing result as a corresponding local general draft. When executing the corresponding processing steps based on the second dialog text to obtain the second editing result corresponding to the second dialog text, the processing unit 42 is specifically configured to: obtain a second editing link corresponding to the second dialog text; process the local general draft based on the second editing link to obtain an updated local general draft; and generate the corresponding second editing result based on the local general draft. In one embodiment of the present disclosure, after obtaining the second editing link corresponding to the second dialogue text, the processing unit 42 is further used to: obtain complexity information corresponding to the second editing link; determine the target execution end based on the complexity information, and the target execution end includes the cloud or the terminal; when the processing unit 42 processes the local general draft based on the second editing link to obtain an updated local general draft, it is specifically used to: place the local general draft on the target execution end and execute the second editing link to obtain a processing result; and obtain an updated local general draft based on the processing result.In one embodiment of the present disclosure, when obtaining an updated local universal draft based on the processing result, the processing unit 42 is specifically configured to: obtain processing results corresponding to at least two second editing links; and fuse the at least two processing results at the terminal to obtain an updated local universal draft. In one embodiment of the present disclosure, the processing unit 42 is further configured to: obtain the execution priority corresponding to each editing capability item in the second editing link; and when processing the local general draft based on the second editing link to obtain an updated local general draft, the processing unit 42 is specifically configured to: iteratively process the local general draft by sequentially invoking the corresponding editing capability items based on the execution priority to obtain the updated local general draft. The interaction unit 41, processing unit 42, and display unit 43 are sequentially connected. The conversation-based video editing device 4 provided in this embodiment can implement the technical solutions of the above-mentioned method embodiments. The implementation principles and technical effects are similar and will not be further described in this embodiment. Figure 18 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. As shown in Figure 18, the electronic device 5 includes: a processor 51 and a memory 52 in communication with the processor 51; the memory 52 stores computer-executable instructions; the processor 51 executes the computer-executable instructions stored in the memory 52 to implement the conversation-based video editing method of the embodiments shown in Figures 2 to 15. Optionally, the processor 51 and the memory 52 are connected via a bus 53. The relevant explanations can be understood by referring to the corresponding descriptions and effects of the steps in the embodiments corresponding to Figures 2-15 , and will not be elaborated upon here. The present embodiment provides a computer-readable storage medium storing computer-executable instructions. When executed by a processor, the computer-executable instructions are used to implement the conversation-based video editing method provided in any of the embodiments corresponding to Figures 2-15 of the present disclosure. The present embodiment provides a computer program product including a computer program. When executed by a processor, the computer program implements the conversation-based video editing method provided in any of the embodiments corresponding to Figures 2-15 of the present disclosure. To implement the above-mentioned embodiments, the present embodiment also provides an electronic device. Referring to Figure 19 , a schematic diagram of an electronic device 900 suitable for implementing the present embodiment is shown. The electronic device 900 may be a terminal device or a server.Terminal devices may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG19 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure. As shown in FIG19 , electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which may perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 902 or programs loaded from a storage device 908 into a random access memory (RAM) 903. RAM 903 also stores various programs and data required for the operation of electronic device 900. The processing device 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904. Typically, the following devices may be connected to the I / O interface 905: an input device 906, such as a touch screen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; an output device 907, such as a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 908, such as a magnetic tape, hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or wired to exchange data. Although FIG19 illustrates an electronic device 900 with various devices, it should be understood that not all illustrated devices are required to be implemented or present. More or fewer devices may alternatively be implemented or present. In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart.In such an embodiment, the computer program can be downloaded and installed from a network via communication device 909, or installed from storage device 908, or installed from ROM 902. When the computer program is executed by processing device 901, the aforementioned functions defined in the method of the embodiment of the present disclosure are performed. It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Furthermore, in the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), or any suitable combination thereof. The computer-readable medium may be included in the electronic device described above, or it may exist separately and not incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the one or more programs cause the electronic device to perform the methods described in the above embodiments. Computer program code for carrying out operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages.The program code may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the accompanying drawings. For example, two blocks shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, as well as combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified functions or operations, or may be implemented using a combination of dedicated hardware and computer instructions. The units described in the embodiments of this disclosure may be implemented via software or hardware. The names of the units do not, in some cases, limit the units themselves. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses." The functions described above may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like. In the context of this disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.Machine-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination thereof. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In a first aspect, according to one or more embodiments of the present disclosure, a conversation-based video editing method is provided, comprising: after loading user material, obtaining first conversation text input by a user via a conversation control, wherein the first conversation text is used to represent an editing requirement for the user material; based on the first conversation text, invoking an editing capability item to edit the user material to obtain a first editing result, wherein the editing capability item is used to perform processing steps to implement the editing requirement; and displaying the first editing result. According to one or more embodiments of the present disclosure, the first editing result includes a target video generated after editing the user material, and / or an execution result corresponding to the editing capability item; displaying the first editing result includes at least one of the following: displaying the execution result through the dialogue control; playing the target video in a video playback interface for playing the user material. According to one or more embodiments of the present disclosure, the execution result includes at least two editing templates corresponding to the editing capability item; after displaying the editing result through the dialogue control, the method further includes: setting one of the at least two editing templates as a target editing template in response to a selection operation input by the user; editing the target video based on the target editing template to obtain an updated first editing result. According to one or more embodiments of the present disclosure, before loading the user material, the method further includes: registering at least one to-be-selected editing capability item to an editing capability library; calling the editing capability item based on the first dialogue text to edit the user material to obtain the first editing result includes: searching the editing capability library based on keywords in the first dialogue text to obtain at least two to-be-selected editing capability items; based on content features of the first dialogue text, The at least two to-be-selected editing capability items are searched to obtain a target editing capability item; and the target editing capability item is invoked to edit the user material to obtain a first editing result.According to one or more embodiments of the present disclosure, invoking an editing capability item based on the first dialogue text to edit the user material and obtain a first editing result includes: processing the first dialogue text using a language model to obtain a first editing link, wherein the first editing link represents an editing process for achieving the editing requirement and includes at least two sequentially invoked editing capability items; and executing the first editing link to edit the user material and obtain the first editing result. According to one or more embodiments of the present disclosure, processing the first dialogue text using a language model to obtain the first editing link includes: generating target prompt words for the language model based on the first dialogue text; processing the target prompt words using the language model to obtain an initial editing link and dynamic parameters, wherein the initial editing link includes a processing template corresponding to the processing steps constituting the editing process, and the dynamic parameters are template parameters for the processing template; and configuring the initial editing link using the dynamic parameters to generate the first editing link. According to one or more embodiments of the present disclosure, processing the target prompt word using the language model to obtain an initial editing link includes: processing the target prompt word using the language model to obtain search information, where the search information represents a search range corresponding to an editing capability item; searching for editing capability items based on the search information to obtain corresponding initial editing capability items; and obtaining the initial editing link based on the initial editing capability items. According to one or more embodiments of the present disclosure, the method further includes: obtaining description information of the user material, where the description information is used to represent content features of the user material; and processing the first dialogue text using the language model to obtain a first editing link includes: processing the first dialogue text and the description information using the language model to obtain the first editing link. According to one or more embodiments of the present disclosure, processing the first dialogue text using a language model to obtain a first editing link includes: searching a cache module based on the first dialogue text to obtain a search result, wherein the cache module stores at least one set of corresponding model input information and model output information; based on the search result, if the first dialogue text matches the corresponding model input information, determining the first editing link as the first dialogue text based on the model output information corresponding to the model input information; and if the first dialogue text does not match the corresponding model input information, processing the first dialogue text using a language model to obtain the first editing link, wherein the model input information includes historical text input into the language model, and the model output information includes search information or a historical editing link corresponding to the historical text.According to one or more embodiments of the present disclosure, the model input information includes the first dialogue text and the description information of the user material, and the retrieval result includes a first sub-retrieval result and a second sub-retrieval result; searching the cache module based on the first dialogue text to obtain the retrieval result includes: obtaining a first retrieval threshold corresponding to the first dialogue text and a second retrieval threshold corresponding to the description information, wherein the first retrieval threshold is greater than the second retrieval threshold; searching the first dialogue text based on the first retrieval threshold and the cache module to obtain a first sub-retrieval result; searching the description information based on the second retrieval threshold and the cache module to obtain a second sub-retrieval result. According to one or more embodiments of the present disclosure, the first dialogue text includes modality information, the modality information being used to characterize a material type used in editing a user material. The invoking an editing capability item to edit the user material based on the first dialogue text to obtain a first editing result includes: generating a corresponding target prompt word based on the modality information; obtaining an editing capability item corresponding to a target material type using a language model and the target prompt word; and editing the user material by invoking the editing capability item corresponding to the target material type to obtain the first editing result. According to one or more embodiments of the present disclosure, obtaining an editing capability item corresponding to a target material type using a language model and the target prompt word includes: generating a target material matching the target material type using a language model and the target prompt word; and generating the editing capability item based on the target material. According to one or more embodiments of the present disclosure, after displaying the first editing result, the method further includes: obtaining a second dialogue text input by the user through the dialogue control, where the second dialogue text is used to represent an editing requirement for the first editing result, and the first editing result includes a target video generated after editing the user material; based on the second dialogue text, executing corresponding processing steps to obtain a second editing result corresponding to the second dialogue text; wherein the processing steps include at least one of the following: calling at least one editing capability item to edit the target video; and undoing the editing effect corresponding to the at least one editing capability item.According to one or more embodiments of the present disclosure, the method further includes: storing the first editing result as a corresponding local general draft; performing corresponding processing steps based on the second dialogue text to obtain a second editing result corresponding to the second dialogue text, including: obtaining a second editing link corresponding to the second dialogue text; processing the local general draft based on the second editing link to obtain an updated local general draft; and generating the corresponding second editing result based on the local general draft. According to one or more embodiments of the present disclosure, after obtaining the second editing link corresponding to the second dialogue text, the method further includes: obtaining complexity information corresponding to the second editing link; determining a target execution end based on the complexity information, the target execution end including a cloud or a terminal; processing the local general draft based on the second editing link to obtain an updated local general draft, including: placing the local general draft on the target execution end and executing the second editing link to obtain a processing result; and obtaining the updated local general draft based on the processing result. According to one or more embodiments of the present disclosure, obtaining an updated local universal draft based on the processing result includes: obtaining processing results corresponding to at least two second editing links; and fusing the at least two processing results on a terminal to obtain an updated local universal draft. According to one or more embodiments of the present disclosure, the method further includes: obtaining an execution priority corresponding to each editing capability item in the second editing link; and processing the local universal draft based on the second editing link to obtain an updated local universal draft includes: iteratively processing the local universal draft based on the execution priority, sequentially invoking the corresponding editing capability items to obtain the updated local universal draft. In a second aspect, according to one or more embodiments of the present disclosure, a dialogue-based video editing device is provided, comprising: an interaction unit, configured to obtain, through a dialogue control, a first dialogue text input by a user after loading a user material, wherein the first dialogue text is used to represent an editing requirement for the user material; a processing unit, configured to invoke an editing capability item based on the first dialogue text to edit the user material to obtain a first editing result, wherein the editing capability item is used to execute a processing step for realizing the editing requirement; and a display unit, configured to display the first editing result.According to one or more embodiments of the present disclosure, the first editing result includes a target video generated after editing the user material and / or an execution result corresponding to the editing capability item; the display unit is specifically configured to perform at least one of the following: display the execution result via the dialogue control; or play the target video within a video playback interface for playing the user material. According to one or more embodiments of the present disclosure, the execution result includes at least two editing templates corresponding to the editing capability item; after displaying the editing result via the dialogue control, the system further includes: the interaction unit is further configured to: set one of the at least two editing templates as a target editing template in response to a selection operation input by a user; and the processing unit is further configured to: edit the target video based on the target editing board to obtain an updated first editing result. According to one or more embodiments of the present disclosure, before loading the user material, the processing unit is further configured to: register at least one candidate editing capability item in an editing capability library; and when invoking an editing capability item to edit the user material based on the first dialogue text to obtain a first editing result, the processing unit is further configured to: search the editing capability library based on keywords in the first dialogue text to obtain at least two candidate editing capability items; search the at least two candidate editing capability items based on content features of the first dialogue text to obtain a target editing capability item; and invoke the target editing capability item to edit the user material to obtain a first editing result. According to one or more embodiments of the present disclosure, the processing unit is further configured to: process the first dialogue text using a language model to obtain a first editing link, wherein the first editing link represents an editing process for achieving the editing requirement and includes at least two sequentially invoked editing capability items; and edit the user material by executing the first editing link to obtain the first editing result.According to one or more embodiments of the present disclosure, when the processing unit processes the first dialogue text using a language model to obtain a first editing link, the processing unit is specifically configured to: generate a target prompt word for the language model based on the first dialogue text; process the target prompt word using the language model to obtain an initial editing link and dynamic parameters, wherein the initial editing link includes a processing template corresponding to the processing steps constituting the editing process, and the dynamic parameters are template parameters for the processing template; and configure the initial editing link using the dynamic parameters to generate the first editing link. According to one or more embodiments of the present disclosure, when the processing unit processes the target prompt word using the language model to obtain an initial editing link, the processing unit is specifically configured to: process the target prompt word using the language model to obtain search information representing a search range corresponding to an editing capability item; perform an editing capability item search based on the search information to obtain a corresponding initial editing capability item; and obtain the initial editing link based on the initial editing capability item. According to one or more embodiments of the present disclosure, the processing unit is further configured to: obtain description information of the user material, the description information being used to characterize content features of the user material; and when processing the first dialogue text using a language model to obtain a first edit link, the processing unit is specifically configured to: process the first dialogue text and the description information using a language model to obtain the first edit link. According to one or more embodiments of the present disclosure, when processing the first dialogue text using a language model to obtain the first edit link, the processing unit is specifically configured to: search a cache module based on the first dialogue text to obtain a search result, the cache module storing at least one set of corresponding model input information and model output information; based on the search result, if the first dialogue text matches the corresponding model input information, then, based on the model output information corresponding to the model input information, determine that the first edit link is the first edit link; and if the first dialogue text does not match the corresponding model input information, then, process the first dialogue text using a language model to obtain the first edit link, wherein the model input information includes historical text input into the language model, and the model output information includes search information or historical edit links corresponding to the historical text.According to one or more embodiments of the present disclosure, the model input information includes the first dialogue text and the description information of the user material, and the retrieval result includes a first sub-retrieval result and a second sub-retrieval result; the processing unit searches the cache module based on the first dialogue text to obtain the retrieval result, including: obtaining a first retrieval threshold corresponding to the first dialogue text and a second retrieval threshold corresponding to the description information, wherein the first retrieval threshold is greater than the second retrieval threshold; based on the first retrieval threshold and the cache module, searching the first dialogue text to obtain a first sub-retrieval result; based on the second retrieval threshold and the cache module, searching the description information to obtain a second sub-retrieval result. According to one or more embodiments of the present disclosure, the first dialogue text includes modal information, and the modal information is used to characterize the material type of the material used in the process of editing the user material; the processing unit is specifically configured to: generate a corresponding target prompt word based on the modal information; obtain an editing capability item corresponding to the target material type through a language model and the target prompt word; and edit the user material by calling the editing capability item corresponding to the target material type to obtain a first editing result. According to one or more embodiments of the present disclosure, when the processing unit obtains the editing capability item corresponding to the target material type through a language model and the target prompt word, it is specifically configured to: generate a target material matching the target material type through the language model and the target prompt word; and generate the editing capability item based on the target material. According to one or more embodiments of the present disclosure, after the first editing result is displayed, the interaction unit is further used to: obtain a second dialogue text input by the user through the dialogue control, the second dialogue text is used to represent the editing requirements for the first editing result, and the first editing result includes a target video generated after editing the user material; the processing unit is further used to: execute corresponding processing steps based on the second dialogue text to obtain a second editing result corresponding to the second dialogue text; wherein the processing steps include at least one of the following: calling at least one editing capability item to edit the target video; and canceling the editing effect corresponding to the at least one editing capability item.According to one or more embodiments of the present disclosure, the processing unit is further configured to: store the first editing result as a corresponding local general draft; when the processing unit executes corresponding processing steps based on the second conversation history to obtain a second editing result corresponding to the second conversation text, the processing unit is further configured to: obtain a second editing link corresponding to the second conversation text; process the local general draft based on the second editing link to obtain an updated local general draft; and generate the corresponding second editing result based on the local general draft. According to one or more embodiments of the present disclosure, after obtaining the second editing link corresponding to the second conversation text, the processing unit is further configured to: obtain complexity information corresponding to the second editing link; determine a target execution end based on the complexity information, the target execution end including a cloud or a terminal; and when the processing unit processes the local general draft based on the second editing link to obtain an updated local general draft, the processing unit is further configured to: place the local general draft on the target execution end and execute the second editing link to obtain a processing result; and obtain an updated local general draft based on the processing result. According to one or more embodiments of the present disclosure, when obtaining an updated local universal draft based on the processing results, the processing unit is specifically configured to: obtain processing results corresponding to at least two second editing links; and merge the at least two processing results on the terminal to obtain an updated local universal draft. According to one or more embodiments of the present disclosure, the processing unit is further configured to: obtain the execution priority corresponding to each editing capability item in the second editing link; and when processing the local universal draft based on the second editing link to obtain an updated local universal draft, the processing unit is specifically configured to: iteratively process the local universal draft by sequentially invoking the corresponding editing capability items based on the execution priority to obtain an updated local universal draft. In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory; the memory storing computer-executable instructions; and the at least one processor executing the computer-executable instructions stored in the memory, such that the at least one processor performs the dialogue-based video editing method described in the first aspect and various possible designs of the first aspect. In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the dialogue-based video editing method described in the first aspect and various possible configurations of the first aspect is implemented.In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program. When executed by a processor, the computer program implements the dialogue-based video editing method described in the first aspect and various possible designs of the first aspect. The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by specific combinations of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the scope of the above-mentioned disclosure. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure. Furthermore, while various operations are depicted in a specific order, this should not be construed as requiring that these operations be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while the above discussion includes several specific implementation details, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Although the subject matter has been described using language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

Claims 1. A video editing method based on dialogue, comprising: After the user material is loaded, a first dialogue text input by the user is obtained through a dialogue control, where the first dialogue text is used to represent an editing requirement for the user material; Based on the first dialogue text, an editing capability item is called to edit the user material to obtain a first editing result, wherein the editing capability item is used to execute a processing step to realize the editing requirement; and the first editing result is displayed.

2. The method according to claim 1, wherein: The first editing result includes a target video generated after editing the user material, and / or an execution result corresponding to the editing capability item; displaying the first editing result includes at least one of the following: displaying the execution result through the dialogue control; playing the target video in the video playback interface used to play the user material.

3. The method according to claim 2, wherein: The execution result includes at least two editing templates corresponding to the editing capability item; After displaying the editing result through the dialogue control, the method further includes: in response to a selection operation input by a user, setting one of the at least two editing templates as a target editing template; and editing the target video based on the target editing template to obtain an updated first editing result.

4. The method according to any one of claims 1 to 3, wherein: Before loading the user material, the method further includes: registering at least one editing capability item to be selected to an editing capability library; based on the first dialogue text, calling the editing capability item to edit the user material to obtain a first editing result, including: based on keywords in the first dialogue text, searching the editing capability library to obtain at least two editing capability items to be selected; based on content features of the first dialogue text, searching the at least two editing capability items to be selected to obtain a target editing capability item; calling the target editing capability item to edit the user material to obtain a first editing result.

5. The method according to claim 1, wherein: Based on the first dialogue text, calling an editing capability item to edit the user material to obtain a first editing result, including: processing the first dialogue text through a language model to obtain a first editing link, wherein the first editing link is used to represent an editing process for realizing the editing requirement, and the first editing link includes at least two editing capability items called in order; 29 The user material is edited by executing the first editing link to obtain a first editing result.

6. The method according to claim 5, wherein: The step of processing the first dialogue text through a language model to obtain a first editing link includes: generating a target prompt word for the language model according to the first dialogue text; processing the target prompt word through the language model to obtain an initial editing link and dynamic parameters, wherein the initial editing link includes a processing template corresponding to the processing steps constituting the editing process, and the dynamic parameters are template parameters for the processing template; and configuring the initial editing link through the dynamic parameters to generate the first editing link.

7. The method according to claim 6, wherein: The step of processing the target prompt word through the language model to obtain an initial editing link includes: processing the target prompt word through the language model to obtain search information, wherein the search information represents a search range corresponding to an editing capability item; searching for editing capability items based on the search information to obtain corresponding initial editing capability items; and obtaining the initial editing link based on the initial editing capability items.

8. The method according to claim 5, further comprising: Acquire description information of the user material, where the description information is used to characterize content features of the user material; The step of processing the first dialogue history text through a language model to obtain a first editing link includes: processing the first dialogue text and the description information through a language model to obtain a first editing link.

9. The method according to claim 5, wherein: The processing of the first dialogue text through a language model to obtain a first editing link includes: searching a cache module based on the first dialogue text to obtain a search result, wherein the cache module stores at least one set of corresponding model input information and model output information; according to the search result, if the first dialogue text hits the corresponding model input information, determining it as the first editing link according to the model output information corresponding to the model input information; if the first dialogue text does not hit the corresponding model input information, processing the first dialogue text through a language model to obtain a first editing link, wherein the model input information includes historical text input into the language model, and the model output information includes search information or historical editing links corresponding to the historical text.

10. The method according to claim 9, wherein: The model input information includes the first dialogue text and the description information of the user material, and the retrieval result includes a first sub-retrieval result and a second sub-retrieval result; the retrieval of the cache module based on the first dialogue text to obtain the retrieval result includes: obtaining a first retrieval threshold corresponding to the first dialogue text and a second retrieval threshold corresponding to the description information, wherein the first retrieval threshold is greater than the second retrieval threshold; based on the first retrieval threshold and the cache module, searching the first dialogue text to obtain a first sub-retrieval result; 30 search result; based on the second search threshold and the cache module, search the description information to obtain a second sub-search result.

11. The method according to claim 1, wherein: The first dialogue text includes modality information, where the modality information is used to represent the material type of the material used in the process of editing the user material; The calling of the editing capability item to edit the user material based on the first dialogue text to obtain a first editing result includes: generating a corresponding target prompt word based on the modality information; Obtain an editing capability item corresponding to a target material type through a language model and the target prompt word; and edit the user material by calling the editing capability item corresponding to the target material type to obtain a first editing result.

12. The method according to claim 11, wherein: The obtaining of the editing capability item corresponding to the target material type through the language model and the target prompt word includes: generating a target material matching the target material type through the language model and the target prompt word; and generating the editing capability item based on the target material.

13. The method according to any one of claims 1 to 12, wherein: After displaying the first editing result, the method further includes: obtaining a second dialogue text input by the user through the dialogue control, the second dialogue text being used to represent the editing requirements for the first editing result, the first editing result including a target video generated after editing the user material; based on the second dialogue text, executing corresponding processing steps to obtain a second editing result corresponding to the second dialogue text; wherein the processing steps include at least one of the following: calling at least one editing capability item to edit the target video; and canceling the editing effect corresponding to the at least one editing capability item.

14. The method according to claim 13, further comprising: storing the first editing result as a corresponding local general draft; The executing corresponding processing steps based on the second dialogue text to obtain a second editing result corresponding to the second dialogue text includes: obtaining a second editing link corresponding to the second dialogue text; processing the local general draft based on the second editing link to obtain an updated local general draft; and generating a corresponding second editing result based on the local general draft.

15. The method according to claim 14, wherein: After obtaining the second editing link corresponding to the second dialogue text, the method further includes: Obtaining complexity information corresponding to the second editing link; determining a target execution end according to the complexity information, wherein the target execution end includes a cloud or a terminal; processing the local general draft based on the second editing link to obtain an updated local general draft, including: placing the local general draft on the target execution end and executing the second editing link to obtain a processing result; obtaining an updated local general draft based on the processing result 16. The method according to claim 15, wherein: The obtaining an updated local general draft based on the processing result includes: obtaining processing results corresponding to at least two second editing links; and fusing at least two of the processing results at the terminal to obtain an updated local general draft.

17. The method according to claim 14, further comprising: Obtaining the execution priority corresponding to each editing capability item in the second editing link; The processing of the local general draft based on the second editing link to obtain an updated local general draft includes: based on the execution priority, sequentially calling corresponding editing capability items to iteratively process the local general draft to obtain an updated local general draft.

18. A conversation-based video editing device, comprising: an interaction unit, configured to obtain, after loading the user material, a first dialogue text input by the user through a dialogue control, wherein the first dialogue text is used to represent an editing requirement for the user material; a processing unit configured to call an editing capability item to edit the user material based on the first dialogue text to obtain a first editing result, wherein the editing capability item is used to execute a processing step to realize the editing requirement; A display unit is configured to display the first editing result.

19. An electronic device, comprising a processor and a memory, wherein: The memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor executes the dialogue-based video editing method as described in any one of claims 1 to 17.

20. A computer-readable storage medium storing computer-executable instructions, wherein: When the processor executes the computer-executable instructions, the dialogue-based video editing method according to any one of claims 1 to 17 is implemented.

Citation Information

Patent Citations

  • Video-based image generation method, video-based image display method, video-based image generation device, video-based image display equipment and storage medium

    CN111726676A

  • Video editing method, video editing device, electronic equipment and readable storage medium

    CN114430499A

  • Artificial intelligence device, and method for operating artificial intelligence device

    WO2022260188A1

Cited By

  • Generative intelligent model-based text travel resource digital reconstruction system and method

    CN121581059A