Video clip method, device, apparatus and medium

CN122802727APending Publication Date: 2026-09-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510346822.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

但是,人工手动处理的方式对操作人员具有较高的门槛,难以满足人们对大量视频进行剪辑的需求

Benefits of technology

[0008]根据本公开实施例提供的视频的剪辑方案,通过获取用于剪辑的原始视频、引导信息以及参考素材,获取多个参考编辑元素,利用目标大模型,基于参考编辑元素、原始视频、引导信息和参考素材,生成自然语言形式的剪辑数据,并基于剪辑数据,获取目标视频。从而不仅提高了对视频进行剪辑的效率,满足了人们对大量视频进行剪辑的需求,还提高了对视频进行剪辑的效果,提升了用户体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802727A_ABST
    Figure CN122802727A_ABST
Patent Text Reader

Abstract

The disclosure provides a video clipping method, device, equipment and medium, a specific embodiment of the method includes: obtaining an original video, guide information and reference materials for clipping; obtaining a plurality of reference editing elements; using a target large model, based on the reference editing elements, the original video, the guide information and the reference materials, generating clipping data in natural language form; based on the clipping data, obtaining a target video. This embodiment not only improves the efficiency of video clipping, meets the demand of people for a large number of video clipping, but also improves the effect of video clipping and enhances the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of video processing technology, and more particularly to video editing methods, apparatus, devices, and media. Background Technology

[0002] With the continuous development of network technology and digital media technology, video is increasingly used in people's lives, providing entertainment, bringing convenience, and adding more fun. To improve videos, editing is usually required to obtain more vivid and expressive content from the original footage. Current techniques typically involve manually cutting video segments from the original video, recombining these segments, and adding various materials and editing elements. However, this manual process has a high skill ceiling and cannot meet the demand for editing large volumes of video. Therefore, a video editing solution is needed. Summary of the Invention

[0003] Embodiments of this disclosure describe a video editing method, apparatus, device, and medium.

[0004] According to a first aspect, a video editing method is provided, the method comprising: acquiring an original video, guiding information, and reference materials for editing; acquiring multiple reference editing elements; using a target large model, generating editing data in natural language form based on the reference editing elements, the original video, the guiding information, and the reference materials; and acquiring a target video based on the editing data.

[0005] According to a second aspect, a video editing apparatus is provided, the apparatus comprising: a first acquisition unit configured to acquire an original video, guiding information, and reference materials for editing; a second acquisition unit configured to acquire multiple reference editing elements; a generation unit configured to generate editing data in natural language form based on the reference editing elements, the original video, the guiding information, and the reference materials using a target large model; and a conversion unit configured to acquire a target video based on the editing data.

[0006] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform any of the methods described in the first aspect.

[0007] According to a fourth aspect, an electronic device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in any one of the first aspects.

[0008] According to the video editing scheme provided in this disclosure, by acquiring the original video, guidance information, and reference materials for editing, multiple reference editing elements are obtained. Using a target large model, editing data in natural language form is generated based on the reference editing elements, the original video, the guidance information, and the reference materials. Based on the editing data, the target video is then obtained. This not only improves the efficiency of video editing and meets the demand for editing large amounts of video, but also improves the editing effect and enhances the user experience.

[0009] Furthermore, since this embodiment obtains reference editing elements based on keywords in the guidance information, the selection of reference editing elements is more in line with the user's needs, thereby improving the effect of video editing and enhancing the user experience.

[0010] Furthermore, in this embodiment, the target large model first selects a more suitable target editing element from the reference editing elements, and then generates editing data based on the target editing element, thereby decoupling the selection of editing elements and the generation of editing data, which further improves the effect of video editing.

[0011] Furthermore, in this embodiment, during the model training phase, sample video templates are obtained from the candidate video templates as sample training data to train the target large model, making the selection of training data more accurate and thus improving the performance of the target large model.

[0012] Furthermore, since the third editing element is very likely unrelated to the keywords in the guiding information, during the model training phase, a third editing element unrelated to the keywords in the guiding information is added to the reference editing elements. This allows the target large model to better learn the differences between editing elements unrelated to keywords and editing elements related to keywords, thereby improving the performance of the target large model. Attached Figure Description

[0013] Figure 1 This disclosure is a schematic diagram illustrating an application scenario of video editing according to an exemplary embodiment;

[0014] Figure 2 This is a schematic diagram of an exemplary system architecture for applying embodiments of this disclosure;

[0015] Figure 3 This is a flowchart illustrating a video editing method according to an exemplary embodiment of the present disclosure;

[0016] Figure 4 This is a block diagram illustrating a video editing apparatus according to an exemplary embodiment of the present disclosure;

[0017] Figure 5 This is a schematic block diagram of an electronic device provided in some embodiments of this disclosure. Detailed Implementation

[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations of the technical solutions disclosed herein, based on the prompt message.

[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0022] The technical solutions provided in this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the relevant invention and not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0023] With the continuous development of network and digital media technologies, video is increasingly used in people's lives, providing entertainment, bringing convenience, and adding more fun. To improve videos, people usually need to edit the original footage to obtain more vivid and expressive videos. In related technologies, this typically involves manually cutting video segments from the original video, then recombining these segments, adding various materials and editing elements to complete the editing process. However, manual processing has a high skill ceiling and is difficult to meet the demand for editing large volumes of video.

[0024] In other related technologies, different video templates can be pre-set for different scenarios and themes, and users can input the video to be edited. Video processing applications or platforms can match suitable video templates to the video to be edited and add the video to the template, thus obtaining the edited video. However, pre-set video templates have certain limitations, resulting in poor processing effects on the video to be edited.

[0025] This disclosure provides a video editing solution that acquires the original video, guiding information, and reference materials for editing, obtains multiple reference editing elements, utilizes a target large model, and generates editing data in natural language form based on the reference editing elements, original video, guiding information, and reference materials. Based on this editing data, the target video is then obtained. This not only improves the efficiency of video editing, meeting the demand for editing large volumes of videos, but also enhances the editing effect and improves the user experience.

[0026] See Figure 1 This diagram illustrates an application scenario of video editing according to an exemplary embodiment. The process shown can be applied to scenarios involving video editing using a model, as well as to scenarios involving model training. The application stage and the model training stage of video editing are described below.

[0027] like Figure 1 As shown, in the video editing application stage, users can input the original video, guidance information, and reference materials into the video editing client via their terminal device. First, the video editing client can extract keywords from the guidance information, obtaining one or more keywords. Then, using the extracted keywords, it searches from one or more template libraries containing video templates to obtain reference video templates related to the keywords, and retrieves the editing elements from these reference video templates. Furthermore, using the extracted keywords, it searches from one or more element libraries containing editing elements to obtain editing elements related to the keywords. The editing elements from the searched reference video templates and the searched editing elements related to the keywords can be used as reference editing elements.

[0028] Then, the video editing client can input the original video, guidance information, reference materials, and reference editing elements into the large model M. The large model M can first select at least one target editing element from multiple reference editing elements, and then generate editing data in natural language form based on the target editing element, the original video, guidance information, and reference materials. Finally, the video editing client can convert the editing data into the target video.

[0029] During the model training phase, sample video templates can be selected from a pre-built pool of candidate video templates. The original video corresponding to the sample video template is obtained as the raw video; the descriptive information corresponding to the sample video template is obtained as guiding information; and the template material corresponding to the sample video template is obtained as reference material.

[0030] Next, in addition to retrieving editable elements associated with keywords from the template and element libraries based on the guiding information, editable elements not associated with keywords can also be randomly selected. These keyword-associated editable elements, unassociated editable elements, and the editable elements corresponding to the sample video template are grouped together as reference editable elements. Then, the large model M generates clip data, and the target video is obtained based on the clip data. Loss1 is calculated based on the target video and the sample video template. The target editable element selected by the large model M from the reference editable elements is then retrieved, and loss Loss2 is calculated based on the target editable element and the editable element corresponding to the sample video template. The large model M is then updated based on Loss1 and Loss2, thus completing the training of the large model M.

[0031] It should be noted that, in Figure 1 In this embodiment, the video editing application stage is described using the example of the video editing client directly performing video editing. In other embodiments, the video editing client can also transmit the original video, guidance information, and reference materials to the video editing server deployed on the service platform via the network. The video editing server then generates editing data based on the original video, guidance information, and reference materials, obtains the target video based on the editing data, and transmits the target video to the video editing client via the network to provide the target video to the user. See details below. Figure 2 Example.

[0032] Figure 2 This is a schematic diagram of an exemplary system architecture for applying embodiments of this disclosure.

[0033] like Figure 2 As shown, system architecture 200 may include terminal device 202, network 203, and server 204. It should be understood that... Figure 2 The number or type of terminal devices, networks, and servers shown are merely illustrative. Depending on implementation needs, there can be any number or type of terminal devices, networks, and servers.

[0034] Network 203 is a medium used to provide communication links between terminal devices and servers. Network 203 can include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0035] The terminal device 202 is equipped with a video editing client. The terminal device 202 can interact with the server via network 203 to receive or send requests or information. The terminal device 202 can be various electronic devices, including but not limited to smartphones, tablets, laptops, desktop computers, and smart wearable devices.

[0036] Server 204 deploys a video editing server. Server 204 can store, analyze, and process received data, and can also send control commands or requests to terminal devices or other servers. The server can provide video editing services in response to user service requests. It is understood that a single server can provide one or more services, and the same service can be provided by multiple servers.

[0037] based on Figure 2 In the system architecture shown in this embodiment, user 201 can input original video, guidance information, and reference materials for editing via terminal device 202. Terminal device 202 can transmit the original video, guidance information, and reference materials to server 204 via network 203. After receiving the original video, guidance information, and reference materials, server 204 can generate editing data based on the original video, guidance information, and reference materials, and obtain the target video based on the editing data. Finally, server 204 can return the target video to terminal device 202 via network 203, allowing user 201 to view and save the target video via terminal device 202.

[0038] The present disclosure will now be described in detail with reference to specific embodiments.

[0039] Figure 3 This is a flowchart illustrating a video editing method according to an exemplary embodiment. The method can be applied to a video editing client or a video editing server. In this embodiment, the video editing client is installed on a terminal device, which may include, but is not limited to, mobile terminal devices such as smartphones, smart wearable devices, tablets, laptops, and desktop computers. The video editing server is deployed in a service platform, which can be any device, server, or device cluster with computing and processing capabilities. It should be noted that this method can be applied to the video editing application stage or the model training stage, and the method includes the following steps:

[0040] like Figure 3 As shown, in step 301, the original video, guidance information, and reference materials for editing are obtained.

[0041] In this embodiment, during the video editing application stage, the user can input the original video, guidance information, and reference materials through a terminal device. The original video can be a video to be edited that the user has pre-downloaded or filmed on the terminal device. The guidance information can be text-based information entered by the user to guide the target model in video editing; for example, the guidance information could be "Edit a video with cool yet cute elements." Another example is "Edit a vacation vlog video." Reference materials can be materials selected by the user that can be added to the video; for example, reference materials can be images, background music, audio, etc., that can be inserted into the video.

[0042] During the model training phase, a video template for training the model can be randomly selected from a pre-built pool of candidate video templates and used as a sample video template. The original video corresponding to this sample video template is then obtained as the raw video. The descriptive information corresponding to this sample video template is also obtained as guiding information. Finally, the template material corresponding to this sample video template is obtained as reference material.

[0043] Specifically, the candidate video templates can be pre-built templates that users can refer to when editing videos. A video template can include the original video, descriptive information, template materials, and the edited effect video. The original video corresponding to the video template can be the object being edited. The descriptive information can be, for example, the title or a brief introduction of the video template. The template materials can be, for example, images inserted into the video, background music, and audio. The effect video is the video obtained after editing the original video using the template materials. Because this embodiment uses sample video templates from the candidate video templates as sample training data to train the target large model during the model training phase, the selection of training data is more accurate, thereby improving the performance of the target large model.

[0044] In step 302, multiple reference editing elements are obtained.

[0045] In this embodiment, the reference editing element can be a candidate editing element. Editing elements can be used to add effects to the video; for example, editing elements can include, but are not limited to, transitions, animations, stickers, filters, and special effects. When editing the video, a target editing element can be selected from the reference editing elements and added to the video effects.

[0046] Specifically, in both the video editing application stage and the model training stage, keywords from the guidance information can be extracted, and multiple editing elements associated with those keywords can be obtained as reference editing elements. The guidance information can contain one or more keywords. If there are multiple keywords, editing elements associated with each keyword can be obtained separately. Because this embodiment obtains reference editing elements based on keywords from the guidance information, the selection of reference editing elements better meets the user's needs, thereby improving the video editing effect and enhancing the user experience.

[0047] In one implementation, reference video templates related to keywords can be searched from a template library containing multiple video templates, and editing elements from these reference video templates can be obtained as reference editing elements. The template library can be a database stored in public resources on the server side or a database stored in local resources on the client side. The template library may include multiple pre-built video templates for user reference. Specifically, different template libraries have different search interfaces. The search interfaces of at least one template library can be obtained first, and then reference video templates related to keywords can be searched from each template library using each search interface. Finally, editing elements from the reference video templates are obtained as reference editing elements.

[0048] For example, you can directly retrieve the titles or descriptions of each video template in the template library through a search interface, then match these titles or descriptions with keywords, and determine the relevant reference video templates based on the matching degree. Alternatively, you can pre-set corresponding tags for each video template in the template library, retrieve the tags for each video template through a search interface, and determine the relevant reference video templates based on these tags.

[0049] In another implementation, keyword-related editing elements can be retrieved from an element library containing multiple editing elements, serving as reference editing elements. The element library can be a database stored in public resources on the server side or a database stored in local resources on the client side. The element library can include multiple pre-defined editing elements, such as transitions, animations, stickers, filters, and effects.

[0050] Specifically, different element libraries may have different search interfaces. One approach is to first obtain the search interfaces for at least one element library, and then use these interfaces to search for relevant editing elements from each element library as reference editing elements. For example, one can pre-set at least one tag for each editing element in the element library, obtain the corresponding tags for each editing element through the search interface, and determine the relevant reference editing elements based on the tags corresponding to each editing element.

[0051] In another implementation, a subset of reference editing elements can be obtained based on at least one template library, and another subset of reference editing elements can be obtained based on at least one element library. It is understood that this embodiment does not limit the specific method of obtaining the reference editing elements.

[0052] In addition, during the model training phase, the reference edit elements can include not only the first edit element associated with the keywords in the guidance information, but also the second edit element corresponding to the sample video template. The acquired first and second edit elements can be used together as reference edit elements. Furthermore, a third edit element can be randomly acquired, thus defining the first, second, and third edit elements as multiple reference edit elements.

[0053] Specifically, during the model training phase, the reference edit elements include not only the first and second edit elements, but also a third edit element. This third edit element can be an edit element included in a video template randomly selected from the template library, or an edit element included in a video template randomly selected from the element library. Since the third edit element is highly likely to be unrelated to the keywords in the guidance information, adding this unrelated third edit element to the reference edit elements during model training allows the target large model to better learn the differences between keyword-unrelated and keyword-related edit elements, thereby improving the performance of the target large model.

[0054] In step 303, using the target large model, editing data in natural language form is generated based on reference editing elements, the original video, guidance information, and reference materials. In step 304, the target video is obtained based on the editing data.

[0055] In this embodiment, a target large model can be used to generate editing data based on reference editing elements, the original video, guidance information, and reference materials. The target large model can be a large-scale language model involving natural language processing. Specifically, reference editing elements, the original video, guidance information, and reference materials can be input into the target large model. The target large model can first select target editing elements from the reference editing elements, and then generate and output editing data based on the target editing elements, the original video, guidance information, and reference materials. The editing data can be in natural language form and can be converted into a target video. Because in this embodiment, the target large model first selects more suitable target editing elements from the reference editing elements, and then further generates editing data based on the target editing elements, the selection of editing elements and the generation of editing data are decoupled, further improving the video editing effect.

[0056] In one implementation, an intermediate format can be pre-created. Data in this intermediate format, after rendering, yields a video file. Conversion rules between the intermediate format and natural language data are established, enabling mutual conversion between the two. After the target large model outputs the edited data, it can be converted back to the intermediate format based on the pre-established conversion rules to obtain a transition file. This transition file is then rendered to obtain the target video. In another implementation, other generative models can be further utilized to generate the target video based on the edited data. It is understood that any other reasonable method can be used to convert the edited data into the target video; this embodiment is not limited in this regard.

[0057] This disclosure provides a video editing method that acquires the original video, guiding information, and reference materials for editing, obtains multiple reference editing elements, utilizes a target large model, and generates editing data in natural language form based on the reference editing elements, the original video, the guiding information, and the reference materials. Based on this editing data, the target video is then obtained. This not only improves the efficiency of video editing, meeting the demand for editing large amounts of video, but also enhances the editing effect and improves the user experience.

[0058] It should be noted that although the operations of the methods of this disclosure embodiment are described in a specific order in the above embodiments, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0059] Corresponding to the aforementioned video editing method embodiments, this disclosure also provides embodiments of a video editing apparatus.

[0060] like Figure 4 As shown, Figure 4 This disclosure is a block diagram of a video editing apparatus according to an exemplary embodiment. The apparatus may include: a first acquisition unit 401, a second acquisition unit 402, a generation unit 403, and a conversion unit 404.

[0061] The first acquisition unit 401 is configured to acquire the original video, guidance information, and reference materials for editing.

[0062] The second acquisition unit 402 is configured to acquire multiple reference editing elements.

[0063] The generation unit 403 is configured to utilize the target large model to generate editing data in natural language form based on reference editing elements, the original video, guiding information, and reference materials.

[0064] The conversion unit 404 is configured to acquire the target video based on the clip data.

[0065] In some implementations, the second acquisition unit 402 is configured to: extract keywords from the guidance information, acquire multiple editing elements associated with the keywords, and use them as multiple reference editing elements.

[0066] In other embodiments, the second acquisition unit 402 acquires multiple editing elements associated with the keyword as multiple reference editing elements by searching for reference video templates related to the keyword from a template library that includes multiple video templates, and acquiring the editing elements in the reference video templates as reference editing elements.

[0067] In other embodiments, the second acquisition unit 402 acquires multiple editing elements associated with the keyword as multiple reference editing elements by acquiring editing elements related to the keyword from an element library that includes multiple editing elements as reference editing elements.

[0068] In other embodiments, the generation unit 403 is configured to: input reference editing elements, original video, guidance information and reference materials into the target large model, the target large model selects target editing elements from the reference editing elements, and generates editing data based on the target editing elements, original video, guidance information and reference materials.

[0069] In other embodiments, the device is used for the training process of a target large model, and the first acquisition unit 401 is configured to: acquire sample video templates from a plurality of candidate video templates, acquire the original video corresponding to the sample video template as the original video, acquire the descriptive information corresponding to the sample video template as guidance information, and acquire the template material corresponding to the sample video template as reference material.

[0070] In other embodiments, the second acquisition unit 402 is configured to: extract keywords from the guidance information, acquire a first editing element associated with the keywords, acquire a second editing element corresponding to the sample video template, and obtain multiple reference editing elements based on the first and second editing elements.

[0071] In other embodiments, the second acquisition unit 402 obtains multiple reference editing elements based on the first and second editing elements in the following manner: randomly acquires a third editing element, and determines the first, second, and third editing elements as multiple reference editing elements.

[0072] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the embodiments of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0073] The following is for reference. Figure 5 , Figure 5 This is a schematic block diagram of an electronic device provided for some embodiments of this disclosure. The electronic device 920 is, for example, suitable for implementing the video editing method provided in the embodiments of this disclosure. The electronic device 920 can be a terminal device, etc., and can be used to implement a client or server. The electronic device 920 can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. It should be noted that... Figure 5 The illustrated electronic device 920 is merely an example and does not impose any limitation on the functionality and scope of use of the embodiments of this disclosure.

[0074] like Figure 5As shown, the electronic device 920 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 921, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 922 or a program loaded from a storage device 928 into a random access memory (RAM) 923. The RAM 923 also stores various programs and data required for the operation of the electronic device 920. The processing unit 921, ROM 922, and RAM 923 are interconnected via a bus 924. An input / output (I / O) interface 925 is also connected to the bus 924.

[0075] Typically, the following devices can be connected to I / O interface 925: input devices 926 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 927 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 928 including, for example, magnetic tapes, hard disks, etc.; and communication devices 929. Communication device 929 allows electronic device 920 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 5 An electronic device 920 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 920 may alternatively implement or have more or fewer devices. Figure 5 Each box shown can represent a device or multiple devices as needed.

[0076] According to embodiments of this disclosure, the video editing method described above can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the video editing method described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 929, or installed from a storage device 928, or installed from a ROM 922. When the computer program is executed by a processing device 921, the functions defined in the video editing method provided by embodiments of this disclosure can be implemented.

[0077] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the methods provided in this disclosure.

[0078] It should be noted that the computer-readable medium described in the embodiments of this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the embodiments of this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.

[0079] Computer program code for performing the operations of embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0080] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for storage media and computing devices are basically similar to the method embodiments, so they are described more simply; relevant parts can be referred to the descriptions of the method embodiments.

[0081] Those skilled in the art will recognize that the functions described in the embodiments of this disclosure in one or more of the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.

[0082] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the embodiments of this disclosure. It should be understood that the above descriptions are merely specific implementations of the embodiments of this disclosure and are not intended to limit the scope of protection of this invention. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of this disclosure should be included within the scope of protection of this invention.

Claims

1. A video editing method, the method comprising: Obtain the original video, guidance information, and reference materials for editing; Retrieve multiple reference editing elements; Using the target large model, based on the reference editing elements, the original video, the guidance information, and the reference materials, edit data in natural language form is generated; Based on the edited data, the target video is obtained.

2. The method according to claim 1, wherein, The process of obtaining multiple reference editing elements includes: Extract keywords from the guidance information; Obtain multiple edit elements associated with the keyword, and use them as multiple reference edit elements.

3. The method according to claim 2, wherein, The step of obtaining multiple edit elements associated with the keyword as multiple reference edit elements includes: Search a template library containing multiple video templates for reference video templates related to the keywords; Obtain the editing elements from the reference video template and use them as the reference editing elements.

4. The method according to claim 2, wherein, The step of obtaining multiple edit elements associated with the keyword as multiple reference edit elements includes: From an element library containing multiple editing elements, obtain editing elements related to the keyword as reference editing elements.

5. The method according to claim 1, wherein, The process of generating editing data in natural language form using a target large model, based on the reference editing elements, the original video, the guiding information, and the reference materials, includes: The reference editing elements, the original video, the guidance information, and the reference materials are input into the target large model; The target large model selects the target editing element from the reference editing elements, and generates the editing data based on the target editing element, the original video, the guidance information, and the reference materials.

6. The method according to claim 1, wherein, The method is used in the training process of the target large model; The acquisition of the original video, guidance information, and reference materials for editing includes: Obtain a sample video template from multiple alternative video templates; Obtain the original video corresponding to the sample video template, and use it as the original video; Obtain the description information corresponding to the sample video template, and use it as the guidance information; Obtain the template material corresponding to the sample video template as the reference material.

7. The method according to claim 6, wherein, The process of obtaining multiple reference editing elements includes: Extract the keywords from the guidance information and obtain the first editing element associated with the keywords; Obtain the second editing element corresponding to the sample video template; Based on the first edit element and the second edit element, the plurality of reference edit elements are obtained.

8. The method according to claim 7, wherein, The process of obtaining the plurality of reference edit elements based on the first edit element and the second edit element includes: Randomly select the third edit element; The first edit element, the second edit element, and the third edit element are identified as the plurality of reference edit elements.

9. A video editing apparatus, the apparatus comprising: The first acquisition unit is configured to acquire the original video, guidance information, and reference materials for editing; The second acquisition unit is configured to acquire multiple reference editing elements; The generation unit is configured to utilize the target large model to generate editing data in natural language form based on the reference editing elements, the original video, the guiding information, and the reference materials; The conversion unit is configured to acquire the target video based on the clipping data.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-8.

11. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-8.