Video editing method and device, electronic equipment and storage medium

By encapsulating the video editing process into tasks and building dependencies, the maintainability and accuracy of automatic editing is solved, and a more efficient video editing process is achieved.

CN120378648APending Publication Date: 2025-07-25CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510596169.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing automatic editing technology has problems with insufficient maintainability and accuracy in financial live broadcasts, especially because the use of GCD technology leads to complex code logic and error-prone.

Method used

The various processes during the video editing process are encapsulated into different tasks according to functions, such as uploading tasks, automatic editing tasks and downloading tasks, and scheduled based on task dependencies, and automatically editing is achieved by calling these tasks.

Benefits of technology

Improve the maintainability and accuracy of the automatic editing process, reduce errors, and improve development efficiency and editing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378648A_ABST
    Figure CN120378648A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video editing method and device, electronic equipment and a storage medium, and is suitable for the field of financial science and technology. The method comprises the following steps: acquiring original video data of a target user; calling an uploading task to upload the original video data to a storage end; under the condition that the uploading task is completed, determining an automatic editing task according to a task dependency relationship corresponding to the uploading task; calling the automatic editing task to obtain the original video data from the storage end and editing the original video data to obtain an edited video; under the condition that the automatic editing task is completed, determining a downloading task according to a task dependency relationship corresponding to the automatic editing task; and calling the downloading task to obtain the edited video. According to the embodiment of the invention, the maintainability and accuracy of automatic editing can be improved. The embodiment of the invention can be used for financial live broadcast slice editing, financial product promotion video editing and other scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, is applicable to the field of fintech, and particularly relates to a video editing method and device, an electronic device, and a storage medium. Background Art

[0002] Video editing technology is a technology that performs operations such as cutting, splicing, and adjusting on video resources. Through video editing technology, a new video can be generated. Video editing technology can be applied to fintech scenarios, and thus can be applied to the processing and optimization of financial live content. For example, in financial live broadcasts, through automatic editing technology, key video segments or core key content can be extracted in real time, providing intelligent and automated support for the extraction of wonderful segments and key content after the live broadcast.

[0003] Currently, for the process nodes that take a long time or require waiting during the automatic editing process, concurrent execution is achieved through the GCD technology. Among them, GCD, that is, Grand Central Dispatch, is a multi-core programming technology used for concurrent programming. However, GCD is a low-level application programming interface provided by C language, which is not object-oriented, and developers need to handle the dependencies and priorities between each process node through logical editing by themselves. With the iteration of business logic, using CGD makes the code logic of the automatic editing program become more and more complex, difficult to maintain, and prone to errors. Therefore, how to improve the maintainability and accuracy of automatic editing has become an urgent technical problem to be solved. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a video editing method and device, an electronic device, and a storage medium, aiming to improve the maintainability and accuracy of automatic editing.

[0005] To achieve the above object, a first aspect of the embodiments of the present application proposes a method, and the method includes:

[0006] Obtain the original video data of the target user;

[0007] Call an upload task to upload the original video data to the storage end;

[0008] In the case where the upload task is completed, determine an automatic editing task according to the task dependency relationship corresponding to the upload task;

[0009] Call the automatic editing task to obtain the original video data from the storage end and perform editing to obtain an edited video;

[0010] In the case where the automatic editing task is completed, determine a download task according to the task dependency relationship corresponding to the automatic editing task;

[0011] Invoke the download task to obtain the clipped video.

[0012] In some embodiments, the method further includes:

[0013] Obtain a first completion progress of the upload task, a second completion progress of the automatic clip task, and a third completion progress of the download task;

[0014] Obtain a first progress weight of the upload task, a second progress weight of the automatic clip task, and a third progress weight of the download task;

[0015] According to the first progress weight, the second progress weight, and the third progress weight, perform a weighted sum of the first completion progress, the second completion progress, and the third completion progress to obtain the total clip progress.

[0016] In some embodiments, obtaining the original video data of the target user includes:

[0017] Invoke an export task to export the original video data according to the video material input by the target user, where the original video data includes material videos, subtitles, dubbing, and special effects corresponding to the videos;

[0018] Correspondingly, invoking the upload task to upload the original video data to the storage end includes:

[0019] In the case where the export task is completed, determine the upload task according to the task dependency corresponding to the export task;

[0020] Invoke the upload task to upload the original video data to the storage end.

[0021] In some embodiments, invoking the upload task to upload the original video data to the storage end includes:

[0022] Invoke the upload task to upload the original video data to the storage end and receive a first video link returned by the storage end, where the first video link is used to jump to the storage address corresponding to the original video data to obtain the original video data.

[0023] In some embodiments, invoking the automatic clip task to obtain the original video data from the storage end and perform a clip to obtain a clipped video includes:

[0024] Send the first video link and the parameters of the automatic clip task to the server, so that the server obtains the original video data according to the first video link and performs an automatic clip task according to the original video data and the parameters of the automatic clip task;

[0025] Receive a second video link returned by the storage end, where the second video link is generated after the server completes the automatic editing task and uploads the edited video to the storage end, and the second video link is used to jump to the storage address corresponding to the edited video to obtain the edited video.

[0026] In some embodiments, the invoking the download task to obtain the edited video includes:

[0027] Invoke the download task to download the edited video according to the second video link;

[0028] When the download task is completed, determine a save task according to the task dependency corresponding to the download task;

[0029] Invoke the save task to save the downloaded edited video locally.

[0030] In some embodiments, after invoking the download task to obtain the edited video, the method further includes:

[0031] Invoke a publishing task to publish the edited video to a video list and generate a request interface, so that a user can access the video list based on the request interface to obtain the edited video.

[0032] To achieve the above object, a second aspect of the embodiments of the present application proposes a video editing device, and the device includes:

[0033] A video acquisition module, configured to acquire original video data of a target user;

[0034] A video upload module, configured to invoke an upload task to upload the original video data to a storage end;

[0035] A task determination module, configured to determine an automatic editing task according to the task dependency corresponding to the upload task when the upload task is completed;

[0036] An automatic editing module, configured to invoke the automatic editing task to obtain the original video data from the storage end and perform editing to obtain an edited video;

[0037] The task determination module is further configured to determine a download task according to the task dependency corresponding to the automatic editing task when the automatic editing task is completed;

[0038] A video download module, configured to invoke the download task to obtain the edited video.

[0039] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0040] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0041] The video editing method, device, electronic device, and storage medium proposed in the present application encapsulate each process in automatic video editing into various tasks according to different functions, such as an upload task, an automatic editing task, a download task, etc. Based on the encapsulated tasks, the tasks are scheduled. By invoking the upload task, the original video data of the target user is uploaded to the storage end, and then based on the task dependency relationship corresponding to the upload task, the subsequent automatic editing task is determined to be executed. By invoking the automatic editing task, the original video data is obtained from the storage end and edited to obtain an edited video. After the editing is completed, based on the task dependency relationship corresponding to the automatic editing task, the subsequent download task is determined to be executed, and the edited video is obtained by invoking the download task. In the above process, by encapsulating each process of automatic editing into corresponding tasks according to functions and then constructing the dependency relationships between the tasks to achieve automatic editing, users can maintain and manage the automatic editing process by constructing and modifying the dependency relationships between the tasks. When an error occurs in automatic editing, it is only necessary to maintain the relevant tasks based on the error occurrence point, which improves the maintainability of the automatic editing process. And based on the task dependency relationship, the execution logic and order of the tasks are arranged, which is not easy to make mistakes, thereby improving the accuracy of automatic editing.

[0042] Other features and advantages of the present application will be described in the subsequent description, and part of them will become obvious from the description or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the description, claims, and drawings. Description of the Drawings

[0043] Figure 1 is a flowchart of the video editing method provided by the embodiments of the present application;

[0044] Figure 2 is Figure 1 a flowchart of steps S101 and S102 in

[0045] Figure 3 is Figure 1 a flowchart of step S103 in

[0046] Figure 4 is Figure 1 the flowchart of step S104 in

[0047] Figure 5 is Figure 1 the flowchart of step S106 in

[0048] Figure 6 is the flowchart of short video clip of financial live broadcast slice provided by the example of the present application;

[0049] Figure 7 is the structural schematic diagram of the video clip device provided by the embodiment of the present application;

[0050] Figure 8 is the hardware structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0051] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0054] First, several nouns involved in the present application are analyzed:

[0055] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.

[0056] Automatic Editing: It refers to the process of automatically completing video editing using artificial intelligence technology through software or tools. Such tools can automatically identify key frames, scene transitions, audio change points, etc. in the video according to preset rules or algorithms, and perform corresponding editing operations, such as deleting invalid segments, adding transition effects, and generating subtitles.

[0057] Transformer Model: It is a deep learning model based on the self-attention mechanism and is widely used in fields such as natural language processing (NLP) and computer vision. The core structure of the Transformer model includes an encoder and a decoder, and each part is stacked by multiple identical layers.

[0058] VideoMAE (Video Masked Autoencoders): It is a model for video self-supervised pre-training, aiming to improve the representation ability of video data through masking and reconstruction. The core idea of this model is to use the masked autoencoder framework to mask and reconstruct pixel blocks in the video sequence, thereby learning the spatio-temporal structure features of the video.

[0059] VideoPrism: It is a general video encoder aiming to handle various video understanding tasks, such as video classification, localization, retrieval, subtitle generation, and question answering, etc. This model is based on the Vision Transformer (ViT) architecture and adopts a spatio-temporal decomposition design to retain the spatial and temporal dimension information in the video.

[0060] Software Development Kit (SDK): It is a collection of tools, libraries, and documents used to assist in the development of specific types of software. It aims to simplify the software development process, reduce the workload of developers, and improve development efficiency.

[0061] In the related art, in the field of fintech, automatic editing technology is usually used to process and optimize financial live broadcast content. For example, in a financial live broadcast, through automatic editing technology, key video segments or core key content can be extracted in real time, providing intelligent and automated support for the extraction of highlight segments and key content after the live broadcast.

[0062] Currently, for the process nodes that take a long time or require waiting during the automatic editing process, concurrent execution is achieved through the GCD technology. Among them, GCD, that is, Grand Central Dispatch, is a multi-core programming technology used for concurrent programming. However, GCD is a low-level application programming interface provided by C language, which is not object-oriented, and developers need to handle the dependencies and priorities between each process node through logical editing by themselves. With the iteration of business logic, using CGD makes the code logic of the automatic editing program become more and more complex, difficult to maintain and prone to errors. Therefore, how to improve the maintainability and accuracy of automatic editing has become a technical problem to be solved urgently.

[0063] The embodiments of the present application provide a video editing method, device, electronic device and storage medium, aiming to improve the maintainability and accuracy of automatic editing.

[0064] The video editing method, device, electronic device and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the video editing method in the embodiments of the present application is described.

[0065] The embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theory, method, technology and application system.

[0066] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0067] The video editing method provided by the embodiments of the present application relates to the technical field of data processing and utilizes artificial intelligence technology. The video editing method provided by the embodiments of the present application can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the video editing method, etc., but is not limited to the above forms.

[0068] The present application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0069] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0070] Figure 1 is an optional flowchart of the video editing method provided by the embodiments of the present application, Figure 1 The method in may include but is not limited to steps S101 to S106.

[0071] Step S101, obtain the original video data of the target user;

[0072] Step S102, call an upload task to upload the original video data to the storage end;

[0073] Step S103, when the upload task is completed, determine an automatic editing task according to the task dependency corresponding to the upload task;

[0074] Step S104, call the automatic editing task to obtain the original video data from the storage end and perform editing to obtain an edited video;

[0075] Step S105, when the automatic editing task is completed, determine a download task according to the task dependency corresponding to the automatic editing task;

[0076] Step S106, call the download task to obtain the edited video.

[0077] In step S101 of some embodiments, the target user refers to a user who needs to perform automatic video editing. The original video data refers to the video data that needs to be provided by the target user and needs to be edited.

[0078] Taking the fintech scenario as an example, when a financial institution is live streaming an insurance product, it is necessary to edit the live content, insert target content into the live content, or generate sliced short videos based on the live content. In this case, the live broadcaster is the target user. For example, a user who conducts a financial live stream can perform automatic editing on the live content to extract key video clips or core key points, providing intelligent and automated support for the extraction of wonderful clips and key points after the live stream. Another example is that when a financial institution conducts online training or product promotion, the video editing method of the present application is used to automatically edit the training video or promotion video, automatically extract key clips such as key points of the explanation and frequently asked questions in customer interactions, and generate short videos.

[0079] In step S102 of some embodiments, the upload task means encapsulating the relevant code or program responsible for the video upload function into a task instance based on NSOperation. Among them, NSOperation is a multi-threaded solution and an abstract class used to encapsulate a single task operation.

[0080] The storage end refers to the server used to provide storage services. Exemplarily, video data can be uploaded to OSS. OSS (Object Storage Service) is an object-based distributed storage solution designed to provide users with the ability to store data in large quantities, securely, with high reliability, and at low cost. OSS stores data in the cloud in the form of objects, and each object contains the data itself, metadata, and a unique identifier, thereby achieving efficient management of large-scale unstructured data.

[0081] In step S103 of some embodiments, the task dependency relationship refers to the preset dependency relationships between various tasks. For example, if task B depends on task A, then task B needs to be executed after task A is completed, and based on the dependency relationship, task B corresponding to it will be automatically executed after task A is completed.

[0082] The automatic video editing task means that the relevant code or program responsible for the automatic video editing function is encapsulated as a task instance based on NSOperation.

[0083] In step S104 of some embodiments, the edited video refers to the video obtained by editing the original video data by executing the automatic video editing task.

[0084] In step S105 of some embodiments, the download task means that the relevant code or program responsible for the video download function is encapsulated as a task instance based on NSOperation.

[0085] In step S106 of some embodiments, saving the edited video to the local refers to saving the edited video in the local storage space of the server or terminal where the video editing method of the present application is deployed for subsequent viewing or calling.

[0086] Steps S101 to S106 shown in the embodiments of the present application encapsulate each process in automatic video editing into respective tasks according to different functions, such as an upload task, an automatic editing task, a download task, etc. Based on the encapsulated tasks, the tasks are scheduled. By invoking the upload task, the original video data of the target user is uploaded to the storage end, and then based on the task dependency relationship corresponding to the upload task, the subsequent automatic editing task is determined to be executed. By invoking the automatic editing task, the original video data is obtained from the storage end and edited to obtain a clipped video. After the editing is completed, based on the task dependency relationship corresponding to the automatic editing task, the subsequent download task is determined to be executed. By invoking the download task, the clipped video obtained through the automatic editing task is obtained. In the above process, by encapsulating each process of automatic editing into corresponding tasks according to functions and then constructing the dependency relationships between the tasks to achieve automatic editing, and at the same time using the storage end to upload and invoke the original video data, it is not necessary for different tasks to directly transfer relevant video data, and the independence between different tasks (such as the upload task and the automatic editing task) is stronger. Thus, the user can conveniently construct and modify the dependency relationships between the tasks to achieve the maintenance and management of the automatic editing process. When an error occurs in the automatic editing, it is only necessary to maintain the relevant tasks based on the error occurrence point, which improves the maintainability of the automatic editing process. And based on the task dependency relationships, the execution logic and order of the tasks are arranged, which is not easy to make mistakes, thereby improving the accuracy of automatic editing.

[0087] By encapsulating each function in the automatic editing process into corresponding tasks and scheduling based on task dependency relationships, if an error occurs during the execution of automatic editing, it is only necessary to go back to the previous task of the task where the error occurred and retry, without having to start the entire process over again, which can better handle possible errors during the error execution process and improve the processing efficiency.

[0088] Moreover, the user can adjust the dependency relationships and priorities between the tasks according to the selection to optimize the overall process of automatic editing according to actual needs. Since the encapsulated tasks can be directly reused, there is no need to re-edit the logical order of the entire process, which can reduce the development time and improve the efficiency.

[0089] Please refer to Figure 2 , in step S101 of some embodiments, step S101 may include but is not limited to step S201:

[0090] Step S201, invoking an export task to export the original video data according to the video material input by the target user, where the original video data includes the material video, the subtitles corresponding to the video, the dubbing, and the special effects.

[0091] Correspondingly, step S102 may include but is not limited to steps S202 and S203:

[0092] Step S202, in the case where the export task is completed, determine the upload task according to the task dependency corresponding to the export task;

[0093] Step S203, call the upload task to upload the original video data to the storage end.

[0094] In step S201 of some embodiments, the export task means encapsulating the relevant code or program responsible for the function of exporting the complete video file into a task instance based on NSOperation. Exemplarily, the export task can export a complete video file including material videos, video subtitles, dubbing, special effects and other data from the album or video material package of the target user by calling the Meishe SDK.

[0095] The original video data is the complete video file exported by calling the export task. In other examples, the video input by the target user can also be directly used as the original video data for subsequent editing.

[0096] In step S202 of some embodiments, through the pre - constructed task dependency, it can be determined that the next task of the export task is the upload task. In some other embodiments, when it is not necessary to upload the original video data, the next task of the export task can also be an automatic editing task, and the automatic editing task directly obtains the complete video data exported based on the export task for editing.

[0097] In the above steps S201 to S203, the video material is pre - processed through the export task to export a complete video file including material videos, video subtitles, dubbing, special effects and other data, which can be more convenient for subsequent editing and improve the editing efficiency. Uploading the exported video data to the storage end through the upload task can facilitate other tasks that need video data to call relevant data from the storage end. When there are multiple tasks that need to call the original video data, directly call it from the storage end, and the upload task only needs to upload the data to the storage end without the need to separately construct the data transmission relationship between the upload task and other tasks, which improves the reusability of the tasks.

[0098] In some embodiments, a save task may also be set between the export task and the upload task. The save task means encapsulating the relevant code or program responsible for the function of saving data files into a task instance based on NSOperation. By calling the save task, the complete video file exported by the export task is saved, and the required file can be directly retrieved later, avoiding repeated calls to the export task. Among them, the task dependency relationships between the export task, the save task, and the upload task can be set as follows: the save task depends on the completion of the export task to trigger execution, and the upload task depends on the completion of the save task to trigger execution; or both the save task and the upload task depend on the completion of the export task to trigger execution, and no task dependency relationship is established between the save task and the upload task.

[0099] Please refer to Figure 3 , in some embodiments, step S103 may include but is not limited to steps S301 to S302:

[0100] Step S301, call the upload task to upload the original video data to the storage end;

[0101] Step S302, receive the first video link returned by the storage end, where the first video link is used to jump to the storage address corresponding to the original video data to obtain the original video data.

[0102] In step S301 of some embodiments, after uploading the original video data to the storage end, the storage end will generate a corresponding link according to the storage location of the original video data, that is, the first video link.

[0103] In step S302 of some embodiments, through the first video link, the storage end storing the original video data can be accessed, and the original video data corresponding to the first video link can be obtained.

[0104] Exemplarily, taking the enterprise collaborative management software (Internet Business Operation System, IBOS) as the storage end as an example, the user first creates a Bucket (storage space) in IBOS, and uploads the original video data to the corresponding Bucket of the IBOS server through the upload task. After the upload is successful, a fileKeys (file name list) will be returned. A basic file link is generated by splicing the IBOS domain name, bucket, and fileKey. Save this link to the background. When entering the page to request resources next time, the background returns an access link (i.e., the first video link) of the IBOS server with a token (token), and this access link can truly access the IBOS resources.

[0105] In the above steps S301 to S302, the original video data is uploaded to the storage end and the first video link returned by the storage end is obtained. When other tasks after the upload task, such as the automatic editing task, need to obtain the original video data, the storage end can be accessed through the first video link to obtain the relevant data. In this way, there is no need to locally store a large amount of original video data, and the transfer of the original video data between different tasks only needs to transfer the corresponding video link, saving storage resources and transmission resources. And when the subsequent automatic editing task or other tasks go wrong, the original video data that has been uploaded to the server can still be directly called through the first video link without having to re-execute the upload task.

[0106] Before step S104 in some embodiments, the video editing method further includes pre-training an automatic editing model, which is used to automatically edit the original video based on the original video data and the user's editing requirements to obtain an edited video, where the user's editing requirements can be carried in the parameters of the automatic editing task. Specifically, the language model can be a model based on transformer, the VideoPrism model, or the VideoPrism model, etc.

[0107] Taking the VideoPrism model as an example, the training process of this model includes:

[0108] The first stage, video-text contrast training: Align the video encoder and the text encoder through contrastive learning to learn the semantic relationship between the video and the text. Use all video-text pairs for contrastive learning, and minimize the symmetric cross-entropy loss to guide the video encoder to learn rich visual semantics. The purpose of this stage is to establish a basic semantic video embedding representation to lay the foundation for subsequent masked video modeling. Exemplarily, a global-local distillation mechanism can also be introduced in this stage to ensure that the model not only pays attention to the appearance information of the video but also can capture the motion information. In addition, a token shuffling scheme is also adopted to further enhance the model's ability to understand the video content. The second stage, masked video modeling: Enhance the model's understanding of the video content by predicting the masked video blocks, especially capturing the dynamic information in the video. In this stage, the model only relies on pure video data instead of relying on the noisy text data used in the first stage. The focus of this stage is to capture more motion information through masked video modeling and use the semantic video embedding obtained in the first stage as a guide to improve the overall performance of the model.

[0109] Please refer to Figure 4 , in some embodiments, step S104 may include but is not limited to steps S401 to S402:

[0110] Step S401: Send the first video link and the parameters of the automatic editing task to the server, so that the server can obtain the original video data according to the first video link and execute the automatic editing task according to the original video data and the parameters of the automatic editing task;

[0111] Step S402: Receive the second video link returned by the storage end. The second video link is generated after the server completes the automatic editing task and uploads the edited video to the storage end. The second video link is used to jump to the storage address corresponding to the edited video to obtain the edited video.

[0112] In step S401 of some embodiments, the server refers to a server on which the pre-trained automatic editing model provided in the above embodiments is deployed, or a platform that can provide video editing services.

[0113] The parameters of the automatic editing task refer to the parameters used to feedback the editing requirements of the target user and instruct the server to perform video editing according to the user's requirements.

[0114] Exemplarily, the parameters of the automatic editing task include: draft name, template configuration, draft position and quantity, which are used to set the organizational structure of video editing, such as the position arrangement of drafts and the selection of editing templates; video length, editing style, rhythm, audio selection parameters; transition effects, these parameters are used to control the transition method between video segments to ensure video smoothness; subtitle parameters, including text size, color, border size, etc., which are used to adjust the display effect of subtitles; dubbing parameters, such as voice actor selection, speech rate, scene setting, etc., which are used to generate dubbing effects; video resolution, editing speed, these parameters affect the visual quality and playback speed of the video; special effect selection parameters, which are used to add corresponding special effects, such as transitions, filters, etc. to the edited video.

[0115] In step S402 of some embodiments, the second video link refers to the link returned by the storage end after the automatic editing task uploads the edited video to the storage end. Through the second video link, the storage end storing the edited video can be accessed, and the edited video corresponding to the second video link can be obtained.

[0116] In the above steps S401 and S402, the server performs video editing according to the first video link and the parameters of the automatic editing task, which can save local computing resources. By using the server for video editing, the editing efficiency can be improved, and the hardware configuration requirements for the client (local) can be reduced. The server uploads the edited video to the storage end, and then obtains the corresponding edited video by returning the second video link through the storage end. By using the storage end as a transfer for uploading and obtaining the edited video, there is no need to directly establish a data transmission connection between the server and the subsequent tasks that need to obtain the edited video, which can make it more convenient to call the server and each task, and improve the maintainability of the video editing process.

[0117] Please refer to Figure 5 , in some embodiments, step S106 may include but is not limited to steps S501 to S503:

[0118] Step S501, call the download task to download the edited video according to the second video link;

[0119] Step S502, when the download task is completed, determine the save task according to the task dependency corresponding to the download task;

[0120] Step S503, call the save task to save the downloaded edited video to the local.

[0121] In step S501 of some embodiments, the download task accesses the storage end through the second video link and downloads the edited video corresponding to the second video link.

[0122] In steps S502 and S503 of some embodiments, the save task is triggered to execute depending on the completion of the download task, and the edited video downloaded by the download task is saved to the local client for the user to view. It should be noted that the save task called after the download task in this embodiment and the save task called after the export task in the above embodiment can reuse the same save task.

[0123] The above steps S501 to S503 realize that after the download task downloads the complete edited video, the edited video is automatically saved to the local through the task dependency, for the convenience of the user to view. Through the save task encapsulated based on NSOperation, it is possible to build task dependencies with other tasks and reuse the same save task for data storage operations when saving is required, without repeatedly writing multiple save nodes as in the prior art.

[0124] In some embodiments, after step S106, it further includes: calling the publish task to publish the edited video to the video list and generating a request interface, so that the user can access the video list based on the request interface to obtain the edited video.

[0125] In this embodiment, the publishing task means encapsulating the relevant code or program responsible for the video publishing function into a task instance based on NSOperation. The publishing task notifies the server by sending a request that the clipped video has been saved locally, and the clipped video can be published to the video list so that other users can obtain the clipped video from the video list through the request interface.

[0126] In some embodiments, the video clipping method of the present application further includes:

[0127] Obtaining the first completion progress of the upload task, the second completion progress of the automatic clipping task, and the third completion progress of the download task;

[0128] Obtaining the first progress weight of the upload task, the second progress weight of the automatic clipping task, and the third progress weight of the download task;

[0129] According to the first progress weight, the second progress weight, and the third progress weight, perform a weighted sum of the first completion progress, the second completion progress, and the third completion progress to obtain the total clipping progress.

[0130] In this embodiment, the first completion progress refers to the progress of uploading the original video data in the upload task, the second completion progress refers to the completion progress of video clipping in the automatic clipping task, and the third completion progress is the download progress of downloading the clipped video from the storage end in the download task.

[0131] The first progress weight refers to the proportion of the progress of the upload task in the overall progress, the second progress weight refers to the proportion of the progress of the automatic clipping task in the overall progress, and the third progress weight refers to the proportion of the progress of the download task in the overall progress.

[0132] After obtaining the total clipping progress, it can be displayed to the user on the front-end interface so that the user can know the progress of the current video clipping.

[0133] Exemplarily, assume that the first progress weight of the uploaded video progress in the total progress is 50%; the second progress weight of the automatic clipping progress in the total progress is 25%; the third progress weight of the download progress in the total progress is 25%. Then, the overall total progress of video clipping = upload task progress * 50% + automatic clipping task progress * 25% + download task progress * 50%.

[0134] Exemplarily, when multiple upload tasks are set for parallel upload, each upload task only needs to care about its own progress. For each upload task, its progress value is set to 0 when it is unfinished, and the progress value is set to 1 when the task is completed. Add up the progress values of all upload tasks and divide by the number of upload tasks to obtain the overall progress of the upload tasks. For example, there are 5 upload tasks, three of which have been completed, with a progress value of 1, and the other two are unfinished, with a progress value of 0. The total progress of the final upload tasks is 3 / 5 = 0.6, which is 60% when converted to a percentage.

[0135] In the above embodiment, different functions are respectively encapsulated as different tasks (such as upload tasks, automatic editing tasks, download tasks, etc.) based on NSOperation. When monitoring the progress of video editing, it is only necessary to determine the progress of each task itself respectively, and perform weighted summation in combination with the preset weights, and then the overall progress of video editing can be obtained. Even if the task dependency relationships of each task are modified based on different video editing requirements, changing the call order and reuse times of the tasks, the total progress of video editing can still be determined based on the same algorithm, without having to modify or adjust the progress monitoring algorithm again.

[0136] Next, a specific overall description of the terminal testing method of the present application will be given through an example. It can be understood that the following embodiments are all for better exemplarily illustrating the terminal testing method of the present application, and no specific limitations are made.

[0137] Please refer to Figure 6 , taking the example of a financial institution editing a live broadcast to generate sliced short videos for financial product promotion in a fintech scenario.

[0138] Before video editing, based on NSOperation, the overall video editing process is respectively encapsulated according to different functions to obtain an export task, a save task, an upload task, an automatic editing task, a download task, and a release task, and the task dependency relationships between each task are constructed.

[0139] First, the client of video editing obtains the live video saved by the financial institution and determines whether it supports exporting to the album. If it supports, the export task is called, and a complete video file including material video, video subtitles, dubbing, and special effects, etc. is exported based on the live video as the original video data.

[0140] According to the task dependency relationship, the client depends on the completion of the export task to trigger the save task, and calls the save task to save the exported original video data.

[0141] Based on the task dependency, the client depends on the completion of the save task to trigger the upload task, calls the upload task to upload the original video data to the storage end, and receives the first video link returned by the storage end. In this example, the upload task is reused multiple times to enable multiple upload tasks to upload the original video data in parallel.

[0142] Based on the task dependency, the client depends on the completion of the upload task to trigger the automatic editing task and calls the automatic editing task. The first video link and the parameters of the automatic editing task are sent to the server so that the server can obtain the original video data based on the first video link and execute the automatic editing task according to the original video data and the parameters of the automatic editing task. After the server completes the video editing to obtain the sliced short video for promoting financial products, it uploads the sliced short video to the storage end, and the storage end generates the second video link. The client receives the second video link returned by the storage end.

[0143] Based on the task dependency, the client depends on the completion of the automatic editing task to trigger the download task, calls the download task to access the storage end according to the second video link, and downloads the sliced short video.

[0144] Based on the task dependency, the client depends on the completion of the download task to trigger the save task, calls the save task, and saves the downloaded sliced short video locally on the client. Among them, the two saves in this example can reuse the same save task.

[0145] Based on the task dependency, the client depends on the completion of the save task to trigger the publish task, calls the publish task, and sends a request to the server to notify it that the sliced short video has been saved locally and the sliced short video can be published to the video list so that other users can obtain the sliced short video from the video list through the request interface.

[0146] In the above example, each process in the automatic video editing is encapsulated into individual tasks according to different functions, such as an upload task, an automatic editing task, a download task, etc. Based on the encapsulated tasks, the tasks are scheduled. By invoking the upload task, the original video data of the target user is uploaded to the storage end, and then based on the task dependency relationship corresponding to the upload task, the subsequent automatic editing task is determined. By invoking the automatic editing task, the original video data is obtained from the storage end and edited to obtain the edited video. After the editing is completed, based on the task dependency relationship corresponding to the automatic editing task, the subsequent download task is determined. By invoking the download task, the edited video obtained through the automatic editing task is acquired. In the above process, by encapsulating each process of automatic editing into corresponding tasks according to functions, and then constructing the dependency relationships between tasks to achieve automatic editing, and at the same time using the storage end to upload and invoke the original video data, it is not necessary to directly transmit relevant video data between different tasks, and the independence between different tasks (such as the upload task and the automatic editing task) is stronger. Thus, users can conveniently construct and modify the dependency relationships between tasks to achieve the maintenance and management of the automatic editing process. When an error occurs in the automatic editing, it is only necessary to maintain the relevant tasks based on the error occurrence point, which improves the maintainability of the automatic editing process. And based on the task dependency relationships, the execution logic and sequence of tasks are arranged, which is not easy to make mistakes, thereby improving the accuracy of automatic editing.

[0147] Please refer to Figure 7 , the embodiment of the present application further provides a video editing device, which can implement the above video editing method. The device includes:

[0148] A video acquisition module, configured to acquire the original video data of the target user;

[0149] A video upload module, configured to invoke the upload task to upload the original video data to the storage end;

[0150] A task determination module, configured to determine the automatic editing task according to the task dependency relationship corresponding to the upload task when the upload task is completed;

[0151] An automatic editing module, configured to invoke the automatic editing task to acquire the original video data from the storage end and perform editing to obtain the edited video;

[0152] The task determination module is further configured to determine the download task according to the task dependency relationship corresponding to the automatic editing task when the automatic editing task is completed;

[0153] A video download module, configured to invoke the download task to acquire the edited video.

[0154] The specific implementation manner of the video clipping device is basically the same as the specific embodiment of the above video clipping method, and will not be elaborated herein.

[0155] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above video clipping method is implemented. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0156] Please refer to Figure 8 , Figure 8 , which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0157] A processor 801, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0158] A memory 802, which can be implemented in the form of a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802, and the processor 801 is used to call and execute the video clipping method of the embodiments of the present application;

[0159] An input / output interface 803, which is used to implement information input and output;

[0160] A communication interface 804, which is used to implement communication interaction between this device and other devices. Communication can be achieved through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);

[0161] A bus 805, which transmits information between various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);

[0162] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 are communicatively connected to each other inside the device through the bus 805.

[0163] The embodiments of the present application further provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above video editing method is implemented.

[0164] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0165] The video editing method, video editing device, electronic device, and storage medium provided by the embodiments of the present application encapsulate each process in automatic video editing into respective tasks according to different functions, such as an upload task, an automatic editing task, a download task, etc. Based on the encapsulated tasks, the tasks are scheduled. By invoking the upload task, the original video data of the target user is uploaded to the storage end, and then based on the task dependency relationship corresponding to the upload task, the subsequent automatic editing task is determined to be executed. By invoking the automatic editing task, the original video data is obtained from the storage end and edited to obtain a clipped video. After the editing is completed, based on the task dependency relationship corresponding to the automatic editing task, the subsequent download task is determined to be executed, and the clipped video is obtained by invoking the download task. In the above process, by encapsulating each process of automatic editing into corresponding tasks according to functions and then constructing the dependency relationships between the tasks to achieve automatic editing, users can maintain and manage the automatic editing process by constructing and modifying the dependency relationships between the tasks. When an error occurs in automatic editing, it is only necessary to maintain the relevant tasks based on the error occurrence point, which improves the maintainability of the automatic editing process. And based on the task dependency relationship, the execution logic and order of the tasks are arranged, which is not easy to make mistakes, thereby improving the accuracy of automatic editing.

[0166] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0167] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0169] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0170] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0171] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0172] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0173] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0174] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0175] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. And the aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0176] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.

Claims

1. A video editing method, characterized in that, The method includes: Obtain the original video data of the target user; Call an upload task to upload the original video data to the storage end; When the upload task is completed, determine an automatic editing task according to the task dependency corresponding to the upload task; Call the automatic editing task to obtain the original video data from the storage end and perform editing to obtain an edited video; When the automatic editing task is completed, determine a download task according to the task dependency corresponding to the automatic editing task; Call the download task to obtain the edited video.

2. The method according to claim 1, wherein The method further includes: Obtain the first completion progress of the upload task, the second completion progress of the automatic editing task, and the third completion progress of the download task; Obtain the first progress weight of the upload task, the second progress weight of the automatic editing task, and the third progress weight of the download task; According to the first progress weight, the second progress weight, and the third progress weight, perform weighted summation on the first completion progress, the second completion progress, and the third completion progress to obtain the total editing progress.

3. The method according to claim 1, characterized in that, The obtaining of the original video data of the target user includes: Call an export task to export the original video data according to the video material input by the target user, where the original video data includes material videos, subtitles corresponding to the videos, dubbing, and special effects; Correspondingly, the calling of the upload task to upload the original video data to the storage end includes: When the export task is completed, determine an upload task according to the task dependency corresponding to the export task; Call the upload task to upload the original video data to the storage end.

4. The method according to claim 1 or 3, characterized in that, The calling of the upload task to upload the original video data to the storage end includes: Call the upload task to upload the original video data to the storage end and receive the first video link returned by the storage end, where the first video link is used to jump to the storage address corresponding to the original video data to obtain the original video data.

5. The method according to claim 4, wherein The calling of the automatic editing task to obtain the original video data from the storage end and perform editing to obtain an edited video includes: Send the first video link and the parameters of the automatic editing task to the server, so that the server obtains the original video data according to the first video link and executes the automatic editing task according to the original video data and the parameters of the automatic editing task; Receive the second video link returned by the storage end, where the second video link is generated after the server completes the automatic editing task and uploads the edited video to the storage end, and the second video link is used to jump to the storage address corresponding to the edited video to obtain the edited video.

6. The method according to claim 5, characterized in that, The calling of the download task to obtain the edited video includes: Call the download task to download the edited video according to the second video link; When the download task is completed, determine a save task according to the task dependency corresponding to the download task; Call the save task to save the downloaded edited video locally.

7. The method according to claim 1, wherein After invoking the download task to obtain the clipped video, the method further includes: Invoking a publishing task to publish the clipped video to a video list and generating a request interface, enabling a user to access the video list based on the request interface to obtain the clipped video.

8. A video editing device, characterized in that, The apparatus includes: A video acquisition module, configured to acquire original video data of a target user; A video upload module, configured to invoke an upload task to upload the original video data to a storage end; A task determination module, configured to determine an automatic clip task according to a task dependency relationship corresponding to the upload task when the upload task is completed; An automatic clip module, configured to invoke the automatic clip task to obtain the original video data from the storage end and perform clipping to obtain a clipped video; The task determination module is further configured to determine a download task according to a task dependency relationship corresponding to the automatic clip task when the automatic clip task is completed; A video download module, configured to invoke the download task to obtain the clipped video.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the video clip method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the video clip method according to any one of claims 1 to 7.