Video editing method and device

The 'thought chain engine' in AI video editing software generates logical editing schemes with interpretable reports, addressing the lack of transparency in current AI editing tools and enabling user-adjustable editing outcomes.

CN120321458APending Publication Date: 2025-07-15SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510525856.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing artificial intelligence video editing software cannot display the video editing logic, resulting in users being unable to understand the editing process and being unable to adjust the video editing targetedly to achieve the desired effect.

Method used

Through the thinking chain engine, the solution to generate video clips and provide interpretability reports, users can understand the editing logic and basis, and make targeted adjustments.

Benefits of technology

Users can clearly understand the video editing process, adjust the editing process as needed, and achieve the desired editing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321458A_ABST
    Figure CN120321458A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video editing method. The video editing method comprises the following steps: acquiring a video material; generating a first scheme of video editing by adopting a thinking chain engine based on the video material, and obtaining logic and basis for generating the first scheme by the thinking chain engine; generating and outputting an interpretability report of the video clip based on the logic and basis; and performing video editing based on the first scheme and the video material. According to the technical scheme provided by the embodiment of the invention, a user can clearly know the logic and basis of intelligent video editing, so that the video editing process can be conveniently and specifically adjusted, and the desired editing effect can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a video editing method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] With the development of social media, people are increasingly inclined to watch videos to acquire knowledge or for entertainment, and the demand for video content is increasing. To improve the efficiency of video editing, video producers are also increasingly inclined to use artificial intelligence video editing software to achieve automatic video editing.

[0003] However, current artificial intelligence video editing software cannot display the editing logic of videos, resulting in users being unable to understand the video editing process and unable to make targeted adjustments to the video editing to achieve the desired effect.

[0004] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention

[0005] Embodiments of the present application provide a video editing method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above-mentioned technical problems.

[0006] One aspect of the embodiments of the present application provides a video editing method, the method comprising: Obtain video materials; Generate a first video editing plan using a thought chain engine based on the video materials, and obtain the logic and basis for the thought chain engine to generate the first plan; Generate and output an interpretability report for video editing based on the logic and basis; Perform video editing based on the first plan and the video materials.

[0007] Optionally, the method further comprises: When receiving an input modification prompt, generate a second video editing plan using a thought chain engine based on the video materials and the modification prompt; Perform video editing based on the second plan and the video materials.

[0008] Optionally, the generating a first video editing plan using a thought chain engine based on the video materials comprises: Perform multimodal analysis on the video materials to obtain the result of multimodal analysis; Generate a first video editing plan using a thought chain engine based on the result of multimodal analysis.

[0009] Optionally, the video material includes video format material, image format material, audio format material, and text format material; Correspondingly, the multi-modal analysis based on the video material to obtain the result of multi-modal analysis includes: Performing visual modal analysis on the video format material and the image format material to obtain a visual modal analysis result; Performing auditory modal analysis on the audio format material to obtain an auditory modal analysis result; Performing text modal analysis on the text format material to obtain a text modal analysis result; Obtaining the multi-modal analysis result based on the visual modal analysis result, the auditory modal analysis result, and the text modal analysis result.

[0010] Optionally, the first scheme for generating a video clip by using a thought chain engine based on the result of the multi-modal analysis includes: Obtaining a description of the input clip requirement; Generating a first scheme for video clip by using a thought chain engine based on the result of the multi-modal analysis and the clip requirement description.

[0011] Optionally, the first scheme for generating a video clip by using a thought chain engine based on the result of the multi-modal analysis and the clip requirement description includes: Using a thought chain engine to match videos in a target knowledge base based on the result of the multi-modal analysis and the clip requirement description to obtain reference videos; Obtaining video clip information of the reference videos; Generating the first scheme based on the video clip information.

[0012] Another aspect of the embodiments of the present application provides a video clip device, and the device includes: An acquisition module, configured to acquire video material; A generation module, configured to generate a first scheme for video clip by using a thought chain engine based on the video material, and acquire the logic and basis for the thought chain engine to generate the first scheme; A reporting module, configured to generate and output an interpretability report for video clip based on the logic and basis; A clip module, configured to perform video clip based on the first scheme and the video material.

[0013] Another aspect of the embodiments of the present application provides a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0014] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0015] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0016] The embodiments of the present application adopting the above technical solutions may include the following advantages: By obtaining video materials, generating a first video editing plan based on the video materials by using a thought chain engine, obtaining the logic and basis for the thought chain engine to generate the first plan, generating an interpretability report for the video editing based on the logic and basis of the first plan, and performing video editing based on the first plan and the video materials, the logic and basis for generating the video editing plan can be obtained through the thought chain engine, and an interpretability report can be generated and output according to the logic and basis of the video editing plan, enabling users to understand the logic and basis of intelligent video editing by viewing the interpretability report, so that the corresponding video editing process can be adjusted targeted when needed to achieve the desired editing effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings exemplarily show the embodiments and form a part of the description, and are used together with the written description of the description to explain the exemplary embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily the same elements.

[0018] Figure 1 Schematically shows the operating environment diagram of the video editing method according to Embodiment 1 of the present application; Figure 2 Schematically shows the flowchart of the video editing method according to Embodiment 1 of the present application; Figure 3 Schematically shows the new process of the video editing method according to Embodiment 1 of the present application; Figure 4 Schematically shows Figure 2 the sub-step flowchart of step S202 in Figure 5 Schematically showsFigure 4 The sub-step flowchart of step S400 in Figure 6 Schematically shows Figure 4 The sub-step flowchart of step S402 in Figure 7 Schematically shows Figure 6 The sub-step flowchart of step S602 in Figure 8 Schematically shows an application example diagram of the video clip method according to the first embodiment of the present application; Figure 9 Schematically shows a block diagram of a video clip device according to the second embodiment of the present application; and Figure 10 Schematically shows a schematic diagram of the hardware architecture of a computer device according to the third embodiment of the present application. Detailed implementation manners

[0019] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0020] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0021] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.

[0022] First, the following provides the term explanations involved in the present application: Chain of Thought (CoT): It is a way of thinking or a method of analyzing problems. It decomposes a complex problem or topic into a series of interrelated links or steps, and reveals the essence and solution of the problem through step-by-step reasoning and analysis.

[0023] Large Language Model (LLM): A type of language model composed of artificial neural networks with a large number of parameters, trained on a large amount of unlabeled text using self-supervised learning or semi-supervised learning.

[0024] Secondly, to facilitate the understanding of the technical solutions provided in the embodiments of the present application by those skilled in the art, the related technologies will be described below: Currently, when artificial intelligence video editing software performs video editing, it is generally a "black box operation" and cannot display the video editing logic, resulting in users being unable to understand the video editing process and unable to adjust the video editing targeted to achieve the desired effect.

[0025] For this reason, the embodiments of the present application provide a video editing technical solution. In this technical solution, a chain of thought engine is used to generate the logic and basis for the video editing solution and generate an interpretability report for video editing, which allows users to clearly understand the process of intelligent video editing, so that they can adjust the corresponding video editing process targeted to achieve the desired editing effect. See the following for details.

[0026] Finally, for the convenience of understanding, an exemplary operating environment is provided below.

[0027] As Figure 1 shown, the environmental schematic diagram includes a service platform 2, a network 4, and a client 6, where: The service platform 2 can be composed of a single or multiple computing devices. The multiple computing devices can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing devices can load virtual machines based on virtual images and / or other data that define specific software (e.g., operating systems, dedicated applications, servers) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices. A hypervisor can be implemented to manage the use of different virtual machines on the same computing device.

[0028] The service platform 2 can be configured to communicate with the client 6, etc. through the network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and combinations thereof, or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.

[0029] The service platform 2 can provide services such as storage, reading, writing, querying, deleting, etc., such as providing video editing services for the client.

[0030] The client 6 can be an electronic device running an operating system such as Windows, Android™, or iOS, such as a smartphone, a tablet device, a laptop computer, a virtual reality device, a gaming device, a set-top box, an in-vehicle terminal, or a smart TV. Based on the above operating systems, various applications can be run, such as an application for running video clips.

[0031] The client 6 can provide / configure a user access page for manipulating the service platform 2 or uploading an object, etc.

[0032] Note that the above devices are exemplary, and in different scenarios or according to different requirements, the number and types of devices can be adjusted.

[0033] The technical solutions of the present application will be introduced below through multiple embodiments. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.

[0034] Embodiment 1 Figure 2 A flowchart of a video clip method according to Embodiment 1 of the present application is schematically shown. It should be noted that the execution subject of the video clip method in the embodiments of the present application can be a client or a service platform. Hereinafter, the service platform will be taken as an example of the execution subject for description.

[0035] As Figure 2 shown, the video clip method may include steps S200 to S206, where: Step S200: Obtain video materials.

[0036] Step S202: Generate a first video clip plan using a thought chain engine based on the video materials, and obtain the logic and basis for the thought chain engine to generate the first plan.

[0037] Step S204: Generate and output an interpretability report for the video clip based on the logic and basis of the first plan.

[0038] Step S206: Perform video clip based on the first plan and the video materials.

[0039] The video editing method provided in this embodiment obtains video material, generates a first video editing plan based on the video material by using a thinking chain engine, obtains the logic and basis for the thinking chain engine to generate the first plan, generates an explainability report for the video editing based on the logic and basis of the first plan, and performs video editing based on the first plan and the video material. The thinking chain engine can be used to generate the logic and basis of the video editing plan, and an explainability report is generated and output based on the logic and basis of the video editing plan, so that users can understand the logic and basis of intelligent video editing by viewing the explainability report, so that they can make targeted adjustments to the corresponding video editing process when necessary to achieve the desired editing effect.

[0040] The following combination Figure 2 , each step in steps S200~S206 and other optional steps are described in detail.

[0041] Step S200 , get the video material.

[0042] Video materials may include video, audio, images, text or special effects, etc. For example, videos or photos taken by users or captions entered by users.

[0043] Specifically, the user can select the corresponding video material from the terminal device, and after the user completes the selection, it can be uploaded to the service platform through the client, so that the service platform can obtain the video material. Optionally, the client can also provide some video materials, such as transition effects, and the user can select the corresponding video material on the client, and the service platform can obtain the corresponding video material according to the user's selection. Alternatively, the user can also upload the address of the video material through the client, such as the URL address, and the service platform can obtain the video material according to the address uploaded by the client.

[0044] Step S202 , based on the video material, a first plan for video editing is generated by using a thinking chain engine, and the logic and basis for the thinking chain engine to generate the first plan are obtained.

[0045] The chain of thought engine can be pre-trained using a large amount of training video materials. When training the chain of thought engine, the chain of thought engine can be guided to divide the video editing process into several parts according to the general order of video editing. For example, it can be divided into "material analysis → narrative goal matching → shot cutting point calculation → transition recommendation → filter parameter optimization → music matching". Then each part is further broken down into several sub-tasks, and the video editing plan is generated step by step according to the sub-tasks, and the logic and basis of each step are given during the reasoning process. Among them, the basic model of the chain of thought engine can adopt a large model or a large language model, and other models can be combined to complete the reasoning of some sub-tasks. For example, when calculating the shot cutting point, a visual model can be called to complete it. For each part or sub-task, the basic model or the combined model can be appropriately pre-trained or fine-tuned to meet the needs of reasoning.

[0046] After obtaining the video materials, the video materials can be input into the trained chain of thought engine, and the trained chain of thought engine is used to generate the first plan for video editing of the currently input video materials, and the logic and basis for the chain of thought engine to generate the first plan are obtained. Among them, the logic and basis of the plan can include: the input of the current task, the analysis of the current input, the decisions made, and the basis for the decisions. For example, if the current task is music matching, the logic and basis of the plan can be: "According to the analysis of XX video materials, the video is related to the theme of a wedding, so the music matching is romantic music, and the specific music is XX".

[0047] Step S204 , generate and output an explainable report for video editing based on the logic and basis of the first plan.

[0048] Specifically, the service platform can generate an explainable report for video editing according to the logic and basis of the first plan, and then output it to the client for display, so that the user can clearly understand the logic basis of each process of video editing. In practical applications, the explainable report can be generated in real time during the reasoning process of the chain of thought engine and output to the client for display in real time, or after the first plan is completed, the explainable report corresponding to the entire first plan can be output to the client for display.

[0049] Step S206 , perform video editing based on the first plan and the video materials.

[0050] After outputting the explainable report, if the user has no modification opinions on the first plan, the video materials can be correspondingly video-edited according to the first plan, and finally the edited video can be obtained.

[0051] In an alternative embodiment, as Figure 3 shown, the video editing method of the embodiment of the present application may further include: Step S300, in the case of receiving an input modification prompt, generate a second video clip plan using the thought chain engine based on the video material and the modification prompt.

[0052] Step S302, perform video clip based on the second plan and the video material.

[0053] Among them, the modification prompt can be a clear instruction, such as modifying a certain transition effect of the video clip to a specific transition effect; it can also be a tendency index, such as "the background music is relatively noisy, modify it to a more soothing music". The modification prompt can be input by the user after seeing the interpretability report; it can also be input by the user according to the effect of the video clip or the preview effect after using the first plan to clip the video.

[0054] Specifically, if the user is not satisfied with the interpretability report or the effect of the first plan for video clipping, the user can input a modification prompt through the client. In the case of receiving the modification prompt input by the user, the service platform uses the thought chain engine for secondary reasoning according to the video material and the modification prompt to generate a second video clip plan, and then performs video clipping according to the second plan and the video material. Optionally, in the case of receiving a modification prompt, targeted adjustment can be made according to the modification prompt on the basis of the first plan. For example, if the modification prompt is to modify the transition effect, only the corresponding transition effect can be modified on the basis of retaining the first plan. Of course, if the modification prompt also has a collateral impact on other processes of video clipping, the collateral impact needs to be considered and the affected processes need to be re-reasoned.

[0055] In this embodiment, by generating a second video clip plan using the thought chain engine based on the video material and the modification prompt in the case of receiving the input modification prompt, and performing video clipping based on the second plan and the video material, the video clipping can be adjusted according to the user's modification requirements, so as to obtain a clipped video that meets the user's needs and improve the flexibility of video clipping.

[0056] In an alternative embodiment, in step S202, using the thought chain engine to generate a first video clip plan based on the video material, as Figure 4 shown, may include: Step S400, perform multimodal analysis on the video material to obtain the result of multimodal analysis.

[0057] Step S402, generate a first video clip plan using the thought chain engine based on the result of the multimodal analysis.

[0058] Specifically, it is possible to first determine which modalities of information the video material includes, and then for each modality of information, adopt the corresponding modality analysis method to perform modality analysis to obtain the results of each modality analysis. Finally, summarize the results of all modality analyses together to obtain the results of multi-modal analysis. For example, if the video material includes video, it can be determined that the video material includes at least two modalities of information, namely video and audio. Then, perform modality analysis on the video modality information and the audio modality information respectively to obtain the results of video modality analysis and the results of audio modality analysis, and summarize to obtain the results of multi-modal analysis. Then, the results of multi-modal analysis can be input into the thought chain engine, and the thought chain engine is used to make various decisions in the video editing process based on the results of multi-modal analysis, and finally generate the first video editing plan.

[0059] In this embodiment, by performing multi-modal analysis based on the video material to obtain the results of multi-modal analysis, and using the thought chain engine to generate the first video editing plan based on the results of multi-modal analysis, the thought chain engine can perform reasoning based on the results of multi-modal analysis of the video material, which can improve the accuracy and efficiency of thought chain reasoning.

[0060] In an alternative embodiment, the video material includes video format material, image format material, audio format material, and text format material. Correspondingly, in step S400, multi-modal analysis is performed based on the video material to obtain the results of multi-modal analysis, as Figure 5 shown, which may include: Step S500, perform visual modality analysis on the video format material and the image format material to obtain the visual modality analysis results.

[0061] Step S502, perform auditory modality analysis on the audio format material to obtain the auditory modality analysis results.

[0062] Step S504, perform text modality analysis on the text format material to obtain the text modality analysis results.

[0063] Step S506, obtain the results of multi-modal analysis based on the visual modality analysis results, the auditory modality analysis results, and the text modality analysis results.

[0064] The visual modality analysis of video format materials can include content recognition, scene classification, motion analysis, shot change detection, etc. Among them, content recognition can be to use object detection algorithms to identify objects, people, scenes, etc. in video frames; scene classification can be to classify various scenes in video frames through a pre-trained image classification model to determine whether it is an indoor scene, an outdoor scene, an urban landscape or a natural landscape, etc.; motion analysis can be to use optical flow algorithms to analyze the dynamic changes in the video frame, such as the motion direction, speed and trajectory of objects, etc.; shot change detection can detect the position and frequency of shot changes through the differences between adjacent frames, so as to determine information such as shot length. The visual modality analysis of image format materials can include content recognition, image classification, feature extraction, composition analysis, etc. Among them, image classification can be to judge the type of the image through a pre-trained image classification model, such as landscape pictures, portrait pictures, product pictures, etc.; feature extraction can be to extract the feature vectors of the image using a neural network (such as CNN) so as to perform similarity comparison or retrieval of images according to the feature vectors, etc.; composition analysis can be to analyze the composition method of the image, such as whether it follows the rule of thirds, symmetric composition, etc. It can be understood that since video format materials include frame images, the visual modality analysis of video format materials can also include the visual modality analysis of image format materials, that is, the visual modality analysis of video format materials can also include content such as feature extraction and composition analysis.

[0065] The auditory modality analysis of audio format materials can include audio classification, emotion analysis, rhythm analysis, frequency analysis, etc. Among them, audio classification can be to use an audio classification model to judge the type of audio, such as background music, dialogue or environmental sound effects, etc.; emotion analysis can use a trained emotion analysis model (such as a model combining mel spectrogram and MobileViT network) to analyze the emotion conveyed by the audio, such as cheerful, sad, tense or relaxed, etc.; rhythm analysis can be to extract the rhythm information of the audio, such as beats, rhythm speed, to understand the rhythm characteristics of the audio; frequency analysis can be to analyze the frequency components of the audio and determine the energy analysis of different frequency bands, such as high-frequency part, low-frequency part, etc.

[0066] The text modality analysis of text format materials can include text classification, emotion analysis, keyword extraction, entity recognition, etc.; among them, text classification can use text classification algorithms in natural language processing to classify the text to determine whether the text is descriptive, explanatory or critical, etc.; emotion analysis can be to use natural language processing to analyze the emotional tendency expressed by the text, such as positive, negative or neutral, etc.; keyword extraction can be to extract the key information and important words in the text through keyword extraction algorithms to understand the core content of the text; entity recognition can be to identify the entities in the text, such as person names, place names, product names, etc., to facilitate the establishment of entity associations.

[0067] After obtaining the visual modality analysis result, the auditory modality analysis result, and the text modality analysis result, these results can be aggregated together as the multimodal analysis result, or the visual modality analysis result, the auditory modality analysis result, and the text modality analysis result can be combined for further analysis to analyze the correlations between these modalities, obtain the result of the correlation analysis, and then aggregate the visual modality analysis result, the auditory modality analysis result, the text modality analysis result, and the correlation analysis result together as the multimodal analysis result. In addition, in practical applications, when the video material does not include material of a certain format, the analysis of the material of that format can be skipped.

[0068] In this embodiment, by performing visual modality analysis, auditory modality analysis, and text modality analysis on the video material, and obtaining the multimodal analysis result based on the visual modality analysis result, the auditory modality analysis result, and the text modality analysis result, effective multimodal analysis of the video material can be performed, which is beneficial to improving the accuracy of the video editing decision-making of the thinking chain engine.

[0069] In an alternative embodiment, in step S402, based on the result of the multimodal analysis, a first video editing plan is generated using the thinking chain engine, as Figure 6 shown, which may include: Step S600, obtain the input description of the editing requirements.

[0070] Step S602, based on the result of the multimodal analysis and the description of the editing requirements, generate a first video editing plan using the thinking chain engine.

[0071] The description of the editing requirements may be a description of requirements such as the style, effect, duration, background music, and transitions of the video. For example, it may be "produce a travel video showing natural scenery, with a brisk rhythm and a duration of about 3 minutes". Optionally, an input box can be provided on the client side to allow users to input the description of the editing requirements by entering text; or, a voice input interface can be provided, and the description of the editing requirements can be obtained by using speech recognition based on the voice input by the user.

[0072] In the case of obtaining the input description of the editing requirements, the result of the multimodal analysis and the description of the editing requirements can be input into the thinking chain engine, and the thinking chain engine is used to generate a first video editing plan. Among them, the description of the editing requirements can be input into the thinking chain engine as a prompt (prompt), and the thinking chain engine infers the result of the multimodal analysis under the prompt of the description of the editing requirements to generate a first video editing plan. The thinking chain engine can perform natural language processing on the description of the editing requirements to determine its semantics, so as to clarify the specific requirements of the user.

[0073] In this embodiment, by obtaining the input description of the editing requirements, a first video editing scheme is generated using a chain-of-thought engine based on the results of multimodal analysis and the editing requirements, which can generate a video editing scheme in combination with the actual needs of the user, so that the video editing scheme better meets the needs of the user and improves the accuracy of video editing.

[0074] In an alternative embodiment, in step S602 above, a first video editing scheme is generated using a chain-of-thought engine based on the results of multimodal analysis and the description of the editing requirements, as Figure 7 shown, and may include: Step S700, using a chain-of-thought engine to match videos in the target knowledge base based on the results of multimodal analysis and the description of the editing requirements to obtain reference videos.

[0075] Step S702, obtaining the video editing information of the reference videos.

[0076] Step S704, generating a first scheme based on the video editing information.

[0077] Among them, the target knowledge base may include an online knowledge base, an offline knowledge base, a knowledge graph, a material library, etc. The target knowledge base may include a large number of edited videos, specifically including videos using various background music, transition effects, filters, etc. The information of the videos may specifically include information such as video titles, comments, bullet screens, tags, descriptions, popularity, and types.

[0078] Specifically, after inputting the results of multimodal analysis and the description of editing requirements into the chain-of-thought engine, the chain-of-thought engine can match videos in the target knowledge base according to the input information to obtain reference videos. For example, the chain-of-thought engine can match the results of multimodal analysis, the description of editing requirements with information such as the title, comments, bullet screens, and tags of the video, and use the videos with a matching degree reaching the preset threshold as reference videos. When there are multiple matching videos, further screening (such as further screening according to popularity) can be performed to obtain reference videos. When the target knowledge base contains a knowledge graph, the information in the knowledge graph can also be combined to obtain reference videos. After obtaining the reference videos, the reference videos can be analyzed to obtain the video editing information of the reference videos, where the video editing information can include information related to video editing such as the transitions used, background music, filters, and duration. After obtaining the video editing information of the reference videos, a first scheme for video editing of the current video material can be generated based on the obtained video editing information. For example, the transition scheme used in the reference video can be used as the transition scheme for the current video material. When there are multiple reference videos or the video editing information contains multiple different types of information, the final editing decision can be further determined according to other metrics or by combining information such as the results of multimodal analysis and the description of editing requirements. For example, if the target knowledge base is a certain video platform, after matching according to the results of multimodal analysis and the description of editing requirements on this video platform, videos A, B, and C are determined as reference videos; at the same time, the description of editing requirements is a 1-minute video duration, and the video duration of video B is close to the duration described in the editing requirements. Since the transition effect has a great relationship with time, the transition used in video B can be selected as the transition scheme for the current video material; another example is that if the result of multimodal analysis is an emotion of "lively", and the background music of video C is more in line with the "lively" emotion, then the background music in video C can be selected as the background music scheme for the current video material.

[0079] In this embodiment, by using the chain-of-thought engine to match videos in the target knowledge base based on the results of multimodal analysis and the description of editing requirements to obtain reference videos, obtaining the video editing information of the reference videos, and generating a first scheme based on the video editing information, the editing method of existing edited videos can be effectively used as a reference to generate the video editing scheme of the current video material, effectively improving the quality and efficiency of video editing.

[0080] To make this application easier to understand, the following is combined with Figure 8 to provide an exemplary application. As Figure 8 shown, the video editing method can generally include the following content: 1. The material selection module can select videos, pictures, etc. as video materials according to the user's selection, and can also receive the description of editing requirements input by the user as a prompt word; 2. A multimodal analysis unit can be used to perform multimodal analysis on video materials to identify information such as creation themes, audio emotions, and text semantics; 3. The chain-of-thought reasoning engine matches the results of the multimodal analysis unit to the knowledge base to match corresponding materials and relevant reference videos, and analyzes the video editing information of the reference videos; then, based on this information, it conducts reasoning and decision-making on video editing to generate a video editing plan and the logic and basis of the reasoning process; 4. According to the logic and basis of the reasoning process of the chain-of-thought reasoning engine, an explainability report generator is used to generate and output an explainability report for the user to view; 5. The video construction unit can apply transitions, filters, music, storyboards, etc. according to the plan of the chain-of-thought reasoning engine; 6. Preview the video. If the user is satisfied, the video can be exported, or it can be further modified or adjusted.

[0081] In this exemplary application, by obtaining the logic and basis of the video editing process through the chain-of-thought reasoning engine and generating an explainability report through the explainability report generator, users can understand the logic and basis of video editing by viewing the explainability report, which is convenient for targeted adjustment of the corresponding video editing process. At the same time, through multimodal analysis by the multimodal analysis unit, the accuracy of the reasoning of the chain-of-thought reasoning engine can be improved; in addition, by matching videos to the knowledge base and generating a editing plan based on the editing information of the matching videos, it is also possible to effectively draw on editing methods and improve the quality and efficiency of video editing.

[0082] Embodiment 2 Figure 9 The block diagram of the video editing device according to Embodiment 2 of the present application is schematically shown. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 9 shown, the device 800 may include: an acquisition module 810, a generation module 820, a report module 830, and an editing module 840, where: The acquisition module 810 is configured to acquire video materials; The generation module 820 is configured to generate a first video editing plan using a chain-of-thought engine based on the video materials and obtain the logic and basis for the chain-of-thought engine to generate the first plan; The report module 830 is configured to generate and output an explainability report for video editing based on the logic and basis; A clip module 840, configured to perform video clip based on the first scheme and the video material.

[0083] In an alternative embodiment, the apparatus 800 is further configured to: In response to receiving an input modification prompt, generate a second scheme for video clip by using a thought chain engine based on the video material and the modification prompt; Perform video clip based on the second scheme and the video material.

[0084] In an alternative embodiment, the generation module 820 is further configured to: Perform multimodal analysis on the video material to obtain a result of multimodal analysis; Generate a first scheme for video clip by using a thought chain engine based on the result of multimodal analysis.

[0085] In an alternative embodiment, the video material includes video format material, image format material, audio format material, and text format material; Correspondingly, the generation module 820 is further configured to: Perform visual modal analysis on the video format material and the image format material to obtain a visual modal analysis result; Perform auditory modal analysis on the audio format material to obtain an auditory modal analysis result; Perform text modal analysis on the text format material to obtain a text modal analysis result; Obtain the result of multimodal analysis based on the visual modal analysis result, the auditory modal analysis result, and the text modal analysis result.

[0086] In an alternative embodiment, the generation module 820 is further configured to: Obtain an input clip requirement description; Generate a first scheme for video clip by using a thought chain engine based on the result of multimodal analysis and the clip requirement description.

[0087] In an alternative embodiment, the generation module 820 is further configured to: Match videos in a target knowledge base by using a thought chain engine based on the result of multimodal analysis and the clip requirement description to obtain a reference video; Obtain video clip information of the reference video; Generate the first scheme based on the video clip information.

[0088] Embodiment III Figure 10Schematically shown is a hardware architecture diagram of a computer device 10000 suitable for implementing the video editing method according to Embodiment 3 of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 10 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can be communicatively linked to each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disc, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the video editing method. In addition, the memory 10010 may also be used to temporarily store various data that have been output or will be output.

[0089] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0090] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (abbreviated as GSM), Wideband Code Division Multiple Access (abbreviated as WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0091] It should be noted that Figure 10 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0092] In this embodiment, the video clip method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of this application.

[0093] Embodiment 4 The embodiments of this application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the video clip method in the embodiments are implemented.

[0094] In this embodiment, the computer-readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (e.g., SD or DX memories, etc.), random access memories (RAM), static random access memories (SRAM), read-only memories (ROM), electrically erasable programmable read-only memories (EEPROM), programmable read-only memories (PROM), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the video editing method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or will be output.

[0095] Embodiment 5 The embodiment of the present application further provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0096] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device. Thus, they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0097] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A video editing method, characterized in that, The method includes: Obtain video materials; Based on the video materials, use a thought chain engine to generate a first video editing plan, and obtain the logic and basis for the thought chain engine to generate the first plan; Generate and output an interpretability report for video editing based on the logic and basis; Perform video editing based on the first plan and the video materials.

2. The method according to claim 1, wherein The method further includes: In the case of receiving an input modification prompt, use a thought chain engine to generate a second video editing plan based on the video materials and the modification prompt; Perform video editing based on the second plan and the video materials.

3. The method according to claim 1, characterized in that The step of using a thought chain engine to generate a first video editing plan based on the video materials includes: Perform multimodal analysis on the video materials to obtain the results of multimodal analysis; Based on the results of the multimodal analysis, use a thought chain engine to generate a first video editing plan.

4. The method according to claim 3, wherein The video materials include video format materials, image format materials, audio format materials, and text format materials; Correspondingly, the step of performing multimodal analysis on the video materials to obtain the results of multimodal analysis includes: Perform visual modal analysis on the video format materials and the image format materials to obtain visual modal analysis results; Perform auditory modal analysis on the audio format materials to obtain auditory modal analysis results; Perform text modal analysis on the text format materials to obtain text modal analysis results; Obtain the results of the multimodal analysis based on the visual modal analysis results, the auditory modal analysis results, and the text modal analysis results.

5. The method according to claim 3, characterized in that The step of using a thought chain engine to generate a first video editing plan based on the results of the multimodal analysis includes: Obtain an input description of the editing requirements; Based on the results of the multimodal analysis and the description of the editing requirements, use a thought chain engine to generate a first video editing plan.

6. The method according to claim 5, characterized in that, The step of using a thought chain engine to generate a first video editing plan based on the results of the multimodal analysis and the description of the editing requirements includes: Based on the results of the multimodal analysis and the description of the editing requirements, use a thought chain engine to match videos in the target knowledge base to obtain reference videos; Obtain the video editing information of the reference videos; Generate the first plan based on the video editing information.

7. A video editing device, characterized in that, The device includes: An acquisition module for obtaining video materials; A generation module for using a thought chain engine to generate a first video editing plan based on the video materials, and obtaining the logic and basis for the thought chain engine to generate the first plan; A report module for generating and outputting an interpretability report for video editing based on the logic and basis; An editing module for performing video editing based on the first plan and the video materials.

8. A computer device, characterized in that It includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.