Video Editing Method, Device, Electronic Device and Computer-Readable Storage Medium
By introducing video and audio fill functions into video editing tools, the problems of low efficiency and lack of personalization in the prior art are solved, and an efficient and personalized video editing experience is achieved.
Patent Information
- Application Number
- CN202110871543.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-07-30
AI Technical Summary
Existing video editing tools cannot meet the needs of personalized and efficient video editing, resulting in a lack of personalization of videos, difficulty in expressing users' unique opinions and ideas, and inefficient editing.
Through the mutually coordinated video and audio filling functions, users allow them to fill video materials and audio materials in the video template to achieve personalized video editing.
This improves the efficiency and personalization of video editing, and users can flexibly add video and audio materials to meet the needs of expressing personalized ideas and opinions.
Smart Images

Figure CN115695680B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of video processing, and in particular, to a video editing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] Videos, especially short videos, have become an important medium for social and information dissemination on the Internet. Users can create videos through video editing tools to express their ideas and post them on the Internet for various forms of interaction.
[0003] In related technologies, users are supported to create videos using the video template function provided by video editing tools. However, the video editing tools provided based on related technologies cannot meet the personalized and efficient video editing needs.
[0004] For example, the videos created based on the video templates provided by related technologies have a high similarity in plot and logic to the video templates themselves, resulting in a lack of personalization and difficulty in expressing the user's own unique views and ideas. Even secondary editing of the video is required to further add content, which in turn affects the editing efficiency. Summary of the Invention
[0005] Embodiments of the present application provide a video editing method, apparatus, electronic device, and computer-readable storage medium, which can achieve personalized and efficient video editing through the mutually cooperative video and audio filling functions.
[0006] The technical solution of the embodiments of the present application is implemented as follows:
[0007] Embodiments of the present application provide a video editing method, including:
[0008] In response to a video filling trigger operation for a video template, display a first video, where the first video is formed by filling at least one video material in the video template;
[0009] Display at least one audio filling area on the timeline of the first video;
[0010] In response to an audio setting operation, obtain the audio material to be filled in the at least one audio filling area;
[0011] In response to an audio filling trigger operation for the first video, display a second video to replace the first video, where the second video is formed by filling the corresponding audio material in the at least one audio filling area of the first video.
[0012] In the above solution, before displaying the first video in response to a video filling trigger operation for a video template, the method further includes: displaying at least one unfilled segment included in the video template and a video material selection control, where the video material selection control is used to select multiple candidate video materials; in response to a video material selection operation through the video material selection control, highlighting at least one selected video material, and in response to a video filling trigger operation for the video template, filling the at least one selected video material into the unfilled segment corresponding to the video template.
[0013] In the above solution, the displaying of at least one unfilled segment included in the video template includes: displaying multiple candidate video templates; in response to a video template selection operation, displaying the usage entry of the selected video template among the multiple candidate video templates; in response to a trigger operation for the usage entry of the selected video template, displaying at least one unfilled segment included in the selected video template; the method further includes: displaying the introduction information of the unfilled segment, where the introduction information is used to characterize the type of the video material to be filled in the unfilled segment.
[0014] In the above solution, after filling the at least one selected video material into the video template, the method further includes: displaying a replacement video entry; in response to a trigger operation for the replacement video entry, displaying a video selection control, where the video selection control is used to select multiple candidate video materials; in response to a video material selection operation through the video selection control, replacing the video material with at least one selected video material among the multiple candidate video materials.
[0015] In the above solution, after filling the at least one selected video material into the video template, the method further includes: displaying a video cropping entry; in response to a trigger operation for the video cropping entry, displaying a video cropping control, where the video cropping control includes at least one of a video frame cropping control and a time period selection control; in response to a cropping frame setting operation based on the video frame cropping control, cropping the video material to remove the video frame outside the set cropping frame; in response to a time setting operation based on the time period selection control, cropping the video material to remove the video segment outside the set time period.
[0016] In the above solution, after filling the selected at least one video material into the video template, the method further includes: in response to a video material selection operation, displaying that a target video material selected from the at least one video material is in an editing state; displaying a volume adjustment entry, and in response to a triggering operation on the volume adjustment entry, displaying a volume adjustment control; and in response to a setting operation on the volume adjustment control, determining the set volume as the volume of the target video material.
[0017] In the above solution, after obtaining the audio material to be filled in the at least one audio filling area, the method further includes: performing at least one of the following processes: converting the data format of the audio material to conform to the data format of the first video; converting the sampling frequency of the audio material to conform to the sampling frequency of the first video.
[0018] An embodiment of the present application provides a video editing device, including:
[0019] A display module, configured to display a first video in response to a video filling trigger operation on a video template, where the first video is formed after filling at least one video material in the video template;
[0020] The display module is further configured to display at least one audio filling area on the timeline of the first video;
[0021] An obtaining module, configured to obtain the audio material to be filled in the at least one audio filling area in response to an audio setting operation;
[0022] The display module is further configured to display a second video to replace the first video in response to an audio filling trigger operation on the first video, where the second video is formed after filling corresponding audio materials in the at least one audio filling area of the first video.
[0023] In the above solution, the display module is further configured to display an audio setting entry corresponding to the first video; and in response to a triggering operation on the audio setting entry, display the timeline of the first video and display at least one audio filling area on the timeline, where the audio filling area is used to indicate a time period on the timeline where the audio material can be filled.
[0024] In the above solution, the obtaining module is further configured to obtain the at least one audio filling area by at least one of the following methods: obtaining the at least one pre-set audio filling area from the video template; dividing the at least one audio filling area from the timeline according to the playing time of the video material filled in the first video; dividing the timeline according to the plot units of the first video, and determining the audio filling area corresponding to each plot unit.
[0025] In the above solution, the display module is further configured to, in response to an audio filling area selection operation, display that the selected target audio filling area is in an editing state; in response to a time setting operation for the target audio filling area, update and display the target audio filling area based on the start time and end time set on the timeline.
[0026] In the above solution, when the audio setting entry is a recording entry, the display module is further configured to display a recording control corresponding to the target audio filling area, where the target audio filling area is the audio filling area in the editing state among the at least one audio filling area; the apparatus further includes a collection module, configured to start collecting audio in response to a first audio setting operation for turning on the recording control; stop collecting audio in response to a second audio setting operation for turning off the recording control; the apparatus further includes a determination module, configured to use the collected audio as the audio material to be filled in the target audio filling area.
[0027] In the above solution, when the audio setting entry is an audio material selection entry, the display module is further configured to display an audio material selection control, where the audio material selection control is used to select multiple candidate audio materials; the determination module is further configured to, in response to an audio setting operation for selecting through the audio material selection control, use at least one of the selected audio materials among the multiple candidate audio materials as the audio material to be filled in the target audio filling area, where the target audio filling area is the audio filling area in the editing state among the at least one audio filling area.
[0028] In the above solution, the display module is further configured to display an audio re-recording entry; and in response to a trigger operation for the audio re-recording entry, display a recording control for re-collecting; the collection module is further configured to start collecting audio in response to a third audio setting operation for turning on the recording control; stop collecting audio in response to a fourth audio setting operation for turning off the recording control; the determination module is further configured to use the re-collected audio as the audio material to be filled in the target audio filling area, where the target audio filling area is the audio filling area in the editing state among the at least one audio filling area.
[0029] In the above solution, the display module is further configured to display an audio deletion entry; the device further includes a deletion module, configured to, in response to a trigger operation on the audio deletion entry, delete the audio material corresponding to the target audio filling area, where the target audio filling area is the audio filling area in an editing state among the at least one audio filling area.
[0030] In the above solution, the display module is further configured to display a volume adjustment entry; and to, in response to a trigger operation on the volume adjustment entry, display a volume adjustment control; the determination module is further configured to, in response to a setting operation on the volume adjustment control, determine the set volume as the volume of the audio material corresponding to the target audio filling area, where the target audio filling area is the audio filling area in an editing state among the at least one audio filling area.
[0031] In the above solution, the display module is further configured to display a voice change entry; and to, in response to a trigger operation on the voice change entry, display a plurality of candidate voice objects; the device further includes a replacement module, configured to, in response to a voice object selection operation, replace the initial voice object of the audio material corresponding to the target audio filling area with the selected target voice object among the plurality of candidate voice objects, where the target audio filling area is the audio filling area in an editing state among the at least one audio filling area.
[0032] In the above solution, the display module is further configured to display a text recognition entry; and to, in response to a trigger operation on the text recognition entry, display the speech recognition result of the target audio material and use it as the subtitle of the time segment in the second video filled with the target audio material, where the target audio material is the audio material in an editing state among the to-be-filled audio materials.
[0033] In the above solution, the display module is further configured to display at least one to-be-filled segment included in the video template and a video material selection control, where the video material selection control is used to select a plurality of candidate video materials; and to, in response to a video material selection operation through the video material selection control, highlight the selected at least one video material; the device further includes a filling module, configured to, in response to a video filling trigger operation on the video template, fill the selected at least one video material into the to-be-filled segment corresponding to the video template.
[0034] In the above solution, the display module is further configured to display a plurality of candidate video templates; in response to a video template selection operation, display the usage entry of the selected video template among the plurality of candidate video templates; in response to a trigger operation on the usage entry of the selected video template, display at least one unfilled segment included in the selected video template; and display the introduction information of the unfilled segment, where the introduction information is used to characterize the type of video material to be filled in the unfilled segment.
[0035] In the above solution, the display module is further configured to display a replacement video entry; in response to a trigger operation on the replacement video entry, display a video selection control, where the video selection control includes a plurality of candidate video materials; the replacement module is further configured to, in response to a video material selection operation in the video selection control, replace the video material with at least one selected video material among the plurality of candidate video materials.
[0036] In the above solution, the display module is further configured to display a video cropping entry; and in response to a trigger operation on the video cropping entry, display a video cropping control, where the video cropping control includes at least one of a picture cropping control and a time period selection control; the device further includes a cropping module configured to crop the picture outside the set cropping frame from the video material in response to a cropping frame setting operation based on the picture cropping control; and crop the video segment outside the set time period from the video material in response to a time setting operation based on the time period selection control.
[0037] In the above solution, the display module is further configured to, in response to a video material selection operation, display that the selected target video material among the at least one video material is in an editing state; display a volume adjustment entry, and in response to a trigger operation on the volume adjustment entry, display a volume adjustment control; the determination module is further configured to, in response to a setting operation on the volume adjustment control, determine the set volume as the volume of the target video material.
[0038] In the above solution, the display module is further configured to display a plurality of candidate video materials; in response to a video material selection operation, highlight at least one selected video material and display a plurality of candidate video templates matching the at least one video material; the filling module is further configured to, in response to a video template selection operation, fill the at least one video material into the unfilled segment corresponding to the selected video template.
[0039] In the above solution, the device further includes a conversion module configured to perform at least one of the following processes: converting the data format of the audio material to conform to the data format of the first video; converting the sampling frequency of the audio material to conform to the sampling frequency of the first video.
[0040] An embodiment of the present application provides an electronic device, including:
[0041] A memory for storing executable instructions;
[0042] A processor, when executing the executable instructions stored in the memory, implements the video editing method provided by the embodiment of the present application.
[0043] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the video editing method provided by the embodiment of the present application.
[0044] An embodiment of the present application provides a computer program product, the computer program product includes computer-executable instructions, which, when executed by a processor, implement the video editing method provided by the embodiment of the present application.
[0045] The embodiment of the present application has the following beneficial effects:
[0046] In the process of using the video template, by cooperating the functions of video material filling and audio material filling, it is possible to support the flexible addition of corresponding video materials and audio materials according to requirements, which not only improves the video editing efficiency but also fully meets the needs of expressing personalized ideas and viewpoints in the video content. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a schematic diagram of the architecture of the video editing system 100 provided by the embodiment of the present application;
[0048] Figure 2 is a schematic diagram of the structure of the terminal 400 provided by the embodiment of the present application;
[0049] Figure 3 is a schematic diagram of the flow of the video editing method provided by the embodiment of the present application;
[0050] Figure 4 is a schematic diagram of the application scenario of the video editing method provided by the embodiment of the present application;
[0051] Figure 5 is a schematic diagram of the application scenario of the video editing method provided by the embodiment of the present application;
[0052] Figure 6 is a schematic diagram of the application scenario of the video editing method provided by the embodiment of the present application;
[0053] Figure 7 It is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application;
[0054] Figure 8 It is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application;
[0055] Figure 9 It is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application;
[0056] Figure 10 It is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application;
[0057] Figure 11 It is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application;
[0058] Figure 12 It is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application;
[0059] Figure 13 It is a schematic diagram of a model of the video editing system provided by an embodiment of the present application;
[0060] Figure 14 It is a schematic flowchart of the recording process provided by an embodiment of the present application. Detailed implementation manners
[0061] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0062] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0063] In the following description, the terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0065] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0066] 1) Video template, a template for making videos, which may include replaceable resource content and non-replaceable resource content. Among them, the replaceable resource content may include pictures, audio used as background music, text content, etc., and the non-replaceable resource content may include pictures other than some replaceable pictures, the dynamic display effects, sound effects, and text display effects when switching between every two pictures, etc.
[0067] 2) Video material, the material used by users when making videos. The media types of video materials may include pictures and videos. For example, when the video material selected by the user is a video, then the video may include resource content of audio media type and image media type.
[0068] 3) Audio material, the material used by users when making videos. The audio material may be audio recorded by the user in real time or audio pre-stored locally on the terminal.
[0069] 4) In response to, used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more of the executed operations may be real-time or may have a set delay; without special instructions, there is no restriction on the execution order of the multiple executed operations.
[0070] Currently, with the rapid development of the short video industry, users' demands for video editing are getting higher and higher. Users can use video editing tools to make videos and express their ideas.
[0071] In the related art, when making videos, generally, users need to collect video materials or write video copy scripts in advance. Then, users use the native camera or video recording software to shoot, and then perform operations such as editing and cutting to generate the final video. The process of making and generating videos requires users to make detailed inputs and adjustments to video text and add their own recordings, etc. However, the efficiency of making videos in this way is very low. The production cycle of videos generally takes 1-2 days, which is very inconvenient and increases the operation time of users.
[0072] In addition, related technologies also provide a function of using video templates to quickly produce videos. Users only need to add video materials, and the system will automatically add effects such as music and text subtitles preset by the template to the video materials, and the production of the video can be completed in one step. However, this way of quickly producing videos using templates cannot add recordings to the templated videos (it is necessary to produce another video separately after the production of the template video to record the sound, resulting in a relatively long time spent on producing the video), which causes users to be unable to use their own voices to express their views and ideas on the produced videos, nor can it increase the diversity and personalization of the videos.
[0073] In view of this, the embodiments of the present application provide a video editing method, device, electronic device and computer-readable storage medium, which can realize personalized and efficient video editing through the functions of video filling and audio filling that cooperate with each other. The following describes an exemplary application of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (for example, mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), or can be implemented jointly by a server and a terminal. The following will describe an exemplary application when the electronic device is implemented as a terminal.
[0074] See Figure 1 , Figure 1 FIG. is a schematic diagram of the architecture of the video editing system 100 provided by the embodiments of the present application. To support a video editing application, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0075] A client 410 runs on the terminal 400. The client 410 can be a video editing client or a client integrated with video editing functions (such as a social network client, an instant messaging client), etc. The client 410 can send a network request for obtaining a video template to the server 200 through the network 300 (for example, the client 410 can send a request for obtaining a video template to the server 200 only when receiving a user-triggered obtaining instruction; of course, the client 410 can also pre-send a network request for obtaining a video template to the server 200 to obtain the video template in advance and store it locally on the terminal 400, so as to reduce the number of interactions), so that the server 200 sends the video template to the client 410. Then, the client 410 can display a template video (i.e., a video formed after filling video materials in the video template) in response to a video filling operation on the video template, and display at least one audio filling area on the timeline of the template video. Subsequently, the client 410 can obtain the audio materials to be filled in at least one audio filling area in response to an audio setting operation (for example, it can be an audio recorded by the user in real time or an audio pre-stored in the terminal 400). Finally, the client 410 can display the template video with the audio materials filled in the audio filling area in response to an audio filling operation on the template video. In this way, audio can be directly added to the template video, improving the video editing efficiency and meeting the user's need to express their personalized ideas and views through audio.
[0076] In some embodiments, the embodiments of the present application can be implemented with the help of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing.
[0077] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, and application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. For example, the service interaction function between the above-mentioned server 200 and the terminal 400 can be implemented through cloud technology.
[0078] For example, Figure 1The server 200 shown in the figure can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal 400 and the server 200 can be directly or indirectly connected by wired or wireless communication means, and there is no limitation in the embodiments of the present application.
[0079] In some other embodiments, the terminal device 400 can also implement the video editing method provided in the embodiments of the present application by running a computer program, and the computer program can be as Figure 1 shown in the client 410. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a video editing APP or a video playing APP integrated with video editing functions; it can also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded in any APP, where the small program can be controlled by the user to run or close. In short, the above computer program can be any form of application program, module or plug-in.
[0080] The following will describe Figure 1 the structure of the terminal 400 shown in the figure. Refer to Figure 2 , Figure 2 which is a schematic structural diagram of the terminal 400 provided in the embodiments of the present application. Figure 2 The terminal 400 shown in the figure includes: at least one processor 420, a memory 460, at least one network interface 430, and a user interface 440. Each component in the terminal 400 is coupled together through a bus system 450. It can be understood that the bus system 450 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 450 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear description, in Figure 2 all kinds of buses are labeled as the bus system 450.
[0081] The processor 420 may be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0082] The user interface 440 includes one or more output devices 441 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 440 also includes one or more input devices 442, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.
[0083] The memory 460 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 460 optionally includes one or more storage devices that are physically remote from the processor 420.
[0084] The memory 460 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 460 described in the embodiments of the present application is intended to include any suitable type of memory.
[0085] In some embodiments, the memory 460 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are exemplarily described below.
[0086] The operating system 461 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0087] The network communication module 462 is used to reach other computing devices via one or more (wired or wireless) network interfaces 430. Exemplary network interfaces 430 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;
[0088] A presentation module 463 for enabling the presentation of information (e.g., a user interface for operating a peripheral device and displaying content and information) via one or more output devices 441 associated with the user interface 440 (e.g., a display screen, a speaker, etc.);
[0089] An input processing module 464 for detecting and translating one or more user inputs or interactions from one of one or more input devices 442.
[0090] In some embodiments, the video editing device provided by the embodiments of the present application may be implemented in software. Figure 2 Shown is a video editing device 465 stored in the memory 460, which may be software in the form of a program and plug-ins, etc., including the following software modules: a display module 4651, an acquisition module 4652, a collection module 4653, a determination module 4654, a deletion module 4655, a replacement module 4656, a filling module 4657, a cropping module 4658, and a module 4659. These modules are logical, and thus can be combined arbitrarily or further split according to the functions implemented. It should be noted that, for the sake of convenience of description, all the above modules are shown at once, but it should not be considered that the video editing device 465 excludes embodiments that may only include the display module 4651 and the acquisition module 4652. The functions of each module will be described below. Figure 2 In order to facilitate the description, all the above modules are shown at once, but it should not be considered that the video editing device 465 excludes embodiments that may only include the display module 4651 and the acquisition module 4652. The functions of each module will be described below.
[0091] In other embodiments, the video editing device provided by the embodiments of the present application may be implemented in hardware. As an example, the video editing device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the video editing method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0092] Next, the video editing method provided by the embodiments of the present application will be described in conjunction with the exemplary applications and implementations of the terminal device provided by the embodiments of the present application. By way of example, refer to Figure 3 , Figure 3 which is a flowchart of the video editing method provided by the embodiments of the present application, and will be described in conjunction withFigure 3 The steps shown will be described.
[0093] It should be noted that Figure 3 the method shown can be executed by various forms of computer programs running on the terminal 400 shown in Figure 1 It is not limited to the above-mentioned client 410, and can also be the operating system 461, software modules and scripts described above. Therefore, the client described below should not be regarded as a limitation to the embodiments of the present application.
[0094] In step S101, in response to a video filling trigger operation for a video template, a first video is displayed.
[0095] Here, the first video is formed after filling at least one video material in the video template.
[0096] In some embodiments, the user can first select a video template, then select video materials, and then fill the selected video materials into the selected video template. Then, before performing Figure 3 the step S101 shown, the following processing can also be performed: display at least one to-be-filled segment included in the video template and a video material selection control, where the video material selection control is used to select multiple candidate video materials; in response to a video material selection operation through the video material selection control, highlight at least one selected video material (for example, highlight the selected video material), and in response to a video filling trigger operation for the video template, fill at least one selected video material into the to-be-filled segment corresponding to the video template.
[0097] In some other embodiments, following the above example, at least one to-be-filled segment included in the video template can be displayed in the following manner: display multiple candidate video templates; in response to a video template selection operation, display the introduction information and usage entry of the selected video template among the multiple candidate video templates; in response to a trigger operation for the usage entry of the selected video template, display at least one to-be-filled segment included in the selected video template; display the introduction information of the to-be-filled segment, where the introduction message is used to characterize the type of video material to be filled in the to-be-filled segment.
[0098] In some other embodiments, the user can also first select video materials, then select a video template, and then fill the selected video materials into the selected video template. Then, before performing Figure 3Before the step S101 shown, the following processing may also be performed: displaying a plurality of candidate video materials; in response to a video material selection operation, highlighting at least one selected video material and displaying a plurality of candidate video templates that match the at least one selected video material (for example, candidate video templates that match the quantity or type of the video material); in response to a video template selection operation, filling at least one selected video material into a to-be-filled segment corresponding to the selected video template.
[0099] The following takes the example where the user first selects a video template and then selects a video material for illustration.
[0100] Exemplarily, refer to Figure 4 , Figure 4 which is a schematic diagram of an application scenario of the video editing method provided in an embodiment of the present application. As Figure 4 shown, a plurality of candidate video templates are displayed on a page 401, including a video template 402, a video template 403, a video template 404, and a video template 405. When a click operation on the video template 403 displayed on the page 401 is received by the user, it will jump to a details page 406 of the video template 403. An introduction information 407 and a usage entry 408 of the video template 403 are displayed on the details page 406. When a click operation on the usage entry 408 displayed on the details page 406 is received by the user, a video material selection control 409 will pop up. A plurality of candidate video materials (such as pictures, videos, etc. in the local album of the user terminal) are displayed in the video material selection control 409. When a click operation on the video material 410 displayed in the video material selection control 409 is received by the user, the selected video material 410 will be highlighted, for example, the user-selected video material 410 is highlighted. In addition, taking the to-be-filled segment 411 as an example, corresponding introduction information 412 may also be displayed below the to-be-filled segment 411, such as the type of video material that needs to be filled in the to-be-filled segment 411. In this way, when the user uses the video template to upload their own video material, it can be more accurate, improving the video editing efficiency.
[0101] In some embodiments, after filling at least one selected video material into the video template, the following processing may also be performed: displaying a replacement video entry; in response to a trigger operation on the replacement video entry, displaying a video selection control, where the video selection control includes a plurality of candidate video materials; in response to a video material selection operation in the video selection control, replacing the original video material with at least one re-selected video material among the plurality of candidate video materials.
[0102] In some other embodiments, after filling at least one selected video material into a video template, the following processing may further be performed: displaying a video cropping entry; in response to a trigger operation on the video cropping entry, displaying video cropping controls, where the video cropping controls include at least one of a picture cropping control and a time period selection control; in response to a cropping frame setting operation based on the picture cropping control, cropping off the picture outside the set cropping frame from the video material; and in response to a time setting operation based on the time period selection control, cropping off the video segments outside the set time period from the video material.
[0103] In some embodiments, after filling at least one selected video material into a video template, the following processing may further be performed: in response to a video material selection operation, displaying that the selected target video material in at least one video material is in an editing state; displaying a volume adjustment entry, and in response to a trigger operation on the volume adjustment entry, displaying volume adjustment controls; and in response to a setting operation on the volume adjustment controls, determining the set volume as the volume of the target video material.
[0104] It should be noted that in practical applications, the above-mentioned various video editing entries (i.e., the video replacement entry, the video cropping entry, and the volume adjustment entry) can be displayed and used selectively or in combination. For example, each of the above-mentioned video editing entries can be displayed separately, or a unified video editing entry can be displayed first, and after the unified video editing entry is triggered, the specific types of video editing entries can be displayed.
[0105] The following takes the combined display and use of the above-mentioned various video editing entries as an example for illustration.
[0106] Exemplarily, refer to Figure 5 , Figure 5 which is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application. As Figure 5As shown, when the user is not satisfied with the video material imported into the video template, the user can also click on the corresponding video material to finely adjust the video material on the pop-up floating window. The fine adjustment can include replacing the video material, cropping the video material, adjusting the volume of the video material, etc. For example, when the user is not satisfied with the video material 501 imported into the video template, the user can select the video material 501. At this time, the video material 501 will be highlighted (e.g., highlighted), indicating that the video material 501 is in the editing state, and a floating window will also pop up above the video material 501. In the floating window, there are a video replacement entry 502, a video cropping entry 503, and a volume adjustment entry 504. When receiving the click operation of the user on the video replacement entry 502, a video selection control 505 will pop up, and the user can re-select the video material to be used in the popped-up video selection control 505 to replace the video material 501; when receiving the click operation of the user on the video cropping entry 503, it will jump to page 506, where there are a picture cropping control 507 and a time period selection control 508. The user can adjust the picture size of the video material 501 by adjusting the cropping frame of the picture cropping control 507. At the same time, by setting the start time and end time of the time period selection control 508, the required video segment can be cropped from the video material 501; when receiving the click operation of the user on the volume adjustment entry 504, it will jump to page 509, where there is a volume adjustment control 510. The user can set the volume adjustment control 510, for example, by dragging the adjustment button on the volume adjustment control 510 to adjust the volume of the video material 501.
[0107] In step S102, at least one audio filling area is displayed on the timeline of the first video.
[0108] In some embodiments, at least one audio filling area can be displayed on the timeline of the first video in the following manner: display an audio setting entry corresponding to the first video (for example, it can be a recording entry for obtaining the audio recorded by the user in real time; or it can be an audio material selection entry for selecting from multiple audio materials pre-stored in the terminal); in response to the trigger operation on the audio setting entry, display the timeline of the first video and display at least one audio filling area on the timeline, where the audio filling area is used to indicate the time period on the timeline where audio materials can be filled.
[0109] In some other embodiments, before at least one audio filling area is displayed on the time axis, the following processing may also be performed: obtaining at least one audio filling area by at least one of the following methods: obtaining at least one preset audio filling area from a video template; dividing at least one audio filling area from the time axis according to the playing time of the video material filled in the first video; dividing the time axis according to the plot units of the first video and determining the audio filling area corresponding to each plot unit.
[0110] For example, the audio filling area may be preset in the video template, and other areas cannot be filled with audio material. For example, when making a video template, the producer of the video template may preset at least one audio filling area in the video template. For example, assuming that the playing duration of the video template is 20 seconds, the producer can set the time period from the 5th second to the 10th second as an audio filling area, that is, the time period from the 5th second to the 10th second of the video template allows users to fill in audio material. For example, users can fill the audio recorded in real time into the time period from the 5th second to the 10th second of the video template. In this way, by presetting the audio filling area in the video template, the operation burden of users is reduced and the efficiency of video production is improved.
[0111] For example, the audio filling area may also be formed by dividing the time axis according to the video material filled in the first video, and each video material corresponds to an audio filling area. For example, assuming that there are 3 video materials filled in the first video, namely video material 1, video material 2 and video material 3, and the playing time of video material 1 is from the 3rd second to the 5th second, the playing time of video material 2 is from the 7th second to the 9th second, and the playing time of video material 3 is from the 11th second to the 13th second, then 3 corresponding audio filling areas can be divided from the time axis of the first video according to the playing time of these 3 video materials. For example, the first audio filling area may be the time period from the 3rd second to the 5th second, the second audio filling area may be the time period from the 7th second to the 9th second, and the third audio filling area may be the time period from the 11th second to the 13th second. In this way, since the number of audio filling areas corresponds to the number of video materials, that is, for each video material, users can express their views through the corresponding audio, improving the user experience.
[0112] Exemplarily, the audio filling area can also be formed by dividing the timeline according to the plot units of the first video, and each plot unit corresponds to an audio filling area. For example, first, a plot unit recognition model is called to perform plot unit recognition processing on the first video to obtain the plot units included in the first video. Suppose the plot unit recognition model divides the first video into 2 plot units, where the playing time corresponding to the first plot unit is from the 1st second to the 10th second, and the playing time corresponding to the second plot unit is from the 11th second to the 20th second. Then, the corresponding audio filling areas can be determined for these 2 plot units respectively. For example, the time period from the 2nd second to the 8th second can be used as the audio filling area corresponding to the first plot unit, and the time period from the 13th second to the 17th second can be used as the audio filling area corresponding to the second plot unit. In this way, since the number of audio filling areas corresponds to the number of plot units of the first video, that is, for each plot unit, the user can use the corresponding audio to express their views, improving the user experience.
[0113] In some other embodiments, the user can manually adjust the number of audio filling areas displayed on the timeline of the first video, as well as the start time and end time corresponding to each audio filling area. That is, the terminal can also perform the following processing: in response to an audio filling area selection operation, display the selected target audio filling area in an editing state (such as in a highlighted manner); in response to a time setting operation for the target audio filling area, update the display of the target audio filling area based on the start time and end time set on the timeline; in response to an audio filling area addition operation, display a time period selection control on the timeline (including a starting point and an ending point that can be moved on the timeline); in response to a setting operation based on the time period selection control (including the set start time and end time), display the newly added audio filling area.
[0114] Exemplarily, refer to Figure 6 , Figure 6 is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application. As Figure 6 shown, two audio filling areas are displayed in the timeline 601 of the first video, namely the audio filling area 602 and the audio filling area 603. When a selection operation of the user for the audio filling area 603 is received, the audio filling area 603 will be highlighted to indicate that the audio filling area 603 is currently in an editing state. At this time, the user can set the start time and end time corresponding to the audio filling area 603. For example, suppose the initial start time and end time of the audio filling area 603 are the 5th second and the 10th second respectively. After the user sets the start time and end time of the audio filling area 603 on the timeline 601, as Figure 6As shown in the right figure, the start time and end time of the audio filling area 603 are updated to the 8th second and the 11th second.
[0115] Exemplarily, refer to Figure 7 , Figure 7 which is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application. As Figure 7 shown, only an audio filling area 702 and a time period selection control 703 are displayed on the time axis 701 of the first video. When the user needs to add an audio filling area to the time axis 701 of the first video, the time period selection control 703 can be used to set the start time and end time of the audio filling area to be added. After the user sets the start time and end time, a newly added audio filling area will be displayed on the time axis 701 of the first video. For example, as Figure 7 shown by the newly added audio filling area 704.
[0116] In step S103, in response to an audio setting operation, at least one audio material to be filled in the audio filling area is obtained.
[0117] In some embodiments, when the audio setting entry is a recording entry, step S103 can be implemented in the following manner: in response to an audio filling area selection operation, display that the selected target audio filling area is in an editing state; display a recording control corresponding to the target audio filling area (for example, display the recording control in the target audio filling area); in response to a first audio setting operation to turn on (for example, trigger for the first time) the recording control, start collecting audio; in response to a second audio setting operation to turn off (for example, trigger again) the recording control, stop collecting audio; use the collected audio as the audio material to be filled in the target audio filling area. In this way, during the process of making a video using a video template, the user can record their own voice through the recording entry and fill the recorded audio into the audio filling area, so as to realize personalized and efficient video editing through the functions of video filling and audio filling that cooperate with each other.
[0118] Exemplarily, refer to Figure 8 , Figure 8 which is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application. As Figure 8As shown in the figure, there are two audio filling areas displayed on the timeline 801 of the first video, namely the audio filling area 802 and the audio filling area 803. When a selection operation of the user for the audio filling area 803 is received, the audio filling area 803 will be highlighted (i.e., the brightness is higher than that of other audio filling areas) to indicate that the audio filling area 803 is currently in an editing state. At the same time, a recording control 804 corresponding to the audio filling area 803 will be displayed. At this time, the user can click the recording control 804 to record, and the terminal will use the collected audio as the audio material to be filled in the audio filling area 803.
[0119] In some other embodiments, when the audio setting entry is an audio material selection entry, step S103 can be implemented in the following manner: in response to an audio filling area selection operation, display that the selected target audio filling area is in an editing state; display an audio material selection control, where the audio material selection control is used to select multiple candidate audio materials; in response to an audio setting operation of selecting through the audio material selection control, use at least one of the selected audio materials from the multiple candidate audio materials as the audio material to be filled in the target audio filling area. In this way, the user can quickly select the audio materials to be used from multiple candidate audio materials pre-stored locally in the terminal device through the audio material selection entry, improving the efficiency of video editing.
[0120] Exemplarily, refer to Figure 9 , Figure 9 is a schematic diagram of an application scenario of the video editing method provided by the embodiments of the present application. As Figure 9 shown in the figure, there are two audio filling areas displayed on the timeline 901 of the first video, namely the audio filling area 902 and the audio filling area 903. When a selection operation of the user for the audio filling area 903 is received, the audio filling area 903 will be highlighted to indicate that the audio filling area 903 is currently in an editing state. At the same time, an audio material selection control 904 will be displayed in a pop-up window. Multiple candidate audio materials (such as audio materials pre-stored locally in the terminal, including recording 1 to recording 12) are displayed in the audio material selection control 904. At this time, the user can select from the multiple candidate audio materials displayed in the audio material selection control 904. For example, when a click operation of the user for the audio material 905 is received, the audio material 905 will be highlighted to indicate that the audio material 905 will be filled into the audio filling area 903.
[0121] It should be noted that in practical applications, when a user selects multiple audio materials at once in the audio material selection control, these multiple audio materials can be used as the corresponding audio materials to be filled in the respective audio filling areas displayed on the timeline of the first video in sequence. For example, assume that the user selects audio material 1, audio material 2, and audio material 3 at once in the audio material selection control. Then, audio material 1 will be filled into audio filling area 1 displayed on the timeline of the first video, audio material 2 will be filled into audio filling area 2 displayed on the timeline of the first video, and audio material 3 will be filled into audio filling area 3 displayed on the timeline of the first video.
[0122] In some embodiments, the audio data format collected by the audio acquisition device may not match the audio data format required by the client, or the sampling frequency of the audio collected by the audio acquisition device may not match the sampling frequency when the client plays the audio. Then, after obtaining the audio materials to be filled in at least one audio filling area, the following processing can also be performed: converting the data format of the audio material to conform to the data format of the first video; converting the sampling frequency of the audio material to conform to the sampling frequency of the first video. In this way, by converting the data format and sampling frequency of the obtained audio materials, the converted audio materials meet the requirements of the client, so that they can be successfully filled into the audio filling areas displayed on the timeline of the first video.
[0123] In step S104, in response to the audio filling trigger operation for the first video, the second video is displayed to replace the first video.
[0124] Here, the second video is formed after filling the corresponding audio materials in at least one audio filling area of the first video.
[0125] In some embodiments, when displaying the second video to replace the first video, the following processing can also be performed: in response to the audio filling area selection operation, displaying the selected target audio filling area in an editing state; displaying an audio re-recording entry (for example, an audio re-recording entry can be displayed in the target audio filling area); in response to the trigger operation for the audio re-recording entry, displaying a recording control for re-acquiring audio; in response to the third audio setting operation for turning on the recording control, starting to acquire audio; in response to the fourth audio setting operation for turning off the recording control, stopping to acquire audio; using the re-acquired audio as the audio material to be filled in the target audio filling area. In this way, when the user is not satisfied with the original audio material filled in the target audio filling area, they can re-record through the audio re-recording entry and replace the original audio material with the re-recorded audio.
[0126] In some other embodiments, when the second video is displayed to replace the first video, the following processing may also be performed: in response to an audio filling area selection operation, display that the selected target audio filling area is in an editing state; display an audio deletion entry; in response to a trigger operation on the audio deletion entry, delete the audio material corresponding to the target audio filling area. In this way, when the user is not satisfied with the audio material filled in a certain audio filling area, the user can quickly delete it through the audio deletion entry corresponding to this audio filling area.
[0127] In some embodiments, when the second video is displayed to replace the first video, the following processing may also be performed: in response to an audio filling area selection operation, display that the selected target audio filling area is in an editing state; display a volume adjustment entry; in response to a trigger operation on the volume adjustment entry, display a volume adjustment control; in response to a setting operation on the volume adjustment control, determine the set volume as the volume of the audio material corresponding to the target audio filling area. In this way, when the user is not satisfied with the volume of the audio material filled in a certain audio filling area, the user can quickly adjust the volume of the audio material through the volume adjustment entry.
[0128] In some other embodiments, when the second video is displayed to replace the first video, the following processing may also be performed: in response to an audio filling area selection operation, display that the selected target audio filling area is in an editing state; display a voice change entry; in response to a trigger operation on the voice change entry, display a plurality of candidate voice objects; in response to a voice object selection operation, replace the initial voice object of the audio material corresponding to the target audio filling area with the selected target voice object among the plurality of candidate voice objects. In this way, the user can change the initial voice object (i.e., the initial timbre of the audio material) of the audio material according to their own preferences and change the timbre to that of other characters to meet the user's personalized needs.
[0129] It should be noted that in practical applications, the above-mentioned various audio editing entries (i.e., the audio re-recording entry, the audio deletion entry, the volume adjustment entry, and the voice change entry) can be displayed and used selectively, or can be displayed and used in combination. For example, the above-mentioned various audio editing entries can be displayed separately, or a unified audio editing entry can be displayed first, and after the unified audio editing entry is triggered, the specific types of audio editing entries are displayed.
[0130] The following takes the combined display and use of the above-mentioned various audio editing entries as an example for illustration.
[0131] Exemplarily, refer to Figure 10 , Figure 10 is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application. As Figure 10As shown, when the user is not satisfied with the imported audio material, the user can also click on the corresponding audio filling area and make fine adjustments to the audio material on the pop-up floating window. The fine adjustments can include deleting the audio material, adjusting the volume of the audio material, changing the voice object of the audio material, etc. For example, when the user is not satisfied with the audio material filled in the audio filling area 1001, the user can select the audio filling area 1001. At this time, the audio filling area 1001 will be highlighted (for example, displayed in a highlighted manner), indicating that the audio filling area 1001 is currently in an editing state, and a floating window 1002 will also pop up above the audio filling area 1001. In the floating window 1002, there are displayed an audio re-recording entry 1003, a volume adjustment entry 1004, an audio deletion entry 1005, and a voice-changing entry 1006; when receiving a click operation from the user on the audio re-recording entry 1003, a recording control 1007 will be displayed. At this time, the user can click on the recording control 1007 to re-record and use the newly captured audio to replace the original audio material corresponding to the audio filling area 1001; when receiving a click operation from the user on the volume adjustment entry 1004, a volume adjustment control 1008 will be displayed. At this time, the user can set the volume adjustment control 1008, for example, slide the adjustment button on the volume adjustment control 1008, and then use the volume adjusted by the user as the volume of the audio material corresponding to the audio filling area 1001; when receiving a click operation from the user on the audio deletion entry 1005, the audio material corresponding to the audio filling area 1001 will be deleted; when receiving a click operation from the user on the voice-changing entry 1006, a plurality of candidate voice-changing objects 1009 will be displayed, such as uncle voice, lolita voice, alien voice, doll voice, etc. At this time, the user can select from the plurality of candidate voice-changing objects 1009, so as to replace the initial voice object of the audio material corresponding to the audio filling area 1001 with the selected target voice object among the plurality of candidate voice-changing objects 1009. For example, when the user selects the uncle voice, the initial voice object of the audio material corresponding to the audio filling area 1001 will be replaced with the uncle voice.
[0132] In some other embodiments, when displaying the second video to replace the first video, the following processing can also be performed: in response to an audio material selection operation, display that the selected target audio material is in an editing state; display a text recognition entry; in response to a trigger operation on the text recognition entry, display the speech recognition result corresponding to the target audio material and use it as the subtitle for the time segment in the second video where the target audio material is filled. Here, the target audio material is the audio material in the to-be-filled audio materials that is in an editing state. In this way, by performing speech recognition processing on the audio material to generate the corresponding subtitle, the time spent by the user in entering text additionally is reduced, and the efficiency of video editing is improved.
[0133] Exemplarily, refer toFigure 11 , Figure 11 is a schematic diagram of an application scenario of the video editing method provided by an embodiment of the present application. As Figure 11 shown, when the user needs the subtitle text corresponding to the audio material filled in the audio filling area 1101, the audio filling area 1101 can be selected first. At this time, the audio filling area 1101 will be highlighted to indicate that the audio material filled in the audio filling area 1101 is in an editing state. Subsequently, when a click operation on the text recognition entry 1102 is received by the user, speech recognition processing will be performed on the audio material to obtain the speech recognition result 1103 of the audio material filled in the audio filling area 1101. At the same time, the corresponding subtitle 1104 will also be displayed in the first video. Among them, the type of the subtitle 1104 can be a soft subtitle (also called an internal subtitle, a packaged subtitle, a subtitle stream, etc., which embeds the subtitle file into the video stream and is part of the video stream) or a hard subtitle (also called an embedded subtitle, which suppresses the subtitle file and the video stream into the same set of data, like a watermark and cannot be separated).
[0134] It should be noted that when the first video itself has audio, after filling the audio material in the audio filling area displayed on the time axis of the first video, the filled audio material can replace the audio of the first video itself (that is, only play the filled audio material and mute the audio of the first video itself). Of course, the original audio of the first video and the filled audio material can also coexist. For example, the audio of the first video itself and the filled audio material are played through two channels respectively. The embodiments of the present application do not make specific limitations on this.
[0135] The video editing method provided by the embodiments of the present application can further fill the audio material in the audio filling area of the first video after obtaining the first video filled with the video material by using the video template, so as to obtain the second video filled with the audio material, achieving the purpose of directly adding audio during the process of making a video by using the video template, improving the video editing efficiency, and at the same time meeting the user's need to express their personalized ideas and views through audio during the process of making a video by using the video template.
[0136] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.
[0137] An embodiment of the present application provides a video editing method. When editing a video based on a template, a recording scene is created, enabling users to click the recording button in any provided video template to record audio, express their views and ideas in their own voices for the produced video, thereby increasing the diversity of the video. In addition, in this scene, in addition to being able to record their own voices, users can also change the voices of the recorded voices to imitate the voices of different characters, and at the same time, voice recognition processing can be performed on the recorded audio to generate corresponding subtitles. In this way, on the one hand, the diversity and personalization of the video can be increased, and on the other hand, the efficiency of video production can be improved, prompting users to quickly produce videos.
[0138] In the embodiment of the present application, users can record audio within the recordable time range of the video template by clicking the recording button in the video template, thereby improving the user experience. At the same time, the personalization of the view expression and the operation speed of the video can be improved, promoting the production of diverse videos.
[0139] Exemplarily, refer to Figure 4 , Figure 4 which is a schematic diagram of the application scenario of the video editing method provided by the embodiment of the present application. As Figure 4 shown, after the user opens the client, multiple candidate video templates are displayed on the home page of the client. The user selects the video template they want to produce on the home page of the client and clicks to use the template (for example, clicks the "Use Immediately" button), and can select pictures / videos in the local album of the terminal (such as a mobile phone) and import them into the selected video template.
[0140] Exemplarily, refer to Figure 5 , Figure 5 which is a schematic diagram of the application scenario of the video editing method provided by the embodiment of the present application. As Figure 5 shown, after the user imports the picture / video into the video template, if they are not satisfied with the imported picture / video, they can also click on the corresponding picture / video and perform fine-tuning of the video on the pop-up floating window. Among them, the fine-tuning includes replacing the video (that is, replacing the just imported picture / video and importing the picture / video reselected by the user into the video template), cropping the video (adjusting the picture size and content of the picture / video imported into the video template, supporting intercepting a certain period of video material, and also supporting cropping the picture / video according to the template frame size), and adjusting the volume of the video (adjusting the volume of the picture / video imported into the video template). If the user is satisfied with the picture / video imported into the video template, they can directly skip this step.
[0141] Exemplarily, refer to Figure 12 , Figure 12 which is a schematic diagram of the application scenario of the video editing method provided by the embodiment of the present application. AsFigure 12 As shown, when a user wants to express their thoughts on a template video through their voice, they can directly click on the recording entry 1201 displayed on the video template to open the recording function (i.e., when a click operation on the recording entry 1201 is received, the recording button 1202 will be displayed). In addition, a black line frame area 1203 is also displayed on the timeline of the video template. The black line frame area 1203 is the range in the video template where recording can be performed, and the part outside the black line frame area 1203 is not available for the user to record. The user clicks the recording button 1202 to start recording, and during the recording process, ripples 1205 will be displayed around the recording button 1202 to reflect the size of the user's recording voice in real time. In addition, during the recording process, a gray mask 1204 will cover the black line frame area 1203. The covered area represents the part where the user has recorded, and the area not covered by the gray mask 1204 represents the part where the user has not recorded.
[0142] When the user needs to pause the recording, they can click the recording button 1202 again to pause the recording. At this time, the user can directly click the confirmation button to end the recording, or continue to click the recording button 1202 to continue recording the unfinished part.
[0143] After the recording is completed, the client will automatically process the recording and display the real-time progress 1206 of the recording process. After the processing is completed, the user can click the text recognition entry 1207 to obtain the recognition result 1208 obtained by performing speech recognition processing on the recording, and at the same time, the corresponding subtitle text 1209 will also be displayed in the video template.
[0144] For example, refer to Figure 10 , Figure 10 which is a schematic diagram of the application scenario of the video editing method provided by the embodiments of the present application. As Figure 10 shown, after the user finishes recording, they can also select the recording that needs to be edited and make fine adjustments to the recording in the pop-up floating window. The fine adjustments include recording volume adjustment (select the recording editing part, select the volume adjustment function, corresponding to the volume adjustment entry 1004, and the adjustable volume range for volume adjustment can be between 0 - 200), recording voice change (select the recording editing part, select the recording voice change function, corresponding to the voice change entry 1006, and there are various optional voices to choose from, such as uncle voice, lolita voice, alien voice, doll voice, fat person voice, ethereal reverberation voice, etc. After the user selects a voice and clicks the confirmation button, the voice change can be completed), and re-recording (select the recording editing part, select the re-recording function, corresponding to the audio re-recording entry 1003, and the content of the recording that has just been completed can be cleared and re-recorded).
[0145] In the embodiments of the present application, the collection of the user's voice by hardware (such as a microphone, headphones, Bluetooth headphones, noise-canceling headphones, etc.) is used as the input of the voice. At the same time, the collected audio data is sampled and encoded by the hardware. During the recording process, the audio data is intercepted in real time to draw a real-time waveform diagram of the recording. In addition, the collected audio data can be stored.
[0146] Exemplarily, refer to Figure 13 , Figure 13 which is a schematic diagram of the model of the video editing system provided by the embodiments of the present application. As Figure 13 shown, the recording controller serves as the entry of the recording function and has functions such as starting recording, pausing recording, and stopping recording. At the same time, by calling the delegate, the user can conveniently record and process the recording result; the recording synchronization protocol is used to obtain the recording progress and call back the recording result to draw a real-time waveform diagram of the recording based on the recording result; the recording configuration management is used for recording status setting and managing input channels, such as microphones, headphones, etc.; the audio data processing is used for processing conversion operations, such as audio format conversion; the audio synchronization frequency control module is used for controlling the call back frequency, that is, manually setting the call back frequency; the client can process the recording information (such as the audio data collected through headphones) by detecting the recording controller, so as to complete the drawing of the real-time waveform diagram of the recording.
[0147] In addition, since the refresh frequency of the terminal device running the client is much higher than the call frequency of the audio collection device (such as headphones), it will cause an illusion of lag in the user experience. To solve this problem, the embodiments of the present application introduce audio synchronization frequency control. The call back frequency can be set by the user. For example, the recording sampling frequency of 44.1 khz (that is, the recording sampling frequency of the audio collection device, such as the call frequency of headphones) is converted to 48 khz (that is, the recording sampling frequency of the client, such as the refresh frequency of the terminal running the client), so as to ensure the smoothness of the user experience. In addition, since the audio data format requirements set externally may not match the audio data format of the client, during the recording process, the recording controller can perform audio data format conversion through audio data processing, so as to solve the problem that the audio data format requirements set externally do not match the audio data format of the internal software.
[0148] The following describes the recording process.
[0149] Exemplarily, refer to Figure 14 , Figure 14 which is a schematic flow diagram of the recording process provided by the embodiments of the present application. As Figure 14As shown, the client will first determine whether it has the recording permission. If it does not have the recording permission, it cannot record and directly enters the end state. When it is determined that the client has the recording permission, after setting the recording format (the default setting is 44.1 kHz) and the storage path of the recording, it can start recording. During the recording process, the audio acquisition device (such as headphones) can send the acquired audio data to the client through the recording synchronization protocol, so that the client can perform operations such as drawing the real-time waveform diagram of the recording based on the received audio data. When the user leaves the client (for example, enters the background), the recording will stop, that is, the recording ends. At this time, the recording synchronization protocol will send the terminated recording file to the client to notify the client that the recording has ended.
[0150] In addition, due to system limitations, such as being interrupted by an incoming call suddenly, reaching the set duration, resetting the system multimedia service (that is, the system crashes, resulting in the restart of the multimedia service information), losing the system multimedia service (that is, after the system crashes and the multimedia service information restarts, some service information is lost) and other software system problems, as well as changes in the audio input channel (that is, the recording channel) (that is, changes in the audio acquisition device, for example, when recording on a mobile phone, inserting headphones and changing to continue recording through the headphones), etc., will all cause the recording to end (that is, stop recording). This situation is similar to the above, and will end the current recording, and synchronize the status to the client through the recording synchronization protocol module, and the client will decide how to handle it.
[0151] In addition, it should be noted that when the recording is paused, if the audio control state of the system does not change, the recording can be resumed, that is to say, the audio data collected before and after the pause can be written into the same recording file; but when the audio state control changes after the pause, then at this time it will be considered that the current recording ends, and if the recording is started again, it will be a new recording, that is, a new recording file.
[0152] The following uses a complete recording process to illustrate the management of the recording file.
[0153] Preparation stage: Before starting the recording, first, it is necessary to set the storage path of the recording file, including creating a recording folder and the recording file name (for example, r_0.aac);
[0154] The first recording: The process of recording is also the process of storing audio data, and the acquired audio data will be written into the r_0 file;
[0155] Pausing the recording: At this time, no data will be written into the r_0 file;
[0156] Resuming the recording: Continue to write the acquired audio data into the r_0 file;
[0157] Finish recording: That is, end the recording, stop writing data to the r_0 file, and at the same time, return the recording file r_0 to the client through the recording synchronization protocol module;
[0158] Prepare for the second recording: Create a recording file named r_1 in the same folder;
[0159] …(Steps of the first recording)
[0160] Prepare for the third recording: Create a recording file named r_2 in the same folder;
[0161] …
[0162] It can be seen that in the working space, when creating a recording folder, the names of the recording files will increase gradually, such as r_0, r_1, r_2, …. In addition, when the editing is completed, the entire recording folder will be deleted when the deletion operation is performed. Also, when deleting an existing recording file, for example, assuming there are three recording files r_0, r_1, and r_2 now, when deleting the recording file r_2, if recording again next time, the name of the recording file will start from r_3. After the recording is completed, there will be three recording files r_0, r_1, and r_3 in the recording folder. If recording again, the name of the next recording file will be r_4, and so on. Here, after deleting the recording file r_2, the naming does not start from r_2 but from r_3 to avoid the problem of duplicate names of the recording files caused by restoring the recording file r_2 after deleting the recording file r_2.
[0163] In addition, for abnormal situations, since the recording is performed in the path of the editing cache space, thus, even if the recording is interrupted, it can continue to be edited at the position where the recording was interrupted last time based on the cache information in the editing cache space during the next editing.
[0164] The video editing method provided by the embodiments of the present application enables users to flexibly add their own recording content during the process of making videos using video templates, thus avoiding the complex process of having to export the template video separately and then re-importing the final video into the editing software to add recordings and add subtitle text when users make videos using the methods provided by related technologies, improving the efficiency of video production. At the same time, by performing speech recognition processing on the recording to generate corresponding subtitle text, it reduces the time for users to input text content additionally, and can also change the voice of their own recordings according to the preferences of users, changing the voice to that of other characters, enhancing the personalization of users, adding a certain amount of fun to users during recording, and creating a simple scene-based recording atmosphere.
[0165] Next, the exemplary structure of the video editing device 465 provided in the embodiments of the present application implemented as software modules will be further described. In some embodiments, as Figure 2 shown, the software modules in the video editing device 465 stored in the memory 460 may include: a display module 4651, an acquisition module 4652,
[0166] The display module 4651 is configured to display a first video in response to a video filling trigger operation for a video template, where the first video is formed by filling at least one video material in the video template; the display module 4651 is further configured to display at least one audio filling area on the time axis of the first video; the acquisition module 4652 is configured to acquire at least one audio material to be filled in at least one audio filling area in response to an audio setting operation; the display module 4651 is further configured to display a second video to replace the first video in response to an audio filling trigger operation for the first video, where the second video is formed by filling the corresponding audio material in at least one audio filling area of the first video.
[0167] In some embodiments, the display module 4651 is further configured to display an audio setting entry corresponding to the first video; and in response to a trigger operation for the audio setting entry, display the time axis of the first video and display at least one audio filling area on the time axis, where the audio filling area is used to indicate the time period on the time axis where the audio material can be filled.
[0168] In some embodiments, the acquisition module 4652 is further configured to acquire at least one audio filling area by at least one of the following methods: acquire at least one pre-set audio filling area from the video template; divide at least one audio filling area from the time axis according to the playing time of the video material filled in the first video; divide the time axis according to the plot units of the first video and determine the audio filling area corresponding to each plot unit.
[0169] In some embodiments, the display module 4651 is further configured to display that the selected target audio filling area is in an editing state in response to an audio filling area selection operation; and in response to a time setting operation for the target audio filling area, update the display of the target audio filling area based on the start time and end time set on the time axis.
[0170] In some embodiments, when the audio setting entry is a recording entry, the display module 4651 is further configured to display recording controls for a corresponding target audio filling area, where the target audio filling area is the audio filling area in an editing state among at least one audio filling area; the video editing device 465 further includes an acquisition module 4653, configured to start acquiring audio in response to a first audio setting operation of enabling the recording control; and stop acquiring audio in response to a second audio setting operation of disabling the recording control; the video editing device 465 further includes a determination module 4654, configured to use the acquired audio as the audio material to be filled in the target audio filling area.
[0171] In some embodiments, when the audio setting entry is an audio material selection entry, the display module 4651 is further configured to display audio material selection controls, where the audio material selection controls are used to select multiple candidate audio materials; the determination module 4654 is further configured to use at least one selected audio material among the multiple candidate audio materials as the audio material to be filled in the target audio filling area in response to an audio setting operation of selecting through the audio material selection controls, where the target audio filling area is the audio filling area in an editing state among at least one audio filling area.
[0172] In some embodiments, the display module 4651 is further configured to display an audio re-recording entry; and to display recording controls for re-acquisition in response to a triggering operation on the audio re-recording entry; the acquisition module 4653 is further configured to start acquiring audio in response to a third audio setting operation of enabling the recording control; and stop acquiring audio in response to a fourth audio setting operation of disabling the recording control; the determination module 4654 is further configured to use the re-acquired audio as the audio material to be filled in the target audio filling area, where the target audio filling area is the audio filling area in an editing state among at least one audio filling area.
[0173] In some embodiments, the display module 4651 is further configured to display an audio deletion entry; the video editing device 465 further includes a deletion module 4655, configured to delete the audio material corresponding to the target audio filling area in response to a triggering operation on the audio deletion entry, where the target audio filling area is the audio filling area in an editing state among at least one audio filling area.
[0174] In some embodiments, the display module 4651 is further configured to display a volume adjustment entry; and to display volume adjustment controls in response to a triggering operation on the volume adjustment entry; the determination module 4654 is further configured to determine the set volume as the volume of the audio material corresponding to the target audio filling area in response to a setting operation on the volume adjustment controls, where the target audio filling area is the audio filling area in an editing state among at least one audio filling area.
[0175] In some embodiments, the display module 4651 is further configured to display a voice-changing entry; and to display a plurality of candidate voice-producing objects in response to a trigger operation on the voice-changing entry; the video editing device 465 further includes a replacement module 4656 configured to replace an initial voice-producing object of an audio material corresponding to a target audio filling area with a selected target voice-producing object among the plurality of candidate voice-producing objects in response to a voice-producing object selection operation, where the target audio filling area is an audio filling area in an editing state among at least one audio filling area.
[0176] In some embodiments, the display module 4651 is further configured to display a text recognition entry; and to display a speech recognition result of a target audio material in response to a trigger operation on the text recognition entry, and use it as a subtitle for a time segment in the second video filled with the target audio material, where the target audio material is an audio material in an editing state among the audio materials to be filled.
[0177] In some embodiments, the display module 4651 is further configured to display at least one unfilled segment included in a video template and a video material selection control, where the video material selection control is used to select a plurality of candidate video materials; and to highlight the selected at least one video material in response to a video material selection operation through the video material selection control; the video editing device 465 further includes a filling module 4657 configured to fill the selected at least one video material into the unfilled segment corresponding to the video template in response to a video filling trigger operation on the video template.
[0178] In some embodiments, the display module 4651 is further configured to display a plurality of candidate video templates; to display a usage entry of a selected video template among the plurality of candidate video templates in response to a video template selection operation; to display at least one unfilled segment included in the selected video template in response to a trigger operation on the usage entry of the selected video template; and to display introduction information of the unfilled segment, where the introduction information is used to characterize the type of the video material to be filled in the unfilled segment.
[0179] In some embodiments, the display module 4651 is further configured to display a replacement video entry; to display a video selection control in response to a trigger operation on the replacement video entry, where the video selection control includes a plurality of candidate video materials; the replacement module 4656 is further configured to replace a video material with at least one selected video material among the plurality of candidate video materials in response to a video material selection operation in the video selection control.
[0180] In some embodiments, the display module 4651 is further configured to display a video cropping entry; and in response to a trigger operation on the video cropping entry, display video cropping controls, where the video cropping controls include at least one of a picture cropping control and a time period selection control; the video editing device 465 further includes a cropping module 4658, configured to crop the picture outside the set cropping frame from the video material in response to a cropping frame setting operation based on the picture cropping control; and crop the video segment outside the set time period from the video material in response to a time setting operation based on the time period selection control.
[0181] In some embodiments, the display module 4651 is further configured to, in response to a video material selection operation, display that the selected target video material among at least one video material is in an editing state; display a volume adjustment entry, and in response to a trigger operation on the volume adjustment entry, display volume adjustment controls; the determination module 4654 is further configured to, in response to a setting operation on the volume adjustment controls, determine the set volume as the volume of the target video material.
[0182] In some embodiments, the display module 4651 is further configured to display a plurality of candidate video materials; in response to a video material selection operation, highlight the selected at least one video material and display a plurality of candidate video templates matching the at least one video material; the filling module 4657 is further configured to, in response to a video template selection operation, fill the at least one video material into the to-be-filled segment corresponding to the selected video template.
[0183] In some embodiments, the video editing device 465 further includes a conversion module 4659, configured to perform at least one of the following processes: convert the data format of the audio material to conform to the data format of the first video; convert the sampling frequency of the audio material to conform to the sampling frequency of the first video.
[0184] It should be noted that the description of the device in the embodiments of the present application is similar to the implementation of the video editing method in the foregoing text and has similar beneficial effects, so it will not be elaborated here. For the technical details not described in the video editing device provided in the embodiments of the present application, they can be understood according to Figure 3 the description.
[0185] The embodiments of the present application provide a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the video editing method described above in the embodiments of the present application.
[0186] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions, when executed by a processor, cause the processor to execute the method provided by the embodiment of the present application. For example, as Figure 3 shown in the video editing method.
[0187] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0188] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0189] As an example, the executable instructions may or may not correspond to a file in the file system, may be stored as part of a file that stores other programs or data. For example, stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program under discussion, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).
[0190] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.
[0191] In summary, in the process of using a video template in the embodiment of the present application, by cooperating the functions of video material filling and audio material filling, it is possible to support the flexible addition of corresponding video materials and audio materials according to requirements, which not only improves the video editing efficiency but also fully meets the needs of expressing personalized ideas and viewpoints in the content of the video.
[0192] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A video editing method, characterized in that, The method includes: In response to a video filling trigger operation for a video template, display a first video, where the first video is formed by filling at least one video material in the video template; Display an audio setting entry corresponding to the first video; In response to a trigger operation for the audio setting entry, display the timeline of the first video, and display at least one audio filling area on the timeline, where the audio filling area is used to indicate the time period on the timeline where the audio material can be filled; When the audio setting entry is a recording entry, display a recording control for the corresponding target audio filling area, where the target audio filling area is the audio filling area in the at least one audio filling area that is in an editing state; In response to a trigger operation for the recording control for the corresponding target audio filling area, obtain the audio material to be filled in the at least one audio filling area; In response to an audio filling trigger operation for the first video, display a second video to replace the first video, where the second video is formed by filling the corresponding audio material in the at least one audio filling area of the first video.
2. The method according to claim 1, wherein Before displaying at least one audio filling area on the timeline, the method further includes: Obtain the at least one audio filling area by at least one of the following methods: Obtain the at least one pre-set audio filling area from the video template; Divide the at least one audio filling area from the timeline according to the playing time of the video material filled in the first video; Divide the timeline according to the plot units of the first video, and determine the audio filling area corresponding to each plot unit.
3. The method according to claim 1, wherein The method further includes: In response to an audio filling area selection operation, display that the selected target audio filling area is in an editing state; In response to a time setting operation for the target audio filling area, update and display the target audio filling area based on the start time and end time set on the timeline.
4. The method according to claim 1, wherein The step of, in response to a trigger operation for the recording control for the corresponding target audio filling area, obtaining the audio material to be filled in the at least one audio filling area includes: In response to a first audio setting operation to turn on the recording control, start collecting audio; In response to a second audio setting operation to turn off the recording control, stop collecting audio; Use the collected audio as the audio material to be filled in the target audio filling area.
5. The method according to claim 1, wherein When the audio setting entry is an audio material selection entry, the method further includes: Display an audio material selection control, where the audio material selection control is used to select multiple candidate audio materials; In response to an audio setting operation to select through the audio material selection control, use at least one of the selected audio materials from the multiple candidate audio materials as the audio material to be filled in the target audio filling area.
6. The method according to claim 1, characterized in that, When displaying the second video to replace the first video, the method further includes: Display an audio re-recording entry; In response to a trigger operation for the audio rerecording entry, display a recording control for re-collecting audio; In response to a third audio setting operation for enabling the recording control, start collecting audio; In response to a fourth audio setting operation for disabling the recording control, stop collecting audio; Use the audio obtained by re-collecting as the audio material to be filled in the target audio filling area.
7. The method according to claim 1, wherein When a second video is displayed to replace the first video, the method further includes: Display an audio deletion entry; In response to a trigger operation for the audio deletion entry, delete the audio material corresponding to the target audio filling area.
8. The method according to claim 1, wherein When a second video is displayed to replace the first video, the method further includes: Display a volume adjustment entry; In response to a trigger operation for the volume adjustment entry, display a volume adjustment control; In response to a setting operation for the volume adjustment control, determine the set volume as the volume of the audio material corresponding to the target audio filling area.
9. The method according to claim 1, characterized in that, When a second video is displayed to replace the first video, the method further includes: Display a voice-changing entry; In response to a trigger operation for the voice-changing entry, display a plurality of candidate voice objects; In response to a voice object selection operation, replace the initial voice object of the audio material corresponding to the target audio filling area with the selected target voice object among the plurality of candidate voice objects.
10. The method according to claim 1, characterized in that When a second video is displayed to replace the first video, the method further includes: Display a text recognition entry; In response to a trigger operation for the text recognition entry, display the speech recognition result of the target audio material and use it as the subtitle for the time segment in the second video filled with the target audio material, where the target audio material is the audio material in the to-be-filled audio materials that is in the editing state.
11. The method according to claim 1, wherein Before displaying the first video in response to a video filling trigger operation for a video template, the method further includes: Display a plurality of candidate video materials; In response to a video material selection operation, highlight at least one selected video material and display a plurality of candidate video templates matching the at least one video material; In response to a video template selection operation, fill the at least one video material into the to-be-filled segment corresponding to the selected video template.
12. A video editing device, characterized in that, The device includes: A display module, configured to display a first video in response to a video filling trigger operation for a video template, where the first video is formed after filling at least one video material in the video template; The display module is further configured to display an audio setting entry corresponding to the first video; in response to a trigger operation for the audio setting entry, display the timeline of the first video and display at least one audio filling area on the timeline, where the audio filling area is used to indicate the time period on the timeline where the audio material can be filled; An acquisition module, configured to display a recording control for a corresponding target audio filling area when the audio setting entry is a recording entry, where the target audio filling area is the audio filling area in an editing state among the at least one audio filling area; and in response to a triggering operation on the recording control for the corresponding target audio filling area, acquire audio material to be filled in the at least one audio filling area. The display module is further configured to, in response to an audio filling trigger operation on the first video, display a second video to replace the first video, where the second video is formed after filling corresponding audio material in the at least one audio filling area of the first video.
13. An electronic device, characterized in that, The electronic device includes: A memory for storing executable instructions; A processor, configured to implement the video editing method according to any one of claims 1-11 when executing the executable instructions stored in the memory.
14. A computer-readable storage medium, characterized in that, Stored with executable instructions, which are configured to implement the video editing method according to any one of claims 1-11 when being executed by a processor.
15. A computer program product comprising computer-executable instructions, characterized in that, When the computer executable instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Cloud editing method, device and equipment and storage medium
CN111918128A
Video synthesis method and device, electronic equipment and storage medium
CN112291484A
Multimedia file material processing method and device, electronic equipment and storage medium
CN112449231A