Video editing method, electronic equipment and storage medium
By intelligently retaining the original video soundtrack in the 'one-click editing' function of electronic devices, the problem of templated editing content is solved, and better audio effects and user experience is achieved.
Patent Information
- Application Number
- CN202311785110.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-01
AI Technical Summary
When electronic devices perform 'one-click editing', the editing content may be templated, which cannot meet users' needs for the 'one-click editing' function.
By intelligently retaining the original video sound, avoid templating the audio content of edited videos. In the edited video, the video soundtrack and background music can coexist, and the background music automatically decreases the volume during the video soundtrack to highlight the sound elements of the material selected by the user.
It realizes the wonderful original soundtrack that retains video materials in one-click editing videos, improves audio effects, avoids the templateization of audio content, and enhances user experience.
Smart Images

Figure CN120238618A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent terminals, and particularly to a video editing method, an electronic device, and a storage medium. Background Art
[0002] With the advent of the self-media era, video editing has become more and more popular, and more and more ordinary people without editing skills also need to edit their multimedia materials. Thus, the "one-key editing" function has emerged, and users can complete the editing of multimedia materials with one key operation.
[0003] However, when the electronic device automatically edits the multimedia material based on the "one-key editing" function, it may cause the editing content to be templatized and fail to meet the user's usage requirements for the "one-key editing" function. Summary of the Invention
[0004] This application provides a video editing method, an electronic device, and a storage medium. In this method, when the electronic device automatically edits the multimedia material based on the "one-key editing" function, it can intelligently retain the original video sound and avoid the templatization of the audio content of the edited video.
[0005] In a first aspect, an embodiment of this application provides a video editing method. The method includes: the electronic device displays a first interface; the first interface displays the multimedia material selected by the user; the multimedia material includes at least one video material;
[0006] The electronic device displays a second interface in response to a first operation; the second interface includes a first edited video corresponding to the multimedia material; wherein, in a first time period of the first edited video, the audio of the first edited video includes background music corresponding to a target editing template, and the audio of the first edited video does not include the original video sound; in a second time period of the first edited video, the audio of the first edited video includes the first original video sound in at least one video material.
[0007] Exemplarily, the first interface includes a one-key editing control, and the first operation may be an operation on the one-key editing control, such as a click operation.
[0008] Exemplarily, in addition to the video material, the multimedia material selected by the user may also include picture materials, etc.
[0009] In the second interface, the first edited video can be played so that the user can understand the details of the first edited video.
[0010] It should be noted that, in the second time period of the first edited video, the audio of the first edited video may include the original video sound but not the background music, or may include both the original video sound and the background music. It can be understood that the first original video sound is the original sound of the video material retained in the edited video.
[0011] Similarly, in the third period of the first clipped video, the audio of the first clipped video only includes the background music corresponding to the target clip template; in the fourth period of the first clipped video, the audio of the first clipped video includes the first video original sound in at least one video material, and so on. In this embodiment, the number of video original sound segments retained in the clipped video is not limited.
[0012] In this way, when the electronic device automatically clips the multimedia material selected by the user, it can intelligently retain the wonderful original sound of the video material, rather than uniformly adding only background music or retaining all the video original sound, which highlights the sound elements of the material input by the user and improves the audio effect of the one-key clipped video.
[0013] According to the first aspect, in the first period of the first clipped video, the playing volume of the background music is the first volume; in the second period of the first clipped video, the audio of the first clipped video further includes the background music, and the playing volume of the background music is the second volume, and the second volume is less than the first volume.
[0014] Among them, the first volume can be understood as the set volume of the electronic device, or the normal volume, standard volume, etc. of the clipped video. Exemplarily, the ratio of the second volume to the first volume is 50%.
[0015] In this way, during the video original sound playing period, the video original sound can coexist with the background music. At this time, the background music automatically reduces the volume to give way to the video original sound, which more prominently shows the sound elements of the material selected by the user.
[0016] According to the first aspect, or any one of the implementation manners of the above first aspect, in the second period of the first clipped video, the playing volume of the first video original sound is the first volume.
[0017] According to the first aspect, or any one of the implementation manners of the above first aspect, the method further includes: at the first moment of the first clipped video, the first video original sound starts to fade in; at the second moment of the first clipped video, the first video original sound starts to fade out.
[0018] In the embodiment of the present application, in order to ensure the effect of retaining the video original sound and avoid the phenomenon of abrupt playing of the video original sound, the appearance and stop of the video original sound adopt appropriate fade-in and fade-out effects, thereby improving the user's auditory experience.
[0019] According to the first aspect, or any one of the implementation manners of the above first aspect, the method further includes:
[0020] When the first video original sound fades in, the playing volume of the first video original sound increases from zero to the first volume according to the first step length; when the first video original sound fades out, the playing volume of the first video original sound decreases from the first volume to zero according to the second step length.
[0021] Among them, the first step length and the second step length can be equal or not equal.
[0022] The fade-in duration of the first video original sound and the fade-out duration of the video original sound can be equal or not equal. Exemplarily, the fade-in duration of the first video original sound is 1 second, and the fade-out duration of the first video original sound is 1 second.
[0023] According to the first aspect, or any implementation manner of the above first aspect, the first video original sound includes the second video original sound, the starting point of the first video original sound is earlier than the starting point of the second video original sound, and the ending point of the first video original sound is later than the ending point of the second video original sound;
[0024] The ratio of the playing volume at the starting point of the second video original sound to the first volume is greater than the first value, and the ratio of the playing volume at the ending point of the second video original sound to the first volume is greater than the second value.
[0025] Among them, the first video original sound can be understood as the video original sound retained in the edited video, and the second video original sound can be understood as the high-value video original sound. In the case of the fade-in and fade-out of the video original sound, in order to ensure the playing volume of the high-value video original sound, the duration of the video original sound retained in the edited video can be set slightly longer.
[0026] The first value and the second value can be equal or not equal. Exemplarily, both the first value and the second value are 75%.
[0027] In this way, in the case of the fade-in and fade-out of the video original sound, it can be ensured that the playing volume at the starting point of the high-value video original sound is not lower than the set value, and the playing volume at the ending point of the high-value video original sound is not lower than the set value.
[0028] According to the first aspect, or any implementation manner of the above first aspect, the method further includes:
[0029] When the first video original sound starts to fade in, the playing volume of the background music decreases from the first volume to the second volume according to the third step length; when the first video original sound starts to fade out, the playing volume of the background music increases from the second volume to the first volume according to the fourth step length.
[0030] Among them, the third step length and the fourth step length can be equal or not equal.
[0031] The fade - out duration when the background music weakens and the fade - in duration when the background music strengthens can be equal or not equal. Exemplarily, the fade - out duration when the background music weakens is 1 second, and the fade - in duration when the background music strengthens is 1 second.
[0032] In the embodiments of the present application, in order to ensure the effect of the background music avoiding the video original sound and to avoid the phenomenon of sudden volume change when the background music weakens or strengthens, a suitable fade - in and fade - out effect is adopted when the background music weakens or strengthens, so as to improve the user's auditory experience.
[0033] According to the first aspect, or any one of the implementation manners of the above - mentioned first aspect, after the second interface is displayed, the method further includes:
[0034] The electronic device responds to the second operation and displays a third interface; among the multiple video sound setting options in the third interface, the first option is in a selected state, and the first option is used to indicate that the video original sound is intelligently retained in the clipped video.
[0035] Exemplarily, the first option is the "Intelligent Retention" option 4031 as shown in (3) of Figure 12a the figure.
[0036] In the embodiments of the present application, the "Intelligent Retention" option is the default - selected option for video sound setting, that is, the first clipped video is generated by clipping in a way that intelligently retains the video original sound.
[0037] According to the first aspect, or any one of the implementation manners of the above - mentioned first aspect, after the electronic device displays the third interface, the method further includes:
[0038] The electronic device responds to the third operation and displays a fourth interface; in the fourth interface, the first option is in a non - selected state, and the second option is in a selected state, and the second option is used to indicate that all the video original sound is retained in the clipped video;
[0039] The electronic device responds to the fourth operation and displays a fifth interface; the fifth interface includes a second clipped video corresponding to the multimedia material; at any time period of the second clipped video, the audio of the second clipped video includes background music and video original sound.
[0040] Exemplarily, the second option is the "All Retention" option 4032 as shown in (3) of Figure 12a the figure. Exemplarily, the third operation is an operation of clicking the second option, and the fourth operation is an operation of clicking the confirmation option.
[0041] According to the first aspect, or any one of the implementation manners of the above - mentioned first aspect, after the electronic device displays the third interface, the method further includes:
[0042] In response to a fifth operation, the electronic device displays a sixth interface; in the sixth interface, a first option is in an unselected state, and a third option is in a selected state, and the third option is used to indicate that only background music is added to the clipped video.
[0043] In response to a sixth operation, the electronic device displays a seventh interface; the seventh interface includes a third clipped video corresponding to the multimedia material; at any time period of the third clipped video, the audio of the second clipped video only includes background music.
[0044] Exemplarily, the third option is the "only add background music" option 4033 shown in (3) of Figure 12a . Exemplarily, the fifth operation is an operation of clicking the third option, and the sixth operation is an operation of clicking the confirmation option.
[0045] When the user obtains a clipped video finished product based on the one - key clipping function, the video clipping application generates a default video clipping finished product based on the default sound setting option, that is, clips and generates a video clipping finished product in a way of intelligently retaining the original video sound. On this basis, if the user adjusts the video sound setting option, the video clipping application dynamically adjusts the audio information in the video clipping finished product so that the adjusted video clipping finished product meets the user's video clipping requirements. Since the analysis process of the high - value original sound in the high - light video segment is integrated into the analysis process of the high - light video segment, when the user adjusts the video sound setting, the video clipping application can adjust the video sound with low latency and will not bring an obvious delay feeling to the user.
[0046] According to the first aspect, or any one of the above - mentioned implementation manners of the first aspect, in response to a first operation, the electronic device displays a second interface, including:
[0047] In response to a first operation, the electronic device analyzes the multimedia material to determine the theme type corresponding to the multimedia material and at least one high - light video segment; the electronic device analyzes the audio events in at least one high - light video segment to determine at least one segment of first video original sound; the electronic device automatically clips the multimedia material according to the target clipping template matching the theme type, at least one high - light video segment, and at least one segment of first video original sound to generate a first clipped video; the electronic device displays the second interface.
[0048] In the embodiments of the present application, when the video clipping application performs offline analysis on the material selected by the user, it can not only analyze the material theme type and high - light video segments, but also analyze the high - value audio in the high - light video segments. Among them, the process of intelligent recognition of the original video sound can be integrated into the analysis process of the high - light segments of the video material.
[0049] According to the first aspect, or any implementation of the above first aspect, the electronic device analyzes audio events in at least one high-light video clip to determine at least one segment of first video original sound, including:
[0050] The electronic device identifies audio events in at least one high-light video clip; the electronic device filters the audio events according to audio filtering conditions to determine at least one segment of first video original sound.
[0051] In the embodiments of the present application, for each high-light video clip, the electronic device can determine whether each high-light video clip includes high-value audio events and the high-value audio events included in each high-light video clip according to the audio events included therein and the signal quality information of each audio event. It can be understood that high-value audio events refer to audio events with better audio signal quality in high-light video clips, for example, audio events whose audio signal quality meets preset conditions.
[0052] According to the first aspect, or any implementation of the above first aspect, the audio filtering conditions include at least one of the following: audio signal loudness filtering conditions, audio signal signal-to-noise ratio filtering conditions.
[0053] The electronic device can filter out high-value audio events from the respective audio events in the high-light video clip according to information such as the signal loudness and signal-to-noise ratio of the audio events. Exemplarily, if the signal loudness and signal-to-noise ratio of a certain audio event meet the corresponding filtering conditions, this audio event can be retained as a high-value audio event (i.e., the first video original sound).
[0054] According to the first aspect, or any implementation of the above first aspect, the types of audio events include human voice types and non-human voice types; among them, the human voice types include at least one of the following: speech, emotional human voice, singing voice; the non-human voice types include at least one of the following: applause, animal calls, environmental sounds.
[0055] For example, speech can include conversations, monologues, etc., emotional human voices can include cries, laughter, shouts, etc., animal calls can include meows, barks, etc., and environmental sounds can include ocean waves, wind, fireworks, etc.
[0056] According to the first aspect, or any implementation of the above first aspect, the electronic device displays a first interface, including: the electronic device displays the interface of a gallery application; the electronic device responds to a seventh operation and displays the first interface; a video material is played in the first interface.
[0057] As Figure 1a shown in the scenario, the first interface can be, for example, Figure 1aIn the interface shown in (2), the first operation may be an operation of clicking on the "AI One - Key Blockbuster" control 104.
[0058] According to the first aspect, or any implementation manner of the above - mentioned first aspect, the electronic device displays a first interface, including: the electronic device displays the interface of the gallery application; the electronic device responds to the eighth operation and displays the first interface; in the first interface, there are thumbnails of multiple multimedia materials, and one or more thumbnails are in a selected state.
[0059] As Figure 1b shown in the scenario, the first interface may, for example, be Figure 1b the interface shown in (3) in the figure, and the first operation may be an operation of clicking on the "One - Key Movie" control 207.
[0060] In a second aspect, an embodiment of the present application provides an electronic device. The electronic device includes: one or more processors; a memory; and one or more computer programs, where one or more computer programs are stored in the memory, and when the computer programs are executed by one or more processors, the electronic device executes the video - editing method according to the first aspect and any one of the first aspect.
[0061] The second aspect and any implementation manner of the second aspect respectively correspond to the first aspect and any implementation manner of the first aspect. The technical effects corresponding to the second aspect and any implementation manner of the second aspect can refer to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, and will not be elaborated here.
[0062] In a third aspect, an embodiment of the present application provides a computer - readable storage medium. The computer - readable storage medium includes a computer program, and when the computer program runs on the electronic device, the electronic device executes the video - editing method according to the first aspect and any one of the first aspect.
[0063] The third aspect and any implementation manner of the third aspect respectively correspond to the first aspect and any implementation manner of the first aspect. The technical effects corresponding to the third aspect and any implementation manner of the third aspect can refer to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, and will not be elaborated here.
[0064] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is run, the computer executes the video - editing method according to the first aspect or any one of the first aspect.
[0065] The fourth aspect and any implementation manner of the fourth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the fourth aspect and any implementation manner of the fourth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated herein.
[0066] In a fifth aspect, the present application provides a chip, which includes a processing circuit and transceiver pins. Among them, the transceiver pins and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the video editing method as described in the first aspect or any item in the first aspect to control the receiving pin to receive a signal and control the sending pin to send a signal.
[0067] The fifth aspect and any implementation manner of the fifth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the fifth aspect and any implementation manner of the fifth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1a Schematic diagram of an application scenario of one - key automatic video editing shown by way of example;
[0069] Figure 1b Schematic diagram of an application scenario of one - key automatic video editing shown by way of example;
[0070] Figure 2 Schematic diagram of the audio in the automatically edited video clip shown by way of example;
[0071] Figure 3 Schematic diagram of the hardware structure of an electronic device shown by way of example;
[0072] Figure 4 Schematic diagram of the software structure of an electronic device shown by way of example;
[0073] Figure 5 Schematic diagram of the module interaction involved in the video editing method shown by way of example;
[0074] Figure 6 Schematic diagram of the intelligent video original sound retention in the automatically edited video clip shown by way of example;
[0075] Figure 7 Schematic diagram of the fade - in and fade - out of the video original sound when the video original sound and background music co - exist shown by way of example;
[0076] Figure 8 Schematic diagram of the change in the playback volume when the video original sound fades in and fades out in the automatically edited video clip shown by way of example;
[0077] Figure 9 A schematic diagram showing a comparison between the original video soundtrack and the high-value original video soundtrack retained in the automatic editing of a video as an example;
[0078] Figure 10 This is a schematic diagram of the volume change of background music when the original sound of a video fades in and out in an automatic editing of a video;
[0079] Figure 11 A schematic diagram showing the volume change of background music before and after the original sound of a video fades in and out in an exemplary video automatic editing;
[0080] Figure 12a A schematic diagram of a user setting video sound in an exemplary one-key video automatic editing scenario;
[0081] Figure 12b The figure is a schematic diagram showing an exemplary one-key video automatic editing scenario in which a user sets the video sound. DETAILED DESCRIPTION
[0082] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0083] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0084] The terms "first" and "second" in the description and claims of the embodiments of the present application are used to distinguish different objects rather than to describe a specific order of objects. For example, a first target object and a second target object are used to distinguish different target objects rather than to describe a specific order of target objects.
[0085] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0086] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more. For example, a plurality of processing units means two or more processing units; a plurality of systems means two or more systems.
[0087] With the advent of the self-media era, video editing has become more and more popular, and more and more ordinary people without editing skills also need to edit their multimedia materials (including but not limited to video materials, picture materials, etc.). Thus, the "one-key editing" function (or the "one-key video creation" function, the "one-key blockbuster" function, etc.) has emerged. Users only need to perform a one-key operation (such as clicking the "one-key editing" control) on an electronic device (hereinafter, the electronic device is taken as an example of a mobile phone for explanation), and then they can complete the automatic editing of the multimedia material and obtain an automatically edited video (or an automatically created video, etc.).
[0088] In one application scenario, a user can install a third-party video editing application on a mobile phone, import one or more multimedia materials into the third-party video editing application, and use the "one-key editing" function of the third-party video editing application to complete the automatic editing of the multimedia materials.
[0089] In another application scenario, a user can also use the "one-key editing" function provided by the system-level application in the mobile phone to complete the automatic editing of the multimedia materials. In this scenario, the user does not need to install an additional third-party video editing application on the mobile phone, and does not need to perform an import operation on the multimedia materials when automatic editing of the multimedia materials is required. Among them, Figure 1a and Figure 1b respectively exemplarily show application scenarios of automatically editing videos based on the "one-key editing" function provided by the system-level application in the mobile phone.
[0090] Figure 1a In (1) and Figure 1b In (1) exemplarily shows the display interface of the gallery application. Exemplarily, in response to the user's operation of clicking the gallery application icon on the mobile phone, the mobile phone starts the gallery application and displays the gallery application interface 101 as shown in Figure 1a In (1) or the gallery application interface 201 as shown in Figure 1b In (1). In the gallery application interface 101, various pictures (or photos) and videos stored in the gallery application are displayed. Among them, for video editing, the pictures and videos stored in the gallery application can all be used as multimedia materials (or materials to be edited, hereinafter simply referred to as materials) for video editing.
[0091] In one implementation, a user can select a video material for automatic video editing. Exemplarily, if the duration of a certain video material is greater than a preset threshold (e.g., 15 s), then the video material meets the condition for individual editing, and the video editing application can complete the automatic editing operation based on the video material to obtain a finished edited video.
[0092] Continue to refer to Figure 1a As shown in (1) of, in the gallery application interface 101, the user clicks on the thumbnail 102 corresponding to the video material A. In response to the user's operation, the mobile phone plays the video material A, as Figure 1a shown in (2) of. Since the duration of the video material A is greater than the preset threshold and meets the condition for individual editing, in the playback interface 103 of the video material A, an "AI One - Key Blockbuster" control 104 can also be displayed. Continue to refer to Figure 1a shown in (2) of. The user clicks on the "AI One - Key Blockbuster" control 104, which can control the video editing application to perform automatic editing on the video material A. In response to the user's operation, the video editing application in the mobile phone performs automatic editing on the video material A to generate an edited video B corresponding to the video material A, as shown Figure 1aThe interface 105 shown in (3). Optionally, before the mobile phone displays the interface 105, a transition interface for prompting the user that a clipped video B is being generated can also be displayed, such as a transition interface for indicating "material analysis in progress" or "video clipping in progress". In the interface 105, the clipped video B is displayed so that the user can learn about the automatic clipping result of the video material A. In the interface 105, confirmation control 106, template control 107, segment control 109, editing control 110, etc. can also be displayed. When the user needs to store the current clipped video in the gallery application, the user can click the confirmation control 106 to store the current clipped video in the gallery application. When the user needs to adjust the clipping template of the clipped video, the user can click the template control 107 to make the mobile phone display a video clipping template selection interface. Among them, in the video clipping template selection interface, the user can reselect an automatic clipping template for the video material A. When the user needs to adjust the background music of the clipped video, the user can click the music control 108 to make the mobile phone display a video background music editing interface. In the video background music editing interface, the user can edit the background music of the clipped video, such as adjusting the background music volume, changing the background music, etc. When the user needs to adjust some segments of the clipped video, the user can click the segment control 109 to make the mobile phone display a video segment editing interface. Among them, in the video segment editing interface, the user can edit the video segments in the clipped video, such as adjusting the video segments, changing the video segments, etc. When the user needs to edit the clipped video, the user can click the editing control 110 to make the mobile phone display a video editing interface. In the video editing interface, the user can edit the clipped video, such as adding stickers, adjusting the duration, etc.
[0093] In another embodiment, the user can select multiple multimedia materials for automatic video clipping. For example, the user can select multiple picture materials for automatic video clipping, or select multiple video materials for automatic video clipping, or select video materials and picture materials for automatic video clipping.
[0094] Exemplarily, continue to refer to Figure 1b as shown in (1). In the gallery application interface 201, the user long-presses the thumbnail corresponding to any one material, such as clicking on the thumbnail 202 corresponding to the picture material a. In response to the user's operation, the mobile phone displays as Figure 1b the material selection interface 203 shown in (2). Optionally, the user can also trigger the mobile phone to display as Figure 1b the material selection interface 203 shown in (2) by clicking on the selection control, or trigger the mobile phone to display as Figure 1bThe material selection interface 203 shown in (2) of the present embodiment is not limited thereto. Exemplarily, in the material selection interface 203, a to-be-selected identifier 204 is displayed on the thumbnail of each material (including picture materials and video materials). When the user selects a certain material, a selected identifier 205 (such as the checkmark identifier shown in the figure) is displayed in the to-be-selected identifier 204 on the thumbnail of the material, which can be referred to Figure 1b as shown in (3) of the present embodiment. Optionally, the selected identifier 205 can also be a numerical identifier (such as "1", "2", etc.), which is used to indicate the order and quantity of the materials selected by the user, and the present embodiment is not limited thereto. When the materials selected by the user meet the conditions for automatic video editing, for example, when the user selects one or more picture materials, or when the user selects one or more video materials, optionally, when the user selects a video material and the video material meets the conditions for separate editing, the mobile phone pops up a prompt information box 206, which is used to prompt the user that a movie-like video short can be automatically generated based on the selected materials. Among them, a "One-key Movie" control 207 is displayed in the prompt information box 206. Continuing to refer to Figure 1b as shown in (3) of the present embodiment, when the user clicks the "One-key Movie" control 207, the video editing application can be controlled to automatically edit the materials selected by the user. In the example as Figure 1b shown in (3) of the present embodiment, the user selects three video materials for video editing. In response to the user's operation, the video editing application in the mobile phone automatically edits the three video materials selected by the user, and generates a clipped video C corresponding to these video materials, and the interface 208 shown in Figure 1b (4) of the present embodiment is displayed. In the interface 208, the clipped video C is displayed so that the user can know the automatic editing result of the selected video materials. When the user needs to store the current clipped video in the gallery application, the user can click the confirmation control 209 to store the current clipped video in the gallery application. Regarding other controls in the interface 208, reference can be made to the explanation of Figure 1a (3) of the present embodiment, which will not be elaborated herein.
[0095] In the above-mentioned scenario of automatic video editing, when the video editing application automatically edits based on one or more video materials selected by the user, usually the original sound in the video materials is removed and the background music (BGM) is globally substituted. Refer to Figure 2As shown, it is assumed that the clipped video obtained by one-key automatic clipping is composed of 5 high-light video segments spliced together. Usually, the original sounds of these 5 high-light video segments will be turned off, and instead, the background music of the automatic clipping template will be used. In this way, in the video clip obtained by one-key automatic clipping, the wonderful sounds in the video material (i.e., the original video) will be lost, making the audio part of the video clip templatized. The audio parts of multiple video clips obtained by the video clipping application automatically clipping different video materials based on the same clipping template are the same, that is, the background music of the video clips is the same, making the video clips obtained by one-key automatic clipping give users a feeling of sameness, resulting in a poor user experience of the "one-key automatic clipping" function.
[0096] An embodiment of the present application provides a video clipping method. In this method, when the electronic device automatically clips the video material selected by the user, it can intelligently retain the wonderful original sounds of the video material, so that the original video sound and the background music coexist in the clipped video clip, and the sound elements of the user input material are more prominent, improving the audio effect of the one-key clipped video clip.
[0097] Figure 3 The structural schematic diagram of the electronic device 100 is shown. It should be understood that Figure 3 The shown electronic device 100 is only an example of an electronic device, and the electronic device 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. Figure 3 The various components shown in can be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.
[0098] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0099] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0100] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0101] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0102] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0103] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface to implement the function of answering calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0104] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0105] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. While the charging management module 140 charges the battery 142, it can also supply power to the electronic device through the power management module 141. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc.
[0106] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc. In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies.
[0107] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel.
[0108] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc. The ISP is used to process the data fed back by the camera 193. The camera 193 is used to capture static images or videos.
[0109] The video codec is used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0110] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0111] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a Universal Flash Storage (UFS), etc.
[0112] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. For example, music playback, recording, etc.
[0113] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0114] The pressure sensor is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor can be disposed on the display screen 194. The touch sensor, also known as the "touch panel". The touch sensor can be disposed on the display screen 194, and the touch sensor and the display screen 194 form a touch screen, also known as the "touch screen".
[0115] The buttons 190 include a power button, a volume button, etc. The motor 191 can generate a vibration prompt. The indicator 192 can be an indicator light, which can be used to indicate the charging status, power change, messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect a SIM card.
[0116] The software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. The embodiment of the present application takes the Android system of the layered architecture as an example to exemplify the software structure of the electronic device 100.
[0117] Figure 4 It is a software structure block diagram of the electronic device 100 according to an embodiment of the present application.
[0118] The layered architecture of the electronic device 100 divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely, the application layer, the application framework layer, the Android runtime (Android Runtime) and the system library, and the kernel layer.
[0119] The application layer can include a series of application packages.
[0120] like Figure 4 As shown, the application package may include a gallery application, a video editing application, etc. Among them, the video editing application may be a system-level application or a third-party application. In addition, the application package may also include camera applications, WLAN, Bluetooth, calls, calendars, maps, navigation, music, videos, short messages, and other applications.
[0121] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0122] like Figure 4 As shown, the application framework layer can include a media processing module (or media processing middle platform), a visual analysis module (or visual analysis middle platform), a media data module (or media data middle platform), and a material shelving module (or material shelving middle platform).
[0123] Among them, the media processing module can be used to perform offline analysis of multimedia materials, which may include but is not limited to material theme analysis, analysis of highlight video clips in video materials, analysis of high-value audio in highlight video clips, etc.
[0124] Exemplarily, the media processing module may call the visual analysis module to perform visual analysis on the multimedia material, obtain the visual analysis result of the material, and then identify the theme type, highlight video segment, etc. of the material according to the visual analysis result of the material. The media processing module may also call other modules (such as the intelligent processing module, etc.) to identify the theme type of the material, which is not limited in this embodiment. Among them, the visual analysis module may perform visual analysis on the multimedia material based on a preset visual analysis algorithm (such as the visual atomization analysis algorithm).
[0125] Also exemplarily, the media processing module may analyze the high-value audio in the highlight video segment according to a preset audio recognition algorithm. For example, an audio recognition module (or an intelligent recognition dynamic library for video original sound) is integrated in the media processing module, which is used to identify audio events in the highlight video segment, analyze audio loudness, audio signal-to-noise ratio, etc. Optionally, the audio recognition module may also be set separately, which is not limited in this embodiment.
[0126] Among them, the media data module may be used for the media processing module to read or cache low-quality (or non-high-definition) materials corresponding to the multimedia original materials, that is, for the media processing module to read or cache analysis information corresponding to the multimedia original materials, so as to improve the processing efficiency of the media processing module.
[0127] The material shelving module may be used for the video editing application to download video editing templates and materials related to video editing, etc.
[0128] Continue to refer to Figure 4 As shown, the application framework layer may further include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.
[0129] Among them, the window manager is used to manage window programs. The window manager may obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0130] The content provider is used to store and obtain data, and make these data accessible to application programs. The data may include videos, images, audio, incoming and outgoing calls, browsing history and bookmarks, phone books, etc.
[0131] The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system may be used to build application programs. The display interface may be composed of one or more views. For example, a display interface including a text message notification icon may include a view for displaying text and a view for displaying pictures.
[0132] The telephone manager is used to provide the communication function of the electronic device. For example, the management of call status (including connection, hanging up, etc.).
[0133] The resource manager provides various resources for the application, such as localized strings, icons, pictures, layout files, video files, and so on.
[0134] The notification manager enables the application to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction.
[0135] The Android Runtime includes the core libraries and the virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.
[0136] The core libraries consist of two parts: one is the functional functions that the Java language needs to call, and the other is the core libraries of Android.
[0137] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0138] The system libraries can include multiple functional modules. For example: the surface manager, the media libraries, the 3D graphics processing library (such as OpenGL ES), the 2D graphics engine (such as SGL), etc. Among them, the surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications. The media libraries support the playback and recording of various common audio and video formats, as well as static image files, etc. The media libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The 2D graphics engine is the drawing engine for 2D drawing.
[0139] The kernel layer is the layer between the hardware and the software. The kernel layer at least includes the display driver, the audio driver, the Wi-Fi driver, the sensor driver, etc. Among them, the hardware at least includes the processor, the display screen, the Wi-Fi module, the sensor, etc.
[0140] It can be understood that Figure 4 The layers in the shown software structure and the components included in each layer do not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may include more or fewer layers than shown, and each layer may include more or fewer components, or combine certain components, or split certain components, or have different component arrangements, which are not limited in the present application.
[0141] It is understandable that in order for an electronic device to implement the video editing method in the embodiments of the present application, it includes the corresponding hardware and / or software modules for performing various functions. Combining the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the manner of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.
[0142] As Figure 5 shown is the interaction schematic diagram of each module. Referring to Figure 5 , the process of the video editing method provided by the embodiments of the present application specifically includes:
[0143] S301, in response to a user's one-key editing operation, the gallery application inputs the materials selected by the user into the video editing application.
[0144] Exemplarily, the user's one-key editing operation can be the operation of the user clicking on the "AI One-key Blockbuster" control as shown in (2) of Figure 1a , or the operation of the user clicking on the "One-key Movie" control as shown in (3) of Figure 1b .
[0145] In one implementation, S301 can also be adjusted to: in response to a user's one-key editing operation, the video editing application obtains the materials selected by the user in the gallery application.
[0146] Among them, the materials selected by the user, which can also be referred to as the materials to be edited, may include picture materials and / or video materials selected in the gallery application.
[0147] S302, the video editing application performs an offline analysis on the materials selected by the user.
[0148] In the embodiments of the present application, when the video editing application performs an offline analysis on the materials selected by the user, it can not only analyze the material theme type and high-light video segments, but also analyze the high-value audio in the high-light video segments. Among them, when the video editing application performs an offline analysis on the materials, the original sound intelligent recognition can be integrated into the analysis process of the high-light segments of the video materials.
[0149] Continuing to refer to Figure 5 shown, the video editing application performing an offline analysis on the materials selected by the user can specifically be:
[0150] S3021, the video editing application calls the media processing module to perform an offline analysis on the materials selected by the user.
[0151] In this embodiment, the offline analysis of the material can be implemented relying on the media processing module (or the media processing middleware) in the application framework layer. Optionally, the media processing module provides a call interface, and the video editing application calls the media processing module through this call interface to perform theme type recognition, highlight video segment recognition, etc. on the material selected by the user.
[0152] As an alternative embodiment, the call parameters of the media processing module may further include audio recognition analysis parameters (or intelligent original sound retention analysis parameters). Among them, the parameter value of the audio recognition analysis parameter is used to indicate whether audio recognition analysis of the highlight video segment is required. For example, when the parameter value of the audio recognition analysis parameter is the target value, it is used to indicate that audio recognition analysis of the highlight video segment is required. Optionally, when the video editing application calls the media processing module, the parameter value of the audio recognition analysis parameter is always set to the target value to enable audio recognition analysis of the highlight video segment. Optionally, when the video editing application calls the media processing module, the parameter value of the audio recognition analysis parameter is set to the target value only when the material to be edited includes video material, to enable audio recognition analysis of the highlight video segment.
[0153] S3022, under the call of the video editing application, the material analysis scheduling decision module in the media processing module calls the visual analysis module to perform visual analysis on the material selected by the user, and obtains the material visual analysis result.
[0154] The visual analysis module provides a call interface, and the media processing module can call the visual analysis module through this call interface to perform visual analysis on the material selected by the user, and obtain the material visual analysis result.
[0155] Under the call of the material analysis scheduling decision module, the visual analysis module can perform visual analysis on the material selected by the user based on a preset visual atomization algorithm, obtain the material visual analysis result, and return the obtained material visual analysis result to the material analysis scheduling decision module.
[0156] Exemplarily, the visual analysis of the material by the visual analysis module may include but is not limited to image classification, object detection, image segmentation, pose estimation, video stream action understanding (such as behavior recognition, temporal action detection, spatio-temporal action detection, etc.), OCR (Optical Character Recognition), etc.
[0157] S3023, the material analysis scheduling decision module in the media processing module identifies the theme type of the material and the highlight video segments in the video material.
[0158] Exemplarily, the theme types may include, but are not limited to, childlike fun, people, food, sports, travel, ancient architecture, night scenes, nature, leisure, dynamic rhythm, lightheartedness, slight sadness, etc.
[0159] The highlight video segment can be understood as the wonderful video segment in the video material, which can refer to the video segment containing wonderful frames (video frames corresponding to wonderful moments). Exemplarily, taking the video material with a birthday theme as an example, the highlight video segment may include the "blowing out candles" video segment, the "cutting the cake" video segment, the "singing the birthday song" video segment, etc. Another exemplarily, taking the video material with a football-playing theme as an example, the highlight video segment may include the "scoring a goal" video segment, the "passing the ball" video segment, the "catching the ball" video segment, etc.
[0160] In an alternative embodiment, the material analysis scheduling decision module may identify the theme type of the material and the highlight video segment in the video material according to the visual analysis result of the material.
[0161] In another alternative embodiment, the material analysis scheduling decision module may identify the highlight video segment in the video material according to the visual analysis result of the material. Among them, a theme recognition module may be integrated in the media processing module for identifying the theme type corresponding to the material. Exemplarily, the material analysis scheduling decision module calls the theme recognition module to identify the theme type corresponding to the material. If the material to be edited is an image material, the material analysis scheduling decision module may input the image material into the theme recognition model to identify the theme of the material; if the material to be edited is a video material, the material analysis scheduling decision module may first perform frame extraction on the video material and input one or more extracted video frames into the theme recognition model to identify the theme of the video material; when comprehensively identifying the themes of multiple materials to be edited, the material analysis scheduling decision module may jointly input each image material included in the materials to be edited and the frame extraction results corresponding to each video material into the theme recognition model to identify the themes of the materials to be edited. Among them, regarding the method of frame extraction for video materials, it can be determined according to actual needs and specific scenarios, and the embodiments of the present application do not limit this. For example, the frame extraction method may be to uniformly extract frames from the video material at a certain step size.
[0162] S3024, the material analysis scheduling decision module in the media processing module calls the audio recognition module to analyze the audio in the highlight video segment to obtain an audio recognition result.
[0163] In the embodiments of the present application, the audio recognition module provides a call interface, and the material analysis scheduling decision module may call the audio recognition module to identify and analyze the audio in the highlight video segment according to this call interface to obtain an audio recognition result.
[0164] When the parameter value of the audio recognition analysis parameter indicates that audio recognition analysis of the highlight video segment is required, the material analysis scheduling decision module continues to call the audio recognition module to analyze the audio in the recognized highlight video segment to obtain the audio recognition result. Optionally, the material analysis scheduling decision module may determine whether to call the audio recognition module for audio analysis according to the parameter value of the audio recognition analysis parameter and whether a highlight video segment is recognized. Among them, when the parameter value of the audio recognition analysis parameter is the target value and the material analysis scheduling decision module determines at least one highlight video segment, the material analysis scheduling decision module continues to call the audio recognition module to perform audio recognition analysis on each highlight video segment respectively to determine whether there is video original sound that needs to be retained in each highlight video segment.
[0165] Exemplarily, the audio recognition module may call the original sound intelligent recognition dynamic library to analyze the audio information of the highlight video segment and feedback the audio recognition result to the material analysis scheduling decision module. Optionally, for each highlight video segment, the audio recognition module recognizes the audio events included therein and calculates the signal quality information of each audio event. The signal quality information may include, for example, audio signal loudness and audio signal signal-to-noise ratio and other information.
[0166] In the embodiments of the present application, the audio event types corresponding to high-value audio may include human voice types and non-human voice types. Among them, the human voice types may include, but are not limited to, speech (such as dialogue, monologue, etc.), emotional human voices (such as laughter, crying, shouting angrily, etc.), singing, etc.; the non-human voice types may include, but are not limited to, applause, animal (such as cats, dogs, birds, etc.) calls, environmental sounds (such as ocean wave sounds, wind sounds, fireworks sounds, etc.).
[0167] Given that the duration of the highlight video segment is uncertain, it may be long or short, and the duration of the audio events recognized by the audio recognition module in the highlight video segment is also uncertain, it may be long or short. Among them, the duration of the audio event may be shorter than the total duration of the highlight video segment, or may be equal to the total duration of the highlight video segment.
[0168] S3025, the material analysis scheduling decision module in the media processing module recognizes the high-value audio events included in the highlight video segment according to the audio recognition result, and returns the material analysis result to the video editing application.
[0169] Among them, the material analysis result may include, but is not limited to, the video theme type, the information of the highlight video segment, and the high-value audio events included in the highlight video segment, etc.
[0170] In the embodiments of the present application, for each highlight video clip, the material analysis scheduling decision module may determine whether each highlight video clip includes high-value audio events and the high-value audio events included in each highlight video clip according to the audio events included therein and the signal quality information of each audio event. It can be understood that high-value audio events refer to audio events with better audio signal quality in the highlight video clip, for example, audio events whose audio signal quality meets preset conditions.
[0171] Optionally, the material analysis scheduling decision module may screen out high-value audio events from the various audio events of the highlight video clip according to information such as the signal loudness and signal-to-noise ratio of the audio events. Exemplarily, if the signal loudness and signal-to-noise ratio of a certain audio event meet the high-value audio screening conditions, then this audio event can be retained as a high-value audio event (or high-value video original sound).
[0172] For example, for the highlight video clip x, the audio recognition module recognizes that audio event 1 and audio event 2 are included in this highlight video clip. The material analysis scheduling decision module filters and screens out high-value audio events from audio event 1 and audio event 2 according to the signal loudness and signal-to-noise ratio of audio event 1 and audio event 2. If the audio signal of audio event 1 meets the high-value audio screening conditions, for example, the audio signal loudness value of audio event 1 is greater than or equal to the audio signal loudness threshold, and the audio signal-to-noise ratio of audio event 1 is greater than or equal to the audio signal-to-noise ratio threshold, then audio event 1 is the high-value audio event obtained by filtering and screening. If the audio signal of audio event 2 does not meet the high-value audio screening conditions, for example, the audio signal loudness value of audio event 2 is less than the audio signal loudness threshold, or the audio signal-to-noise ratio of audio event 2 is less than the audio signal-to-noise ratio threshold, then audio event 2 does not belong to high-value audio events. Thus, the high-value audio event corresponding to the highlight video clip x is audio event 1, and then the material analysis scheduling decision module may generate a material analysis result according to the information of audio event 1.
[0173] In the embodiments of the present application, the high-value audio events screened out by the material analysis scheduling decision module from multiple audio events are the video original sounds that need to be retained. It can be understood that given the variable duration of the highlight video clip, which may be long or short, the duration of the high-value audio events recognized by the audio recognition module in the highlight video clip is also not fixed, and may be long or short. Among them, the duration of the high-value audio events may be shorter than the total duration of the highlight video clip, or may be equal to the total duration of the highlight video clip. When the duration of the high-value audio event is equal to the total duration of the highlight video clip, it indicates that all the audio (or all the original sounds) of this highlight video clip are high-value audio that needs to be retained.
[0174] S303. The video editing application selects a video editing template and downloads the template and related editing materials in the material shelving module.
[0175] In an alternative implementation, the video editing application selects a template as the video automatic editing template according to the material theme type. Exemplarily, the video editing application can select multiple templates to be matched in the template library according to the material theme type, calculate the matching scores respectively according to the matching degree between the material content and the templates to be matched, and use the template with the highest matching score as the video automatic editing template. Also exemplarily, after the video editing application selects multiple templates to be matched in the template library according to the material theme type, it can calculate the matching scores of each template to be matched respectively according to information such as the number of materials (such as the sum of the number of highlight video clips and image materials) and / or the duration of the highlight video clips, and use the template with the highest matching score as the video automatic editing template.
[0176] After determining the video automatic editing template, the video editing application can download the corresponding template in the material shelving module, and download the editing materials related to this template, etc. It should be noted that the editing materials related to the template at least include background music.
[0177] S304. The video editing application performs video editing and splicing operations according to the video automatic editing template.
[0178] When the materials selected by the user only include video materials, the video editing application performs editing and splicing operations on the video materials according to the video automatic editing template; when the materials selected by the user include video materials and picture materials, the video editing application performs editing and splicing operations on the video materials and picture materials according to the video automatic editing template.
[0179] The following takes the video editing application performing editing and splicing on video materials according to the video automatic editing template as an example for explanation. Since in the material offline analysis stage, the material analysis result not only includes highlight video clips, but also includes high-value audio clips in the highlight video clips, in the video editing and splicing stage, the video editing application needs to perform editing and splicing on the highlight video clips, background music, and high-value audio clips (i.e., video original sound) according to the editing template. That is, in terms of video, the video editing application needs to splice the highlight video clips; in terms of audio, the video editing application needs to splice the background music and the high-value video original sound.
[0180] When the video editing application splices audio, during the period of playing the original video sound, the high-value original video sound can coexist with the background music. That is, when the high-value original video sound is played, the background music automatically reduces its volume to give way to the original video sound. The high-value original video sound can also exist alone. That is, when the high-value original video sound is played, the background music automatically stops playing. This embodiment does not make any limitations in this regard.
[0181] The following takes "when the high-value audio segment is played, the background music automatically reduces its volume to give way to the original video sound" as an example to explain the audio splicing.
[0182] Figure 6 An exemplary audio splicing example during video automatic editing is shown. As Figure 6 shown, the edited video (or the edited video finished product) is spliced by 5 high-light video segments. Part of the high-value audio in the high-light video segment 2 needs to be retained, and all the audio in the high-light video segment 4 is high-value audio that needs to be retained. Continuing to refer to Figure 6 , in the edited video finished product, the audio part played in the t1 period is the background music, the audio part played in the t2 period is the weakened background music and the retained original video sound, the audio part played in the t3 period is the background music, the audio part played in the t4 period is the weakened background music and the retained original video sound, and the audio part played in the t5 period is the background music. In the t2 period and the t4 period, the audio part of the video finished product is no longer just the background music, but the background music and the high-value original video sound coexist. Among them, when the background music and the high-value original video sound coexist, the background music volume is reduced for playback to give way to the playback of the original video sound, and the sound elements of the materials selected by the user are made more prominent.
[0183] As an optional implementation manner, assume that the set volume of the electronic device (or the normal volume, standard volume, etc. of the edited video) is V. Then, when the audio part only includes the background music, the playback volume of the background music is 100%*V. When the background music and the high-value original video sound coexist, the playback volume of the background music is reduced to N1%*V, where N1 is a positive integer less than 100. Exemplarily, N1 = 50.
[0184] In this way, in the video finished product obtained by one-key automatic editing, the wonderful sounds in the original video will not be lost, and the audio part of the video finished product will not be fixed and templated. As a result, the video finished product will not give users a feeling of sameness, and the user experience of one-key automatic editing is improved.
[0185] In the embodiments of the present application, in order to ensure the effect of retaining the original video sound and avoid the phenomenon of abrupt playback of the original video sound, the appearance and stop of the original video sound can adopt appropriate fade-in and fade-out effects.
[0186] AsFigure 7 As shown, in the edited video, it is assumed that the background music is played before time T1 (the playing volume is 100%*V). Between time T1 and time T4, the background volume is played at a decreasing level (the playing volume is N1%*V) and the original video sound that needs to be retained is played. After time T4, the background music is played (the playing volume is 100%*V). Among them, the original video sound fades in starting from time T1, and the playing volume of the original video sound reaches 100%*V at time T2. The original video sound fades out starting from time T3 until it stops playing at time T4 (such as the playing volume is 0). Among them, the time period from time T1 to time T2 is the fade-in duration of the original video sound, and the time period from time T3 to time T4 is the fade-out duration of the original video sound. During the time period from time T1 to time T2, the playing volume of the original video sound increases from 0%*V to 100%*V according to the set step size 1. During the time period from time T3 to time T4, the playing volume of the original video sound decreases from 100%*V to 0%*V according to the set step size 2. Among them, the set step size 1 and the set step size 2 can be equal or not equal, and this embodiment does not limit this.
[0187] Among them, the fade-in duration of the original video sound and the fade-out duration of the original video sound can be equal or not equal, and this embodiment does not limit this. Exemplarily, the fade-in duration of the original video sound is 1 second, and the fade-out duration of the original video sound is 1 second.
[0188] In the embodiment of the present application, in order to ensure the playing effect of high-value original video sounds, for any segment of high-value original video sounds, the duration of a corresponding segment of retained original video sound in the edited video can be slightly longer. Among them, the starting point of the retained original video sound is earlier than the starting point of the high-value original video sound, and the ending point of the retained original video sound is later than the ending point of the high-value original video sound.
[0189] As an optional implementation manner, when the original video sound fades in, the playing volume at the starting point of the high-value original video sound is greater than or equal to N2%*V. When the original video sound fades out, the playing volume at the ending point of the high-value original video sound is greater than or equal to N3%*V. Both N2 and N3 are positive integers less than 100. Among them, N2 and N3 can be equal or not equal, and this embodiment does not limit this. Exemplarily, N2 = N3 = 75.
[0190] Such as Figure 8As shown in the figure, assume that the playback period of the high-value video original sound is from time T5 (i.e., the starting point of the high-value video original sound) to time T6 (i.e., the ending point of the high-value video original sound). The playback period of the video original sound retained in the edited video can be the period from time T1 to time T4. Time T1 is slightly earlier than time T5, and time T4 is slightly later than time T6. During the period from time T1 to time T4 in the edited video, the background volume with reduced playback and the video original sound to be retained are played. Among them, the video original sound can fade in at time T1, the playback volume of the video original sound reaches 100%*V at time T2, the video original sound starts to fade out at time T3, and stops playing until time T4 (such as the playback volume is 0). The volume of the video original sound at time T1 is 0%*V. During the period from time T1 to time T2, the playback volume of the video original sound increases from 0%*V to 100%*V according to the set step 1, and the volume of the video original sound has increased to more than 75%*V at time T5. The volume of the video original sound at time T3 is 100%*V. During the time period from time T3 to time T4, the playback volume of the video original sound decreases from 100%*V to 0%*V according to the set step 2, and the volume of the video original sound has not decreased to less than 75%*V at time T6.
[0191] As described above, since the duration of the high-light video segment is uncertain, the duration of the high-value video original sound in the high-light video segment is also uncertain. Therefore, the high-value video original sound in a high-light video segment may be part of the video original sound in the high-light video segment, such as Figure 9 the high-value video original sound in the high-light video segment X1 shown in (1) below, and the high-value video original sound in a high-light video segment may also be all of the video original sound in the high-light video segment, such as Figure 9 the high-value video original sound in the high-light video segment X2 shown in (2) below.
[0192] To ensure the fade-in and fade-out effects of the video original sound, and to ensure that the playback volume at the starting point of the high-value video original sound is not lower than N2%*V, and the playback volume at the ending point of the high-value video original sound is not lower than N3%*V, the time difference between the starting point of the video original sound retained in the edited video and the starting point of the high-value video original sound is t_1, and the time difference between the ending point of the video original sound retained in the edited video and the starting point of the high-value video original sound is t_2. Among them, t_1 and t_2 can be equal or not equal, and this embodiment does not make any limitations in this regard.
[0193] In the case shown in Figure 9 below (1), the time difference between the starting point of the high-value video original sound and the starting point of the high-light video segment X1 is greater than or equal to t_1, and the time difference between the ending point of the high-value video original sound and the ending point of the high-light video segment X1 is greater than t_2. At this time, the retained video original sound all belongs to the video original sound of the high-light video segment X1.
[0194] In the situation shown in (1) below Figure 9 the time difference between the starting point of the high-value video original sound and the starting point of the highlight video segment X2 is less than t_1, and the time difference between the ending point of the high-value video original sound and the ending point of the highlight video segment X2 is less than t_2. At this time, the reserved video original sound does not completely belong to the video original sound of the highlight video segment X2. In the reserved video original sound, in addition to the video original sound corresponding to the highlight video segment X2, it also includes the video original sound before the highlight video segment X2 in the video material and the video original sound after the highlight video segment X2 in the video material.
[0195] It can be understood that the reserved video original sound is a continuous video original sound, that is, the video original sound before the highlight video segment X2, the video original sound corresponding to the highlight video segment X2, and the video original sound after the highlight video segment X2 are a continuous video original sound.
[0196] In an alternative embodiment, if the time difference between the starting point of the high-value video original sound and the starting point of the video material is less than t_1, the playback volume of the starting point of the high-value video original sound can also be ensured to be not lower than N2% * V by increasing the set step length 1.
[0197] In an alternative embodiment, if the time difference between the ending point of the high-value video original sound and the ending point of the video material is less than t_2, the playback volume of the ending point of the high-value video original sound can also be ensured to be not lower than N3% * V by delaying the start point of the fade-out of the video original sound and increasing the set step length 2.
[0198] In the embodiments of the present application, in order to ensure the avoidance effect of the background music and prevent the background music from weakening or strengthening abruptly, the background music can be weakened by fading out (i.e., reducing the volume). For example, the playback volume of the background music is reduced from 100% * V to N1% * V according to the set step length 3. Correspondingly, the background music can be increased by fading in (i.e., increasing the volume). For example, the playback volume of the background music is increased from N1% * V to 100% * V according to the set step length 4. Among them, the set step length 3 and the set step length 4 can be equal or not equal, and this embodiment does not limit this.
[0199] Among them, the fade-out duration when the background music weakens and the fade-in duration when the background music strengthens can be equal or not equal, and this embodiment does not limit this. Exemplarily, the fade-out duration when the background music weakens is 1 second, and the fade-in duration when the background music strengthens is 1 second.
[0200] Among them, the fade-out start time when the background music weakens can be the same as the fade-in start time of the video original sound, or different from the fade-in start time of the video original sound. For example, the fade-out start time when the background music weakens can be slightly earlier than the fade-in start time of the video original sound. This embodiment does not limit this. The fade-in start time when the background music strengthens can be the same as the fade-out start time of the video original sound, or different from the fade-out start time of the video original sound. For example, the fade-in start time when the background music strengthens can be slightly later than the fade-out start time of the video original sound. This embodiment does not limit this.
[0201] Among them, the fade-out duration when the background music weakens and the fade-in duration of the video original sound can be equal or unequal. The fade-in duration when the background music strengthens and the fade-out duration of the video original sound can be equal or unequal. This embodiment does not limit this.
[0202] The following takes "the fade-out start time when the background music weakens is the same as the fade-in start time of the video original sound, the fade-out duration when the background music weakens is equal to the fade-in duration of the video original sound, the fade-in start time when the background music strengthens is the same as the fade-out start time of the video original sound, and the fade-in duration when the background music strengthens is equal to the fade-out duration of the video original sound" as an example for explanation.
[0203] As Figure 10 shown, in the edited video, it is assumed that the background music is played before time T7 (the playing volume is 100%*V). Between time T7 and time T10, the weakened background volume (the playing volume is N1%*V) and the video original sound to be retained are played. After time T10, the background music is played (the playing volume is 100%*V). Among them, the video original sound starts to fade in at time T7, and the playing volume of the video original sound reaches 100%*V at time T8. The background music weakens in a fade-out form at time T7, and the playing volume of the background music drops to N1%*V at time T8. The video original sound starts to fade out at time T9, and the video original sound stops playing at time T10 (such as the playing volume is 0). The background music strengthens in a fade-in form at time T9, and the playing volume of the background music rises to 100%*V at time T10. Among them, the time period from time T7 to time T8 is the fade-in duration of the video original sound and also the fade-out duration when the background music weakens. The time period from time T9 to time T10 is the fade-out duration of the video original sound and also the fade-in duration when the background music strengthens. In the time period from time T7 to time T8, the playing volume of the video original sound rises from 0%*V to 100%*V according to the set step 1, and the playing volume of the background music drops from 100%*V to N1%*V (such as 50%*V) according to the set step 3. In the time period from time T9 to time T10, the playing volume of the video original sound drops from 100%*V to 0%*V according to the set step 2, and the playing volume of the background music rises from N1%*V to 100%*V according to the set step 4.
[0204] The following takes "the fade-out start time of the background music is earlier than the fade-in start time of the video original sound, and the fade-in start time of the background music is later than the fade-out start time of the video original sound" as an example for explanation. Exemplarily, the fade-out end time of the background music when it weakens is the fade-in start time of the video original sound, and the fade-in start time of the background music when it strengthens is the fade-out end time of the video original sound.
[0205] As Figure 11 shown, in the edited video, assume that the background music is played before time T11 (the playing volume is 100% * V). At time T11, the background music weakens in a fade-out form, that is, in the time period from T11 to T12, the playing volume of the background music gradually weakens, and at time T12, the playing volume of the background music drops to N1% * V (such as 50% * V). The video original sound starts to fade in at time T12, and the playing volume of the video original sound reaches 100% * V at time T13. Among them, the time period from T11 to T12 is the fade-out duration when the background music weakens, and the playing volume of the background music drops from 100% * V to N1% * V (such as 50% * V) according to the set step size 3; the time period from T12 to T13 is the fade-in duration of the video original sound, and the playing volume of the video original sound rises from 0% * V to 100% * V according to the set step size 1. Continuing to refer to Figure 11 , the video original sound starts to fade out at time T14, and the video original sound stops playing at time T15 (such as the playing volume is 0). At time T15, the background music strengthens in a fade-in form, that is, in the time period from T15 to T16, the playing volume of the background music gradually strengthens, and at time T16, the playing volume of the background music rises to 100% * V. Among them, the time period from T14 to T15 is the fade-out duration of the video original sound, and the playing volume of the video original sound drops from 100% * V to 0% * V according to the set step size 2; the time period from T15 to T16 is the fade-in duration when the background music strengthens, and the playing volume of the background music rises from N1% * V (such as 50% * V) to 100% * V according to the set step size 4.
[0206] In the scenario of retaining the video original sound mentioned above, the effect of retaining the video original sound can be achieved by dynamically adjusting the volume of the video original sound. For example, in the segments where the video original sound does not need to be retained, the playing volume of the video original sound is set to 0% * V, and in the segments where the video original sound needs to be retained, the playing volume of the video original sound is set to 100%. Among them, when the video original sound starts to play in a fade-in form, the playing volume of the video original sound rises from 0% * V to 100% * V according to the set step size 1, and when the video original sound stops playing in a fade-out form, the playing volume of the video original sound drops from 100% * V to 0% * V according to the set step size 2.
[0207] In a possible implementation, it is also possible to extract the corresponding high-value audio from the video material and add the corresponding high-value audio according to the playing moment, so as to achieve the effect of retaining the original video sound. Optionally, each added high-value original video sound starts playing in a fade-in form and stops playing in a fade-out form. Among them, when playing or about to play each high-value original video sound, the playing volume of the background music is reduced to achieve avoidance of the original video sound. Optionally, when the background music starts to avoid, the playing volume of the background music can be reduced in a fade-out form, and when the background music stops avoiding, the playing volume of the background music can be restored in a fade-in form.
[0208] S305, the video editing application synthesizes and edits the video into a finished product.
[0209] The video editing application sets the relevant visual effects of the edited video and synthesizes the relevant special effects of the edited video according to the currently determined video automatic editing template, such as setting the transition method between high-light video segments, setting the filter method of the edited video into a finished product, adding special effects, adding subtitles, etc., to obtain the final edited video into a finished product.
[0210] It should be noted that during the process of the video editing application synthesizing and editing the video into a finished product, if some synthesis effects (such as frame interpolation processing, audio rhythm point analysis processing, etc.) need to be implemented relying on specific algorithms, the media processing module in the application framework layer can be called to implement them.
[0211] Among them, S304 and S305 can also be integrated into one operation step to complete, and the embodiment does not limit the division of the video editing and synthesis steps.
[0212] So far, the video editing application has completed the automatic editing operation on the user-selected material and obtained the edited video into a finished product. At this time, the user can view the effect of the video automatic editing to see if it meets the user's needs. Among them, the user can adjust the automatic editing template so that the video editing application can complete the automatic editing operation on the material again based on the user's re-selected editing template.
[0213] Figure 12a and Figure 12b Exemplarily shows an application scenario. As Figure 12a shown in (1) below, after the user selects multiple video materials in the gallery application, clicks the "One-key Movie" control 401. In response to the user operation, the video editing application calls the media processing module to perform offline analysis on the multiple materials selected by the user, analyzes and obtains the material theme type, high-light video segments, and high-value audio information contained in the high-light video segments, and then selects a matching editing template for video editing and effect synthesis operations to obtain the video editing into a finished product, as Figure 12a shown in (2) below. In Figure 12aIn the interface shown in (2), the user can view the finished video clip. Continuing to refer to Figure 12a As shown in (2), the interface also includes a video sound setting control 402. When the user clicks on the video sound setting control 402, in response to this user operation, the mobile phone can display an interface as shown in (3) of 12a. In the interface shown in (3) of 12a, a video sound setting prompt box 403 is displayed, and video sound setting options are shown in the video sound setting prompt box 403, such as a "Smart Retain" option 4031, an "All Retain" option 4032, and a "Only Add Background Music" option 4033. Among them, the "Smart Retain" option 4031 is the default selected option, which is used to indicate that the finished video clip is generated based on the method of smartly retaining the original video sound, that is, smartly retaining the high-value original video sound in the finished video clip; the "All Retain" option 4032 is used to indicate that the finished video clip is generated based on the method of retaining all the original sounds of the video segments, that is, retaining all the original sounds of each highlight video segment in the finished video clip; the "Only Add Background Music" option 4033 is used to indicate that the finished video clip is generated based on the method of only adding background music, that is, not retaining any original video sound in the finished video clip.
[0214] Since the "Smart Retain" option 4031 is the default selected option for video sound setting, therefore Figure 12a the finished video clip shown in the interface shown in (2) is generated by editing based on the method of smartly retaining the original video sound. That is, Figure 12a in the finished video clip shown in the interface shown in (2), the background music and the high-value original video sound coexist, which can more prominently display the sound elements of the user-input materials. If the user needs to retain all the original sounds in the video segment, the user can select the "All Retain" option 4032 in the interface as shown in Figure 12a (3), and then click the "OK" option 4034; if the user needs to only add background music and not retain any original video sound, the user can select the "Only Add Background Music" option 4033 in the interface as shown in Figure 12a (3), and then click the "OK" option 4034.
[0215] Taking the example that the user needs to retain all the original video sounds, the user selects the "All Retain" option 4032 in the interface as shown in Figure 12a (3), and then clicks the "OK" option 4034. In response to the user operation, the video clip application adjusts the sound in the finished video clip of the edited video, that is, retains all the original video sounds of each highlight video segment on the basis of the background music, and displays an interface as shown in Figure 12b (1). In Figure 12bIn the video clip shown in (1), not only the background music is played globally, but also the original video sound is played globally. At this time, if the user clicks on the video sound setting control 402, in response to the user's operation, the mobile phone can display the interface as shown in Figure 12b (2). In this interface, the "Keep All" option 4032 is selected, indicating that the current video clip is generated based on the method of keeping all the original sounds of the video segments. If the user only needs to add background music in the video clip, after the user selects the "Add Background Music Only" option 4033 and then clicks the "OK" option 4034, the user can view the effect of the video clip after the video sound is adjusted.
[0216] The video clip method provided by the embodiments of the present application supports the user to flexibly select between the "Intelligent Retention" option, the "Keep All" option, and the "Add Background Music Only" option for the video sound setting (or the setting related to the original video sound in video clip). After the user clicks the "One - key Movie" option or the "One - key Clip" option, the video clip application generates a default video clip based on the default sound setting option. On this basis, if the user adjusts the video sound setting option, the video clip application dynamically adjusts the audio information in the video clip to make the adjusted video clip meet the user's video clip requirements. Among them, when the video clip application generates the video clip for the first time, it can obtain the intelligently retained original video sound and all the original video sounds. Therefore, when the user flexibly selects between the "Intelligent Retention" option, the "Keep All" option, and the "Add Background Music Only" option, the video clip application dynamically adjusts the audio information in the video clip without bringing an obvious delay feeling to the user. Since the analysis process of the high - value original sounds in the highlight video segments is integrated into the analysis process of the highlight video segments, even if the default sound setting option in the video clip application is the "Keep All" option or the "Add Background Music Only" option, the video clip application can adjust the video sound with low latency when the user adjusts the video sound setting.
[0217] In Figure 12a and Figure 12b the shown scenarios, the video sound setting (such as "Intelligent Retention" or "Keep All" or "Add Background Music Only") takes effect globally in the video clip.
[0218] As an alternative embodiment, the video sound settings (such as "intelligent retention", "all retention", or "only add background music") can also take effect partially (i.e., not globally) in the edited video clip, which means that the user can make separate settings for different parts of the edited video clip (such as the opening part, the ending part, or a certain highlight video segment). For example, for the opening part of the edited video clip, the user sets the video sound to "all retention"; for the ending part of the edited video clip, the user sets the video sound to "only add background music"; for a certain highlight video segment in the edited video clip, the user sets the video sound to "intelligent retention". In this way, the user can set the sound of the edited video clip according to actual needs, which is highly personalized. Moreover, since the analysis process of the high-value original sound in the highlight video segment is integrated into the analysis process of the highlight video segment, the video editing application can also adjust the sound in the edited video clip with low latency for the flexible adjustment of the sound of individual segments in the video by the user, thus enhancing the user experience.
[0219] It should be noted that for the edited video clip generated based on the one-key automatic editing method provided in the embodiments of the present application, the user can also flexibly adjust its opening, ending, transitions, etc., which will not be elaborated here.
[0220] This embodiment also provides a computer storage medium, in which computer instructions are stored. When the computer instructions run on an electronic device, the electronic device is enabled to execute the above-related method steps to implement the video editing method in the above embodiments.
[0221] This embodiment also provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above-related steps to implement the video editing method in the above embodiments.
[0222] In addition, the embodiments of the present application also provide a device, which may specifically be a chip, a component, or a module. The device may include a processor and a memory connected to each other. The memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory to enable the chip to execute the video editing method in each of the above method embodiments.
[0223] Among them, the electronic device (such as a mobile phone, etc.), the computer storage medium, the computer program product, or the chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, which will not be elaborated here.
[0224] From the description of the above embodiments, those skilled in the art can understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0225] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0226] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.
Claims
1. A video editing method, characterized in that, Including: Display a first interface; The multimedia materials selected by the user are displayed in the first interface; the multimedia materials include at least one video material; In response to a first operation, display a second interface; the first clipped video corresponding to the multimedia material is included in the second interface; Wherein, in a first period of the first clipped video, the audio of the first clipped video includes background music corresponding to a target clip template and does not include the video original sound in the video material; in a second period of the first clipped video, the audio of the first clipped video includes the first video original sound in the at least one video material.
2. The method according to claim 1, characterized in that, In the first period of the first clipped video, the playing volume of the background music is a first volume; In the second period of the first clipped video, the audio of the first clipped video further includes the background music, and the playing volume of the background music is a second volume, and the second volume is less than the first volume.
3. The method according to claim 2, characterized in that, In the second period of the first clipped video, the playing volume of the first video original sound is the first volume.
4. The method according to claim 2 or 3, characterized in that, It further includes: At the first moment of the first clipped video, the first video original sound starts to fade in; At the second moment of the first clipped video, the first video original sound starts to fade out.
5. The method according to claim 4, wherein It further includes: When the first video original sound fades in, the playing volume of the first video original sound rises from zero to the first volume according to a first step length; When the first video original sound fades out, the playing volume of the first video original sound decreases from the first volume to zero according to a second step length.
6. The method according to claim 5, wherein The first video original sound includes a second video original sound, the starting point of the first video original sound is earlier than the starting point of the second video original sound, and the ending point of the first video original sound is later than the ending point of the second video original sound; The ratio of the playing volume at the starting point of the second video original sound to the first volume is greater than a first value, and the ratio of the playing volume at the ending point of the second video original sound to the first volume is greater than a second value.
7. The method according to claim 4, wherein It further includes: When the first video original sound starts to fade in, the playing volume of the background music decreases from the first volume to the second volume according to a third step length; When the first video original sound starts to fade out, the playing volume of the background music increases from the second volume to the first volume according to a fourth step length.
8. The method according to claim 1, characterized in that, After displaying the second interface, it further includes: In response to a second operation, display a third interface; Wherein, a plurality of video sound setting options are included in the third interface, and in the plurality of video sound setting options, a first option is in a selected state, and the first option is used to indicate that the video original sound is intelligently retained in the clipped video.
9. The method according to claim 8, wherein After displaying the third interface, it further includes: In response to a third operation, display a fourth interface; in the fourth interface, the first option is in a non-selected state, and a second option is in a selected state, and the second option is used to indicate that all the video original sounds are retained in the clipped video; In response to a fourth operation, display a fifth interface; the fifth interface includes a second clipped video corresponding to the multimedia material; at any time period of the second clipped video, the audio of the second clipped video includes the background music and the video original sound; Alternatively, after displaying the third interface, it further includes: In response to a fifth operation, display a sixth interface; in the sixth interface, the first option is in an unselected state and the third option is in a selected state, and the third option is used to indicate that only the background music is added to the clipped video; In response to a sixth operation, display a seventh interface; the seventh interface includes a third clipped video corresponding to the multimedia material; at any time period of the third clipped video, the audio of the second clipped video only includes the background music.
10. The method according to claim 1, wherein The step of displaying the second interface in response to the first operation includes: In response to the first operation, analyze the multimedia material to determine the theme type corresponding to the multimedia material and at least one highlight video segment; Analyze the audio events in the at least one highlight video segment to determine at least one segment of first video original sound; Automatically clip the multimedia material according to the target clip template matching the theme type, the at least one highlight video segment, and the at least one segment of first video original sound to generate a first clipped video; Display the second interface.
11. The method according to claim 10, wherein Analyzing the audio events in the at least one highlight video segment to determine at least one segment of first video original sound includes: Identifying the audio events in the at least one highlight video segment; Filter the audio events according to the audio filtering conditions to determine the at least one segment of first video original sound.
12. The method according to claim 11, wherein The audio filtering conditions include at least one of the following: an audio signal loudness filtering condition, an audio signal signal-to-noise ratio filtering condition.
13. The method according to claim 11, wherein The types of the audio events include a human voice type and a non-human voice type; wherein, The human voice type includes at least one of the following: speech, emotional human voice, singing; The non-human voice type includes at least one of the following: applause, animal calls, environmental sounds.
14. The method according to claim 1, characterized in that, The step of displaying the first interface includes: Display the interface of the gallery application; In response to a seventh operation, display the first interface; play a video material in the first interface; Alternatively, the step of displaying the first interface includes: Display the interface of the gallery application; In response to an eighth operation, display the first interface; the first interface includes thumbnails of a plurality of multimedia materials, and one or more of the thumbnails are in a selected state.
15. An electronic device, characterized in that, It includes: One or more processors; A memory; And one or more computer programs, wherein the one or more computer programs are stored on the memory, and when the computer programs are executed by the one or more processors, the electronic device executes the video clipping method according to any one of claims 1-14.
16. A computer-readable storage medium, comprising a computer program, characterized in that, When the computer program runs on the electronic device, the electronic device executes the video clipping method according to any one of claims 1-14.
Citation Information
Patent Citations
Audio processing method and device, electronic equipment and storage medium
CN112637632A
Video shooting method and electronic equipment
CN112887584A
Video playing method, terminal and storage medium
CN113766331A
Song concatenation transition method, terminal and storage medium
CN115762581A
Video score processing method, electronic equipment and computer readable storage medium
CN117119266A