Video editing method, electronic equipment and storage medium
By automatically generating ending and opening copy that fits with the material content in electronic devices, the problem of templated editing content is solved, the personalized effect of video editing is improved, and the user experience is enhanced.
Patent Information
- Application Number
- CN202311785828.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-12-22
AI Technical Summary
In the prior art, when electronic devices automatically edit multimedia materials based on the "one-click editing" function, it is easy to cause the editing content to be templated and cannot meet the personalized needs of users. Especially in the end and opening parts of the video, there is a lack of copy description that fits the content of the material.
When electronic devices edit videos, they automatically generate a matching copy based on the content of the end credits and display it at the end credits. At the same time, they generate a copy based on the content of the opening credits, adjust the size of the material screen to display the copy, and enhance the personalized effect of the video.
Improve the effect of automatically editing videos into endings and openings, avoiding templates, and improving user experience.
Smart Images

Figure CN120238619A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent terminals, and particularly to a video editing method, an electronic device, and a storage medium. Background Art
[0002] With the advent of the self-media era, video editing has become more and more popular, and more and more ordinary people without editing skills also need to edit their multimedia materials. Thus, the "one-key editing" function has emerged, and users can complete the editing of multimedia materials through one-key operation.
[0003] However, when the electronic device automatically edits the multimedia material based on the "one-key editing" function, it may cause the edited content to be templated and unable to meet the user's usage requirements for the "one-key editing" function. Summary of the Invention
[0004] This application provides a video editing method, an electronic device, and a storage medium. In this method, when the electronic device automatically edits the multimedia material selected by the user, it generates an end-credit copy that fits the end-credit material, thereby overcoming the defect of templated end-credit content in the edited video.
[0005] In a first aspect, an embodiment of this application provides a video editing method. The method includes: the electronic device displays a first interface; the first interface displays the multimedia material selected by the user; the electronic device responds to a first operation and displays a second interface; the second interface includes a first edited video corresponding to the multimedia material; when playing the end-credit video segment of the first edited video, at least one first video frame is displayed, and the first video frame includes the playing picture of the end-credit material and a first copy related to the content of the end-credit material.
[0006] Exemplarily, the first interface includes a one-key editing control, and the first operation may be an operation on the one-key editing control, such as a click operation.
[0007] Exemplarily, the multimedia material may include video material and / or picture material, etc. Among them, the last material selected by the user in sequence is the end-credit material.
[0008] In the second interface, the first edited video can be played to enable the user to understand the details of the first edited video.
[0009] In this way, when the electronic device automatically edits the multimedia material selected by the user, it can automatically generate an end-credit copy according to the content of the end-credit material, and display the end-credit copy and the end-credit material together in the end-credit picture. In this way, the defect of templated end-credit in the edited video can be overcome, and the end-credit effect of the video edited automatically is improved.
[0010] According to the first aspect, in the first video frame, the size of the playback screen of the end-credit material is the target size, and the target size is smaller than the size of the screen of the first clipped video. Correspondingly, when the electronic device plays the end-credit video segment of the first clipped video and before the electronic device displays at least one first video frame, the method further includes: the size of the playback screen of the end-credit material gradually shrinks from the size of the screen of the first clipped video to the target size.
[0011] Among them, the size of the playback screen of the end-credit material gradually shrinks from the size of the screen of the first clipped video to the target size in equal proportion. Equal proportion can be understood as the aspect ratio remains unchanged.
[0012] During a certain period of the end-credit video segment of the first clipped video, the size of the playback screen of the end-credit material gradually shrinks from the size of the screen of the first clipped video to the target size, leaving a display area for the first piece of text.
[0013] According to the first aspect, or any one of the implementation manners of the above first aspect, the aspect ratio direction of the first clipped video is horizontal; in the first video frame, the playback screen of the end-credit material and the first piece of text are arranged side by side.
[0014] Exemplarily, when the aspect ratio of the first clipped video is 16:9, the aspect ratio direction of the first clipped video is horizontal.
[0015] According to the first aspect, or any one of the implementation manners of the above first aspect, the size of the playback screen of the end-credit material gradually shrinks from the size of the screen of the first clipped video to the target size, including:
[0016] The size of the playback screen of the end-credit material gradually shrinks from the size of the playback screen of the first clipped video to the left to the first size; the first size is smaller than the size of the screen of the first clipped video; in the first video frame, the first piece of text is displayed on the right side of the playback screen of the end-credit material.
[0017] According to the first aspect, or any one of the implementation manners of the above first aspect, the aspect ratio direction of the first clipped video is vertical; in the first video frame, the playback screen of the end-credit material and the first piece of text are arranged one above the other.
[0018] Exemplarily, when the aspect ratio of the first clipped video is 4:3, the aspect ratio direction of the first clipped video is vertical.
[0019] According to the first aspect, or any one of the implementation manners of the above first aspect, the size of the playback screen of the end-credit material gradually shrinks from the size of the screen of the first clipped video to the target size, including:
[0020] The size of the playback screen of the end-credit material gradually shrinks downward from the size of the playback screen of the first clipped video to a second size; the second size is smaller than the screen size of the first clipped video; in the first video frame, the first copywriting is displayed above the playback screen of the end-credit material.
[0021] According to the first aspect, or any one of the implementation manners of the above first aspect, after the size of the playback screen of the end-credit material gradually shrinks from the screen size of the first clipped video to the target size, the method further includes:
[0022] In the first video frame, the first copywriting is gradually displayed according to a first animation effect.
[0023] Among them, the first animation effect can be, for example, word-by-word display.
[0024] According to the first aspect, or any one of the implementation manners of the above first aspect, before the electronic device displays the second interface, it further includes: the electronic device determines the end-credit material; the electronic device generates multiple pieces of end-credit alternative copywriting according to the end-credit material; the electronic device screens out the first copywriting from the multiple pieces of end-credit alternative copywriting.
[0025] According to the first aspect, or any one of the implementation manners of the above first aspect, the electronic device generates multiple pieces of end-credit alternative copywriting according to the end-credit material, including:
[0026] When the end-credit material is a picture material, the electronic device performs text generation from pictures on the picture material to obtain multiple pieces of end-credit alternative copywriting; when the end-credit material is a video material, the electronic device performs frame extraction on the video material to obtain a second video frame, and performs text generation from pictures on the second video frame to obtain multiple pieces of end-credit alternative copywriting.
[0027] According to the first aspect, or any one of the implementation manners of the above first aspect, the electronic device screens out the first copywriting from the multiple pieces of end-credit alternative copywriting, including:
[0028] The electronic device screens out the first copywriting from the multiple pieces of end-credit alternative copywriting according to a copywriting screening strategy;
[0029] Among them, the copywriting screening strategy includes at least one of the following: text character count limit condition, punctuation mark limit condition, duplicate removal processing, text style limit condition; the customized style corresponding to the style limit condition supports user settings.
[0030] According to the first aspect, or any one of the implementation manners of the above first aspect, after the electronic device displays the second interface, it further includes:
[0031] The electronic device displays a third interface in response to a second operation; the second interface includes a second clipped video corresponding to the multimedia material; the clip template of the second clipped video has a different frame orientation from the clip template of the first clipped video;
[0032] In the first clipped video, the playback screen of the end-credit material and the first piece of text are arranged in a first direction, and in the second clipped video, the playback screen of the end-credit material and the first piece of text are arranged in a second direction; wherein, the first direction is the left-right direction or the up-down direction, and the second direction corresponds to the up-down direction or the left-right direction.
[0033] Exemplarily, the second operation can be, for example, that after the user clicks on the template control 107, the user reselects a video-clipping template to trigger the electronic device to re-clip the video into a finished product.
[0034] In this way, if the aspect ratio direction of the automatically clipped template reselected by the user changes, then in the video clip formed by the video-clipping application after regeneration, the layout of the reduced material image in the material end-credit image and the end-credit text is adjusted accordingly.
[0035] According to the first aspect, or any implementation manner of the above first aspect, the method further includes: when the electronic device plays the opening video segment of the first clipped video, at least one third video frame is displayed, and a second piece of text related to the opening material is included in the third video frame.
[0036] In this way, when the electronic device automatically clips the multimedia material selected by the user, it can automatically generate an opening text based on the content of the opening material, and display the opening text in the opening screen, thereby overcoming the defect of template-based content in the opening of the video clip formed by automatic clipping, and improving the opening effect of the video clip obtained by automatic clipping.
[0037] According to the first aspect, or any implementation manner of the above first aspect, in the third video frame, the second piece of text is gradually displayed according to a second animation effect.
[0038] Wherein, the second animation effect can be, for example, word-by-word display.
[0039] According to the first aspect, or any implementation manner of the above first aspect, the method further includes:
[0040] When the electronic device plays the target video segment of the first clipped video, the electronic device displays at least one fourth video frame, and a target text corresponding to the target video segment and a target audio corresponding to the target text are included in the fourth video frame.
[0041] When the electronic device automatically clips the multimedia material selected by the user, it can generate a paragraph text (which can be abbreviated as caption) corresponding to at least one video segment, and add the paragraph text and the audio corresponding to the paragraph text to the video clip formed by the video clipping, thereby improving the effect of the video clip formed by the one-key clipping function and improving the user experience.
[0042] According to the first aspect, or any implementation of the above first aspect, before the electronic device displays the second interface, it further includes: the electronic device analyzes the constraint conditions of the editing template of the first clipped video; the electronic device generates a target text according to the constraint conditions and the video clips; the electronic device generates a target audio according to the target text and determines the insertion positions of the target text and the target audio. The insertion positions of the target text and the target audio are used for editing and generating the first clipped video.
[0043] Exemplarily, the constraint conditions of the editing template may include, but are not limited to, the word count constraint of the paragraph text corresponding to each video clip, the duration constraint of the voiceover, etc.
[0044] According to the first aspect, or any implementation of the above first aspect, for the electronic device to generate a target audio according to the target text, it includes: the electronic device selects a target voice; the electronic device generates a target audio according to the target voice.
[0045] Among them, when the electronic device generates an audio that matches the paragraph text according to the target audio voice, the audio duration needs to meet the constraint conditions of the video editing template.
[0046] According to the first aspect, or any implementation of the above first aspect, for the electronic device to display the first interface, it includes: the electronic device displays the interface of the gallery application; the electronic device displays the first interface in response to a seventh operation; a video material is played in the first interface.
[0047] As Figure 1a shown in the scenario, the first interface may, for example, be the interface shown in Figure 1a (2) therein, and the first operation may be an operation of clicking on the "AI One-Click Blockbuster" control 104.
[0048] According to the first aspect, or any implementation of the above first aspect, for the electronic device to display the first interface, it includes: the electronic device displays the interface of the gallery application; the electronic device displays the first interface in response to an eighth operation; the first interface includes thumbnails of multiple multimedia materials, and one or more thumbnails are in a selected state.
[0049] As Figure 1b shown in the scenario, the first interface may, for example, be the interface shown in Figure 1b (3) therein, and the first operation may be an operation of clicking on the "One-Click Movie" control 207.
[0050] In a second aspect, embodiments of the present application provide an electronic device. The electronic device includes: one or more processors; a memory; and one or more computer programs, where the one or more computer programs are stored on the memory, and when the computer programs are executed by the one or more processors, the electronic device is caused to execute the video clip method according to the first aspect and any one of the first aspect.
[0051] The second aspect and any implementation manner of the second aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the second aspect and any implementation manner of the second aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0052] In a third aspect, embodiments of the present application provide a computer-readable storage medium. The computer-readable storage medium includes a computer program, and when the computer program runs on an electronic device, the electronic device is caused to execute the video clip method according to the first aspect and any one of the first aspect.
[0053] The third aspect and any implementation manner of the third aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the third aspect and any implementation manner of the third aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0054] In a fourth aspect, embodiments of the present application provide a computer program product, including a computer program, and when the computer program is run, the computer is caused to execute the video clip method according to the first aspect or any one of the first aspect.
[0055] The fourth aspect and any implementation manner of the fourth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the fourth aspect and any implementation manner of the fourth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated here.
[0056] In a fifth aspect, the present application provides a chip, which includes a processing circuit and transceiver pins. Among them, the transceiver pins and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the video clip method according to the first aspect or any one of the first aspect to control the receive pin to receive a signal and control the transmit pin to transmit a signal.
[0057] The fifth aspect and any implementation manner of the fifth aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the fifth aspect and any implementation manner of the fifth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, which will not be elaborated herein. Description of the Drawings
[0058] Figure 1a Schematic diagram of an application scenario of one-key automatic video editing shown by way of example;
[0059] Figure 1b Schematic diagram of an application scenario of one-key automatic video editing shown by way of example;
[0060] Figure 2 Schematic diagram of the end of a video automatically edited into a finished piece shown by way of example;
[0061] Figure 3 Schematic diagram of the hardware structure of an electronic device shown by way of example;
[0062] Figure 4 Schematic diagram of the software structure of an electronic device shown by way of example;
[0063] Figure 5 Schematic diagram of the module interaction of a video editing method shown by way of example;
[0064] Figure 6 Schematic diagram of the end of a video automatically edited into a finished piece shown by way of example;
[0065] Figure 7a Schematic diagram of the end animation of a horizontal format in a video automatically edited into a finished piece shown by way of example;
[0066] Figure 7b Schematic diagram of the end animation of a vertical format in a video automatically edited into a finished piece shown by way of example;
[0067] Figure 8a Schematic diagram of the start screen of a video automatically edited into a finished piece shown by way of example;
[0068] Figure 8b Schematic diagram of the start screen of a video automatically edited into a finished piece shown by way of example;
[0069] Figure 9 Schematic diagram of the module interaction of a video editing method shown by way of example;
[0070] Figure 10 Schematic diagram of the captioning and voiceover of video clips in an automatically edited finished piece shown by way of example. Detailed Description of the Invention
[0071] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0072] The term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0073] The terms "first", "second", etc. in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, rather than to describe the specific order of the target objects.
[0074] In the embodiments of the present application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0075] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of systems refers to two or more systems.
[0076] With the advent of the self-media era, video editing has become more and more popular, and more and more ordinary people without editing skills also need to edit their multimedia materials (including but not limited to video materials, picture materials, etc.). Therefore, the "one-key editing" function (or the "one-key video creation" function, the "one-key blockbuster" function, etc.) has emerged. Users only need to perform a one-key operation (such as clicking the "one-key editing" control) on an electronic device (hereinafter, the electronic device is taken as an example of a mobile phone for explanation), and then the automatic editing of the multimedia materials can be completed to obtain an automatically edited video (or an automatically created video, etc.).
[0077] In an application scenario, a user can install a third-party video editing application on a mobile phone, import one or more multimedia materials into the third-party video editing application, and use the "one-key editing" function of the third-party video editing application to complete the automatic editing of the multimedia materials.
[0078] In another application scenario, users can also use the "one - click editing" function provided by system - level applications in the mobile phone to complete the automatic editing of multimedia materials. In this scenario, users do not need to install additional third - party video editing applications in the mobile phone, nor do they need to perform the import operation of multimedia materials when they need to automatically edit multimedia materials. Among them, Figure 1a and Figure 1b respectively exemplarily show the application scenarios of automatically editing videos based on the "one - click editing" function provided by system - level applications in the mobile phone.
[0079] Figure 1a In (1) and Figure 1b In (1) exemplarily shows the display interface of the gallery application. Exemplarily, in response to the user's operation of clicking on the gallery application icon in the mobile phone, the mobile phone starts the gallery application and displays the gallery application interface 101 as shown in Figure 1a In (1) or the gallery application interface 201 as shown in Figure 1b In (1). In the gallery application interface 101, various pictures (or photos) and videos stored in the gallery application are displayed. Among them, for video editing, the pictures and videos stored in the gallery application can all be used as multimedia materials (or materials to be edited, hereinafter simply referred to as materials) for video editing.
[0080] In one implementation, users can select a video material for automatic video editing. Exemplarily, if the duration of a certain video material is greater than a preset threshold (for example, 15s), then the video material meets the condition for separate editing, and the video editing application can complete the automatic editing operation based on this video material to obtain a clipped video finished product.
[0081] Continuing to refer to Figure 1a As shown in (1), in the gallery application interface 101, the user clicks on the thumbnail 102 corresponding to the video material A. In response to the user's operation, the mobile phone plays the video material A, as shown in Figure 1a In (2). Since the duration of the video material A is greater than the preset threshold and meets the condition for separate editing, in the playback interface 103 of the video material A, there can also be a display of the "AI one - click blockbuster" control 104. Continuing to refer to Figure 1a As shown in (2), when the user clicks on the "AI one - click blockbuster" control 104, it can control the video editing application to perform automatic editing on the video material A. In response to the user's operation, the video editing application in the mobile phone performs automatic editing on the video material A and generates a clipped video B corresponding to the video material A, as shown in Figure 1aThe interface 105 shown in (3). Optionally, before the mobile phone displays the interface 105, a transition interface for prompting the user that a clipped video B is being generated can also be displayed, such as a transition interface for indicating "material analysis in progress" or "video clipping in progress". In the interface 105, the clipped video B is displayed so that the user can learn the automatic clipping result of the video material A. In the interface 105, confirmation controls 106, template controls 107, segment controls 109, editing controls 110, etc. can also be displayed. When the user needs to store the current clipped video in the gallery application, the user can click the confirmation control 106 to store the current clipped video in the gallery application. When the user needs to adjust the clipping template of the clipped video, the user can click the template control 107 to make the mobile phone display a video clipping template selection interface. Among them, in the video clipping template selection interface, the user can reselect an automatic clipping template for the video material A. When the user needs to adjust the background music of the clipped video, the user can click the music control 108 to make the mobile phone display a video background music editing interface. In the video background music editing interface, the user can edit the background music of the clipped video, such as adjusting the background music volume or replacing the background music. When the user needs to adjust some segments of the clipped video, the user can click the segment control 109 to make the mobile phone display a video segment editing interface. Among them, in the video segment editing interface, the user can edit the video segments in the clipped video, such as adjusting the video segments or replacing the video segments. When the user needs to edit the clipped video, the user can click the editing control 110 to make the mobile phone display a video editing interface. In the video editing interface, the user can edit the clipped video, such as adding stickers or adjusting the duration.
[0082] In another embodiment, the user can select multiple multimedia materials for automatic video clipping. For example, the user can select multiple picture materials for automatic video clipping, or select multiple video materials for automatic video clipping, or select video materials and picture materials for automatic video clipping.
[0083] Exemplarily, continue to refer to Figure 1b as shown in (1) above, in the gallery application interface 201, the user long-presses the thumbnail corresponding to any one material, such as clicks the thumbnail 202 corresponding to the picture material a. In response to the user operation, the mobile phone displays as Figure 1b the material selection interface 203 shown in (2) above. Optionally, the user can also trigger the mobile phone to display as Figure 1b the material selection interface 203 shown in (2) above by clicking the selection control, or trigger the mobile phone to display as Figure 1bThe material selection interface 203 shown in (2) of this embodiment is not limited herein. Exemplarily, in the material selection interface 203, a selectable identifier 204 is displayed on the thumbnail of each material (including picture materials and video materials). When the user selects a certain material, a selected identifier 205 (such as the numerical identifier shown in the figure) is displayed in the selectable identifier 204 on the thumbnail of this material, which can be referred to Figure 1b as shown in (3) of this embodiment. Optionally, the selected identifier 205 can also be a tick identifier (such as "√"), which is not limited in this embodiment. It should be noted that when the user selects multiple materials, the sorting of these materials in the video clip is determined according to the order of user selection. Taking Figure 1b the selected identifier 205 shown in (3) of this embodiment as an example, the value corresponding to the selected identifier 205 is the order of the material being selected. When the materials selected by the user meet the conditions for automatic video editing, for example, when the user selects one or more picture materials, or when the user selects one or more video materials, optionally, when the user selects a video material and this video material meets the conditions for separate editing, the mobile phone pops up a prompt message box 206 to prompt the user that a movie-like video short can be automatically generated based on the selected materials. Among them, a "One-key Movie" control 207 is displayed in the prompt message box 206. Continuing to refer to Figure 1b as shown in (3) of this embodiment, when the user clicks the "One-key Movie" control 207, the video editing application automatically edits the materials selected by the user. In the example shown in Figure 1b as shown in (3) of this embodiment, the user selects three video materials and one picture material for video editing. In response to the user's operation, the video editing application in the mobile phone automatically edits the four selected materials by the user to generate a clipped video C corresponding to these video materials, and the interface 208 shown in Figure 1b as shown in (4) of this embodiment is displayed. In the interface 208, the clipped video C is displayed so that the user can know the automatic editing result of the selected video materials. When the user needs to store the current clipped video in the gallery application, the user can click the confirmation control 209 to store the current clipped video in the gallery application. Regarding other controls in the interface 208, reference can be made to the explanation of Figure 1a as shown in (3) of this embodiment, which will not be elaborated herein.
[0084] In the above-mentioned scenario of automatic video editing, the end of the video clip is usually a common end screen, for example Figure 2 the end screen shown in (1) of this embodiment, the main body of the end screen is the text "THE END" and the text "Thank you for watching", and for another example Figure 2The end - screen image shown in (2). The main subject of the end - screen image is the "LOGO" pattern and the text "Please pay attention", etc. For the end - screens of such edited videos, the main subjects are similar, and the differences may only lie in different text fonts, different "LOGO" patterns, etc., resulting in a poor user experience.
[0085] To improve the user experience of the one - key automatic editing function, an embodiment of this application provides a video editing method. In this method, when the electronic device automatically edits the multimedia materials selected by the user, it can generate text (or copywriting) according to the content of the end - screen materials, and generate the end - part of the edited video according to the text, so that the end - part of the edited video contains a text description that fits the material content, enhancing the difference of the end - part of the edited video and avoiding giving the user a feeling of sameness for the end - parts of the automatically edited video clips, thereby improving the user experience of the one - key automatic editing function.
[0086] In this method, when the electronic device automatically edits the multimedia materials selected by the user, it can also generate text according to the content of the start - screen materials, and generate the start - part of the edited video according to the text, so that the start - part of the edited video contains a text description that fits the material content, thereby enhancing the start - screen effect of the automatically edited video.
[0087] Optionally, in this method, when the electronic device automatically edits the multimedia materials selected by the user, it can also generate captions and voiceovers according to at least one video clip, and add them to the matching positions in the edited video. Among them, the video clip can be generated from video materials or from picture materials, and this embodiment does not limit this.
[0088] Figure 3 Shows a schematic structural diagram of the electronic device 100. It should be understood that Figure 3 The shown electronic device 100 is only an example of an electronic device, and the electronic device 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. Figure 3 The various components shown can be implemented in hardware, software, or a combination of hardware and software, including one or more signal - processing and / or application - specific integrated circuits.
[0089] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0090] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0091] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0092] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0093] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0094] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 may be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface to implement the function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0095] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only for illustrative purposes and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods in the above embodiments.
[0096] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. While the charging management module 140 charges the battery 142, it can also supply power to the electronic device through the power management module 141. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc.
[0097] The wireless communication function of the electronic device 100 can be implemented through Antenna 1, Antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc. In some embodiments, Antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and Antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies.
[0098] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel.
[0099] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc. The ISP is used to process the data fed back by the camera 193. The camera 193 is used to capture static images or videos.
[0100] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0101] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0102] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0103] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc.
[0104] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0105] The pressure sensor is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, the pressure sensor can be disposed on the display screen 194. The touch sensor, also known as the "touch panel". The touch sensor can be disposed on the display screen 194, and the touch sensor and the display screen 194 form a touch screen, also known as the "touch screen".
[0106] The keys 190 include a power-on key, a volume key, etc. The motor 191 can generate a vibration prompt. The indicator 192 can be an indicator light, which can be used to indicate the charging state, the change in battery power, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect the SIM card.
[0107] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 100.
[0108] Figure 4 It is the software structure block diagram of the electronic device 100 in the embodiments of the present application.
[0109] The layered architecture of the electronic device 100 divides software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android Runtime and system libraries, and the kernel layer.
[0110] The application layer may include a series of application packages.
[0111] As Figure 4 shown, the application packages may include a gallery application, a video editing application, etc. Among them, the video editing application may be a system-level application or a third-party application. In addition, the application packages may also include applications such as a camera application, WLAN, Bluetooth, call, calendar, map, navigation, music, video, and short message.
[0112] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0113] As Figure 4 shown, the application framework layer may include a media processing module (or media processing center), an intelligent processing module (or intelligent processing center), a visual analysis module (or visual analysis center), and a material shelving module (or material shelving center).
[0114] Among them, the media processing module can be used for offline analysis of multimedia materials, and may include but not limited to material theme analysis, highlight video segment analysis in video materials, generating alternative copywriting for the opening and / or ending materials, generating captions and voiceovers for video segments, etc.
[0115] Exemplarily, the media processing module can call the visual analysis module to perform visual analysis on the multimedia material, obtain the material visual analysis result, and identify and analyze the highlight video segment according to the material visual analysis result. Among them, the visual analysis module can perform visual analysis on the multimedia material based on a preset visual analysis algorithm (such as a visual atomization analysis algorithm).
[0116] Exemplarily, the media processing module can call the intelligent processing module to identify the video theme type. Among them, the intelligent processing module performs semantic understanding on the video frames or pictures extracted by the media processing module, generates a description text, and then identifies the theme type of the video according to the description text.
[0117] Exemplarily, the media processing module may call the intelligent processing module to generate alternative copywriting for the opening material and / or alternative copywriting for the ending material. Among them, the intelligent processing module may perform image-to-text processing on relevant video frames or images to obtain alternative copywriting for the opening material and / or alternative copywriting for the ending material. As the name implies, image-to-text means generating matching text (or copywriting) based on an image.
[0118] Exemplarily, the media processing module may also call the intelligent processing module to generate caption and voiceover information for a video clip and determine the insertion position information of the caption and voiceover information in the edited video. Among them, the intelligent processing module may generate paragraph text that meets template constraint conditions (such as duration, number of words, etc.) based on relevant video frames and / or pictures, and call a TTS (Text To Speech) service (TTS function module) to generate corresponding audio. The tone of the audio (such as male voice or female voice) may be determined according to relevant query information (such as the gender of the owner of the electronic device, etc.).
[0119] In this embodiment, the media processing module may read or cache low-quality (or non-high-definition) materials corresponding to the multimedia original materials in the media data module (or media data center), and may also read or cache analysis information corresponding to the multimedia original materials in the media data module to improve the processing efficiency of the media processing module.
[0120] The material shelving module can be used for a video editing application to download video editing templates and materials related to video editing, etc.
[0121] Continue to refer to Figure 4 As shown, the application framework layer may also include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.
[0122] Among them, the window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0123] The content provider is used to store and obtain data and make this data accessible to application programs. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, phone books, etc.
[0124] The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build application programs. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon may include a view for displaying text and a view for displaying pictures.
[0125] The phone manager is used to provide the communication function of the electronic device. For example, the management of call status (including answering, hanging up, etc.).
[0126] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on.
[0127] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction.
[0128] The Android Runtime includes the core libraries and the virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.
[0129] The core libraries contain two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.
[0130] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0131] The system libraries can include multiple functional modules. For example: the surface manager, the media libraries, the 3D graphics processing library (such as: OpenGL ES), the 2D graphics engine (such as: SGL), etc. Among them, the surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications. The media libraries support the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The 2D graphics engine is the drawing engine for 2D drawing.
[0132] The kernel layer is the layer between the hardware and the software. The kernel layer at least includes a display driver, an audio driver, a Wi-Fi driver, a sensor driver, etc. Among them, the hardware at least includes a processor, a display screen, a Wi-Fi module, a sensor, etc.
[0133] It can be understood that Figure 4The layers in the software structure shown and the components included in each layer do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer layers than shown, and each layer may include more or fewer components, or combine certain components, or split certain components, or have different component arrangements, which are not limited in the present application.
[0134] It can be understood that in order for the electronic device to implement the video editing method in the embodiments of the present application, it includes the corresponding hardware and / or software modules for performing various functions. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described function for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.
[0135] Scenario 1
[0136] In this scenario, when the electronic device automatically edits the multimedia materials selected by the user, it can generate text (or copywriting) according to the content of the end material, and / or generate text according to the content of the start material, and accordingly generate the end part and the start part of the edited video, so that the end and start of the edited video contain text descriptions that fit the material content, thereby enriching the content of the end and start of the edited video and avoiding the templatization of the end and start of the edited video.
[0137] Such as Figure 5 shown is the interaction schematic diagram of each module. Referring to Figure 5 , the process of the video editing method provided by the embodiments of the present application specifically includes:
[0138] S301, in response to the user's one-key editing operation, the gallery application inputs the materials selected by the user into the video editing application.
[0139] Exemplarily, the user's one-key editing operation may be the operation of the user clicking the "AI One-key Blockbuster" control as shown in (2) of Figure 1a , or the operation of the user clicking the "One-key Movie" control as shown in (3) of Figure 1b .
[0140] In one implementation, S301 can also be adjusted to: in response to the user's one-key editing operation, the video editing application obtains the materials selected by the user in the gallery application.
[0141] Among them, the materials selected by the user, which can also be called the materials to be edited, may include picture materials and / or video materials selected in the gallery application.
[0142] S302. The video editing application performs offline analysis on the materials selected by the user.
[0143] In the embodiment of the present application, when the video editing application performs offline analysis on the materials selected by the user, it can not only analyze the material theme type and highlight video segments, but also generate alternative copy corresponding to the end material and / or the start material.
[0144] Among them, the end material can be a video material or a picture material, and the type of the end material is not limited in this embodiment. The start material can be a video material or a picture material, and the type of the start material is not limited in this embodiment.
[0145] When the number of materials selected by the user for video editing is one, then this material selected by the user is both the start material and the end material; when the number of materials selected by the user for video editing is multiple, the first material selected by the user, for example Figure 1b the material with the selection mark 205 being "1" in the interface shown in (3) in the figure, the last material selected by the user, for example Figure 1b the material with the selection mark 205 being "4" in the interface shown in (3) in the figure (assuming the user has selected a total of four materials).
[0146] Continue to refer to Figure 5 As shown, the video editing application performing offline analysis on the materials selected by the user can specifically be:
[0147] S3021. The video editing application calls the media processing module to perform offline analysis on the materials selected by the user.
[0148] In this embodiment, the material offline analysis can be implemented relying on the media processing module (or called the media processing middle platform) in the application framework layer. Optionally, the media processing module provides a call interface, and the video editing application calls the media processing module through this call interface to analyze the materials selected by the user, including but not limited to identifying the video theme type, identifying the highlight video segments, generating alternative copy for the end material and / or the start material, etc.
[0149] In an alternative embodiment, the call parameters of the media processing module may include the end-credit material copy analysis parameter and the start-credit material copy analysis parameter. Among them, the parameter value of the end-credit material copy analysis parameter is used to indicate whether it is necessary to generate alternative copies corresponding to the end-credit material, and the parameter value of the start-credit material copy analysis parameter is used to indicate whether it is necessary to generate alternative copies corresponding to the start-credit material. Exemplarily, the parameter values of the end-credit material copy analysis parameter and the start-credit material copy analysis parameter can be set by default, or can be preset by the user according to usage requirements, or can be flexibly set by the video editing application according to the number of materials selected by the user. For example, when the number of materials selected by the user for video editing is one, the video editing application sets the parameter value of the end-credit material copy analysis parameter to a target value so as to generate alternative copies corresponding to the end-credit material. When the number of materials selected by the user for video editing is multiple, the video editing application sets the parameter value of the end-credit material copy analysis parameter to target value 1 and sets the parameter value of the start-credit material copy analysis parameter to target value 2 so as to generate alternative copies corresponding to the end-credit material and to generate alternative copies corresponding to the start-credit material. Among them, target value 1 and target value 2 can be the same or different, and this embodiment does not limit this. Of course, the video editing application can only set the parameter value of the start-credit material copy analysis parameter to the target value so as to generate alternative copies corresponding to the start-credit material, and this embodiment does not limit this.
[0150] S3022, under the call of the video editing application, the material analysis scheduling decision module in the media processing module calls the visual analysis module to perform visual analysis on the materials selected by the user to obtain the material visual analysis result.
[0151] The visual analysis module provides a call interface through which the media processing module can call the visual analysis module to perform visual analysis on the materials selected by the user to obtain the material visual analysis result.
[0152] Under the call of the material analysis scheduling decision module, the visual analysis module can perform visual analysis on the materials selected by the user based on a preset visual atomization algorithm to obtain the material visual analysis result and return the obtained material visual analysis result to the material analysis scheduling decision module.
[0153] Exemplarily, the visual analysis of the materials by the visual analysis module may include but is not limited to image classification, object detection, image segmentation, pose estimation, video stream action understanding (such as behavior recognition, temporal action detection, spatio-temporal action detection, etc.), OCR (Optical Character Recognition), etc.
[0154] S3023, the material analysis and scheduling decision-making module in the media processing module analyzes the highlight video clips in the video material.
[0155] When the material selected by the user includes video material, the media processing module can analyze the highlight video clips in the video material based on the visual analysis results of the material by the visual analysis module.
[0156] Among them, the highlight video clips can be understood as the wonderful video clips in the video material, which can refer to the video clips containing wonderful frames (video frames corresponding to wonderful moments). Exemplarily, taking a video material with a birthday theme as an example, the highlight video clips can include "blowing out candles" video clip, "cutting the cake" video clip, "singing the birthday song" video clip, etc. Another exemplarily, taking a video material with a football theme as an example, the highlight video clips can include "scoring a goal" video clip, "passing the ball" video clip, "catching the ball" video clip, etc.
[0157] S3024, the material analysis and scheduling decision-making module in the media processing module calls the intelligent processing module to identify the video theme type, and generate alternative copywriting for the opening material and / or alternative copywriting for the ending material.
[0158] Exemplarily, the theme type can include but is not limited to childlike fun, people, food, sports, travel, ancient architecture, night view, nature, leisure, dynamic rhythm, lively and cheerful, a little sad, etc.
[0159] When the material selected by the user (i.e., the material to be edited) includes video material, the media processing module can perform frame extraction on the video material to obtain at least one video frame image. Optionally, the media processing module can perform frame extraction on the highlight video clips to obtain at least one video frame image. Among them, the video frame images obtained by the media processing module through frame extraction can be used for video theme type identification. When the material selected by the user includes picture material, these picture materials can be directly used for video theme type identification.
[0160] Exemplarily, the media processing module inputs at least one video frame image and / or picture material into the intelligent processing module for video theme type recognition. Optionally, the intelligent processing module can perform video theme type recognition on the material to be edited according to a preset theme type recognition model. If the material to be edited is video material, the intelligent processing module inputs at least one video frame image obtained by frame extraction processing by the media processing module into the theme type recognition model to perform theme recognition on the video material; if the material to be edited is picture material, the intelligent processing module inputs the picture material into the theme type recognition model to perform theme recognition on the material; when performing comprehensive theme recognition on multiple materials to be edited, the intelligent processing module can jointly input each picture material included in the materials to be edited and the frame extraction results (i.e., video frame images obtained by frame extraction operations) corresponding to each video material into the theme recognition model to perform theme recognition on the materials to be edited. Among them, regarding the method of frame extraction for video material, it can be determined according to actual needs and specific scenarios, and the embodiments of the present application do not limit this. For example, the frame extraction method can be to uniformly extract frames from the video material at a certain step size.
[0161] Also exemplarily, after the media processing module inputs at least one video frame image and / or picture material into the intelligent processing module, the intelligent processing module first performs semantic understanding processing on at least one video frame image and / or picture material to obtain corresponding description texts, and then performs video theme type recognition according to the description texts. Optionally, the intelligent processing module can perform video theme type recognition on the material to be edited according to a preset theme type recognition model. If the material to be edited is video material, the intelligent processing module performs semantic understanding processing on at least one video frame image obtained by frame extraction processing by the media processing module to obtain description texts corresponding to the at least one video frame image respectively, and then inputs these description texts into the theme type recognition model to perform theme recognition on the video material; if the material to be edited is picture material, the intelligent processing module performs semantic understanding processing on the picture material to obtain a description text corresponding to the picture material, and then inputs the description text into the theme type recognition model to perform theme recognition on the material; when performing comprehensive theme recognition on multiple materials to be edited, the intelligent processing module can perform semantic understanding processing on at least one video frame image and the picture material respectively to obtain description texts corresponding to the at least one video frame image and the picture material respectively, and then jointly input these description texts into the theme type recognition model to perform theme recognition on the materials to be edited.
[0162] Taking the example that the material analysis scheduling decision module in the media processing module calls the intelligent processing module to generate alternative copywriting for the end credits material, the material analysis scheduling decision module first determines whether the last material to be edited selected by the user in sequence is video material or picture material.
[0163] When the last clip material to be selected by the user in sequence is a video material, the material analysis scheduling decision module performs frame extraction on it to obtain a video frame image. Optionally, when the last clip material to be selected by the user in sequence is a video material, the material analysis scheduling decision module performs frame extraction on the highlight video segment corresponding to the video material to obtain a video frame image. The material analysis scheduling decision module inputs the video frame image into the intelligent processing module, and the intelligent processing module performs image-to-text processing on the video frame image to obtain multiple pieces of copy corresponding to the video frame image. These alternative copies are the alternative copies for the end-credit materials. Among them, the copy corresponding to the video frame image can be understood as the copy used to describe the content of the video frame image.
[0164] Exemplarily, the intelligent processing module uses a preset image-to-text neural network model to process the video frame image. For example, the video frame image is input into the image-to-text neural network model, and the image-to-text neural network model outputs multiple alternative copies corresponding to the video frame image. Regarding the neural network or algorithm adopted by the image-to-text neural network model, etc., reference can be made to the existing technology and will not be elaborated here.
[0165] As an alternative implementation manner, the intelligent processing module performs image-to-text processing on the video frame image according to the identified video theme type to obtain multiple pieces of copy corresponding to the video frame image. These alternative copies are the alternative copies for the end-credit materials. Exemplarily, the intelligent processing module uses a preset image-to-text neural network model to process the video frame image. For example, the video frame image and the video theme type are input into the image-to-text neural network model, and the image-to-text neural network model outputs multiple alternative copies corresponding to the video frame image.
[0166] When the last clip material to be selected by the user in sequence is a picture material, the material analysis scheduling decision module inputs the picture material into the intelligent processing module, and the intelligent processing module performs image-to-text processing on the picture material to obtain multiple pieces of copy corresponding to the picture material. These alternative copies are the alternative copies for the end-credit materials. Among them, the copy corresponding to the picture material can be understood as the copy used to describe the content of the picture.
[0167] Exemplarily, the intelligent processing module uses a preset image-to-text neural network model to process the picture material. For example, the picture material is input into the image-to-text neural network model, and the image-to-text neural network model outputs multiple alternative copies corresponding to the picture material.
[0168] Taking the material analysis scheduling decision module in the media processing module calling the intelligent processing module to generate alternative copies for the opening-credit materials as an example, the material analysis scheduling decision module first determines whether the first clip material to be selected by the user in sequence is a video material or a picture material.
[0169] When the first material to be edited selected by the user in sequence is a video material, the material analysis scheduling decision module performs frame extraction processing on it to obtain a video frame image. Optionally, when the first material to be edited selected by the user in sequence is a video material, the material analysis scheduling decision module performs frame extraction processing on the highlight video clip corresponding to the video material to obtain a video frame image. The material analysis scheduling decision module inputs the video frame image into the intelligent processing module, and the intelligent processing module performs image-to-text processing on the video frame image to obtain multiple texts corresponding to the video frame image. These alternative texts are the alternative texts for the title material.
[0170] When the first material to be edited selected by the user in sequence is a picture material, the material analysis, scheduling and decision-making module inputs the picture material into the intelligent processing module. The intelligent processing module performs image-to-text processing on the picture material to obtain multiple texts corresponding to the picture material. These alternative texts are the alternative texts for the opening material.
[0171] As an optional implementation, the material analysis scheduling decision module can call the intelligent processing module to generate the title material alternative copy based on all the materials to be edited. Taking the case where the materials to be edited include video materials and picture materials, the material analysis scheduling decision module extracts frames from the video materials to obtain video frame images corresponding to each video material. The intelligent processing module comprehensively generates the title material alternative copy based on these video frame images and each picture material.
[0172] S3025, the material analysis scheduling decision module in the media processing module returns the material analysis result to the video editing application.
[0173] The material analysis results may include but are not limited to the video theme type, information about highlight video clips, and multiple alternative texts for opening materials and / or multiple alternative texts for ending materials.
[0174] S303, the video editing application selects a video editing template, and downloads the template and editing-related materials in the material shelf module.
[0175] Given that there may be differences in the frame orientation (or frame layout direction) of the materials, for example, the frame orientation of some image materials and / or video materials is vertical (or vertical), and the frame orientation of some image materials and / or video materials is horizontal (or horizontal), the video editing application can refer to the frame orientation of the material when selecting the frame orientation of the video editing template.
[0176] For example, the video editing application separately counts the number of vertical materials and the number of horizontal materials in the materials to be edited. If the number of vertical materials is more than the number of horizontal materials, it selects the video automatic editing template with a vertical aspect ratio. If the number of vertical materials is less than the number of horizontal materials, it selects the video automatic editing template with a horizontal aspect ratio. If the number of vertical materials is equal to the number of horizontal materials, it selects the video automatic editing template according to the aspect ratio direction of the first material selected by the user. If the aspect ratio direction of the first material is horizontal, it selects the video automatic editing template with a horizontal aspect ratio; otherwise, it selects the video automatic editing template with a vertical aspect ratio.
[0177] For another example, the video editing application separately counts the duration ratio of vertical materials and the duration ratio of horizontal materials in the materials to be edited. If the duration ratio of vertical materials is more than the duration ratio of horizontal materials, it selects the video automatic editing template with a vertical aspect ratio. If the duration ratio of vertical materials is less than the duration ratio of horizontal materials, it selects the video automatic editing template with a horizontal aspect ratio. If the duration ratio of vertical materials is equal to the duration ratio of horizontal materials, it selects the video automatic editing template according to the aspect ratio direction of the first material selected by the user. If the aspect ratio direction of the first material is horizontal, it selects the video automatic editing template with a horizontal aspect ratio; otherwise, it selects the video automatic editing template with a vertical aspect ratio.
[0178] After determining the aspect ratio direction of the video automatic editing template required for editing, the video editing application selects a template that matches the aspect ratio direction as the video automatic editing template according to the material theme type. Exemplarily, the video editing application can select multiple templates to be confirmed that match the aspect ratio direction in the template library according to the material theme type, and calculate the matching scores respectively according to the matching degree between the material content and the templates to be confirmed, and use the template with the highest matching score as the video automatic editing template. Another example is that after the video editing application selects multiple templates to be confirmed that match the aspect ratio direction in the template library according to the material theme type, it can calculate the matching scores of each template to be confirmed respectively according to information such as the number of materials (such as the sum of the number of highlight video clips and image materials) and / or the duration of the highlight video clips, and use the template with the highest matching score as the video automatic editing template.
[0179] After selecting the video automatic editing template, the video editing application can download the corresponding template in the material shelving module, as well as download the editing materials related to this template, etc. It should be noted that the editing materials related to the template at least include background music.
[0180] S304, the video editing application selects the end text in the alternative end text of the materials, and / or selects the start text in the alternative start text of the materials.
[0181] When the material analysis result includes alternative captions for the end credits of the material, the video editing application selects the end credit caption from multiple alternative captions for the end credits of the material according to a preset screening strategy; when the material analysis result includes alternative captions for the start credits of the material, the video editing application selects the start credit caption from multiple alternative captions for the start credits of the material according to a preset screening strategy.
[0182] In the embodiments of the present application, the video editing application selects the target captions (i.e., end credit captions and start credit captions) for short video editing from multiple alternative captions (including alternative captions for the end credits of the material and alternative captions for the start credits of the material). Among them, the screening strategies for end credit captions and start credit captions can be the same or different, and this embodiment does not make any limitations.
[0183] Exemplarily, the caption screening strategy may include, but is not limited to: text character count limit conditions, punctuation mark limit conditions, style screening conditions, etc. Among them, in the style screening conditions, the caption styles supported by the video editing application may include, but are not limited to, themes such as literary, aesthetic, lively, cute, seasons, and holidays. For example, when the set style corresponding to the style screening condition is the literary style, the video editing application screens alternative captions that meet the literary style from multiple alternative captions. Regarding the set style corresponding to the style screening condition, the user can flexibly set it according to their own needs, so that the end credit caption and / or start credit caption in the video editing finished product meet the user's psychological expectations.
[0184] Exemplarily, for each alternative caption for the end credits of the material, the video editing application respectively determines whether it meets the caption screening strategy, such as whether the caption character count conforms to the text character count limit condition, whether the punctuation marks in the caption meet the punctuation mark limit condition, and whether the style of the caption conforms to the caption style customized by the video editing application. Furthermore, the video editing application can use any alternative caption for the end credits of the material that meets the caption screening strategy as the final end credit caption.
[0185] Exemplarily, for each alternative caption for the end credits of the material, the video editing application scores it according to the caption screening strategy and the corresponding scoring rules, and uses the alternative caption for the end credits of the material with the highest score as the final end credit caption.
[0186] After determining the end credit caption, the video editing application can also design the font of the end credit caption according to the customized style, so that the video editing application can perform animation rendering according to the corresponding design when synthesizing the video finished product.
[0187] The same applies to the method of selecting the final title copy from multiple alternative title copy texts for the video clip, and this embodiment will not elaborate on this. Before the video clip application filters the title copy and the end credit copy, the video clip application can also perform deduplication on multiple alternative title copy texts and multiple alternative end credit copy texts to avoid duplication of the title copy and the end credit copy. The video clip application can also filter multiple alternative title copy texts and multiple alternative end credit copy texts based on the number of times the copy is selected to avoid the same copy being used multiple times.
[0188] This embodiment does not limit the sequence of S303 and S304.
[0189] S305, the video clip application synthesizes the video clip into a complete video according to the video clip template.
[0190] When the materials selected by the user only include video materials, the video clip application performs editing and splicing operations on the video materials according to the automatic video editing template; when the materials selected by the user only include picture materials, the video clip application performs editing and splicing operations on the picture materials according to the automatic video editing template; when the materials selected by the user include video materials and picture materials, the video clip application performs editing and splicing operations on the video materials and picture materials according to the automatic video editing template.
[0191] In this embodiment, if the video clip application determines the end credit copy, the end credit copy is added to the end part of the video clip. In the end screen of the video clip, the last material (such as a video material or a picture material) selected by the user and the end credit copy are displayed.
[0192] Before the end credit copy is displayed, the screen size of the last material selected by the user gradually shrinks to reserve a display area for the display of the end credit copy in the end screen. Optionally, during the process of the screen size of the material gradually shrinking, the aspect ratio of the screen remains unchanged. Exemplarily, the screen size of the last material selected by the user can be approximately reduced to half of the screen size of the video clip. All the text in the end credit copy can be displayed on the end screen at the same time, or can be gradually displayed on the end screen according to a preset animation effect, and this embodiment does not limit this.
[0193] In addition, it should be noted that during the process of the screen size of the last material selected by the user gradually shrinking, and after the screen size of the last material shrinks to the target size, the playback of the material screen is continuous. Among them, taking the video material as an example, the playback of the material screen is the playback of the video segment. Taking the picture material as an example, the playback of the material screen can be the playback of the animation effect corresponding to the picture material (related to the automatic editing template).
[0194] Figure 6An exemplary end screen in a finished video clip. As Figure 6 shown, in this end screen, it includes the picture 401 of the last material selected by the user, and the end screen copywriting 402 that matches the last material. Among them, the end screen copywriting 402 can be used to describe the last material selected by the user. Continuing to refer to Figure 6 , the size of the picture 401 is reduced from the picture size of the finished video clip. During the continuous reduction of the picture, the aspect ratio of the material picture remains unchanged.
[0195] When the aspect ratio direction of the automatic editing template is horizontal (such as a 16:9 picture), in the end screen of the finished video clip, the reduced material picture and the end screen copywriting are distributed horizontally. For example, the reduced material picture is on the left and the end screen copywriting is on the right (as shown in the example of Figure 6 ), or for another example, the reduced material picture is on the right and the end screen copywriting is on the left. This embodiment does not make any limitations in this regard.
[0196] When the aspect ratio direction of the automatic editing template is vertical (such as a 4:3 picture), in the end screen of the finished video clip, the reduced material picture and the end screen copywriting are distributed vertically. For example, the reduced material picture is on the lower side and the end screen copywriting is on the upper side, or for another example, the reduced material picture is on the upper side and the end screen copywriting is on the lower side. This embodiment does not make any limitations in this regard.
[0197] In the embodiments of the present application, the aspect ratio direction of the automatic editing template has no constraint on generating the end screen copywriting, and only affects the end screen animation playback and the layout method of the elements in the end screen.
[0198] Taking the aspect ratio direction of the finished video clip as horizontal as an example, Figure 7a An exemplary schematic diagram of the end segment of the finished video clip is shown. As Figure 7a shown in (1) in it, the last material starts to be played at the moment t1 of the finished video clip. At this time, the picture size of this material is the same as the picture size of the finished video clip. At the moment t2 of the finished video clip, the playing picture of this material gradually shrinks to the left, from the picture size of the finished video clip to size 1, which can be referred to as shown in (2) in Figure 7a . Continuing to refer to Figure 7a shown in (3) in it, after the picture size of the material shrinks to size 1, the end screen copywriting "All things stand towards the sun" corresponding to this material is displayed on the right side of the material picture. Among them, the end screen copywriting corresponding to this material can complete the display process through an animation effect of gradually showing. This embodiment does not make any limitations on the animation effect of the end file display.
[0199] Taking the aspect ratio direction of the finished video clip as vertical as an example, Figure 7b An exemplary schematic diagram of the end segment of the finished video clip is shown. AsFigure 7b As shown in (1), the last material starts to play at time t3 when the video clip is finalized. At this time, the picture size of this material is the same as that of the finalized video clip. At time t4 when the video clip is finalized, the playing picture of this material gradually shrinks downward, from the picture size of the finalized video clip to size 2, which can be referred to Figure 7b as shown in (2). Continuing to refer to Figure 7b as shown in (3), after the picture size of the material shrinks to size 2, the end text "Midsummer Comes in Due Time" corresponding to this material is displayed on the upper side of the material picture. Among them, the end text corresponding to this material can complete the display process through an animation effect of gradual display. This embodiment does not limit the animation effect of the end file display.
[0200] In this embodiment, if the video editing application determines the start text, the start text is added to the start picture of the finalized video clip. Given that the automatic editing templates are different, the start segment of the finalized video clip can be the pre-video slice or the video segment corresponding to the first material selected by the user. This embodiment does not limit this.
[0201] Among them, the start text can be displayed in the upper half area of the start picture, or in the center area of the start picture, or in the lower half area of the start picture. This embodiment does not limit this. In this embodiment, the display position, display form, display animation effect, etc. of the start text can be determined according to the relevant settings of the automatic editing template. This embodiment does not limit this.
[0202] Taking the start segment of the finalized video clip being the pre-video slice as an example, Figure 8a an exemplary schematic diagram of the start segment of the finalized video clip is shown. As Figure 8a shown in (1), at time t5 when the video clip is finalized, the start text gradually starts to be displayed on the pre-video slice picture until the start text is fully displayed on the pre-video slice picture, as shown in (1) of 8a. Among them, the display animation effect, display position, display form, etc. of the start text are related to the relevant settings of the automatic editing template.
[0203] Taking the start segment of the finalized video clip being the video segment corresponding to the first material selected by the user as an example, Figure 8b an exemplary schematic diagram of the start segment of the finalized video clip is shown. As Figure 8b shown in (1), at time t6 when the video clip is finalized, the start text gradually starts to be displayed on the video segment picture until the start text is fully displayed on the video segment picture, as shown in (1) of 8b. Among them, the display animation effect, display position, display form, etc. of the start text are related to the relevant settings of the automatic editing template.
[0204] For other operations in video clip synthesis (such as splicing, adding audio, rendering, etc.), reference can be made to the existing technology and will not be elaborated here.
[0205] After the video clip application generates a video clip, the electronic device displays the video clip, as shown in Figure 1a in (3) or Figure 1b in (4). If the user needs to adjust the template of the video clip, the user can click on the template option 107 as shown in Figure 1a in (3) and reselect an automatic editing template for the material to be edited according to their own needs. In this way, the video clip application can regenerate the video clip according to the automatic editing template selected by the user.
[0206] Among them, the user's re-selection of the automatic editing template will not affect the title copywriting and / or end copywriting determined by the video editing template. Assume that the aspect ratio direction of the automatic editing template reselected by the user changes. Then, in the video clip regenerated by the video clip application, the layout of the reduced material image and the end copywriting in the end image of the material is adjusted accordingly. For example, before the user reselects the automatic editing template, the video clip application generates a video clip 1 based on the automatic editing template with a horizontal aspect ratio. In the end image of video clip 1, the reduced material image and the end copywriting are distributed horizontally, for example, the reduced material image is on the left and the end copywriting is on the right. Assume that the aspect ratio direction of the automatic editing template reselected by the user is vertical. Then, the video clip application generates a video clip 2 based on the automatic editing template with a vertical aspect ratio. In the end image of video clip 2, the reduced material image and the end copywriting are distributed vertically, for example, the reduced material image is at the bottom and the end copywriting is at the top. It can be understood that the end copywriting in the end image of video clip 1 is the same as the end copywriting in the end image of video clip 2.
[0207] Taking the title material and end material in the video clip as examples above, the video editing template calls the media processing module to generate multiple alternative copywritings for the corresponding materials, and selects one copywriting from these multiple alternative copywritings and adds it to the corresponding segment of the video clip. For the intermediate materials in the video clip, the video editing template also calls the media processing module to generate multiple alternative copywritings for the corresponding materials, and selects one copywriting from these multiple alternative copywritings and adds it to the corresponding video segment of the video clip, which will not be elaborated here.
[0208] Scenario Two
[0209] In this scenario, when the electronic device automatically edits the multimedia materials selected by the user, it can generate paragraph text (which can be simply referred to as caption) corresponding to at least one video segment (a video segment corresponding to a video material or a picture material) according to the video theme type, and add the paragraph text and the audio corresponding to the paragraph text to the video clip to form a video, thereby improving the effect of the video clip formed by the one-key editing function and enhancing the user experience.
[0210] As Figure 9 shown is the interaction schematic diagram of each module. Referring to Figure 9 , the process of the video editing method provided by the embodiment of the present application specifically includes:
[0211] S501, in response to the user's one-key editing operation, the gallery application inputs the materials selected by the user into the video editing application.
[0212] S502, the video editing application performs offline analysis on the materials selected by the user.
[0213] Continuing to refer to Figure 9 shown, the video editing application performing offline analysis on the materials selected by the user can specifically be:
[0214] S5021, the video editing application calls the media processing module to perform offline analysis on the materials selected by the user.
[0215] In this embodiment, the offline analysis of the materials can rely on the media processing module in the application framework layer (or
[0216] S5022, under the call of the video editing application, the material analysis scheduling decision module in the media processing module calls the visual analysis module to perform visual analysis on the materials selected by the user to obtain the material visual analysis result.
[0217] S5023, the material analysis scheduling decision module in the media processing module analyzes the highlight video segments in the video material.
[0218] S5024, the material analysis scheduling decision module in the media processing module calls the intelligent processing module to identify the video theme type.
[0219] S5025, the intelligent processing module performs semantic understanding on the video frame images and / or pictures input by the media processing module to generate corresponding description texts.
[0220] S5026, the intelligent processing module identifies the video theme type according to the description texts corresponding to the video frame images and / or pictures, and returns the identified theme type to the media processing module.
[0221] S5027, the material analysis and scheduling decision-making module in the media processing module returns the material analysis result to the video editing application.
[0222] Among them, the material analysis result may include but is not limited to the video theme type, information of high-light video clips, etc.
[0223] S503, the video editing application selects a video editing template and downloads the template and related editing materials in the material shelving module.
[0224] S504, the video editing application determines the caption and voiceover for at least one video clip according to the video editing template.
[0225] The caption can be understood as a text paragraph used to describe a video clip, and the voiceover refers to an audio segment that matches the caption.
[0226] Among them, the video editing application can select one of the video clips according to the preset information and determine the caption and voiceover for this video clip, or can determine the caption and voiceover for each video clip respectively. This embodiment does not make a limitation on this.
[0227] Continue to refer to Figure 9 As shown, the video editing application determines the caption and voiceover for at least one video clip according to the video editing template, which can be specifically:
[0228] S5041, the video editing application calls the caption and voiceover module in the media processing module to generate the caption and voiceover for at least one video clip.
[0229] Exemplarily, the caption and voiceover module in the media processing module provides an external call interface, and the call parameters may include but are not limited to video clip information, information of the video editing template selected by the video editing application, etc.
[0230] S5042, the caption and voiceover module in the media processing module parses the constraint conditions of the video editing template and calls the intelligent processing module to generate the caption and voiceover for at least one video clip.
[0231] After the video editing application selects the video editing template, the duration of each video clip in the edited video finished product is fixed. Subsequently, the caption and voiceover module can reverse infer the word limit of the paragraph text corresponding to each video clip, the duration limit of the voiceover, etc. according to the duration of each video clip.
[0232] Exemplarily, the constraint conditions of the video editing template may include but are not limited to the word constraint of the paragraph text corresponding to each video clip, the duration constraint of the voiceover, etc.
[0233] In this embodiment, when the caption and voiceover module calls the intelligent processing module to generate the caption and voiceover for at least one video clip, it can send the constraint conditions of the video editing template to the intelligent processing module, or use the constraint conditions of the video editing template as call parameters to pass to the intelligent processing module.
[0234] Optionally, when the caption and voiceover module calls the intelligent processing module to generate the caption and voiceover for at least one video clip, it can also pass the information of at least one video clip to the intelligent processing module.
[0235] S5043, the intelligent processing module generates the paragraph text for at least one video clip.
[0236] Under the call of the video editing application, the intelligent processing module generates the paragraph text corresponding to at least one video clip with reference to the constraint conditions of the video editing template. For example, the intelligent processing module can call a relevant neural network model to generate the paragraph text corresponding to at least one video clip. Exemplarily, the paragraph text can be the subtitles of the video material, the text described by the user for the video clip, or the voiceover description corresponding to the video clip, etc. All in all, the paragraph text corresponding to a certain video clip is used to describe this video clip, and the content and style of the paragraph text are not limited in this embodiment.
[0237] S5044, the intelligent processing module calls the TTS service to generate the audio corresponding to the paragraph text, and returns the caption and voiceover of at least one video clip to the caption and voiceover module in the media processing module.
[0238] In this embodiment, the intelligent processing module first determines the target audio timbre, and generates the audio matching the paragraph text according to the target audio timbre. Among them, when the intelligent processing module generates the audio matching the paragraph text according to the target audio timbre, the duration of the audio matching the paragraph text needs to meet the constraint conditions of the video editing template. Optionally, the voiceover duration corresponding to a certain video clip can be equal to the duration of the video clip.
[0239] Exemplarily, the intelligent processing module can call other intelligent modules or determine the gender of the owner of the electronic device through relevant big data, and use the audio timbre corresponding to the gender of the owner as the target audio timbre. For example, when the gender of the owner is male, the intelligent processing module uses the male timbre as the target audio timbre. For another example, if the intelligent processing module cannot determine the gender of the owner, the intelligent processing module can default the female timbre as the target audio timbre.
[0240] Another exemplarily, the intelligent processing module can also determine the popular timbres (such as authorized star timbres, funny timbres, etc.) by querying big data, etc., and use any one of them, or a timbre screened according to a preset strategy, as the target audio timbre.
[0241] S5045. The caption and voiceover module in the media processing module returns the caption and voiceover of the video clip and the insertion position to the video editing application.
[0242] The caption and voiceover module determines the caption insertion position (or addition position) and the voiceover insertion position for each video clip, and transmits the caption and voiceover of each video clip to the video editing application.
[0243] Exemplarily, the caption insertion position may only include the caption insertion start position, or may include both the caption insertion start position and the caption insertion end position. This embodiment does not limit this.
[0244] Exemplarily, the voiceover insertion position may only include the voiceover insertion start position, or may include both the voiceover insertion start position and the voiceover insertion end position. This embodiment does not limit this.
[0245] In an alternative embodiment, the caption insertion start position and the voiceover insertion start position of a certain video clip are the same as the start position of this video clip.
[0246] S305. The video editing application synthesizes a video editing finished product according to the video editing template.
[0247] During the process of the video editing application synthesizing the video editing finished product, if the caption and voiceover have been generated for a certain video clip, the corresponding paragraph text is inserted according to the caption insertion position corresponding to this video clip, and the corresponding voiceover audio is inserted according to the voiceover insertion position corresponding to this video clip.
[0248] Optionally, the voiceover audio of the video clip can coexist with the background music of the video editing finished product. Exemplarily, in the video editing finished product, when the voiceover audio is played, the playing volume of the background music can be automatically reduced to highlight the voiceover audio more.
[0249] Figure 10 Exemplarily shows the automatic editing effect of the caption and voiceover for video clip X. As Figure 10 shown, during the playing period of video clip X, the corresponding paragraph caption is automatically added, and the corresponding voiceover audio for the paragraph caption is automatically added. Among them, the voiceover audio of video clip X can coexist with the background music of the video editing finished product.
[0250] Regarding other operations involved in the video editing synthesis process (such as splicing, rendering, etc.), reference can be made to the existing technology and will not be elaborated here. For the parts not explained in detail in this process, reference can be made to the description of Scenario 1 above and will not be elaborated here.
[0251] In the one-key automatic video editing method provided in this application, the solution of automatically adding an end text and / or a start text that fits the video content in the edited video can coexist with the solution of automatically adding captions and voiceovers to video clips described in Scenario 2. This embodiment will not elaborate on this.
[0252] In addition, for the function of automatically adding an end text that fits the video content in the one-key automatic editing scenario, the settings option of the video editing application may include a corresponding switch control 1. The user can set the on / off state of the switch control 1 according to personal needs to turn on or off the function of automatically adding an end text that fits the video content in the one-key automatic editing scenario.
[0253] Similarly, for the function of automatically adding a start text that fits the video content in the one-key automatic editing scenario, the settings option of the video editing application may include a corresponding switch control 2. The user can set the on / off state of the switch control 2 according to personal needs to turn on or off the function of automatically adding a start text that fits the video content in the one-key automatic editing scenario.
[0254] Similarly, for the function of automatically adding captions and voiceovers to video clips in the one-key automatic editing scenario, the settings option of the video editing application may include a corresponding switch control 3. The user can set the on / off state of the switch control 3 according to personal needs to turn on or off the function of automatically adding captions and voiceovers to video clips in the one-key automatic editing scenario.
[0255] This embodiment also provides a computer storage medium. Computer instructions are stored in this computer storage medium. When the computer instructions run on an electronic device, the electronic device is enabled to execute the above relevant method steps to implement the video editing method in the above embodiment.
[0256] This embodiment also provides a computer program product. When this computer program product runs on a computer, the computer is enabled to execute the above relevant steps to implement the video editing method in the above embodiment.
[0257] In addition, an embodiment of this application also provides a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected to each other. Among them, the memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory so that the chip executes the video editing method in each of the above method embodiments.
[0258] Among them, the electronic device (such as a mobile phone, etc.), computer storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.
[0259] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0260] In several embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0261] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A video editing method, characterized in that, including: displaying a first interface; the multimedia material selected by the user is displayed in the first interface; responding to a first operation, displaying a second interface; the second interface includes a first clipped video corresponding to the multimedia material; when playing the end video segment of the first clipped video, displaying at least one first video frame, where the first video frame includes the playing picture of the end material and a first copy related to the content of the end material.
2. The method according to claim 1, wherein The size of the playing picture is the target size, and the target size is smaller than the picture size of the first clipped video; when playing the end video segment of the first clipped video and before displaying the at least one first video frame, the method further includes: the size of the playing picture of the end material gradually shrinks from the picture size of the first clipped video to the target size.
3. The method according to claim 2, characterized in that, The aspect ratio direction of the first clipped video is horizontal; in the first video frame, the playing picture of the end material and the first copy are arranged side by side horizontally.
4. The method according to claim 2, wherein The size of the playing picture of the end material gradually shrinks from the picture size of the first clipped video to the target size, including: the size of the playing picture of the end material gradually shrinks from the playing picture size of the first clipped video to the left to a first size; the first size is smaller than the picture size of the first clipped video; in the first video frame, the first copy is displayed on the right side of the playing picture of the end material.
5. The method according to claim 2, wherein The aspect ratio direction of the first clipped video is vertical; in the first video frame, the playing picture of the end material and the first copy are arranged one above the other vertically.
6. The method according to claim 5, characterized in that, The size of the playing picture of the end material gradually shrinks from the picture size of the first clipped video to the target size, including: the size of the playing picture of the end material gradually shrinks from the playing picture size of the first clipped video downwards to a second size; the second size is smaller than the picture size of the first clipped video; in the first video frame, the first copy is displayed above the playing picture of the end material.
7. The method according to claim 2, wherein after the size of the playing picture of the end material gradually shrinks from the picture size of the first clipped video to the target size, it further includes: in the first video frame, the first copy is gradually displayed according to a first animation effect.
8. The method according to claim 1, characterized in that, before displaying the second interface, it further includes: determining the end material; generating multiple alternative end copies according to the end material; screening out the first copy from the multiple alternative end copies.
9. The method according to claim 8, wherein generating multiple alternative end copies according to the end material, including: when the end material is a picture material, performing text generation from image processing on the picture material to obtain the multiple alternative end copies; when the end material is a video material, performing frame extraction on the video material to obtain a second video frame, and performing text generation from image processing on the second video frame to obtain the multiple alternative end copies.
10. The method according to claim 8, wherein screening out the first copy from the multiple alternative end copies, including: screening out the first copy from the multiple alternative end copies according to a copy screening strategy; Among them, the copywriting screening strategy includes at least one of the following: text character count limit condition, punctuation mark limit condition, deduplication process, text style limit condition; the customized style corresponding to the style limit condition supports user settings.
11. The method according to any one of claims 1 to 10, characterized in that, After displaying the second interface, it further includes: In response to a second operation, a third interface is displayed; the second interface includes a second clipped video corresponding to the multimedia material; the aspect ratio direction of the clipping template of the second clipped video is different from that of the first clipped video; In the first clipped video, the playback screen of the end material and the first copywriting are arranged in a first direction, and in the second clipped video, the playback screen of the end material and the first copywriting are arranged in a second direction; wherein, the first direction is the left-right direction or the up-down direction, and the second direction corresponds to the up-down direction or the left-right direction.
12. The method according to any one of claims 1 to 10, characterized in that, It further includes: When playing the opening video segment of the first clipped video, at least one third video frame is displayed, and in the third video frame, a second copywriting related to the opening material is included.
13. The method according to claim 12, wherein In the third video frame, the second copywriting is gradually displayed according to a second animation effect.
14. The method according to any one of claims 1 to 10, characterized in that, It further includes: When playing a target video segment of the first clipped video, at least one fourth video frame is displayed, and the fourth video frame includes a target text corresponding to the target video segment, and a target audio corresponding to the target text is played.
15. The method according to claim 14, characterized in that Before displaying the second interface, it further includes: Analyzing the constraint conditions of the clipping template of the first clipped video; Generating the target text according to the constraint conditions and the video segment; Generating the target audio according to the target text, and determining the insertion positions of the target text and the target audio.
16. The method according to claim 15, characterized in that, Generating the target audio according to the target text includes: Selecting a target voice; Generating the target audio according to the target voice.
17. An electronic device, characterized in that, It includes: One or more processors; A memory; And one or more computer programs, wherein the one or more computer programs are stored on the memory, and when the computer programs are executed by the one or more processors, the electronic device executes the video clipping method according to any one of claims 1-16.
18. A computer-readable storage medium, comprising a computer program, characterized in that, When the computer program runs on the electronic device, the electronic device executes the video clipping method according to any one of claims 1-16.
Citation Information
Patent Citations
Text generation method and device
CN110851622A
Picture copywriting generation method, apparatus and device, and storage medium
CN112256902A
Playing processing method and device, equipment and storage medium
CN113766295A
Small video automatic production method based on high-definition camera acquisition
CN114897749A
Video generation method and device based on AI and electronic equipment
CN115460459A