Video clip method, electronic device, and storage medium
By automatically generating intro and outro text that fits multimedia materials within electronic devices, the problem of templated editing content is solved, enhancing the personalization of video editing and the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2026-03-27
AI Technical Summary
When existing electronic devices automatically edit multimedia materials, the edited content tends to become templated, failing to meet users' personalized needs, especially in the video's opening and closing credits, where there is a lack of differentiation and user experience.
Electronic devices automatically generate end credits that match the end credits based on the multimedia content selected by the user, and display them gradually during playback to enhance the end credits effect; similarly, they generate opening credits based on the opening credits to enhance the opening credits effect.
By automatically generating opening and closing credits that fit the content, the template-based nature of video editing is overcome, enhancing the personalization and user experience of automatically edited videos.
Smart Images

Figure CN120238619B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent terminals, and in particular to a video editing method, an electronic device, and a storage medium. BACKGROUND
[0002] With the advent of the era of self-media, video editing is becoming more and more popular, and more and more ordinary people without editing skills also need to edit their multimedia materials. As a result, the "one-click editing" function has emerged. Users can complete the editing of multimedia materials through one-click operation.
[0003] However, when the electronic device automatically edits the multimedia materials based on the "one-click editing" function, it may lead to the templateization of the edited content, which cannot meet the user's demand for the "one-click editing" function. SUMMARY
[0004] The present application provides a video editing method, an electronic device, and a storage medium. In the method, when the electronic device automatically edits the multimedia materials selected by the user, a tail-end script that matches the tail-end material is generated, thereby overcoming the defect of templateization of the tail-end content of the edited video.
[0005] In a first aspect, an embodiment of the present application provides a video editing method. The method includes: displaying, by an electronic device, a first interface; the first interface displays multimedia materials selected by a user; in response to a first operation, the electronic device displays a second interface; the second interface includes a first edited video corresponding to the multimedia materials; when playing a tail-end video segment of the first edited video, at least one first video frame is displayed, the first video frame includes a playing picture of a tail-end material and a first script related to the content of the tail-end material.
[0006] For example, the first interface includes a one-click editing control, and the first operation can be an operation on the one-click editing control, such as a click operation.
[0007] For example, the multimedia materials can include video materials and / or picture materials, etc. The last material selected by the user in sequence is the tail-end material.
[0008] In the second interface, the first edited video can be played to enable the user to understand the details of the first edited video.
[0009] In this way, when the electronic device automatically edits the multimedia materials selected by the user, the tail-end script can be automatically generated according to the content of the tail-end material, and the tail-end script and the tail-end material can be displayed together in the tail-end picture. In this way, the defect of templateization of the tail-end of the edited video can be overcome, and the tail-end effect of the edited video can be improved.
[0010] According to the first aspect, in the first video frame, the size of the playing picture of the end material is a target size, and the target size is smaller than the size of the picture of the first clip video. Correspondingly, when the electronic device plays the end video segment of the first clip video, and before the electronic device displays at least one first video frame, the method further includes: gradually reducing the size of the playing picture of the end material from the size of the picture of the first clip video to the target size.
[0011] The size of the playing picture of the end material is gradually reduced to the target size in a proportional manner from the size of the picture of the first clip video. The proportional manner can be understood as that the aspect ratio is unchanged.
[0012] In a certain period of the end video segment of the first clip video, the size of the playing picture of the end material is gradually reduced to the target size from the size of the picture of the first clip video, so as to leave a display area for the first text.
[0013] According to the first aspect, or any one of the implementation manners of the first aspect, the aspect direction of the first clip video is horizontal; and in the first video frame, the playing picture of the end material and the first text are arranged side by side.
[0014] For example, when the aspect ratio of the first clip video is 16:9, the aspect direction of the first clip video is horizontal.
[0015] According to the first aspect, or any one of the implementation manners of the first aspect, gradually reducing the size of the playing picture of the end material from the size of the picture of the first clip video to the target size includes:
[0016] The size of the playing picture of the end material is gradually reduced to a first size from the size of the playing picture of the first clip video to the left; the first size is smaller than the size of the picture of the first clip video; and in the first video frame, the first text is displayed to the right of the playing picture of the end material.
[0017] According to the first aspect, or any one of the implementation manners of the first aspect, the aspect direction of the first clip video is vertical; and in the first video frame, the playing picture of the end material and the first text are arranged above and below.
[0018] For example, when the aspect ratio of the first clip video is 4:3, the aspect direction of the first clip video is vertical.
[0019] According to the first aspect, or any one of the implementation manners of the first aspect, gradually reducing the size of the playing picture of the end material from the size of the picture of the first clip video to the target size includes:
[0020] The size of the playing picture of the end-credits material is gradually reduced from the size of the playing picture of the first video clip to a second size, and the second size is smaller than the size of the playing picture of the first video clip; and the first text is displayed on the upper side of the playing picture of the end-credits material in the first video frame.
[0021] According to the first aspect, or any one of the implementations of the first aspect, after the size of the playing picture of the end-credits material is gradually reduced from the size of the playing picture of the first video clip to the target size, the method further includes:
[0022] The first text is gradually displayed according to a first animation effect in the first video frame.
[0023] The first animation effect may be, for example, word-by-word display.
[0024] According to the first aspect, or any one of the implementations of the first aspect, before the electronic device displays the second interface, the method further includes: the electronic device determines the end-credits material; the electronic device generates a plurality of end-credits candidate texts according to the end-credits material; and the electronic device screens the first text from the plurality of end-credits candidate texts.
[0025] According to the first aspect, or any one of the implementations of the first aspect, the electronic device generates the plurality of end-credits candidate texts according to the end-credits material, including:
[0026] When the end-credits material is picture material, the electronic device performs picture-to-text processing on the picture material to obtain the plurality of end-credits candidate texts; and when the end-credits material is video material, the electronic device performs frame extraction processing on the video material to obtain second video frames, and performs picture-to-text processing on the second video frames to obtain the plurality of end-credits candidate texts.
[0027] According to the first aspect, or any one of the implementations of the first aspect, the electronic device screens the first text from the plurality of end-credits candidate texts, including:
[0028] The electronic device screens the first text from the plurality of end-credits candidate texts according to a text screening strategy;
[0029] The text screening strategy includes at least one of the following: a text character number limit condition, a punctuation symbol limit condition, a de-duplication processing, and a text style limit condition; and a customized style corresponding to the style limit condition is supported to be set by a user.
[0030] According to the first aspect, or any one of the implementations of the first aspect, after the electronic device displays the second interface, the method further includes:
[0031] The electronic device displays a third interface in response to a second operation; the second interface includes a second video clip corresponding to the multimedia material; and a frame direction of a clip template of the second video clip is different from a frame direction of a clip template of the first video clip.
[0032] In the first cut video, the playing picture of the end material and the first text are arranged in a first direction, and in the second cut video, the playing picture of the end material and the first text are arranged in a second direction; the first direction is a left-right direction or an up-down direction, and the second direction is a corresponding up-down direction or left-right direction.
[0033] For example, the second operation can be that after the user clicks the template control 107, the user reselects a video cutting template to trigger the electronic device to re-cut the video into a film.
[0034] In this way, if the aspect ratio of the automatically cut template reselected by the user changes, the layout of the reduced material picture in the end picture of the material and the end text in the video cutting application is adjusted accordingly in the video cutting film re-generated.
[0035] According to the first aspect or any one of the implementations of the first aspect, the method further includes: displaying at least one third video frame including second text related to the beginning material when the electronic device plays a beginning video clip of the first cut video.
[0036] In this way, when the electronic device automatically cuts the multimedia material selected by the user, the beginning text can be automatically generated according to the content of the beginning material and displayed in the beginning picture, thereby overcoming the defect of the template of the beginning content of the video cutting film and improving the beginning effect of the video film obtained by automatic cutting.
[0037] According to the first aspect or any one of the implementations of the first aspect, the second text is gradually displayed in the third video frame according to a second animation effect.
[0038] For example, the second animation effect can be word-by-word display.
[0039] According to the first aspect or any one of the implementations of the first aspect, the method further includes:
[0040] When the electronic device plays a target video clip of the first cut video, the electronic device displays at least one fourth video frame including target text corresponding to the target video clip and plays target audio corresponding to the target text.
[0041] When the electronic device automatically cuts the multimedia material selected by the user, the passage text (which can be referred to as a script) corresponding to at least one video clip can be generated, and the passage text and the audio corresponding to the passage text are added to the video cutting film, thereby improving the effect of the cutting video film obtained by the one-key cutting function and improving the user experience.
[0042] According to a first aspect, or any of the implementations of the first aspect, before the electronic device displays the second interface, the method further includes: parsing, by the electronic device, a constraint condition of the clip template of the first clip video; generating, by the electronic device, a target text according to the constraint condition and the video clips; generating, by the electronic device, a target audio according to the target text, and determining an insertion position of the target text and the target audio. The insertion position of the target text and the target audio is used for generating the first clip video.
[0043] For example, the constraint condition of the clip template can include, but is not limited to, a word number constraint of the paragraph text corresponding to each video clip, a time length constraint of the dubbing, etc.
[0044] According to the first aspect, or any of the implementations of the first aspect, generating, by the electronic device, the target audio according to the target text includes: selecting, by the electronic device, a target voice tone; and generating, by the electronic device, the target audio according to the target voice tone.
[0045] When the electronic device generates the audio matching the paragraph text according to the target voice tone, the audio time length needs to meet the constraint condition of the video clip template.
[0046] According to the first aspect, or any of the implementations of the first aspect, displaying, by the electronic device, the first interface includes: displaying, by the electronic device, an interface of a gallery application; displaying, by the electronic device, the first interface in response to a seventh operation; and playing, by the electronic device, a video material in the first interface.
[0047] As shown in the scenario of FIG. 2(2), the first interface can be, for example, the interface shown in FIG. 2(2), and the first operation can be an operation of clicking the AI one-key blockbuster control 104. Figure 1a Figure 1a According to the first aspect, or any of the implementations of the first aspect, displaying, by the electronic device, the first interface includes: displaying, by the electronic device, an interface of a gallery application; displaying, by the electronic device, the first interface in response to an eighth operation; and displaying, by the electronic device, a plurality of thumbnail images of multimedia materials in the first interface, one or more of the thumbnail images being in a selected state.
[0048] As shown in the scenario of FIG. 2(3), the first interface can be, for example, the interface shown in FIG. 2(3), and the first operation can be an operation of clicking the one-key movie control 207.
[0049] According to the first aspect, or any of the implementations of the first aspect, displaying, by the electronic device, the first interface includes: displaying, by the electronic device, an interface of a gallery application; displaying, by the electronic device, the first interface in response to an eighth operation; and displaying, by the electronic device, a plurality of thumbnail images of multimedia materials in the first interface, one or more of the thumbnail images being in a selected state. Figure 1b Figure 1b
[0050] In a second aspect, an embodiment of the present application provides an electronic device. The electronic device includes one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory and, when the computer programs are executed by the one or more processors, cause the electronic device to perform the video clip method in the first aspect and any one of the implementations of the first aspect.
[0051] The second aspect and any one of the implementations of the second aspect correspond to the first aspect and any one of the implementations of the first aspect respectively. For details of the technical effects of the second aspect and any one of the implementations of the second aspect, reference can be made to the technical effects of the first aspect and any one of the implementations of the first aspect, which will not be described here.
[0052] In a third aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium includes a computer program, when the computer program is run on an electronic device, causes the electronic device to perform the video clip method in the first aspect and any one of the implementations of the first aspect.
[0053] The third aspect and any one of the implementations of the third aspect correspond to the first aspect and any one of the implementations of the first aspect respectively. For details of the technical effects of the third aspect and any one of the implementations of the third aspect, reference can be made to the technical effects of the first aspect and any one of the implementations of the first aspect, which will not be described here.
[0054] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, when the computer program is run, causes a computer to perform the video clip method in the first aspect or any one of the implementations of the first aspect.
[0055] The fourth aspect and any one of the implementations of the fourth aspect correspond to the first aspect and any one of the implementations of the first aspect respectively. For details of the technical effects of the fourth aspect and any one of the implementations of the fourth aspect, reference can be made to the technical effects of the first aspect and any one of the implementations of the first aspect, which will not be described here.
[0056] In a fifth aspect, the present application provides a chip, including processing circuitry, transceiver pins. Wherein the transceiver pins and the processing circuitry communicate with each other through internal connection paths, the processing circuitry executes the video clip method in the first aspect or any one of the implementations of the first aspect to control the receiving pin to receive the signal, to control the sending pin to send the signal.
[0057] The fifth aspect and any kind of implementation manner of the fifth aspect correspond to the first aspect and any kind of implementation manner of the first aspect respectively. The technical effects corresponding to the fifth aspect and any kind of implementation manner of the fifth aspect can refer to the technical effects corresponding to the first aspect and any kind of implementation manner of the first aspect, which will not be described here again. BRIEF DESCRIPTION OF DRAWINGS
[0058] Figure 1a An application scenario schematic diagram of one-key video automatic clipping is exemplarily shown;
[0059] Figure 1b An application scenario schematic diagram of one-key video automatic clipping is exemplarily shown;
[0060] Figure 2 A film end schematic diagram of video automatic clipping into a film is exemplarily shown;
[0061] Figure 3 A hardware structure schematic diagram of an electronic device is exemplarily shown;
[0062] Figure 4 A software structure schematic diagram of an electronic device is exemplarily shown;
[0063] Figure 5 A module interaction schematic diagram of a video clipping method is exemplarily shown;
[0064] Figure 6 A film end schematic diagram of video automatic clipping into a film is exemplarily shown;
[0065] Figure 7a A film end animation schematic diagram of a horizontal frame in video automatic clipping into a film is exemplarily shown;
[0066] Figure 7b A film end animation schematic diagram of a vertical frame in video automatic clipping into a film is exemplarily shown;
[0067] Figure 8a A film head picture schematic diagram in video automatic clipping into a film is exemplarily shown;
[0068] Figure 8b A film head picture schematic diagram in video automatic clipping into a film is exemplarily shown;
[0069] Figure 9 A module interaction schematic diagram of a video clipping method is exemplarily shown;
[0070] Figure 10 A text and voice matching schematic diagram in automatic clipping into a film is exemplarily shown. DETAILED DESCRIPTION
[0071] With reference to the drawings and the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are some of the embodiments of the present application, but not all of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall within the scope of the present application.
[0072] The term "and / or" used herein is merely used to describe an association relationship of associated objects, and indicates that three relationships can exist, for example, A and / or B can represent three cases of A existing alone, A and B existing together, and B existing alone.
[0073] The terms "first" and "second" and the like in the specification and claims of the embodiments of the present application are used to distinguish different objects, and are not used to describe a specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, and are not used to describe a specific order of the target objects.
[0074] In the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design presented as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or advantageous than other embodiments or designs. Rather, the use of "exemplary" or "for example" is intended to present concepts in a concrete manner.
[0075] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more. For example, a plurality of processing units means two or more processing units; a plurality of systems means two or more systems.
[0076] With the advent of the era of self-media, video editing is becoming more and more popular, and more and more ordinary people without editing skills also need to edit their multimedia materials (including but not limited to video materials, picture materials, etc.). Therefore, the "one-key editing" function (or "one-key film" function, "one-key blockbuster" function, etc.) has emerged as the times require. The user only needs to perform a one-key operation (for example, an operation of clicking a "one-key editing" control) in an electronic device (hereinafter, the electronic device is taken as a mobile phone for explanation and description) to complete automatic editing of the multimedia materials and obtain an automatically edited video (or an automatically edited film, etc.).
[0077] In one application scenario, the user can install a third-party video editing application in the mobile phone, import one or more multimedia materials into the third-party video editing application, and use the "one-key editing" function of the third-party video editing application to complete automatic editing of the multimedia materials.
[0078] In another application scenario, the user can also use the "one-click clipping" function provided by the system-level application in the mobile phone to complete the automatic clipping of the multimedia material. In this scenario, the user does not need to additionally install a third-party video clipping application in the mobile phone, and does not need to perform an import operation on the multimedia material when needing to automatically clip the multimedia material. Among them, Figure 1a and Figure 1b respectively exemplarily show the application scenario of completing automatic clipping of a video based on the "one-click clipping" function provided by the system-level application in the mobile phone.
[0079] Figure 1a (1) and Figure 1b (1) exemplarily show the display interface of the gallery application. Exemplarily, in response to the user's operation of clicking the gallery application icon in the mobile phone, the mobile phone starts the gallery application and displays the gallery application interface 101 as shown in Figure 1a (1) or the gallery application interface 201 as shown in Figure 1b (1). In the gallery application interface 101, each picture (or photo) and video stored in the gallery application is displayed. Among them, for video clipping, the pictures and videos stored in the gallery application can all be used as multimedia materials (or to-be-clipped materials, hereinafter referred to as materials) for video clipping.
[0080] In an embodiment, the user can select a video material for automatic video clipping. Exemplarily, if the time length of a certain video material is greater than a preset threshold (for example, 15s), the video material meets the condition for individual clipping, and the video clipping application can complete the automatic clipping operation according to the video material to obtain a clipped video.
[0081] Continuing to refer to Figure 1a (1), in the gallery application interface 101, the user clicks the thumbnail 102 corresponding to the video material A. In response to the user's operation, the mobile phone plays the video material A, as shown in Figure 1a (2). Since the time length of the video material A is greater than the preset threshold, it meets the condition for individual clipping, and therefore, in the play interface 103 of the video material A, the "AI one-click blockbuster" control 104 can also be displayed. Continuing to refer to Figure 1a (2), the user clicks the "AI one-click blockbuster" control 104 to control the video clipping application to automatically clip the video material A. In response to the user's operation, the video clipping application in the mobile phone automatically clips the video material A to generate a clipped video B corresponding to the video material A, which is displayed as Figure 1aThe interface 105 is shown in FIG. 3 (3). Optionally, before displaying the interface 105, the phone can display a transition interface to prompt the user that the clip video B is being generated, such as a transition interface indicating "material analyzing", "video clipping". In the interface 105, the clip video B is displayed to inform the user of the automatic clipping result of the video material A. In the interface 105, a confirmation control 106, a template control 107, a segment control 109, an edit control 110, etc. can also be displayed. When the user needs to store the current clip video into the gallery application, the user can click the confirmation control 106 to store the current clip video into the gallery application. When the user needs to adjust the clipping template of the clip video, the user can click the template control 107 to make the phone display a video clipping template selection interface. In the video clipping template selection interface, the user can reselect an automatic clipping template for the video material A. When the user needs to adjust the background music of the clip video, the user can click the music control 108 to make the phone display a video background music editing interface. In the video background music editing interface, the user can edit the background music of the clip video, such as adjusting the volume of the background music, replacing the background music, etc. When the user needs to adjust the segments of the clip video, the user can click the segment control 109 to make the phone display a video segment editing interface. In the video segment editing interface, the user can edit the video segments of the clip video, such as adjusting the video segments, replacing the video segments, etc. When the user needs to edit the clip video, the user can click the edit control 110 to make the phone display a video editing interface. In the video editing interface, the user can edit the clip video, such as adding stickers, adjusting the duration, etc.
[0082] In another implementation, the user can select multiple multimedia materials for automatic video clipping, such as selecting multiple picture materials for automatic video clipping, or selecting multiple video materials for automatic video clipping, or selecting video materials and picture materials for automatic video clipping.
[0083] For example, continuing to refer to Figure 1b In the gallery application interface 201, the user long-presses a thumbnail corresponding to any one material, such as clicking the thumbnail 202 corresponding to the picture material a, as shown in FIG. 3 (1). In response to the user operation, the phone displays a material selection interface 203, as shown in FIG. 3 (2). Figure 1b Optionally, the user can also trigger the phone to display the material selection interface 203 by clicking the selection control, or by other operation manners. Figure 1b Optionally, the user can also trigger the phone to display the material selection interface 203 by clicking the selection control, or by other operation manners. Figure 1bThe material selection interface 203 is not limited in the present embodiment. For example, in the material selection interface 203, a to-be-selected mark 204 is displayed on the thumbnail of each material (including picture material and video material). When a user selects a certain material, a selected mark 205 (for example, a number mark as shown in the figure) is displayed in the to-be-selected mark 204 on the thumbnail of the material. For details, refer to Figure 1b The selected mark 205 is not limited in the present embodiment. For example, the selected mark 205 can also be a check mark (such as “√”). It should be noted that when a user selects multiple materials, the order of the materials in the video clip is determined according to the order of the user selection. For example, as shown in Figure 1b The value corresponding to the selected mark 205 is the order of the material selection. When the user-selected material meets the automatic video clipping condition, for example, when the user selects one or more picture materials, or when the user selects one or more video materials, optionally, when the user selects a video material, the video material meets the condition of separate clipping, the mobile phone pops up a prompt information box 206 to prompt the user that a movie-like video short can be automatically generated based on the selected material. In the prompt information box 206, a “one-key movie” control 207 is displayed. Continue to refer to Figure 1b The user clicks the “one-key movie” control 207, and the video clipping application automatically clips the user-selected material. For example, as shown in Figure 1b In the example shown in (3), the user selects three video materials and one picture material for video clipping. In response to the user operation, the video clipping application in the mobile phone automatically clips the four user-selected materials to generate a clipping video C corresponding to the video materials, and displays an interface 208 as shown in Figure 1b In the interface 208, the clipping video C is displayed to enable the user to know the automatic clipping result of the selected video material. When the user needs to store the current clipping video into the gallery application, the user can click a confirmation control 209 to store the current clipping video into the gallery application. For details of other controls in the interface 208, refer to the explanation of the interface 207 in (2), which will not be repeated here. Figure 1a The explanation of the interface 208 in (3) will not be repeated here.
[0084] In the above automatic video clipping scenario, the end of the video clip is usually a general end picture, for example Figure 2 The end picture shown in (1), the main body of the picture is the text “THE END” and the text “Thank you for watching”, for example Figure 2The end screen shown in the middle (2) is an end screen with a picture subject of a "LOGO" pattern and text "Please pay attention" and the like. Such a cut video end screen has a picture subject that is similar, and the difference can only be in different text fonts, different "LOGO" patterns, and the like, and the user experience is not good.
[0085] To improve the user experience of the one-key automatic clipping function, an embodiment of the present application provides a video clipping method. In this method, when the electronic device automatically clips the multimedia material selected by the user, a text (or script) can be generated according to the content of the end material, and the end part of the cut video is generated according to the text, so that the end part of the cut video contains a text description that matches the content of the material, improving the difference of the end part of the cut video, avoiding the feeling of uniformity of the end part of the automatically cut video into a film, and thus improving the user experience of the one-key automatic clipping function.
[0086] In this method, when the electronic device automatically clips the multimedia material selected by the user, a text can also be generated according to the content of the end material, and the end part of the cut video is generated according to the text, so that the end part of the cut video contains a text description that matches the content of the material, thus improving the end effect of the automatically cut video into a film.
[0087] Optionally, in this method, when the electronic device automatically clips the multimedia material selected by the user, a text and a voiceover can also be generated according to at least one video segment, and added to the matching position in the cut video into a film. The video segment can be generated from video material or from picture material, and this embodiment does not limit it.
[0088] Figure 3 The structure of the electronic device 100 is shown. It should be understood that, Figure 3 The electronic device 100 shown is only an example of an electronic device, and the electronic device 100 can have more or fewer components than those shown in the figure, can combine two or more components, or can have a different component configuration. Figure 3 The various components shown in the middle can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.
[0089] The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor, a gyroscope sensor, a barometric sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0090] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.
[0091] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0092] The memory in the processor 110 can also be provided for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thus improving the efficiency of the system.
[0093] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0094] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In some embodiments, the audio module 170 can be coupled with the wireless communication module 160 through a PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface, enabling the function of answering a phone call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0095] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation of the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection methods or a combination of multiple interface connection methods.
[0096] The charging management module 140 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger. The charging management module 140 can charge the battery 142 and also supply power to the electronic device through the power management module 141. The power management module 141 is configured to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193 and the wireless communication module 160, etc.
[0097] The wireless communication function of the electronic device 100 can be implemented by the antennas 1, 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc. In some embodiments, the antenna 1 and the mobile communication module 150 of the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 160 are coupled, so that the electronic device 100 can communicate with the network and other devices through wireless communication technology.
[0098] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel.
[0099] The electronic device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc. The ISP is used to process the data fed back by the camera 193. The camera 193 is used to capture still images or videos.
[0100] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0101] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, music, video, etc. Files are saved in the external memory card.
[0102] The internal memory 121 can be used to store computer executable program codes including instructions. The processor 110 performs various function applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program (such as a sound playing function, an image playing function, etc.) required by a function, etc. The data storage area can store data (such as audio data, a phone book, etc.) created during the use of the electronic device 100, etc. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0103] The electronic device 100 can realize audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, the application processor, etc. For example, music playing, recording, etc.
[0104] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be arranged in the processor 110, or part of the functions of the audio module 170 can be arranged in the processor 110.
[0105] The pressure sensor is used to sense a pressure signal, and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor can be arranged in the display screen 194. The touch sensor, also known as a "touch panel". The touch sensor can be arranged in the display screen 194, and the touch sensor and the display screen 194 form a touch screen, also known as a "touch screen".
[0106] The keys 190 include a power-on key, a volume key, etc. The motor 191 can generate a vibration prompt. The indicator 192 can be an indicator light, which can be used to indicate a charging state, a power change, and can also be used to indicate a message, a missed call, a notification, etc. The SIM card interface 195 is used to connect a SIM card.
[0107] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes an Android system with a layered architecture as an example to exemplarily illustrate the software structure of the electronic device 100.
[0108] Figure 4 is a software structure block diagram of the electronic device 100 of the embodiment of the present application.
[0109] The layered architecture of the electronic device 100 divides software into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, the application layer, the application framework layer, the Android runtime and system library, and the kernel layer.
[0110] The application layer can include a series of application packages.
[0111] As shown in Figure 4 , the application package can include a gallery application, a video clip application, etc. Among them, the video clip application can be a system-level application or a third-party application. In addition, the application package can also include camera applications, WLAN, Bluetooth, calls, calendars, maps, navigation, music, video, short messages, etc.
[0112] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications of the application layer. The application framework layer includes some pre-defined functions.
[0113] As shown in Figure 4 , the application framework layer can include a media processing module (or media processing middleware), a wisdom processing module (or wisdom processing middleware), a visual analysis module (or visual analysis middleware), and a material shelving module (or material shelving middleware).
[0114] The media processing module can be used for offline analysis of multimedia materials, which can include but is not limited to material theme analysis, highlight video segment analysis in video materials, generating alternative scripts for opening and / or closing materials, generating video segment scripts, etc.
[0115] For example, the media processing module can call the visual analysis module to perform visual analysis on the multimedia material to obtain a material visual analysis result, and identify the highlight video segment based on the material visual analysis result. The visual analysis module can perform visual analysis on the multimedia material based on a pre-set visual analysis algorithm (e.g., visual atomization analysis algorithm).
[0116] For example, the media processing module can call the wisdom processing module to identify the video theme type. The wisdom processing module performs semantic understanding on the video frames or pictures extracted by the media processing module to generate a description text, and then identifies the theme type of the video based on the description text.
[0117] Exemplarily, the media processing module can invoke the wisdom processing module to generate the opening material candidate script and / or the closing material candidate script. The wisdom processing module can perform image-to-text processing on the relevant video frames or images to obtain the opening material candidate script and / or the closing material candidate script. As the name suggests, image-to-text refers to generating matching text (or script) from images.
[0118] Exemplarily, the media processing module can also invoke the wisdom processing module to generate the video clip's caption and dubbing information, and determine the insertion position information of the caption and dubbing information in the edited video. The wisdom processing module can generate a paragraph of text that meets the template constraint conditions (such as time length, word count, etc.) according to the relevant video frames and / or images, and invoke a TTS (Text To Speech) service (TTS function module) to generate the corresponding audio. The timbre of the audio (such as male timbre, female timbre) can be determined according to the relevant query information (such as the gender of the owner of the electronic device, etc.).
[0119] In this embodiment, the media processing module can read or cache low-quality (or non-high-definition) material corresponding to the multimedia original material in the media data module (or media data middle station), and can also read or cache analysis information corresponding to the multimedia original material in the media data module to improve the processing efficiency of the media processing module.
[0120] The material shelf module can be used for video editing application to download video editing templates and materials related to video editing, etc.
[0121] Continuing to refer to Figure 4 The application framework layer can also include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.
[0122] The window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, and take screenshots, etc.
[0123] The content provider is used to store and obtain data, and make the data accessible to application programs. The data can include videos, images, audios, dialed and received calls, browsing history and bookmarks, phone books, etc.
[0124] The view system includes visual controls, such as controls for displaying text, controls for displaying images, etc. The view system can be used to build application programs. A display interface can be composed of one or more views. For example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying images.
[0125] The telephony manager is used to provide the communication function of the electronic device. For example, the management of the call status (including the connection, hang-up, etc.).
[0126] The resource manager provides various resources for the application program, such as the localized string, the icon, the picture, the layout file, the video file, etc.
[0127] The notification manager enables the application program to display the notification information in the status bar, which can be used to convey the message of the notification type and can automatically disappear after a short stay without the user interaction.
[0128] The Android Runtime includes the core library and the virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.
[0129] The core library contains two parts: one part is the function function that the java language needs to call, and the other part is the core library of the Android.
[0130] The application program layer and the application framework layer run in the virtual machine. The virtual machine executes the java file of the application program layer and the application framework layer into the binary file. The virtual machine is used to execute the management of the object life cycle, the stack management, the thread management, the security and the exception management, and the garbage collection, etc.
[0131] The system library can include multiple function modules. For example: the surface manager, the media library, the three-dimensional graphics processing library (for example: OpenGL ES), the 2D graphics engine (for example: SGL), etc. Among them, the surface manager is used to manage the display subsystem and provides the fusion of the 2D and 3D layers for multiple application programs. The media library supports the playback and recording of multiple commonly used audio, video formats, and static image files, etc. The media library can support multiple audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. The three-dimensional graphics processing library is used to realize the three-dimensional graphics drawing, the image rendering, the synthesis, and the layer processing, etc. The 2D graphics engine is the drawing engine of the 2D drawing.
[0132] The kernel layer is the layer between the hardware and the software. The kernel layer at least contains the display driver, the audio driver, the Wi-Fi driver, the sensor driver, etc. Among them, the hardware at least includes the processor, the display screen, the Wi-Fi module, the sensor, etc.
[0133] It can be understood that, Figure 4The layers in the illustrated software structure and the components contained in each layer do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer layers than illustrated, and each layer can include more or fewer components, or some components can be combined, or some components can be split, or different component arrangements, which are not limited in the present application.
[0134] It can be understood that, in order to implement the video clipping method in the embodiments of the present application, the electronic device contains the corresponding hardware and / or software modules for executing various functions. The algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application of the technical solution and the design constraints. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered beyond the scope of the present application.
[0135] Scenario One
[0136] In this scenario, when the electronic device automatically clips the multimedia material selected by the user, a text (or script) can be generated according to the content of the end material, and / or a text can be generated according to the content of the beginning material, and a corresponding end part and beginning part of the clipped video are generated, so that the end and beginning of the clipped video contain text descriptions that match the content of the material, thereby enriching the content of the end and beginning of the clipped video and avoiding the template of the end and beginning of the clipped video.
[0137] As Figure 5 shown is an interaction diagram of each module. Referring to Figure 5 , the flow of the video clipping method provided by the embodiments of the present application specifically includes:
[0138] S301, in response to a one-key clipping operation of the user, the gallery application inputs the material selected by the user to the video clipping application.
[0139] For example, the one-key clipping operation of the user can be an operation of the user clicking the "AI one-key blockbuster" control as shown in (2) of Figure 1a , or can be an operation of the user clicking the "one-key movie" control as shown in (3) of Figure 1b .
[0140] In an implementation manner, S301 can also be adjusted to: in response to the one-key clipping operation of the user, the video clipping application acquires the material selected by the user in the gallery application.
[0141] The user-selected material, which can also be referred to as material to be cut, can include picture material and / or video material selected in the gallery application.
[0142] In S302, the video cutting application performs offline analysis on the user-selected material.
[0143] In the embodiments of the present application, when the video cutting application performs offline analysis on the user-selected material, the video cutting application can analyze not only the theme type of the material and the highlight video clip, but also generate candidate scripts corresponding to the end-of-clip material and / or the start-of-clip material.
[0144] The end-of-clip material can be video material or picture material, and the embodiments of the present application do not limit the type of the end-of-clip material. The start-of-clip material can be video material or picture material, and the embodiments of the present application do not limit the type of the start-of-clip material.
[0145] When the number of the user-selected materials for video cutting is one, the user-selected material is both the start-of-clip material and the end-of-clip material. When the number of the user-selected materials for video cutting is more than one, the first user-selected material, for example, the material with the identification 205 of "1" in the interface shown in (3) of Figure 1b the last user-selected material, for example, the material with the identification 205 of "4" in the interface shown in (3) of Figure 1b the last user-selected material, for example, the material with the identification 205 of "4" in the interface shown in (3) of
[0146] Continuing to refer to (1) shown above, the offline analysis of the user-selected material by the video cutting application can specifically include: Figure 5
[0147] In S3021, the video cutting application calls the media processing module to perform offline analysis on the user-selected material.
[0148] In the embodiments, the offline analysis of the material can be implemented by relying on the media processing module (or media processing middle platform) in the application program framework layer. Optionally, the media processing module provides a calling interface, and the video cutting application calls the media processing module to analyze the user-selected material through the calling interface, including but not limited to identifying the theme type of the video, identifying the highlight video clip, generating the candidate script of the end-of-clip material and / or the candidate script of the start-of-clip material, etc.
[0149] In an optional implementation, the calling parameter of the media processing module can include a tail material script analysis parameter and a head material script analysis parameter. The parameter value of the tail material script analysis parameter is used to indicate whether the alternative script corresponding to the tail material needs to be generated, and the parameter value of the head material script analysis parameter is used to indicate whether the alternative script corresponding to the head material needs to be generated. For example, the parameter value of the tail material script analysis parameter and the parameter value of the head material script analysis parameter can be a default setting, can be pre-set by the user according to the use demand, or can be flexibly set by the video clipping application according to the number of materials selected by the user. For example, when the number of materials selected by the user for video clipping is one, the video clipping application sets the parameter value of the tail material script analysis parameter to a target value, so as to enable the generation of the alternative script corresponding to the tail material. When the number of materials selected by the user for video clipping is multiple, the video clipping application sets the parameter value of the tail material script analysis parameter to a target value 1 and sets the parameter value of the head material script analysis parameter to a target value 2, so as to enable the generation of the alternative script corresponding to the tail material and the generation of the alternative script corresponding to the head material. The target value 1 and the target value 2 can be the same or different, which is not limited in the embodiment. Of course, the video clipping application can only set the parameter value of the head material script analysis parameter to a target value, so as to enable the generation of the alternative script corresponding to the head material, which is not limited in the embodiment.
[0150] S3022, under the calling of the video clipping application, the material analysis scheduling decision module in the media processing module calls the visual analysis module to perform visual analysis on the material selected by the user, to obtain a material visual analysis result.
[0151] The visual analysis module provides a calling interface, through which the media processing module can call the visual analysis module to perform visual analysis on the material selected by the user, to obtain a material visual analysis result.
[0152] Under the calling of the material analysis scheduling decision module, the visual analysis module can perform visual analysis on the material selected by the user based on a preset visual atomization algorithm, to obtain a material visual analysis result, and return the material visual analysis result to the material analysis scheduling decision module.
[0153] For example, the visual analysis of the visual analysis module on the material can include but is not limited to image classification, target detection, image segmentation, pose estimation, video stream action understanding (such as behavior recognition, time sequence action detection, space-time action detection, etc.), OCR (Optical Character Recognition, optical character recognition), etc.
[0154] S3023, the highlight video clip in the video material is analyzed by the material analysis and scheduling decision module in the media processing module.
[0155] When the user-selected material includes video material, the media processing module can analyze the highlight video clip in the video material based on the visual analysis result of the visual analysis module on the material.
[0156] The highlight video clip can be understood as a highlight video clip in the video material, which can refer to a video clip containing highlight frames (video frames corresponding to highlight moments). For example, in a birthday theme video material, the highlight video clip can include a "blowing candles" video clip, a "cutting cake" video clip, a "singing birthday song" video clip, etc. For example, in a video material with a theme of kicking a ball, the highlight video clip can include a "goal" video clip, a "passing" video clip, a "receiving" video clip, etc.
[0157] S3024, the material analysis and scheduling decision module in the media processing module calls the wisdom processing module to identify the video theme type, and generates a clip head material candidate script and / or a clip tail material candidate script.
[0158] For example, the theme type can include but is not limited to childlike fun, character, food, sports, travel, ancient architecture, night scene, nature, leisure, dynamic rhythm, light and happy, slightly sad, etc.
[0159] When the user-selected material (i.e. the material to be edited) includes video material, the media processing module can perform frame extraction processing on the video material to obtain at least one video frame image. Optionally, the media processing module can perform frame extraction processing on the highlight video clip to obtain at least one video frame image. The video frame image obtained by the media processing module through frame extraction can be used for video theme type identification. When the user-selected material includes picture material, the picture material can be directly used for video theme type identification.
[0160] For example, the media processing module inputs the at least one video frame image and / or the picture material into the wisdom processing module for video theme type identification. Optionally, the wisdom processing module can identify the video theme type of the material to be cut according to a preset theme type identification model. If the material to be cut is a video material, the wisdom processing module inputs the at least one video frame image obtained by the frame extraction processing of the media processing module into the theme type identification model to identify the theme of the video material; if the material to be cut is a picture material, the wisdom processing module inputs the picture material into the theme type identification model to identify the theme of the material; when comprehensively identifying the themes of multiple materials to be cut, the wisdom processing module can jointly input each picture material and the frame extraction result (i.e., the video frame image obtained by the frame extraction operation) corresponding to each video material included in the material to be cut into the theme identification model to identify the theme of the material to be cut. The frame extraction method for the video material can be determined according to actual needs and specific scenarios, and the embodiments of the present application do not limit this. For example, the frame extraction method can be uniform frame extraction of the video material at a certain step length.
[0161] For another example, after the media processing module inputs the at least one video frame image and / or the picture material into the wisdom processing module, the wisdom processing module first performs semantic understanding processing on the at least one video frame image and / or the picture material to obtain corresponding description texts, and then performs video theme type identification according to the description texts. Optionally, the wisdom processing module can identify the video theme type of the material to be cut according to a preset theme type identification model. If the material to be cut is a video material, the wisdom processing module performs semantic understanding processing on the at least one video frame image obtained by the frame extraction processing of the media processing module to obtain description texts corresponding to the at least one video frame image, and then inputs the description texts into the theme type identification model to identify the theme of the video material; if the material to be cut is a picture material, the wisdom processing module performs semantic understanding processing on the picture material to obtain a description text corresponding to the picture material, and then inputs the description text into the theme type identification model to identify the theme of the material; when comprehensively identifying the themes of multiple materials to be cut, the wisdom processing module can perform semantic understanding processing on the at least one video frame image and the picture material respectively to obtain description texts corresponding to the at least one video frame image and the picture material respectively, and then jointly input the description texts into the theme identification model to identify the theme of the material to be cut.
[0162] Taking the case that the material analysis scheduling decision module in the media processing module calls the wisdom processing module to generate a trailer material candidate script as an example, the material analysis scheduling decision module first determines whether the last material to be cut selected by the user in order is a video material or a picture material.
[0163] When the last one of the to-be-edited materials selected by the user in sequence is a video material, the material analysis scheduling decision module performs frame extraction processing on the video material to obtain a video frame image. Optionally, when the last one of the to-be-edited materials selected by the user in sequence is a video material, the material analysis scheduling decision module performs frame extraction processing on a highlight video clip corresponding to the video material to obtain a video frame image. The material analysis scheduling decision module inputs the video frame image into the wisdom processing module, and the wisdom processing module performs image-to-text processing on the video frame image to obtain a plurality of text cases corresponding to the video frame image, and the text cases are the end-credit material candidate text cases. The text case corresponding to the video frame image can be understood as a text case used to describe the content of the video frame image.
[0164] For example, the wisdom processing module processes the video frame image by using a preset image-to-text neural network model, such as inputting the video frame image into the image-to-text neural network model, and the image-to-text neural network model outputs a plurality of candidate text cases corresponding to the video frame image. The neural network or algorithm used by the image-to-text neural network model can refer to the prior art, and will not be described herein.
[0165] As an optional implementation, the wisdom processing module performs image-to-text processing on the video frame image according to the identified video theme type to obtain a plurality of text cases corresponding to the video frame image, and the text cases are the end-credit material candidate text cases. For example, the wisdom processing module processes the video frame image by using a preset image-to-text neural network model, such as inputting the video frame image and the video theme type into the image-to-text neural network model, and the image-to-text neural network model outputs a plurality of candidate text cases corresponding to the video frame image.
[0166] When the last one of the to-be-edited materials selected by the user in sequence is a picture material, the material analysis scheduling decision module inputs the picture material into the wisdom processing module, and the wisdom processing module performs image-to-text processing on the picture material to obtain a plurality of text cases corresponding to the picture material, and the text cases are the end-credit material candidate text cases. The text case corresponding to the picture material can be understood as a text case used to describe the content of the picture.
[0167] For example, the wisdom processing module processes the picture material by using a preset image-to-text neural network model, such as inputting the picture material into the image-to-text neural network model, and the image-to-text neural network model outputs a plurality of candidate text cases corresponding to the picture material.
[0168] Taking the case that the material analysis scheduling decision module in the media processing module calls the wisdom processing module to generate the opening-credit material candidate text cases as an example, the material analysis scheduling decision module first determines whether the first to-be-edited material selected by the user in sequence is a video material or a picture material.
[0169] When the first selected material is a video material, the material analysis and scheduling decision module performs frame extraction on the video material to obtain a video frame image. Alternatively, when the first selected material is a video material, the material analysis and scheduling decision module performs frame extraction on a highlight video clip corresponding to the video material to obtain a video frame image. The material analysis and scheduling decision module inputs the video frame image into the wisdom processing module, and the wisdom processing module performs image-to-text processing on the video frame image to obtain a plurality of text cases corresponding to the video frame image, which are the selected text cases for the opening material.
[0170] When the first selected material is a picture material, the material analysis and scheduling decision module inputs the picture material into the wisdom processing module, and the wisdom processing module performs image-to-text processing on the picture material to obtain a plurality of text cases corresponding to the picture material, which are the selected text cases for the opening material.
[0171] As an optional implementation, the material analysis and scheduling decision module can call the wisdom processing module to generate the selected text cases for the opening material according to all the selected materials. For example, when the selected materials include video materials and picture materials, the material analysis and scheduling decision module performs frame extraction on the video materials to obtain video frame images corresponding to each video material, respectively. The wisdom processing module generates the selected text cases for the opening material according to the video frame images and the picture materials.
[0172] S3025, the material analysis and scheduling decision module in the media processing module returns the material analysis result to the video editing application.
[0173] The material analysis result can include, but is not limited to, the video theme type, the information of the highlight video clip, and a plurality of selected text cases for the opening material and / or a plurality of selected text cases for the ending material.
[0174] S303, the video editing application selects a video editing template and downloads the template and the related materials in the material shelving module.
[0175] Since the aspect ratio of the materials (or the aspect page direction) can be different, for example, the aspect ratio of some picture materials and / or video materials is vertical (or vertical direction), and the aspect ratio of some picture materials and / or video materials is horizontal (or horizontal direction), when the video editing application selects the aspect ratio of the video editing template, the aspect ratio of the materials can be referred to.
[0176] For example, the video editing application counts the number of vertical materials and the number of horizontal materials in the materials to be edited, respectively. If the number of vertical materials is more than the number of horizontal materials, a vertical-frame video automatic editing template is selected. If the number of vertical materials is less than the number of horizontal materials, a horizontal-frame video automatic editing template is selected. If the number of vertical materials is equal to the number of horizontal materials, a video automatic editing template is selected according to the frame direction of the first material selected by the user. If the frame direction of the first material is horizontal, a horizontal-frame video automatic editing template is selected. Otherwise, a vertical-frame video automatic editing template is selected.
[0177] For another example, the video editing application counts the time length proportion of vertical materials and the time length proportion of horizontal materials in the materials to be edited, respectively. If the time length proportion of vertical materials is more than the time length proportion of horizontal materials, a vertical-frame video automatic editing template is selected. If the time length proportion of vertical materials is less than the time length proportion of horizontal materials, a horizontal-frame video automatic editing template is selected. If the time length proportion of vertical materials is equal to the time length proportion of horizontal materials, a video automatic editing template is selected according to the frame direction of the first material selected by the user. If the frame direction of the first material is horizontal, a horizontal-frame video automatic editing template is selected. Otherwise, a vertical-frame video automatic editing template is selected.
[0178] After determining the frame direction of the video automatic editing template required for editing, the video editing application selects a template matching the frame direction according to the theme type of the materials as the video automatic editing template. For example, the video editing application can select a plurality of to-be-confirmed templates matching the frame direction according to the theme type of the materials in the template library, and calculate a matching score according to the matching degree between the materials and the to-be-confirmed templates, respectively. The template with the highest matching score is selected as the video automatic editing template. For another example, after the video editing application selects a plurality of to-be-confirmed templates matching the frame direction according to the theme type of the materials in the template library, the video editing application can calculate a matching score of each to-be-confirmed template according to the number of materials (such as the sum of the number of highlight video clips and image materials) and / or the time length of the highlight video clips, and select the template with the highest matching score as the video automatic editing template.
[0179] After selecting the video automatic editing template, the video editing application can download the corresponding template in the material shelving module, and download the editing materials related to the template. It should be noted that the editing materials related to the template at least include background music.
[0180] In S304, the video editing application selects an ending material in the ending material candidate scripts, and / or selects a starting material in the starting material candidate scripts.
[0181] When the end-clip candidate texts are included in the material analysis result, the video editing application selects an end-clip text from the plurality of end-clip candidate texts according to a preset filtering strategy. When the start-clip candidate texts are included in the material analysis result, the video editing application selects a start-clip text from the plurality of start-clip candidate texts according to the preset filtering strategy.
[0182] In the embodiments of the present application, the video editing application selects a target text (i.e., an end-clip text or a start-clip text) for the short video clip from the plurality of candidate texts (including end-clip candidate texts and start-clip candidate texts). The filtering strategies for the end-clip text and the start-clip text can be the same or different, and the embodiments are not limited in this regard.
[0183] For example, the text filtering strategy can include, but is not limited to, a text character number limit condition, a punctuation symbol limit condition, a style filtering condition, and the like. In the style filtering condition, the supported text styles of the video editing application can include, but are not limited to, literary, beautiful, lively, cute, seasonal, holiday, and the like. For example, when the set style corresponding to the style filtering condition is a literary style, the video editing application filters the end-clip candidate texts that meet the literary style from the plurality of candidate texts. The set style corresponding to the style filtering condition can be set flexibly by the user according to the user's own needs, so that the end-clip text and / or the start-clip text in the video clip meets the user's psychological expectation.
[0184] For example, for each end-clip candidate text, the video editing application determines whether it meets the text filtering strategy, such as whether the text character number meets the text character number limit condition, whether the punctuation symbol in the text meets the punctuation symbol limit condition, and whether the style of the text meets the text style customized by the video editing application. Furthermore, the video editing application can select any end-clip candidate text that meets the text filtering strategy as the final end-clip text.
[0185] For example, for each end-clip candidate text, the video editing application scores it according to the text filtering strategy and the corresponding scoring rule, and selects the end-clip candidate text with the highest score as the final end-clip text.
[0186] After determining the end-clip text, the video editing application can also design the font of the end-clip text according to the customized style, so that the video editing application can perform animation rendering according to the corresponding design when synthesizing the video clip.
[0187] The same is true for the method of screening the final opening text from a plurality of opening text material candidate texts. The present embodiment will not repeat the details. Before the video editing application screens the opening text and the ending text, the video editing application can also perform a deduplication process on the plurality of opening text material candidate texts and the plurality of ending text material candidate texts to avoid duplication of the opening text and the ending text. The video editing application can also filter the plurality of opening text material candidate texts and the plurality of ending text material candidate texts according to the number of times the texts are selected to avoid the same text being used multiple times.
[0188] The present embodiment does not limit the sequence of S303 and S304.
[0189] S305, the video editing application synthesizes a video clip according to the video editing template.
[0190] When the user-selected material only includes video material, the video editing application performs a video editing and splicing operation on the video material according to the automatic video editing template; when the user-selected material only includes picture material, the video editing application performs a video editing and splicing operation on the picture material according to the automatic video editing template; when the user-selected material includes video material and picture material, the video editing application performs a video editing and splicing operation on the video material and the picture material according to the automatic video editing template.
[0191] In the present embodiment, if the video editing application determines the ending text, the ending text is added to the ending part of the video clip. In the ending picture of the video clip, the last material (such as video material or picture material) selected by the user and the ending text are displayed.
[0192] Before the ending text is displayed, the picture size of the last material selected by the user gradually decreases to reserve a display area for the display of the ending text in the ending picture. Optionally, during the gradual decrease of the picture size of the material, the picture aspect ratio remains unchanged. For example, the picture size of the last material selected by the user can be reduced to about half of the picture size of the video clip. All the text in the ending text can be displayed on the ending picture at the same time, or can be gradually displayed on the ending picture according to a preset animation effect. The present embodiment does not limit this.
[0193] In addition, it should be noted that during the gradual decrease of the picture size of the last material selected by the user, and after the picture size of the last material is reduced to the target size, the playback of the material picture is continuous. For example, for video material, the playback of the material picture is the playback of the video segment. For example, for picture material, the playback of the material picture can be the playback of the animation effect (related to the automatic editing template) corresponding to the picture material.
[0194] Figure 6An example is shown as the end credits of a video clip. Figure 6 As shown, the end credits screen includes a frame 401 showing the last clip selected by the user, and end credits text 402 matching the last clip. The end credits text 402 can be used to describe the last clip selected by the user. (Continue to refer to...) Figure 6 The size of the 401 image is derived from the reduced size of the video clip. As the image continues to shrink, the aspect ratio of the source image remains unchanged.
[0195] When the automatic editing template is set to landscape orientation (e.g., 16:9), the scaled-down footage and the end credits are positioned horizontally in the final cut of the video, for example, the scaled-down footage is on the left and the end credits are on the right (e.g., the scaled-down footage is on the left and the end credits are on the right). Figure 6 (As shown in the example), for instance, the scaled-down footage is on the right and the end credits are on the left. This embodiment does not limit this.
[0196] When the aspect ratio of the automatic editing template is vertical (such as a 4:3 screen), in the end credits of the video edit, the scaled-down footage and the end credits text are distributed vertically. For example, the scaled-down footage is on the bottom and the end credits text is on the top, or the scaled-down footage is on the top and the end credits text is on the bottom. This embodiment does not limit this.
[0197] In this embodiment, the aspect ratio of the automatic editing template does not constrain the generation of end credits text, but only affects the playback of the end credits animation and the layout of elements in the end credits screen.
[0198] Taking the horizontal aspect ratio of a video edit as an example, Figure 7a An exemplary diagram is shown of the end credits segment of a video clip. For example... Figure 7a As shown in (1), at time t1 when the video is edited into a final cut, the last clip begins to play, and at this time, the image size of the clip is the same as the image size of the final cut. At time t2 when the video is edited into a final cut, the playback image of the clip gradually shrinks to the left, from the image size of the final cut to size 1, which can be seen by referring to... Figure 7a As shown in (2). Continue to refer to Figure 7a As shown in (3), after the image size of the material is reduced to size 1, the corresponding end credits text "All things stand facing the sun" is displayed on the right side of the image. The end credits text corresponding to the material can be displayed through a gradually unfolding animation effect. This embodiment does not limit the animation effect displayed in the end credits file.
[0199] Taking the vertical aspect ratio of the video clip as an example, Figure 7b An exemplary diagram is shown of the end credits segment of a video clip. For example...Figure 7b As shown in (1), at time t3 when the video is edited into a final cut, the last clip begins to play, and at this time, the image size of the clip is the same as the image size of the final cut. At time t4 when the video is edited into a final cut, the playback image of the clip gradually shrinks downwards, from the image size of the final cut to size 2, which can be seen by referring to... Figure 7b As shown in (2). Continue to refer to Figure 7b As shown in (3), after the image size of the material is reduced to size 2, the corresponding end credits text "Summer Arrives in Time" is displayed on the upper side of the image. The end credits text corresponding to the material can be displayed through a gradually unfolding animation effect. This embodiment does not limit the animation effect displayed in the end credits file.
[0200] In this embodiment, if the video editing application determines the intro text, the intro text is added to the intro screen of the edited video. Given that automatic editing templates vary, the intro segment of the edited video can be a slice of the video before it appears, or it can be a video segment corresponding to the first clip selected by the user; this embodiment does not limit this.
[0201] The opening title text can be displayed in the upper half, center, or lower half of the opening sequence; this embodiment does not limit this. In this embodiment, the display position, format, and animation effects of the opening title text can be determined based on the relevant settings of the automatic editing template; this embodiment does not limit this as well.
[0202] Taking the opening sequence of a video clip as an example (which is the first slice of the video), Figure 8a An exemplary diagram is shown of the opening sequence of a video clip. For example... Figure 8a As shown in Figure (1), at the moment t5 when the video is edited into a complete clip, the opening title text gradually begins to appear on the front slice of the video until the entire opening title text is displayed on the front slice of the video, as shown in Figure (1) of 8a. The display animation effects, display position, display format, etc. of the opening title text are related to the relevant settings of the automatic editing template.
[0203] Taking the opening sequence of a video edit as an example, which is the video segment corresponding to the first clip selected by the user, Figure 8b An exemplary diagram is shown of the opening sequence of a video clip. For example... Figure 8b As shown in Figure (1), at the moment t6 when the video is edited into a complete clip, the opening title text gradually begins to appear on the video clip screen until the entire opening title text is displayed on the video clip screen, as shown in Figure (1) of 8b. The display animation effects, display position, display format, etc. of the opening title text are related to the relevant settings of the automatic editing template.
[0204] As to other operations in the video clip synthesis (such as splicing, adding audio, rendering, etc.), reference can be made to the existing technology, which will not be described here.
[0205] After the video clip application generates the video clip piece, the electronic device displays the video clip piece, as shown in (3) of Figure 1a or (4) of Figure 1b If the user needs to adjust the template of the video clip piece, the template option 107 shown in (3) of Figure 1a may be clicked, and an automatic clip template is selected for the to-be-clip material according to the user's own needs. In this way, the video clip application can regenerate the video clip piece according to the automatic clip template selected by the user.
[0206] The user's re-selection of the automatic clip template will not affect the opening text and / or the ending text determined by the video clip template. Assuming that the aspect ratio of the automatic clip template re-selected by the user is changed, the layout of the reduced material picture and the ending text in the ending picture of the material in the video clip piece regenerated by the video clip application is adjusted accordingly. For example, before the user re-selects the automatic clip template, the video clip application generates a video clip piece 1 based on an automatic clip template with a horizontal aspect ratio. In the ending picture of the video clip piece 1, the reduced material picture and the ending text are distributed left and right, for example, the reduced material picture is on the left side and the ending text is on the right side. Assuming that the aspect ratio of the automatic clip template re-selected by the user is vertical, the video clip application generates a video clip piece 2 based on an automatic clip template with a vertical aspect ratio. In the ending picture of the video clip piece 2, the reduced material picture and the ending text are distributed up and down, for example, the reduced material picture is on the lower side and the ending text is on the upper side. It can be understood that the content of the ending text in the ending picture of the video clip piece 1 is the same as that of the ending text in the ending picture of the video clip piece 2.
[0207] The above takes the opening material and the ending material in the video clip piece as an example. The video clip template calls the media processing module to generate a plurality of material candidate texts corresponding to the opening material and the ending material, and selects one of the plurality of material candidate texts to add to the corresponding segment of the video clip piece. For the intermediate material in the video clip piece, the video clip template also calls the media processing module to generate a plurality of material candidate texts corresponding to the intermediate material, and selects one of the plurality of material candidate texts to add to the corresponding video segment of the video clip piece, which will not be described here.
[0208] Scenario Two
[0209] In the present scenario, when the electronic device automatically clips the multimedia material selected by the user, the paragraph text (which can be referred to as a caption) corresponding to at least one video segment (a video segment corresponding to a video material or a picture material) can be generated according to the video theme type, and the paragraph text and the audio corresponding to the paragraph text are added to the video clip to improve the effect of the clip video obtained by the one-key clipping function, and improve the user experience.
[0210] As Figure 9 shown is an interaction diagram of each module. Referring to Figure 9 , the flow of the video clipping method provided by the embodiments of the present application specifically includes:
[0211] S501, in response to a one-key clipping operation of the user, the gallery application inputs the material selected by the user into the video clipping application.
[0212] S502, the video clipping application performs offline analysis on the material selected by the user.
[0213] Continuing to refer to Figure 9 shown, the offline analysis of the video clipping application on the material selected by the user can specifically be:
[0214] S5021, the video clipping application calls the media processing module to perform offline analysis on the material selected by the user.
[0215] In the present embodiment, the offline analysis of the material can rely on the media processing module (or
[0216] S5022, under the call of the video clipping application, the material analysis scheduling decision module in the media processing module calls the visual analysis module to perform visual analysis on the material selected by the user, and obtains the visual analysis result of the material.
[0217] S5023, the material analysis scheduling decision module in the media processing module analyzes the highlight video segment in the video material.
[0218] S5024, the material analysis scheduling decision module in the media processing module calls the wisdom processing module to identify the video theme type.
[0219] S5025, the wisdom processing module performs semantic understanding on the video frame image and / or picture input by the media processing module, and generates the corresponding description text.
[0220] S5026, the wisdom processing module identifies the video theme type according to the description text corresponding to the video frame image and / or picture, and returns the identified theme type to the media processing module.
[0221] S5027, the material analysis result is returned to the video clip application by the material analysis scheduling decision module in the media processing module.
[0222] The material analysis result can include, but is not limited to, a video theme type, information of a highlight video segment, and the like.
[0223] S503, the video clip application selects a video clip template, and downloads the template and clip-related materials in the material shelving module.
[0224] S504, the video clip application determines the text-to-speech for at least one video segment according to the video clip template.
[0225] The text-to-speech can be understood as a text paragraph used to describe a video segment, and the speech-to-text refers to a piece of audio matched with the text-to-speech.
[0226] The video clip application can select one of the video segments according to preset information, and determine the text-to-speech for the video segment, or determine the text-to-speech for each video segment respectively, which is not limited in the embodiment.
[0227] Continuing to refer to Figure 9 The video clip application determines the text-to-speech for at least one video segment according to the video clip template, which can be specifically:
[0228] S5041, the video clip application calls the text-to-speech module in the media processing module to generate the text-to-speech for at least one video segment.
[0229] Exemplarily, the text-to-speech module in the media processing module provides an external calling interface, and the calling parameters can include, but are not limited to, video segment information, information of the video clip template selected by the video clip application, and the like.
[0230] S5042, the text-to-speech module in the media processing module analyzes the constraint conditions of the video clip template, and calls the wisdom processing module to generate the text-to-speech for at least one video segment.
[0231] After the video clip application selects the video clip template, the duration of each video segment in the video clip is also fixed. Then, the text-to-speech module can inversely deduce the word limit of the paragraph text corresponding to each video segment, the duration limit of the speech, and the like according to the duration of each video segment.
[0232] Exemplarily, the constraint conditions of the video clip template can include, but are not limited to, the word constraint of the paragraph text corresponding to each video segment, the duration constraint of the speech, and the like.
[0233] In this embodiment, when the text-to-speech module calls the intelligent processing module to generate the text-to-speech of the at least one video segment, the constraint condition of the video editing template can be sent to the intelligent processing module or be passed to the intelligent processing module as a calling parameter.
[0234] Optionally, when the text-to-speech module calls the intelligent processing module to generate the text-to-speech of the at least one video segment, information of the at least one video segment can also be passed to the intelligent processing module.
[0235] S5043, the intelligent processing module generates the paragraph text corresponding to the at least one video segment.
[0236] Under the calling of the video editing application, the intelligent processing module refers to the constraint condition of the video editing template to generate the paragraph text corresponding to the at least one video segment. For example, the intelligent processing module can call a related neural network model to generate the paragraph text corresponding to the at least one video segment. Illustratively, the paragraph text can be a subtitle of the video material, a text describing the video segment by the user, or a voice-over description corresponding to the video segment, etc. In summary, the paragraph text corresponding to a video segment is used to describe the video segment, and the content and style of the paragraph text are not limited in this embodiment.
[0237] S5044, the intelligent processing module calls a TTS service to generate the audio corresponding to the paragraph text and returns the text-to-speech of the at least one video segment to the text-to-speech module in the media processing module.
[0238] In this embodiment, the intelligent processing module first determines the target audio timbre and generates the audio matching the paragraph text according to the target audio timbre. When the intelligent processing module generates the audio matching the paragraph text according to the target audio timbre, the duration of the audio matching the paragraph text needs to meet the constraint condition of the video editing template. Optionally, the duration of the text-to-speech corresponding to a video segment can be equal to the duration of the video segment.
[0239] Illustratively, the intelligent processing module can call other intelligent modules or determine the gender of the owner of the electronic device through related big data and take the audio timbre corresponding to the gender of the owner of the electronic device as the target audio timbre. For example, when the gender of the owner of the electronic device is male, the intelligent processing module takes the male timbre as the target audio timbre. For another example, if the intelligent processing module cannot determine the gender of the owner of the electronic device, the intelligent processing module can take the female timbre as the target audio timbre by default.
[0240] For another example, the intelligent processing module can also determine a popular timbre (such as an authorized star timbre, a funny timbre, etc.) through querying big data, etc. and take any one of the timbres or a timbre selected according to a preset strategy as the target audio timbre.
[0241] S5045, the text-to-speech module in the media processing module returns the text-to-speech and insertion position of the video clip to the video clip application.
[0242] The text-to-speech module determines the text insertion position (or addition position) and the speech insertion position of each video clip, and transmits the text-to-speech of each video clip to the video clip application.
[0243] For example, the text insertion position can include only a text insertion start position, or both a text insertion start position and a text insertion end position, which is not limited in the embodiment.
[0244] For example, the speech insertion position can include only a speech insertion start position, or both a speech insertion start position and a speech insertion end position, which is not limited in the embodiment.
[0245] In an optional embodiment, the text insertion start position and the speech insertion start position of a certain video clip are the same as the start position of the video clip.
[0246] S305, the video clip application synthesizes the video clip into a video according to the video clip template.
[0247] During the process of synthesizing the video clip into a video by the video clip application, if the text-to-speech has been generated for a certain video clip, the corresponding passage text is inserted according to the text insertion position corresponding to the video clip, and the corresponding speech audio is inserted according to the speech insertion position corresponding to the video clip.
[0248] Optionally, the speech audio of the video clip can coexist with the background music of the video. For example, when the speech audio is played in the video, the volume of the background music can be automatically reduced to highlight the speech audio.
[0249] Figure 10 For example, the text-to-speech automatic editing effect for the video clip X is shown. As shown in Figure 10 During the playing period of the video clip X, the passage text corresponding to the video clip X is automatically added, and the speech audio corresponding to the passage text is automatically added. The speech audio of the video clip X can coexist with the background music of the video.
[0250] For other operations (such as splicing, rendering, etc.) involved in the video clip synthesis process, please refer to the prior art, which will not be described here. For the details not explained in the process, please refer to the description of scenario one in the foregoing, which will not be described here.
[0251] In the one-key automatic video clipping method provided in the present application, the solution of automatically adding end-credit text and / or beginning-credit text that is consistent with the video content in the video clip in scene one can coexist with the solution of automatically adding text and dubbing for the video clip in the video clip in scene two, and the present embodiment will not repeat the description.
[0252] In addition, for the function of automatically adding end-credit text that is consistent with the video content in the one-key automatic clipping scene, the setting option of the video clipping application can include a switch control 1 corresponding thereto, and the user can set the on-off state of the switch control 1 according to personal needs to turn on or turn off the function of automatically adding end-credit text that is consistent with the video content in the one-key automatic clipping scene.
[0253] Similarly, for the function of automatically adding beginning-credit text that is consistent with the video content in the one-key automatic clipping scene, the setting option of the video clipping application can include a switch control 2 corresponding thereto, and the user can set the on-off state of the switch control 2 according to personal needs to turn on or turn off the function of automatically adding beginning-credit text that is consistent with the video content in the one-key automatic clipping scene.
[0254] Similarly, for the function of automatically adding text and dubbing for the video clip in the one-key automatic clipping scene, the setting option of the video clipping application can include a switch control 3 corresponding thereto, and the user can set the on-off state of the switch control 3 according to personal needs to turn on or turn off the function of automatically adding text and dubbing for the video clip in the one-key automatic clipping scene.
[0255] The present embodiment also provides a computer storage medium, which stores computer instructions, and when the computer instructions run on an electronic device, the electronic device executes the related method steps to implement the video clipping method in the above embodiment.
[0256] The present embodiment also provides a computer program product, which, when running on a computer, causes the computer to execute the related steps to implement the video clipping method in the above embodiment.
[0257] In addition, the present embodiment also provides a device, which can be a chip, a component or a module. The device can include a processor and a memory connected to each other. The memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory to make the chip execute the video clipping method in the above method embodiments.
[0258] The electronic device (such as a mobile phone), the computer storage medium, the computer program product or the chip provided by the embodiment are used to execute the corresponding method provided above, and the beneficial effects achieved by the electronic device (such as a mobile phone), the computer storage medium, the computer program product or the chip can refer to the beneficial effects of the corresponding method provided above, and details are not described herein.
[0259] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0260] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other ways. For example, the apparatus embodiments described above are merely illustrative, for example, the division of modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between the devices or units, which can be electrical, mechanical or other forms.
[0261] The above-described embodiments are merely used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A video clip method characterized by, The method comprises: displaying a first interface; the first interface displays multimedia material selected by a user; in response to a first operation, displaying a second interface; the second interface includes a first clip video corresponding to the multimedia material; when playing a trailer video segment of the first clip video, displaying at least one first video frame, the first video frame including a playing picture of a trailer material and a first text related to the content of the trailer material, the first text being a text corresponding to the trailer picture material, and the first text being used to describe the content of the trailer picture.
2. The method of claim 1, wherein, The size of the playing picture is a target size, which is smaller than the size of the playing picture of the first clip video; when playing the trailer video segment of the first clip video, and before displaying the at least one first video frame, the method further comprises: the size of the playing picture of the trailer material is gradually reduced from the size of the playing picture of the first clip video to the target size.
3. The method of claim 2, wherein, The aspect ratio of the first clip video is horizontal; in the first video frame, the playing picture of the trailer material and the first text are arranged side by side.
4. The method of claim 2, wherein, The size of the playing picture of the trailer material is gradually reduced from the size of the playing picture of the first clip video to the target size, comprising: the size of the playing picture of the trailer material is gradually reduced from the size of the playing picture of the first clip video to a first size on the left; the first size is smaller than the size of the playing picture of the first clip video; in the first video frame, the first text is displayed on the right side of the playing picture of the trailer material.
5. The method of claim 2, wherein, The aspect ratio of the first clip video is vertical; in the first video frame, the playing picture of the trailer material and the first text are arranged vertically.
6. The method of claim 5, wherein, The size of the playing picture of the trailer material is gradually reduced from the size of the playing picture of the first clip video to the target size, comprising: the size of the playing picture of the trailer material is gradually reduced from the size of the playing picture of the first clip video to a second size on the bottom; the second size is smaller than the size of the playing picture of the first clip video; in the first video frame, the first text is displayed on the top side of the playing picture of the trailer material.
7. The method of claim 2, wherein, After the size of the playing picture of the trailer material is gradually reduced from the size of the playing picture of the first clip video to the target size, the method further comprises: in the first video frame, the first text is gradually displayed according to a first animation effect.
8. The method of claim 1, wherein, Before displaying the second interface, the method further comprises: determining trailer material; generating a plurality of trailer candidate texts according to the trailer material; selecting the first text from the plurality of trailer candidate texts.
9. The method of claim 8, wherein, Generating a plurality of trailer candidate texts according to the trailer material, comprising: when the trailer material is a picture material, performing picture-to-text processing on the picture material to obtain the plurality of trailer candidate texts; when the trailer material is a video material, performing frame extraction processing on the video material to obtain a second video frame, and performing picture-to-text processing on the second video frame to obtain the plurality of trailer candidate texts.
10. The method of claim 8, wherein, Selecting the first text from the plurality of trailer candidate texts, comprising: According to the script screening strategy, the first script is screened from the plurality of trailer candidate scripts; The script screening strategy at least includes one of the following: a text character number limit condition, a punctuation mark limit condition, a de-duplication processing, and a text style limit condition; and a customized style corresponding to the style limit condition supports user setting.
11. The method according to any one of claims 1 to 10, characterized in that, After displaying the second interface, further comprising: In response to a second operation, a third interface is displayed; the second interface includes a second clip video corresponding to the multimedia material; the clip template of the second clip video is different from the clip template of the first clip video in the frame direction; In the first clip video, the playing picture of the trailer material and the first script are arranged in a first direction, and in the second clip video, the playing picture of the trailer material and the first script are arranged in a second direction; wherein the first direction is a left-right direction or an up-down direction, and the second direction is a corresponding up-down direction or left-right direction.
12. The method according to any one of claims 1 to 10, characterized in that, Further comprising: When playing a trailer video segment of the first clip video, at least one third video frame is displayed, and the third video frame includes a second script related to the trailer material.
13. The method of claim 12, wherein, In the third video frame, the second script is gradually displayed according to a second animation effect.
14. The method according to any one of claims 1 to 10, characterized in that, Further comprising: When playing a target video segment of the first clip video, at least one fourth video frame is displayed, and the fourth video frame includes a target text corresponding to the target video segment, and a target audio corresponding to the target text is played.
15. The method of claim 14, wherein, Before displaying the second interface, further comprising: Parsing a constraint condition of a clip template of the first clip video; According to the constraint condition, the target text is generated according to the video segment; According to the target text, the target audio is generated, and the insertion position of the target text and the target audio is determined.
16. The method of claim 15, wherein, According to the target text, the target audio is generated, comprising: Selecting a target tone; According to the target tone, the target audio is generated.
17. An electronic device, comprising: Comprising: One or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored on the memory, and when the computer programs are executed by the one or more processors, the electronic device executes the video clip method as claimed in any one of claims 1-16.
18. A computer readable storage medium comprising a computer program, characterized in that, When the computer programs run on the electronic device, the electronic device executes the video clip method as claimed in any one of claims 1-16.
Citation Information
Patent Citations
Text generation method and device
CN110851622A
Video generation method and device based on AI and electronic equipment
CN115460459A
Video dubbing method and related device, electronic equipment and storage medium
CN117177024A