Video generation method, device, electronic device and storage medium

By displaying the background audio's card points on the video editing interface and automatically adjusting the position and duration of the video clips, the tedious problem of manually aligning the video clips and background audio card points in the existing technology is solved, and efficient automatic alignment and rhythm matching of video generation are achieved.

CN120281993BActive Publication Date: 2025-09-09BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510704479.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-09
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In the existing video generation process, users need to manually align the card points of the video clip and the background audio, which is cumbersome and affects the generation efficiency.

Method used

Multiple card points of background audio are displayed on the video editing interface, and the position and duration of the video clip are automatically adjusted according to the changes in the card point density, so as to achieve automatic alignment of the video clip and the background audio.

Benefits of technology

This simplifies the video generation process, improves operational efficiency, ensures that the rhythm of video clips matches the background audio, and reduces manual operation steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281993B_ABST
    Figure CN120281993B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video generation method, device, electronic device, and storage medium, belonging to the field of multimedia technology. The method includes: displaying multiple first card points of background audio on a video editing interface, and displaying multiple video clips at positions indicated by the time intervals between the multiple first card points; displaying multiple second card points of the background audio in response to a card point switching operation, wherein the density of the multiple second card points is different from the density of the multiple first card points; updating at least one of the positions and durations of the multiple video clips based on the multiple second card points; and generating a video based on the updated multiple video clips and the background audio. The above method not only meets the user's requirements for video rhythm, but is also simple to operate and does not require manual alignment by the user, thereby improving video generation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of multimedia technology, and in particular to a video generation method, device, electronic device, and storage medium. Background Art

[0002] With the development of multimedia technology, more and more users are accustomed to sharing things through videos. Before publishing a video, users usually generate the video through a video editing application.

[0003] During the video generation process, users usually need to manually add background audio to the video, manually add check points on the audio track, manually mark the highlights in the video clip on the video track, and finally manually align the highlights in the video clip with the check points to generate a video that matches the rhythm of the video image and the background audio.

[0004] However, the above technical solution requires the user to perform multiple manual operations, which is cumbersome and affects the efficiency of video generation. Summary of the Invention

[0005] The present disclosure provides a video generation method, device, electronic device, and storage medium that not only meet the user's requirements for video rhythm, but also are easy to operate and do not require manual alignment by the user, thereby improving video generation efficiency. The technical solution of the present disclosure is as follows.

[0006] According to one aspect of an embodiment of the present disclosure, a video generation method is provided, including:

[0007] On the video editing interface, multiple first card points of the background audio are displayed, and multiple video clips are displayed at positions indicated by time intervals between the multiple first card points, where each first card point indicates a playback moment of the background audio;

[0008] In response to the card point switching operation, displaying a plurality of second card points of the background audio, wherein the density of the plurality of second card points is different from the density of the plurality of first card points, and each second card point is used to indicate a playback moment of the background audio;

[0009] Based on the plurality of second checkpoints, displaying that at least one of the positions and durations of the plurality of video segments is updated;

[0010] A video is generated based on the updated plurality of video clips and the background audio.

[0011] According to another aspect of the present disclosure, there is provided a video generating apparatus, including:

[0012] The display unit is configured to display a plurality of first card points of the background audio on a video editing interface, and display a plurality of video clips at positions indicated by time intervals between the plurality of first card points, wherein each first card point indicates a playback moment of the background audio;

[0013] The display unit is further configured to display a plurality of second card points of the background audio in response to a card point switching operation, wherein a density of the plurality of second card points is different from a density of the plurality of first card points, and each second card point is used to indicate a playback moment of the background audio;

[0014] The display unit is further configured to display an update of at least one of positions and durations of the plurality of video segments based on the plurality of second card points;

[0015] The generating unit is configured to generate a video based on the updated plurality of video clips and the background audio.

[0016] In some embodiments, the display unit is configured to perform at least one of the following:

[0017] For any video segment among the multiple video segments, when the time interval corresponding to the video segment changes from a first time interval to a second time interval, the position of displaying the video segment is updated from the position indicated by the first time interval to the position indicated by the second time interval, where the first time interval is the time interval corresponding to the video segment between the multiple first card points, and the second time interval is the time interval corresponding to the video segment between the multiple second card points;

[0018] If the duration of the first time interval is different from the duration of the second time interval, the duration of displaying the video clip is updated from the duration of the first time interval to the duration of the second time interval.

[0019] In some embodiments, the display unit is configured to execute, for any video clip among the multiple video clips, when the duration of the first time interval is shorter than the duration of the second time interval, updating the playback speed of the video clip from the first speed to the second speed, the time consumed to complete the playback of the video clip based on the first speed is the duration of the first time interval, and the time consumed to complete the playback of the video clip based on the second speed is the length of the second time interval, and the second speed is slower than the first speed.

[0020] In some embodiments, the display unit is configured to perform any of the following:

[0021] For any video segment among the multiple video segments, when the duration of the first time interval is longer than the duration of the second time interval, displaying that the video segment is updated to a target video segment, where the target video segment is a portion of the video segment, the duration of the video segment is the duration of the first time interval, and the duration of the target video segment is the duration of the second time interval;

[0022] For any video clip among the multiple video clips, when the duration of the first time interval is longer than the duration of the second time interval, the playback speed of the video clip is updated from the first speed to the third speed, the time consumed for the video clip to be played back based on the first speed is the duration of the first time interval, and the time consumed for the video clip to be played back based on the third speed is the length of the second time interval, and the third speed is faster than the first speed.

[0023] In some embodiments, for any video segment among the multiple video segments, if the number of time intervals between the multiple second card points is less than the number of the multiple video segments, the number of second time intervals corresponding to the video segment is one, and the duration of the video segment is less than or equal to the duration of the second time interval;

[0024] When the number of time intervals between the multiple second card points is greater than the number of the multiple video clips, the number of second time intervals corresponding to the video clips is at least one, and the duration of the video clip is equal to the total duration of the at least one second time interval.

[0025] In some embodiments, for any video segment among the multiple video segments, when the video segment meets a preset condition, the duration of the second time interval corresponding to the video segment reaches a duration threshold, where the preset condition is a condition that must be met for the display duration of the video segment to meet the standard;

[0026] When the video segment does not meet the preset condition, the duration of the second time interval corresponding to the video segment is lower than the duration threshold.

[0027] In some embodiments, the preset condition includes any one of the following:

[0028] The object in the video clip is in motion;

[0029] At least one of the value of the motion vector in the video clip, the color complexity, the number of scenes, the number of subjects, the number of words in the subtitles, and the weight of the video clip reaches a corresponding preset value.

[0030] In some embodiments, the display unit is further configured to display a card point switching control in the video editing interface, where the card point switching control is used to switch between card points of different densities; in response to a triggering operation on the card point switching control, display multiple card point modes, where the density of card points indicated by different card point modes is different; and in response to a selection operation on any card point mode among the multiple card point modes, display multiple second card points indicated by the card point mode.

[0031] In some embodiments, the display unit is further configured to display card point prompt information when at least one material among the multiple video clips and the background audio changes, and the card point prompt information is used to prompt whether to reallocate based on the changed material; in response to a confirmation operation in the card point prompt information, at least one of the positions and durations of the multiple video clips is displayed to be updated based on the card point of the background audio after the material changes.

[0032] In some embodiments, the background audio includes at least one of the following:

[0033] background audio that matches the style of the plurality of video clips;

[0034] Background audio whose heat meets the heat threshold;

[0035] Background audio having a duration matching the total duration of the plurality of video clips;

[0036] Background audio having a theme matching the themes of the plurality of video clips.

[0037] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, the electronic device including:

[0038] one or more processors;

[0039] a memory for storing program codes executable by the processor;

[0040] The processor is configured to execute the program code to implement the above-mentioned video generation method.

[0041] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When a program code in the computer-readable storage medium is executed by a processor of an electronic device, the electronic device is enabled to perform the above-mentioned video generating method.

[0042] According to another aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program / instruction, which implements the above-mentioned video generation method when executed by a processor.

[0043] The solution provided by the embodiments of the present disclosure can, in the process of generating a video based on multiple video clips and background audio, reallocate the video clips when the card point density of the background audio changes, and automatically allocate the multiple video clips to the positions indicated by the time intervals between the updated card points, so that the video clips and the latest time intervals are aligned with each other in terms of duration. That is, the multiple second card points in the background audio are automatically aligned with the multiple video clips. This not only ensures that the video clips can be switched at the positions of the second card points in the subsequently generated video, meeting the user's requirements for video rhythm, but also is simple to operate. There is no need for manual alignment by the user. The rhythm of the subsequently generated video can be controlled by simply switching the card point density, which is conducive to improving video generation efficiency.

[0044] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0046] Figure 1 The figure is a schematic diagram showing an implementation environment of a video generation method according to an exemplary embodiment.

[0047] Figure 2 The figure is a flowchart of a video generating method according to an exemplary embodiment.

[0048] Figure 3 The figure is a flowchart of another method for generating a video according to an exemplary embodiment.

[0049] Figure 4 FIG. 4 is a schematic diagram showing a first card point according to an exemplary embodiment.

[0050] Figure 5 FIG. 4 is a schematic diagram showing a switching card point density according to an exemplary embodiment.

[0051] Figure 6 is a flowchart of yet another video generating method according to an exemplary embodiment.

[0052] Figure 7 The figure is a schematic diagram showing a method of allocating video segments to time intervals according to an exemplary embodiment.

[0053] Figure 8 FIG. 2 is a schematic diagram showing another method of allocating video segments to time intervals according to an exemplary embodiment.

[0054] Figure 9The figure is a schematic diagram showing a method of generating prompt information according to an exemplary embodiment.

[0055] Figure 10 The figure is a schematic diagram showing a method of trimming an audio segment according to an exemplary embodiment.

[0056] Figure 11 The figure is a schematic diagram showing a method of reallocating video segments according to an exemplary embodiment.

[0057] Figure 12 The figure is a block diagram of a video generating apparatus according to an exemplary embodiment.

[0058] Figure 13 It is a block diagram of a terminal according to an exemplary embodiment.

[0059] Figure 14 The figure is a block diagram of a server according to an exemplary embodiment. DETAILED DESCRIPTION

[0060] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0061] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0062] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the video clips and background audio involved in this disclosure were obtained with full authorization.

[0063] Figure 1 FIG. 1 is a schematic diagram showing an implementation environment of a video generation method according to an exemplary embodiment. Taking the electronic device as a terminal as an example, see Figure 1The implementation environment specifically includes: a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected via wired or wireless communication, which is not limited in the present disclosure.

[0064] Terminal 101 is at least one of a smartphone, a smartwatch, a desktop computer, a laptop, an MP3 player, an MP4 player, and a portable computer. Terminal 101 runs an application that supports video generation. This application can be an editing application or a multimedia application, which is not limited in the present embodiment. A user can log in to this application through terminal 101 to access the services provided by the application. For example, a user can generate a video through this application on terminal 101. During the video generation process, terminal 101 can align the multiple video clips and background audio imported by the user with the card points in the background audio to generate a video that matches the video frame and the card points.

[0065] Terminal 101 generally refers to one of multiple terminals. This embodiment uses terminal 101 as an example. Those skilled in the art will appreciate that the number of terminals may be greater or lesser. For example, there may be a few terminals, or dozens, hundreds, or even more. This embodiment does not limit the number or device type of terminals.

[0066] Server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 is used to provide background services for applications that support video generation. Server 102 can be connected to terminal 101 via a wireless network or a wired network. Server 102 can match background audio for multiple video clips imported by the user and identify the card points in the background audio so that terminal 101 can subsequently align the multiple video clips with the card points in the background audio to generate a video. In some embodiments, the number of the above-mentioned servers can be more or less, and this is not limited in the embodiments of the present disclosure. Of course, server 102 also includes other functional servers to provide more comprehensive and diverse services.

[0067] Figure 2 is a flow chart of a video generation method according to an exemplary embodiment. Figure 2 ,The video generation method is applied in the terminal and includes the following steps.

[0068] In step 201, the terminal displays multiple first card points of the background audio on the video editing interface, and displays multiple video clips at positions indicated by time intervals between the multiple first card points, where each first card point is used to indicate a playback moment in the background audio.

[0069] In the disclosed embodiment, during the video generation process, the terminal displays a video editing interface to facilitate subsequent video generation within the video editing interface. In response to a video import operation within the video editing interface, the terminal displays the imported multiple video clips and background audio within the video editing interface. The disclosed embodiment does not limit the number of multiple video clips or their display locations.

[0070] The background audio can be audio imported by the user in the video editing interface, or it can be audio automatically matched to multiple video clips, which is not limited in the present embodiment. While displaying the background audio, the terminal can also display multiple first card points of the background audio. The multiple first card points can be the playback moments of the rhythm points of the background audio, or any playback moments in the background audio, which is not limited in the present embodiment. The multiple first card points are user-defined, or they can be determined by rhythm recognition of the background audio, which is not limited in the present embodiment.

[0071] There is a time interval between adjacent first card points. Multiple video clips can be assigned to the time intervals between the multiple first card points specified by the user. That is, the multiple video clips can be manually displayed at the locations indicated by the time intervals between the multiple first card points through user operation. Alternatively, the multiple video clips can be automatically assigned to the time intervals between the multiple first card points by a device. For example, a terminal or server can assign video clips from multiple video clips to the time intervals between the multiple first card points. This embodiment of the disclosure is not limited in this regard. After the assignment is completed, the terminal displays the video clips assigned to each time interval at the locations indicated by the time intervals between the multiple first card points. This embodiment of the disclosure does not limit the method of assigning (matching) the video clips to the time intervals.

[0072] In step 202, in response to the card point switching operation, the terminal displays multiple second card points of the background audio, the density of the multiple second card points is different from the density of the multiple first card points, and each second card point is used to indicate a playback moment in the background audio.

[0073] In the disclosed embodiment, the time intervals between card points match those between video segments. Accordingly, the card point density can determine the rhythm of the video segment, including the playback rhythm within the video segment and the switching rhythm between video segments. Users can adjust the density of background audio card points based on their desired video rhythm. In response to a card point switching operation, the terminal replaces the multiple first card points in the video editing interface with the multiple second card points indicated by the card point switching operation. The density of the multiple second card points is different from the density of the multiple first card points. Accordingly, the multiple second card points differ from the multiple first card points in multiple dimensions, such as number, location (i.e., the indicated playback time), and time interval.

[0074] In step 203, the terminal displays, based on the plurality of second checkpoints, that at least one of the positions and durations of the plurality of video segments is updated.

[0075] In the disclosed embodiment, after switching to multiple second card points, the multiple video clips can be reallocated to the positions indicated by the time intervals between the multiple second card points. The terminal then displays the multiple video clips at the positions indicated by the time intervals between the multiple second card points. Because the multiple second card points differ from the multiple first card points in terms of number, position (i.e., indicated playback time), and time interval, at least one of the position and duration of the multiple video clips will change.

[0076] In step 204, the terminal generates a video based on the updated multiple video clips and background audio.

[0077] In an embodiment of the present disclosure, in response to a video generation operation in a video editing interface, the terminal synthesizes the video segments allocated at the time intervals between the plurality of second card points with the background audio to generate a video. The generated video sequentially plays the video segments allocated at each time interval between the plurality of second card points in the order of the plurality of second card points.

[0078] The disclosed embodiments provide a video generation method. In the process of generating a video based on multiple video clips and background audio, when the card point density of the background audio changes, the video clips can be reallocated, and the multiple video clips can be automatically allocated to the positions indicated by the time intervals between the updated card points, so that the video clips are aligned with the latest time intervals. That is, the multiple second card points in the background audio are automatically aligned with the multiple video clips. This not only ensures that the video clips can be switched at the positions of the second card points in the subsequently generated video, which meets the user's requirements for the video rhythm, but also is simple to operate. There is no need for manual alignment by the user. The rhythm of the subsequently generated video can be controlled by switching the card point density, which is conducive to improving the efficiency of video generation.

[0079] In some embodiments, based on the plurality of second card points, at least one of the positions and durations of the plurality of video segments is updated, including at least one of the following:

[0080] For any video segment among the multiple video segments, when the time interval corresponding to the video segment changes from a first time interval to a second time interval, the position of the displayed video segment is updated from the position indicated by the first time interval to the position indicated by the second time interval, where the first time interval is the time interval corresponding to the multiple first card points in the video segment, and the second time interval is the time interval corresponding to the multiple second card points in the video segment;

[0081] If the duration of the first time interval is different from the duration of the second time interval, the duration of the displayed video segment is updated from the duration of the first time interval to the duration of the second time interval.

[0082] The solution provided by the embodiment of the present disclosure can reallocate each video clip among multiple video clips to the position indicated by the time interval between the updated card points, so that the video clips are aligned with the latest time interval. That is, the time intervals between multiple second card points in the background audio are automatically aligned with the multiple video clips. This not only ensures that video clips can be switched at the positions of the second card points in subsequently generated videos, meeting the user's requirements for video rhythm, but also is simple to operate. There is no need for manual alignment by the user. The rhythm of subsequently generated videos can be controlled by simply switching the card point density, which is conducive to improving video generation efficiency.

[0083] In some embodiments, if the duration of the first time interval is different from the duration of the second time interval, updating the duration of the displayed video segment from the duration of the first time interval to the duration of the second time interval includes:

[0084] For any video clip among multiple video clips, when the duration of the first time interval is shorter than the duration of the second time interval, the playback speed of the displayed video clip is updated from the first speed to the second speed, the time consumed for the video clip to be played back based on the first speed is the duration of the first time interval, and the time consumed for the video clip to be played back based on the second speed is the length of the second time interval, and the second speed is slower than the first speed.

[0085] The solution provided by the embodiment of the present disclosure is that after the video clip is assigned to the matching time interval, if the duration of the video clip is less than the duration of the matching time interval, the playback speed of the video clip can be slowed down so that the time consumed to complete the playback of the video clip based on the slower playback speed is equal to the duration of the matching time interval. That is, by moderately slowing down the playback speed of the video clip, the playback duration of the video clip is aligned with the duration of the time interval, ensuring that the corresponding video clip can be displayed in each time interval, and video clip switching can be achieved at the card point, which is conducive to improving the rhythm matching of the generated video, that is, it is conducive to improving the quality of the generated video.

[0086] In some embodiments, if the duration of the first time interval is different from the duration of the second time interval, updating the duration of the displayed video segment from the duration of the first time interval to the duration of the second time interval includes any of the following:

[0087] For any video segment among the multiple video segments, when the duration of the first time interval is longer than the duration of the second time interval, the displayed video segment is updated to a target video segment, where the target video segment is a portion of the video segment, the duration of the video segment is the duration of the first time interval, and the duration of the target video segment is the duration of the second time interval;

[0088] For any video clip among multiple video clips, when the duration of the first time interval is longer than the duration of the second time interval, the playback speed of the video clip is updated from the first speed to the third speed, the time consumed for the video clip to be played back based on the first speed is the duration of the first time interval, and the time consumed for the video clip to be played back based on the third speed is the duration of the second time interval, and the third speed is faster than the first speed.

[0089] The solution provided by the embodiment of the present disclosure is that after the video clips are assigned to the matching time intervals, if the duration of the video clips is longer than the duration of the matching time interval, it is possible to obtain part of the video clips or increase the playback speed of the video clips so that the time consumed to complete the playback of the adjusted video clips is equal to the duration of the matching time interval, that is, by appropriately cropping or increasing the playback speed of the video clips, the playback duration of the video clips is aligned with the duration of the time interval, thereby ensuring that the corresponding video clips can be displayed in each time interval and realizing video clip switching at the card point, which is beneficial to improving the rhythm matching of the generated video, that is, it is beneficial to improving the quality of the generated video.

[0090] In some embodiments, for any video segment among the multiple video segments, when the number of time intervals between the multiple second card points is less than the number of the multiple video segments, the number of second time intervals corresponding to the video segment is one, and the duration of the video segment is less than or equal to the duration of the second time interval;

[0091] When the number of time intervals between the multiple second card points is greater than the number of the multiple video clips, the number of second time intervals corresponding to the video clips is at least one, and the duration of the video clip is equal to the total duration of the at least one second time interval.

[0092] The solution provided by the embodiment of the present disclosure can automatically assign at least one video clip to a matching time interval when the time intervals between background audio card points are small and the number of video clips is large, so that at least one video clip can be played within each time interval; and can automatically assign each video clip to at least one matching time interval when the time intervals between background audio card points are large and the number of video clips is small, so that each video clip can be displayed within at least one matching time interval; this not only ensures that each video clip has time to be displayed, but also ensures that video clip switching can be achieved at the position of the card point in the subsequently generated video. Compared with directly cutting certain video clips, the video generated by this solution has better integrity, that is, it can improve the video quality.

[0093] In some embodiments, for any video segment among the plurality of video segments, when the video segment satisfies a preset condition, the duration of the second time interval corresponding to the video segment reaches a duration threshold, where the preset condition is a condition that must be met for the display duration of the video segment to meet the standard;

[0094] When the video segment does not meet the preset condition, the duration of the second time interval corresponding to the video segment is lower than the duration threshold.

[0095] The solution provided by the embodiment of the present disclosure can automatically assign video clips to time intervals with longer durations when the video clips meet preset conditions, and can automatically assign video clips to time intervals with shorter durations when the video clips do not meet the preset conditions. This not only ensures that video clip switching can be achieved at the card point position in the subsequently generated video, thereby improving the rhythm matching of the generated video, but also ensures that video clips that need to be displayed for a long time have sufficient time to be displayed, and video clips that do not need to be displayed for a long time have a shorter display time, so that users can fully understand the information in each video clip without having to watch video clips with less information for a long time, which is beneficial to improving the accuracy and efficiency of information transmission in the generated video, thereby improving the quality of the generated video.

[0096] In some embodiments, the preset condition includes any of the following:

[0097] The objects in the video clip are in motion;

[0098] At least one of the value of the motion vector in the video clip, the color complexity, the number of scenes, the number of subjects, the number of words in the subtitles, and the weight of the video clip reaches a corresponding preset value.

[0099] The solution provided by the embodiment of the present disclosure can reflect to a certain extent that the video clip contains more information or is of higher importance because the objects in the video clip are in motion, or at least one of the numerical value of the motion vector, color complexity, number of scenes, number of subjects, number of words in the subtitles, and weight of the video clip in the video clip reaches a corresponding preset value. In this case, the video clip is automatically assigned to a time interval with a longer duration, which not only ensures that the video clip can be switched at the card point in the subsequently generated video and improves the rhythm matching of the generated video, but also allows video clips that need to be displayed for a long time to have sufficient time to be displayed, so that users can fully understand the information in each video clip, which is beneficial to improving the accuracy and efficiency of information transmission in the generated video, thereby improving the quality of the generated video.

[0100] In some embodiments, in response to the card point switching operation, displaying multiple second card points of the background audio includes:

[0101] Display the card point switching control in the video editing interface, which is used to switch card points of different densities;

[0102] In response to a triggering operation on a card point switching control, multiple card point modes are displayed, and the density of card points indicated by different card point modes is different;

[0103] In response to a selection operation on any one of the plurality of card point modes, a plurality of second card points indicated by the card point mode are displayed.

[0104] The solution provided by the embodiment of the present disclosure can provide users with multiple card point modes. Different card point modes indicate different card point densities, which makes it convenient for users to switch card point modes to adjust the density of card points of background audio according to their desired video rhythm. The operation is simple and can improve video generation efficiency.

[0105] In some embodiments, the method further comprises:

[0106] When at least one of the multiple video clips and background audio materials changes, a card point prompt is displayed, and the card point prompt is used to prompt whether to re-match the card points based on the changed material;

[0107] In response to a confirmation operation in the card point prompt information, at least one of the positions and durations of the multiple video clips is updated based on the card points of the background audio after the material changes.

[0108] The solution provided by the embodiments of the present disclosure displays a card point prompt when a video clip or background audio is changed, prompting the user whether to re-match the card points based on the changed material. If the user allows re-matching, the solution automatically re-assigns video clips to the current time intervals, or re-determines new card points and assigns video clips to the time intervals between the new card points. This eliminates the need for manual allocation by the user, simplifies the operation, and improves video generation efficiency.

[0109] In some embodiments, the background audio includes at least one of the following:

[0110] Background audio that matches the style of multiple video clips;

[0111] Background audio whose heat meets the heat threshold;

[0112] Background audio with a duration that matches the combined duration of the multiple video clips;

[0113] Background audio with a theme that matches the themes of multiple video clips.

[0114] The solution provided by the embodiments of the present disclosure can select background audio that matches the style of multiple video clips, or whose popularity meets the popularity threshold, or matches the duration of multiple video clips, or matches the theme of multiple video clips. This is beneficial to ensuring the compatibility between the background audio and multiple video clips, that is, it can improve the accuracy of the background audio, determine suitable background audio for multiple video clips, and thus improve the quality of the generated video.

[0115] above Figure 2 The following is only a basic process of the present disclosure. The solution provided by the present disclosure is further described based on a specific implementation method. Figure 3 FIG. 1 is a flow chart of another video generation method according to an exemplary embodiment. Assume that the electronic device is provided as a terminal as an example. Figure 3 , the method includes the following steps.

[0116] In step 301, the terminal displays multiple first card points of the background audio on the video editing interface, and displays multiple video clips at positions indicated by time intervals between the multiple first card points, where each first card point is used to indicate a playback moment in the background audio.

[0117] In the disclosed embodiment, a video track is displayed in the video editing interface. The video track is used to accommodate imported video clips. Accordingly, the terminal displays the imported multiple video clips in the video track in the video editing interface. In this case, the multiple video clips are arranged and displayed in the video track in the order in which they were imported.

[0118] The video editing interface also displays an audio track. The audio track is used to accommodate background audio added to a video clip. In response to adding audio to the audio track, the terminal displays the background audio and multiple first card points of the background audio in the audio track. The terminal then displays multiple video clips at positions indicated by the time intervals between the multiple first card points. The multiple video clips can be displayed at positions indicated by user-specified time intervals or automatically at positions indicated by device-specified time intervals, which is not limited in this embodiment of the present disclosure.

[0119] For example, Figure 4 FIG. 1 is a schematic diagram showing a first card point according to an exemplary embodiment. Figure 4 , a video track 401, an audio track 402, and a video generation control 403 are displayed in the video editing interface. The video track 401 includes multiple video clips, namely video clip 1, video clip 2, video clip 3, and video clip 4. Due to the limitation of the screen size, only video clip 1 and video clip 2 are currently displayed in the video track 401. Each video clip contains multiple frames of images. In response to the triggering operation of the video generation control 403, the terminal displays the background audio 404 and multiple first card points 405 of the background audio 404 in the audio track 402. Then, the terminal displays multiple video clips at the positions indicated by the time intervals between the multiple first card points 405. Figure 4 As shown, the terminal displays video clip 1 at the position indicated by the first card point interval, displays video clip 2 at the position indicated by the second card point interval, displays video clip 3 at the position indicated by the third card point interval, and displays video clip 4 at the position indicated by the fourth card point interval.

[0120] The background audio may be determined based on a plurality of video clips. In some embodiments, the background audio may be at least one of the following.

[0121] First, the background audio is background audio that matches the style of multiple video clips. The style of the background audio can be determined based on at least one of the rhythm of the background audio and the speech within the background audio. The style of the video clip can be determined based on at least one of the subtitles, subject matter, and background within the video clip. The disclosed embodiments do not limit the method for determining the style.

[0122] The second item is background audio whose popularity meets a popularity threshold. The popularity of background audio can be, for example, the number of times the background audio has been used, the number of times it has been collected, the number of likes, or the number of likes for videos containing the background audio, though this disclosure is not limited to this. Optionally, this solution uses audio that has reached a word threshold as background audio for multiple video clips.

[0123] Third, the background audio must match the duration of the video clips. Duration matching means the difference between the duration of the background audio and the total duration of the video clips is less than a preset difference. This embodiment of the disclosure does not limit the size of this preset difference.

[0124] Fourth, background audio is background audio whose theme matches the themes of multiple video clips. The theme of the background audio can be determined based on at least one of the rhythm of the background audio and the speech (such as lyrics) in the background audio. The theme of a video clip can be determined based on at least one of the subtitles, subject matter, and background in the video clip. This disclosed embodiment does not limit the method for determining the theme.

[0125] The solution provided by the embodiments of the present disclosure can select background audio that matches the style of multiple video clips, or whose popularity meets the popularity threshold, or matches the duration of multiple video clips, or matches the theme of multiple video clips. This is beneficial to ensuring the compatibility between the background audio and multiple video clips, that is, it can improve the accuracy of the background audio, determine suitable background audio for multiple video clips, and thus improve the quality of the generated video.

[0126] In the embodiment of the present disclosure, the multiple first card points displayed by the terminal can be multiple card points (rhythm points) directly identified based on background audio, or can be card points selected from multiple directly identified card points. The embodiment of the present disclosure is not limited to this.

[0127] In some embodiments, the terminal obtains multiple initial card points of background audio. The multiple initial card points are obtained based on background audio recognition. The terminal then displays multiple first card points from the multiple initial card points. The time interval between adjacent first card points (card point interval) reaches an interval threshold. For example, the time interval between adjacent first card points reaches 0.1 seconds. The solution provided by the disclosed embodiment can filter out card points that are relatively close to each other among the multiple initial card points, ensuring that there is sufficient time between the filtered first card points. This ensures that when video clips are subsequently allocated based on the multiple first card points, each video clip has sufficient time to be displayed, thereby improving the quality of the generated video.

[0128] In the process of obtaining multiple first checkpoints from multiple initial checkpoints, the terminal may first determine a base checkpoint from the multiple initial checkpoints. Then, using the base checkpoint as a reference, the terminal may determine multiple first checkpoints from the multiple initial checkpoints. Specifically, the terminal may determine the checkpoint closest to the base checkpoint from the multiple initial checkpoints, with a time interval that reaches an interval threshold. Then, using the base checkpoint as a reference, the terminal may search further away from the base checkpoint to determine another checkpoint closest to the base checkpoint, with a time interval that reaches the interval threshold. This cycle repeats until all initial checkpoints have been traversed, and the multiple checkpoints retrieved based on the base checkpoint are the multiple first checkpoints.

[0129] The basic checkpoint may be the first initial checkpoint among multiple initial checkpoints, or the initial checkpoint with the strongest rhythm intensity among multiple initial checkpoints, or may be specified by the user from multiple initial checkpoints, which is not limited in the embodiments of the present disclosure.

[0130] After determining multiple first card points of the background audio, the terminal can assign multiple video clips to the time intervals between adjacent card points, so that the corresponding video clip is displayed during the time intervals between each first card point. The terminal can assign one video clip to a certain time interval, or one video clip to multiple time intervals, or multiple video clips to a certain time interval, etc., which is not limited in the present embodiment.

[0131] In step 302, in response to the card point switching operation, the terminal displays multiple second card points of the background audio, the density of the multiple second card points is different from the density of the multiple first card points, and each second card point is used to indicate a playback moment in the background audio.

[0132] In the embodiment of the present disclosure, when a card point switching operation occurs, the terminal replaces multiple first card points in the video editing interface with multiple second card points indicated by the card point switching operation. The card point switching operation can be a global switch of all card points or a single switch of individual card points, which is not limited in the embodiment of the present disclosure.

[0133] In the process of displaying multiple second card points of the background audio, the terminal displays a card point switching control in the video editing interface, and the card point switching control is used to switch card points of different densities. In response to the triggering operation of the card point switching control, the terminal displays multiple card point modes. The density of the card points indicated by different card point modes is different. In response to the selection operation of any card point mode among the multiple card point modes, the terminal displays multiple second card points indicated by the card point mode. The solution provided by the embodiment of the present disclosure can provide users with multiple card point modes, and the card point densities indicated by different card point modes are different, which is convenient for users to switch card point modes to adjust the density of the card points of the background audio according to their desired video rhythm. The operation is simple and can improve the efficiency of video generation.

[0134] For example, Figure 5 FIG. 1 is a schematic diagram showing a switching card point density according to an exemplary embodiment. Figure 5 , the terminal displays three card point modes in the card point switching panel 501. Different card point modes indicate different card point densities. The card point density indicated by card point mode 2 is greater than the card point density indicated by card point mode 1. In response to the selection of card point mode 2, the terminal previews multiple second card points 502 in the card point switching panel 501. The number of second card points 502 is greater than the number of first card points 503.

[0135] In the disclosed embodiment, the multiple second checkpoints displayed by the terminal can be multiple checkpoints (rhythm points) directly identified based on background audio, or can be checkpoints selected from multiple directly identified checkpoints. This disclosure is not limited to this. The principle of the terminal selecting the second checkpoints is similar to the principle of selecting the first checkpoint in step 301 and will not be further described here.

[0136] In step 303, for any video clip among the multiple video clips, when the time interval corresponding to the video clip changes from a first time interval to a second time interval, the position at which the terminal displays the video clip is updated from the position indicated by the first time interval to the position indicated by the second time interval, where the first time interval is the time interval corresponding to the video clip between multiple first card points, and the second time interval is the time interval corresponding to the video clip between multiple second card points.

[0137] In the disclosed embodiment, a time interval exists between adjacent ones of the plurality of second checkpoints. The terminal can allocate video segments from the plurality of video segments to the time intervals between the plurality of second checkpoints. After the allocation is complete, the terminal displays the video segments allocated to each time interval at the positions indicated by the time intervals between the plurality of second checkpoints. The position of each video segment is updated from the position indicated by the first time interval to the position indicated by the second time interval.

[0138] When allocating video clips to time intervals, the terminal can allocate video segments based on the video attributes of the video segments and the length of the time interval, allocating video frequency bands whose video attributes match the length of the time interval. Video attributes include video status and video duration. Video status indicates whether a video segment is dynamic. Dynamic video refers to a video in which objects in the video are in motion. If the objects in the video segment are stationary, the video segment is static. For example, if a video segment contains only one image, the video segment is static, but this is not limited in the present embodiment.

[0139] Alternatively, the process of identifying checkpoints and allocating (or matching) video segments for the time intervals between the checkpoints may also be performed by the server, and the terminal only needs to display the result, which is not limited in this embodiment of the present disclosure.

[0140] After determining multiple second card points of the background audio, the terminal can assign multiple video clips to the time intervals between adjacent card points, so that the corresponding video clip is displayed during the time intervals between each second card point. The terminal can assign one video clip to a certain time interval, or one video clip to multiple time intervals, or multiple video clips to a certain time interval, etc., which is not limited in the present embodiment.

[0141] For example, see Figure 5 The terminal displays video clip 1 at the positions indicated by the first two time intervals between the multiple second card points 502, displays video clip 2 at the positions indicated by the third and fourth time intervals between the multiple second card points 502, displays video clip 3 at the position indicated by the fifth time interval between the multiple second card points 502, and displays video clip 4 at the position indicated by the sixth time interval between the multiple second card points 502.

[0142] For details on how to allocate multiple video clips, see Figure 6 Step 603, step 604 and step 605 in the illustrated embodiment.

[0143] In step 304, if the duration of the first time interval is different from the duration of the second time interval, the duration of the video clip displayed by the terminal is updated from the duration of the first time interval to the duration of the second time interval.

[0144] In the embodiment of the present disclosure, after allocating the video segments to the matching time intervals, the terminal can also perform time alignment on the matching video segments and time intervals. The embodiment of the present disclosure does not limit the method of time alignment.

[0145] In some embodiments, for any video clip among multiple video clips, if the duration of the first time interval is shorter than the duration of the second time interval, the playback speed of the video clip displayed by the terminal is updated from the first speed to the second speed. The time consumed for the video clip to be played back at the first speed is the duration of the first time interval. The time consumed for the video clip to be played back at the second speed is the duration of the second time interval. The second speed is slower than the first speed. The solution provided by the embodiment of the present disclosure can slow down the playback speed of the video clip after the video clip is assigned to the matching time interval. If the duration of the video clip is less than the duration of the matching time interval, the playback speed of the video clip can be slowed down so that the time consumed for the video clip to be played back at the slower playback speed is equal to the duration of the matching time interval. That is, by appropriately slowing down the playback speed of the video clip, the playback duration of the video clip is aligned with the duration of the time interval, ensuring that the corresponding video clip can be displayed in each time interval and switching between video clips can be achieved at the card point, which is conducive to improving the rhythm matching of the generated video, that is, it is conducive to improving the quality of the generated video.

[0146] In other embodiments, for any video segment among the multiple video segments, if the duration of the first time interval is longer than the duration of the second time interval, the terminal updates the displayed video segment to a target video segment. The target video segment is a portion of the video segment. The duration of the video segment is the duration of the first time interval. The duration of the target video segment is the duration of the second time interval.

[0147] Alternatively, for any video clip among the multiple video clips, if the first time interval is longer than the second time interval, the terminal updates the playback speed of the video clip displayed by the terminal from the first speed to a third speed. The time it takes to play the video clip at the first speed is the duration of the first time interval. The time it takes to play the video clip at the third speed is the duration of the second time interval. The third speed is faster than the first speed.

[0148] The solution provided by the embodiment of the present disclosure is that after the video clips are assigned to the matching time intervals, if the duration of the video clips is longer than the duration of the matching time interval, it is possible to obtain part of the video clips or increase the playback speed of the video clips so that the time consumed to complete the playback of the adjusted video clips is equal to the duration of the matching time interval, that is, by appropriately cropping or increasing the playback speed of the video clips, the playback duration of the video clips is aligned with the duration of the time interval, thereby ensuring that the corresponding video clips can be displayed in each time interval and realizing video clip switching at the card point, which is beneficial to improving the rhythm matching of the generated video, that is, it is beneficial to improving the quality of the generated video.

[0149] If the duration of the first time interval is equal to the duration of the second time interval, then there is no need to adjust the duration of the video segment. Accordingly, step 304 is an optional step.

[0150] In step 305, the terminal generates a video based on the updated multiple video clips and background audio.

[0151] In the disclosed embodiment, in response to a confirmation generation operation in the video editing interface, the terminal directly synthesizes the video segments allocated for the time intervals between the plurality of second card points with the background audio to generate a video. The generated video sequentially plays the video segments allocated for each time interval between the plurality of second card points in the order of the plurality of second card points.

[0152] The disclosed embodiments provide a video generation method. When generating a video based on multiple video clips and background audio, if the density of background audio card points changes, the video clips can be reallocated, automatically assigning the multiple video clips to positions indicated by the updated time intervals between the card points, so that the video clips are aligned with the latest time intervals. That is, multiple second card points in the background audio are automatically aligned with the multiple video clips. This not only ensures that video clips can be switched at the positions of the second card points in subsequently generated videos, meeting user requirements for video rhythm, but also is simple to operate. Manual alignment is not required; the rhythm of subsequently generated videos can be controlled by simply switching the card point density, thereby improving video generation efficiency. Furthermore, since video attributes such as video status and video duration match the length of the time interval between the aligned video clips and the time interval, each video clip can be displayed within a sufficient time interval. Compared to directly adjusting the duration of the video clips based on the time interval, this solution can minimize the risk of some video clips being played too quickly, too slowly, or repeatedly simply to accommodate the time interval, thereby improving the quality of the generated video.

[0153] above Figure 3 The illustrated embodiment only illustrates the basic process of allocating the time intervals between card points to video segments. The following further illustrates the solution provided by the present disclosure based on a specific allocation method. Figure 6 FIG. 1 is a flow chart of another video generation method according to an exemplary embodiment. Assume that the electronic device is provided as a terminal. Figure 6 , the method includes the following steps.

[0154] In step 601, the terminal displays multiple first card points of the background audio on the video editing interface, and displays multiple video clips at positions indicated by time intervals between the multiple first card points, where each first card point is used to indicate a playback moment in the background audio.

[0155] In step 602, in response to the card point switching operation, the terminal displays multiple second card points of the background audio, the density of the multiple second card points is different from the density of the multiple first card points, and each second card point is used to indicate a playback moment in the background audio.

[0156] In the embodiment of the present disclosure, the execution principle of steps 601-602 is the same as that of steps 301-302, and will not be repeated here. The allocation method in step 303 can be further refined into the following steps 603, 604 and 605, etc., and the details can be found in the following steps.

[0157] In step 603, when the number of time intervals between the plurality of second card points is greater than the number of the plurality of video segments, the terminal displays a video segment allocated from the plurality of video segments by at least one time interval at a position indicated by at least one time interval.

[0158] In an embodiment of the present disclosure, a terminal obtains the number of time intervals (i.e., second time intervals) between multiple second card points and the number of multiple video segments. If the number of time intervals is greater than the number of video segments, the terminal can assign each video segment to at least one time interval. Accordingly, the terminal displays each video segment at the location indicated by the at least one assigned time interval. The at least one time interval corresponding to each video segment is continuous. That is, the relationship between the number of video segments and time intervals can be one-to-many. In other words, if the number of time intervals between multiple second card points is greater than the number of multiple video segments, each video segment corresponds to at least one second time interval, and the duration of each video segment is equal to the total duration of the at least one second time interval. Accordingly, for any video segment among the multiple video segments, the duration of the displayed video segment is updated from the duration of the first time interval to the total duration of the corresponding at least one second time interval.

[0159] The solution provided by the embodiment of the present disclosure can automatically assign each video clip to at least one matching time interval when there are many time intervals between the card points of the background audio and fewer video clips, so that each video clip can be displayed within at least one matching time interval. This not only ensures that each video clip has sufficient time to be displayed, but also ensures that video clips can be switched at the card points in subsequently generated videos. Compared with repeatedly playing certain video clips, the video generated by this solution has a better sense of rhythm, that is, it can improve the video quality.

[0160] In the case where the number of time intervals is greater than the number of video clips, the terminal can divide the number of time intervals by the number of multiple video clips to obtain T and N. T is the divisor, N is the remainder, T is a positive integer, and N is a natural number. T and N are used to represent the number of time intervals matched by each video clip. Then, for any video clip among the first N video clips in the multiple video clips, the terminal allocates the video clip to T+1 time intervals. For any video clip other than the first N video clips in the multiple video clips, the terminal allocates the video clip to T time intervals. Accordingly, for any video clip among the first N video clips in the multiple video clips, the terminal displays the video clip at the position indicated by the corresponding T+1 time intervals; for any video clip other than the first N video clips in the multiple video clips, the terminal displays the video clip at the position indicated by the corresponding T time intervals.

[0161] For example, Figure 7 FIG2 is a schematic diagram showing a method of allocating video segments to time intervals according to an exemplary embodiment. Figure 7 , the number of time intervals is 5, and the number of video segments is 3. In this case, T = 1 and N = 2. For the first two video segments (video segment a and video segment b), the terminal assigns both video segments a and b to two time intervals. For video segment c, excluding the first two video segments, the terminal assigns video segment c to one time interval.

[0162] You can also continue to see Figure 5 , the number of time intervals between multiple second card points is 6, and the number of video segments is 4. In this case, T=1, N=2. For the first video segments 1 and 2, the terminal assigns video segments 1 and 2 to 2 time intervals. For video segments 3 and 4 other than the first 2 video segments in the multiple video segments, the terminal assigns video segments 3 and 4 to 1 time interval respectively. The terminal can also change Figure 5 The sorting positions of the four video clips are still allocated in the same manner as above and will not be described in detail.

[0163] After determining the quantitative relationship between video segments and time intervals based on the above method, the terminal can sequentially assign the video segments to the corresponding time intervals according to the time sequence and the order of the video segments. For example, the first video segment can be assigned to the time intervals 1 to T+N, the second video segment can be assigned to the time intervals T+N+1 to 2T+2N, and so on. In this method, the order of the video segments remains unchanged. Alternatively, after determining the quantitative relationship between video segments and time intervals based on the above method, the terminal can also adjust the order of the video segments so that the video segments are assigned to at least one matching time interval.

[0164] In some embodiments, for any video clip among the multiple video clips, if the video clip meets a preset condition, the duration of the second time interval corresponding to the video clip reaches a duration threshold. The preset condition is a condition that must be met for the display duration of the video clip to meet the standard. If the video clip does not meet the preset condition, the duration of the second time interval corresponding to the video clip is less than the duration threshold. The disclosed embodiments provide a solution that automatically assigns a video clip to a longer time interval if the video clip meets the preset condition, and automatically assigns the video clip to a shorter time interval if the video clip does not meet the preset condition. This not only ensures that video clips can be switched at the card point in the subsequently generated video, improving the rhythm matching of the generated video, but also ensures that video clips that require a long display have sufficient time to be displayed, while video clips that do not require a long display have a shorter display time. This allows users to fully understand the information in each video clip without having to watch video clips with less information for a long time, thereby improving the accuracy and efficiency of information transmission in the generated video, thereby improving the quality of the generated video.

[0165] Specifically, if the number of time intervals between the plurality of second checkpoints is greater than the number of video clips, for any video clip from the plurality of video clips, if the video clip meets a preset condition, the terminal displays the video clip at a location indicated by at least one third time interval. The at least one third time interval is a second time interval where the total duration between the plurality of second checkpoints reaches a duration threshold. If the video clip does not meet the preset condition, the terminal displays the video clip at a location indicated by at least one fourth time interval. The at least one fourth time interval is a second time interval where the total duration between the plurality of second checkpoints does not reach the duration threshold.

[0166] The embodiments of the present disclosure do not limit the preset conditions. In some embodiments, the preset conditions may include any of the following.

[0167] First, the object in the video clip is in motion. This indicates that the video clip is dynamic. If the object in the video clip is stationary, it is static. In other words, the terminal can assign dynamic videos to longer time intervals and static videos to shorter time intervals. This approach ensures that dynamic images with high information content have sufficient time to be perceived, while static images with low information content can be quickly switched to maintain attention.

[0168] The second item is that at least one of the following: the motion vector value, color complexity, number of scenes, number of subjects, number of subtitle characters, and the weight of the video clip in the video clip reaches a corresponding preset value. If the motion vector value in the video clip reaches the corresponding preset value, it indicates that the video clip is highly dynamic. In this case, the video clip is assigned to a longer time interval. In other words, the terminal can assign high-dynamic video clips to longer time intervals.

[0169] The solution provided by the embodiment of the present disclosure can reflect to a certain extent that the video clip contains more information or is of higher importance because the objects in the video clip are in motion, or at least one of the numerical value of the motion vector, color complexity, number of scenes, number of subjects, number of words in the subtitles, and weight of the video clip in the video clip reaches a corresponding preset value. In this case, the video clip is automatically assigned to a time interval with a longer duration, which not only ensures that the video clip can be switched at the card point in the subsequently generated video and improves the rhythm matching of the generated video, but also allows video clips that need to be displayed for a long time to have sufficient time to be displayed, so that users can fully understand the information in each video clip, which is beneficial to improving the accuracy and efficiency of information transmission in the generated video, thereby improving the quality of the generated video.

[0170] The weights of video segments can be user-defined, meaning that the terminal can assign user-specified important video segments to longer time intervals. Alternatively, the weights of video segments can be determined based on at least one of color complexity, number of scenes, number of subjects, and number of words in the subtitles. For example, the weight of a video segment can be positively correlated with each of these factors.

[0171] In step 604, when the number of time intervals between the plurality of second card points is less than the number of the plurality of video segments, the terminal displays at least one video segment allocated to the time interval from the plurality of video segments at a position indicated by the time interval for any time interval.

[0172] In an embodiment of the present disclosure, a terminal obtains the number of time intervals between multiple second card points and the number of multiple video segments. If the number of time intervals is less than the number of video segments, the terminal can allocate at least one video segment to each time interval. Accordingly, the terminal displays at the location indicated by each time interval at least one video segment allocated to the time interval from the multiple video segments. In other words, the relationship between the number of video segments and time intervals can be a many-to-one relationship. In other words, for any video segment among the multiple video segments, if the number of time intervals between the multiple second card points is less than the number of video segments, the number of second time intervals corresponding to the video segment is one, and the duration of the video segment is less than or equal to the duration of the second time interval. Each second time interval corresponds to at least one video segment. Accordingly, for any video segment among the multiple video segments, the duration of the displayed video segment is updated from the duration of the first time interval to the entire duration or a portion of the corresponding second time interval.

[0173] The solution provided by the embodiment of the present disclosure can automatically assign at least one video clip to a matching time interval when the time intervals between background audio card points are small and there are many video clips, so that at least one video clip can be played in each time interval. This not only ensures that each video clip has time to be displayed, but also ensures that video clips can be switched at the card points in subsequently generated videos. Compared with directly cutting certain video clips, the video generated by this solution has better integrity, that is, it can improve video quality.

[0174] If the number of time intervals is less than the number of video segments, the terminal can divide the number of video segments by the number of time intervals to obtain P and M. P is the divisor, M is the remainder, P is a positive integer, and M is a natural number. P and M represent the number of video segments that match each time interval between the multiple second checkpoints. Then, for any time interval among the first M time intervals between the multiple second checkpoints, the terminal allocates P+1 video segments to that time interval. For any time interval other than the first M time intervals between the multiple second checkpoints, the terminal allocates P video segments to that time interval. Accordingly, for any time interval among the first M time intervals between the multiple second checkpoints, P+1 video segments are displayed at the position indicated by that time interval. For any time interval other than the first M time intervals between the multiple second checkpoints, the terminal displays P video segments at the position indicated by that time interval.

[0175] For example, Figure 8 FIG is a schematic diagram showing another method of allocating video segments to time intervals according to an exemplary embodiment. Figure 8, the number of time intervals is 3, and the number of video segments is 5. In this case, P = 1 and M = 2. For the first two time intervals, the terminal allocates video segments a and b for the first time interval, and video segments c and d for the second time interval. For the last time interval, the terminal allocates one video segment e.

[0176] After determining the quantitative relationship between video segments and time intervals based on the above method, the terminal can also assign the video segments to the corresponding time intervals in sequence according to the time sequence and the order of the video segments. For example, the 1st to the P+Mth video segments are assigned to the 1st time interval, the P+M+1th to the 2P+2Mth video segments are assigned to the 2nd time interval, and so on. In this method, the order of the video segments remains unchanged. Alternatively, after determining the quantitative relationship between video segments and time intervals based on the above method, the terminal can also adjust the order of the video segments so that the video segments are assigned to matching time intervals.

[0177] In some embodiments, for any video clip among the multiple video clips, if the video clip meets a preset condition, the duration of the second time interval corresponding to the video clip reaches a duration threshold. The preset condition is a condition that must be met for the display duration of the video clip to meet the standard. If the video clip does not meet the preset condition, the duration of the second time interval corresponding to the video clip is less than the duration threshold.

[0178] Specifically, if the number of time intervals between the plurality of second checkpoints is less than the number of the plurality of video segments, for any video segment among the plurality of video segments, if the video segment meets the preset condition, the video segment is displayed at the position indicated by the fifth time interval. The fifth time interval is a second time interval where the duration between the plurality of second checkpoints reaches a duration threshold. If the video segment does not meet the preset condition, the terminal displays the video segment at the position indicated by the sixth time interval. The sixth time interval is a second time interval where the duration between the plurality of second checkpoints does not reach the duration threshold.

[0179] The present embodiment does not limit the preset conditions. For details, please refer to the description of the preset conditions in step 603, which will not be repeated here.

[0180] In step 605, when the number of time intervals between the plurality of second card points is equal to the number of the plurality of video segments, the terminal displays a video segment allocated to the time interval from the plurality of video segments at a position indicated by the time interval for any time interval.

[0181] In an embodiment of the present disclosure, if the number of time intervals is equal to the number of video clips, the terminal can assign each video clip to a time interval. Accordingly, the terminal displays each video clip at the location indicated by the assigned time interval. That is, the relationship between the number of video clips and time intervals can be one-to-one. In other words, for any video clip among the multiple video clips, if the number of time intervals between the multiple second card points is less than or equal to the number of the multiple video clips, the number of second time intervals corresponding to the video clip is one, and the duration of the video clip is equal to the duration of the second time interval. The solution provided by an embodiment of the present disclosure, when the time intervals between the background audio card points are the same as the number of video clips, can automatically assign each video clip to a matching time interval, allowing each video clip to be displayed within a matching time interval. This not only ensures that each video clip has sufficient time to be displayed, but also ensures that video clips can be switched at the card points in subsequently generated videos. Compared to repeatedly playing certain video clips, the video generated by this solution has a better sense of rhythm, that is, it can improve video quality.

[0182] After determining the quantitative relationship between video segments and time intervals based on the above method, the terminal can also assign the video segments to corresponding time intervals in sequence according to the time sequence and the order of the video segments. For example, the first video segment can be assigned to the first time interval, the second video segment can be assigned to the second time interval, and so on. In this method, the order of the video segments remains unchanged. Alternatively, after determining the quantitative relationship between video segments and time intervals based on the above method, the terminal can also adjust the order of the video segments to assign the video segments to a matching time interval. The terminal can also determine the time interval that matches the video segment based on whether the video segment meets preset conditions. The process principle is the same as in steps 603 and 604 and will not be repeated here.

[0183] When matching time intervals for video clips, in addition to the aforementioned quantity matching, the terminal may also perform at least one of duration matching, theme matching, and weight matching to match the time intervals for the video clips. Duration matching refers to ensuring that the difference between the duration of the video clip and the duration of the time interval is less than a preset difference. Theme matching refers to ensuring that the theme of the video clip matches the theme of the audio clip within the time interval. Weight matching refers to ensuring that the weight of the video clip matches the weight of the audio clip within the time interval. The weights of the video clips and the audio clips can be customized by the user and are not limited in this regard in the present disclosed embodiments.

[0184] In some embodiments, for multiple first checkpoints, the terminal may also adopt the above-mentioned method to allocate video segments for the time intervals between adjacent checkpoints in the multiple first checkpoints, which will not be described in detail here.

[0185] In some embodiments, before matching background audio and card points for multiple video clips for the first time, the terminal may display a prompt message to automatically generate a card point video. A card point video is one in which the card points of the background audio in a video are aligned with the switching moments between video clips in the video, so that the video clips can be switched at the card points in the subsequently generated videos.

[0186] For example, Figure 9 FIG. 1 is a schematic diagram showing a method of generating prompt information according to an exemplary embodiment. Figure 9 In response to a trigger operation on the video generation control 901, the terminal displays a generation prompt message 902 and a confirmation generation control 903. The terminal demonstrates the function of the video generation control 901 by generating prompt message 902, namely, that the video generation control 901 can automatically generate a card point video in which the video image and the background audio card points are aligned. In response to a trigger operation on the confirmation generation control 903, the terminal matches the background audio for the video clip and performs card point alignment. The terminal can display the progress of matching the background audio. When the matching is complete, the terminal displays the background audio 905 and multiple first card points 906 of the background audio 905 in the audio track 904. Then, based on the multiple first card points 906, the terminal adjusts the multiple video clips in the video track 907 so that each video clip is displayed at the position indicated by the corresponding time interval.

[0187] In some embodiments, in response to a cropping operation on the audio selected for a video clip in the video editing interface, the terminal replaces the audio with the cropped audio clip. In other words, users can also extract an audio clip from a specific audio source and use it as background audio for generating a video. This approach meets the user's video generation needs, allows users to edit audio as needed, is simple to operate, and can improve video generation efficiency.

[0188] For example, Figure 10 FIG. 1 is a schematic diagram showing a method of trimming an audio segment according to an exemplary embodiment. Figure 10, the terminal displays the matched audio 1002 in the audio track 1001. In response to the triggering operation of the audio editing control 1003, the terminal displays the audio editing panel 1004 in the video editing interface. When audio 2 in the audio editing panel 1004 is selected, in response to the triggering operation of the segment interception control 1005, the terminal displays audio 2 and the interception control 1006 in the audio editing panel 1004. The interception control 1006 is used to intercept the audio segment in the circle. The user can determine the audio segment to be intercepted by sliding the interception control 1006 on audio 2. When the interception is confirmed, the terminal can display the intercepted audio segment and multiple card points of the audio segment in the audio track 1001.

[0189] In step 606, the terminal generates a video based on the video segments and background audio allocated to the time intervals between the plurality of second card points.

[0190] In the embodiment of the present disclosure, the execution principle of step 606 is the same as the execution principle of step 303 and will not be repeated here.

[0191] Embodiments of the present disclosure provide a video generation method. When generating a video based on multiple video clips and background audio, if the density of background audio card points changes, the method can automatically assign multiple video clips to positions indicated by time intervals whose lengths match the video attributes based on the video attributes of the video clips, the length of the time intervals between the changed card points, and the quantitative relationship between the video clips and the time intervals. This allows the matching video clips and time intervals to be aligned with each other. Specifically, multiple first card points in the background audio are automatically aligned with multiple video clips. This not only ensures that each video clip has sufficient time to be displayed, but also ensures that video clips can be switched at the first card points in subsequently generated videos, thereby ensuring the quality of the generated video. Manual alignment is also eliminated, thereby improving video generation efficiency. Furthermore, because the video attributes, such as video status and video duration, of the aligned video clips and time intervals match the length of the time interval, each video clip can be displayed within a sufficient time interval. Compared to directly adjusting the duration of the video clips based on the time interval, this solution can minimize the risk of some video clips being played too quickly, too slowly, or repeatedly simply to accommodate the time interval, thereby improving the quality of the generated video.

[0192] In some embodiments, after a user edits at least one of multiple video clips and background audio, the terminal can detect that the at least one of the multiple video clips and background audio has changed. In this case, the terminal can also reallocate the video clips. Accordingly, if at least one of the multiple video clips and background audio has changed, the terminal displays a card point prompt. The card point prompt prompts whether to re-match the card points based on the changed material. In response to a confirmation operation in the card point prompt, the terminal displays an update of at least one of the positions and durations of the multiple video clips based on the card points of the background audio after the material change. If the background audio remains unchanged (changed) and only the video clips have changed, the card points of the background audio may remain unchanged. For example, after the material change, the card points of the background audio may be the same as multiple first card points or multiple second card points. Alternatively, if the background audio remains unchanged and only the video clips have changed, the card points of the background audio may change, i.e., the card points of the background audio after the material change may be different from those before the material change. If the background audio has changed, the terminal can re-identify the card points of the background audio, i.e., the card points before and after the material change may be different.

[0193] The solution provided by the embodiments of the present disclosure displays a card point prompt when a video clip or background audio is changed, prompting the user whether to re-match the card points based on the changed material. If the user allows re-matching, the solution automatically re-assigns video clips to the current time intervals, or re-determines new card points and assigns video clips to the time intervals between the new card points. This eliminates the need for manual allocation by the user, simplifies the operation, and improves video generation efficiency.

[0194] For example, Figure 11 FIG2 is a schematic diagram showing a method of reallocating video segments according to an exemplary embodiment. Figure 11 If the number of video clips increases, the terminal displays a card point prompt 1101 to prompt whether to reallocate the video clips based on the changed material. In response to a triggering operation on a confirmation control 1002 in the card point prompt 1101, the terminal displays the video clip newly allocated from the six video clips at the position indicated by the time interval between adjacent card points in the plurality of first card points.

[0195] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.

[0196] Figure 12 FIG is a block diagram of a video generating apparatus according to an exemplary embodiment. Figure 12 , the video generating device includes: a display unit 1201 and a generating unit 1202.

[0197] The display unit 1201 is configured to display a plurality of first card points of the background audio on the video editing interface, and display a plurality of video clips at positions indicated by time intervals between the plurality of first card points, wherein each first card point indicates a playback moment of the background audio;

[0198] The display unit 1201 is further configured to display a plurality of second card points of the background audio in response to the card point switching operation, wherein the density of the plurality of second card points is different from the density of the plurality of first card points, and each second card point is used to indicate a playback moment of the background audio;

[0199] The display unit 1201 is further configured to display, based on the plurality of second card points, at least one of the positions and durations of the plurality of video segments being updated;

[0200] The generating unit is configured to generate a video based on the updated plurality of video clips and background audio.

[0201] In some embodiments, the display unit 1201 is configured to perform at least one of the following:

[0202] For any video segment among the multiple video segments, when the time interval corresponding to the video segment changes from a first time interval to a second time interval, the position of the displayed video segment is updated from the position indicated by the first time interval to the position indicated by the second time interval, where the first time interval is the time interval corresponding to the multiple first card points in the video segment, and the second time interval is the time interval corresponding to the multiple second card points in the video segment;

[0203] If the duration of the first time interval is different from the duration of the second time interval, the duration of the displayed video segment is updated from the duration of the first time interval to the duration of the second time interval.

[0204] In some embodiments, the display unit 1201 is further configured to, for any video clip among the multiple video clips, when the duration of the first time interval is shorter than the duration of the second time interval, update the playback speed of the displayed video clip from the first speed to the second speed, the time consumed for the video clip to be played back based on the first speed is the duration of the first time interval, and the time consumed for the video clip to be played back based on the second speed is the length of the second time interval, and the second speed is slower than the first speed.

[0205] In some embodiments, the display unit 1201 is further configured to perform any of the following:

[0206] For any video segment among the multiple video segments, when the duration of the first time interval is longer than the duration of the second time interval, the displayed video segment is updated to a target video segment, where the target video segment is a portion of the video segment, the duration of the video segment is the duration of the first time interval, and the duration of the target video segment is the duration of the second time interval;

[0207] For any video clip among multiple video clips, when the duration of the first time interval is longer than the duration of the second time interval, the playback speed of the video clip is updated from the first speed to the third speed, the time consumed for the video clip to be played back based on the first speed is the duration of the first time interval, and the time consumed for the video clip to be played back based on the third speed is the duration of the second time interval, and the third speed is faster than the first speed.

[0208] In some embodiments, for any video segment among the multiple video segments, when the number of time intervals between the multiple second card points is less than the number of the multiple video segments, the number of second time intervals corresponding to the video segment is one, and the duration of the video segment is less than or equal to the duration of the second time interval;

[0209] When the number of time intervals between the multiple second card points is greater than the number of the multiple video clips, the number of second time intervals corresponding to the video clips is at least one, and the duration of the video clip is equal to the total duration of the at least one second time interval.

[0210] In some embodiments, for any video segment among the plurality of video segments, when the video segment satisfies a preset condition, the duration of the second time interval corresponding to the video segment reaches a duration threshold, where the preset condition is a condition that must be met for the display duration of the video segment to meet the standard;

[0211] When the video segment does not meet the preset condition, the duration of the second time interval corresponding to the video segment is lower than the duration threshold.

[0212] In some embodiments, the preset condition includes any of the following:

[0213] The objects in the video clip are in motion;

[0214] At least one of the value of the motion vector in the video clip, the color complexity, the number of scenes, the number of subjects, the number of words in the subtitles, and the weight of the video clip reaches a corresponding preset value.

[0215] In some embodiments, the display unit 1201 is further configured to display a card point switching control in the video editing interface, where the card point switching control is used to switch between card points of different densities; in response to a triggering operation on the card point switching control, display multiple card point modes, where the density of card points indicated by different card point modes is different; and in response to a selection operation on any of the multiple card point modes, display multiple second card points indicated by the card point mode.

[0216] In some embodiments, the display unit 1201 is further configured to display card point prompt information when at least one material among multiple video clips and background audio changes, and the card point prompt information is used to prompt whether to reallocate based on the changed material; in response to a confirmation operation in the card point prompt information, at least one of the positions and durations of the multiple video clips is updated based on the card point of the background audio after the material changes.

[0217] In some embodiments, the background audio includes at least one of the following:

[0218] Background audio that matches the style of multiple video clips;

[0219] Background audio whose heat meets the heat threshold;

[0220] Background audio with a duration that matches the combined duration of the multiple video clips;

[0221] Background audio with a theme that matches the themes of multiple video clips.

[0222] The disclosed embodiments provide a video generation device. In the process of generating a video based on multiple video clips and background audio, when the card point density of the background audio changes, the video clips can be reallocated, and the multiple video clips can be automatically allocated to the positions indicated by the time intervals between the updated card points, so that the video clips are aligned with the latest time intervals. That is, the multiple second card points in the background audio are automatically aligned with the multiple video clips. This not only ensures that the video clips can be switched at the positions of the second card points in the subsequently generated video, meeting the user's requirements for the video rhythm, but also is simple to operate. There is no need for manual alignment by the user. The rhythm of the subsequently generated video can be controlled by switching the card point density, which is conducive to improving the efficiency of video generation.

[0223] It should be noted that the video generation device provided in the above embodiments, when generating a stuck point video, uses the division of the aforementioned functional units as an example only. In actual applications, the aforementioned functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. Furthermore, the video generation device provided in the above embodiments and the video generation method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0224] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0225] When an electronic device is provided as a terminal, Figure 13 FIG1 is a block diagram of a terminal 1300 according to an exemplary embodiment. Figure 13 The following is a block diagram of a terminal 1300 according to an exemplary embodiment of the present disclosure. Terminal 1300 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 1300 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.

[0226] Typically, the terminal 1300 includes a processor 1301 and a memory 1302 .

[0227] Processor 1301 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 1301 may be implemented in hardware using at least one of the following: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). Processor 1301 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), processes data while awake, while the coprocessor is a low-power processor that processes data while in standby mode. In some embodiments, processor 1301 may integrate a graphics processing unit (GPU), which is responsible for rendering and drawing content displayed on the display. In some embodiments, processor 1301 may also include an artificial intelligence (AI) processor for handling computational operations related to machine learning.

[0228] The memory 1302 may include one or more computer-readable storage media, which may be non-transitory. The memory 1302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1302 is used to store at least one computer program, which is executed by the processor 1301 to implement the video generation method provided in the method embodiment of the present application.

[0229] In some embodiments, terminal 1300 may optionally include a peripheral device interface 1303 and at least one peripheral device. Processor 1301, memory 1302, and peripheral device interface 1303 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1303 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1304, a display screen 1305, a camera assembly 1306, an audio circuit 1307, and a power supply 1308.

[0230] The peripheral device interface 1303 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1301 and the memory 1302. In some embodiments, the processor 1301, the memory 1302, and the peripheral device interface 1303 are integrated on the same chip or circuit board. In other embodiments, any one or two of the processor 1301, the memory 1302, and the peripheral device interface 1303 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0231] RF circuit 1304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. RF circuit 1304 communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 1304 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. In some embodiments, RF circuit 1304 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. RF circuit 1304 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, RF circuit 1304 may also include circuitry related to Near Field Communication (NFC), although this application does not limit this.

[0232] Display screen 1305 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. If display screen 1305 is a touchscreen display, it is also capable of collecting touch signals on or above the surface of display screen 1305. These touch signals can be input as control signals to processor 1301 for processing. Display screen 1305 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 1305, located on the front panel of terminal 1300. In other embodiments, there can be at least two display screens 1305, located on different surfaces of terminal 1300 or in a foldable design. In still other embodiments, display screen 1305 can be a flexible display, located on a curved or foldable surface of terminal 1300. Display screen 1305 can also be configured as a non-rectangular, irregular shape, known as a special-shaped screen. The display screen 1305 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0233] The camera component 1306 is used to capture images or videos. In some embodiments, the camera component 1306 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera component 1306 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0234] The audio circuit 1307 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1301 for processing, or input into the radio frequency circuit 1304 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 1300. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1301 or the radio frequency circuit 1304 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 1307 may also include a headphone jack.

[0235] Power supply 1308 is used to power various components in terminal 1300. Power supply 1308 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1308 includes a rechargeable battery, the rechargeable battery can be wired or wirelessly rechargeable. A wired rechargeable battery is charged via a wired line, while a wireless rechargeable battery is charged via a wireless coil. The rechargeable battery can also support fast charging technology.

[0236] Those skilled in the art will understand that Figure 13 The structure shown in the figure does not constitute a limitation on the terminal 1300, and the terminal 1300 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0237] When an electronic device is provided as a server, Figure 14 This is a block diagram of a server 1400 according to an exemplary embodiment. The server 1400 may vary significantly depending on its configuration or performance. It may include one or more processors (CPUs) 1401 and one or more memories 1402. The memories 1402 store at least one program code, which is loaded and executed by the processor 1401 to implement the video generation methods provided in the various method embodiments described above. Of course, the server may also include components such as a wired or wireless network interface, a keyboard, and input / output interfaces for input and output. The server 1400 may also include other components for implementing device functions, which are not described in detail here.

[0238] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as memory 1302 or memory 1402 including instructions. The instructions may be executed by processor 1301 of terminal 1300 or processor 1401 of server 1400 to implement the above-described video generation method. Alternatively, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, or the like.

[0239] A computer program product includes a computer program / instruction, which implements the above-mentioned video generation method when executed by a processor.

[0240] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0241] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video generation method, characterized in that: The method comprises: On the video editing interface, multiple first card points of the background audio are displayed, and multiple video clips are displayed at first positions indicated by first time intervals between the multiple first card points, where each first card point indicates a playback moment of the background audio; In response to the card point switching operation, displaying a plurality of second card points of the background audio, wherein the density of the plurality of second card points is different from the density of the plurality of first card points, and each second card point is used to indicate a playback moment of the background audio; When the number of second time intervals between the plurality of second card points is greater than the number of the plurality of video segments, displaying a video segment allocated from the plurality of video segments to which the at least one second time interval is assigned at a position indicated by the at least one second time interval; When the number of second time intervals between the plurality of second card points is less than the number of the plurality of video segments, for any second time interval, displaying at least one video segment allocated to the second time interval from the plurality of video segments at a position indicated by the second time interval; A video is generated based on the updated plurality of video clips and the background audio.

2. The video generation method according to claim 1, wherein: The method further comprises: For any video clip among the multiple video clips, when the time interval corresponding to the video clip changes from a first time interval to a second time interval, if the duration of the first time interval is different from the duration of the second time interval, the duration of the video clip displayed is updated from the duration of the first time interval to the duration of the second time interval.

3. The video generation method according to claim 2, characterized in that If the duration of the first time interval is different from the duration of the second time interval, updating the duration of displaying the video clip from the duration of the first time interval to the duration of the second time interval includes: For any video clip among the multiple video clips, when the duration of the first time interval is shorter than the duration of the second time interval, the playback speed of the video clip is updated from the first speed to the second speed, the time consumed for the video clip to be played back based on the first speed is the duration of the first time interval, and the time consumed for the video clip to be played back based on the second speed is the length of the second time interval, and the second speed is slower than the first speed.

4. The video generation method according to claim 2, wherein: If the duration of the first time interval is different from the duration of the second time interval, updating the duration of displaying the video clip from the duration of the first time interval to the duration of the second time interval includes any one of the following: For any video segment among the multiple video segments, when the duration of the first time interval is longer than the duration of the second time interval, displaying that the video segment is updated to a target video segment, where the target video segment is a portion of the video segment, the duration of the video segment is the duration of the first time interval, and the duration of the target video segment is the duration of the second time interval; For any video clip among the multiple video clips, when the duration of the first time interval is longer than the duration of the second time interval, the playback speed of the video clip is updated from the first speed to the third speed, the time consumed for the video clip to be played back based on the first speed is the duration of the first time interval, and the time consumed for the video clip to be played back based on the third speed is the length of the second time interval, and the third speed is faster than the first speed.

5. The video generation method according to claim 2, characterized in that: For any video segment among the multiple video segments, when the number of time intervals between the multiple second card points is less than the number of the multiple video segments, the duration of the video segment allocated to each second time interval is less than or equal to the duration of the second time interval; When the number of time intervals between the plurality of second card points is greater than the number of the plurality of video segments, the duration of the video segment allocated to at least one second time interval is equal to the total duration of the at least one second time interval.

6. The video generation method according to claim 2, wherein: For any video segment among the multiple video segments, when the video segment meets a preset condition, the duration of the second time interval corresponding to the video segment reaches a duration threshold, where the preset condition is a condition that must be met for the display duration of the video segment to meet the standard; When the video segment does not meet the preset condition, the duration of the second time interval corresponding to the video segment is lower than the duration threshold.

7. The video generation method according to claim 6, characterized in that: The preset conditions include any of the following: The object in the video clip is in motion; At least one of the value of the motion vector in the video clip, the color complexity, the number of scenes, the number of subjects, the number of words in the subtitles, and the weight of the video clip reaches a corresponding preset value.

8. The video generation method according to claim 1, wherein: The step of displaying a plurality of second card points of the background audio in response to the card point switching operation includes: Displaying a card point switching control in the video editing interface, wherein the card point switching control is used to switch card points of different densities; In response to a triggering operation on the card point switching control, a plurality of card point modes are displayed, and the density of card points indicated by different card point modes is different; In response to a selection operation on any one of the plurality of card point modes, a plurality of second card points indicated by the card point mode are displayed.

9. The video generation method according to claim 1, wherein: The method further comprises: When at least one of the plurality of video clips and the background audio material is changed, displaying card point prompt information, wherein the card point prompt information is used to prompt whether to re-match the card points based on the changed material; In response to a confirmation operation in the card point prompt information, at least one of the positions and durations of the multiple video clips is updated based on the card points of the background audio after the material changes.

10. The video generation method according to claim 1, characterized in that: The background audio includes at least one of the following: background audio that matches the style of the plurality of video clips; Background audio whose heat meets the heat threshold; Background audio having a duration matching the total duration of the plurality of video clips; Background audio having a theme matching the themes of the plurality of video clips.

11. A video generating device, characterized in that: The device comprises: The display unit is configured to display a plurality of first card points of the background audio on a video editing interface, and display a plurality of video clips at positions indicated by time intervals between the plurality of first card points, wherein each first card point indicates a playback moment of the background audio; The display unit is further configured to display a plurality of second card points of the background audio in response to a card point switching operation, wherein a density of the plurality of second card points is different from a density of the plurality of first card points, and each second card point is used to indicate a playback moment of the background audio; The display unit is further configured to, when the number of second time intervals between the plurality of second card points is greater than the number of the plurality of video segments, display a video segment allocated from the plurality of video segments to the at least one second time interval at a position indicated by the at least one second time interval; and, when the number of second time intervals between the plurality of second card points is less than the number of the plurality of video segments, display at least one video segment allocated from the plurality of video segments to any second time interval at a position indicated by the second time interval. The generating unit is configured to generate a video based on the updated plurality of video clips and the background audio.

12. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the video generation method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video generating method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the video generation method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Processing method, processing device, electronic device and storage medium

    CN110519638A

  • Video generation method and device, storage medium and electronic equipment

    CN116095422A