Video generation method and device, equipment and medium

By extracting target video clips that match the original video beat information and obtaining target video based on material characteristics, the problem of poor coordination between target video audio and picture and limited types of video templates in the prior art is solved, and efficient video creation is achieved.

CN120111275APending Publication Date: 2025-06-06BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311651138.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When the existing video creation platform applies video templates, the audio and picture matching of the target video is poor, and the types of video templates provided are limited, which limits the creative freedom of video creators.

Method used

The target video clips matching the original video beat information from the to-process video and obtain the target video based on the material characteristics of the target video clip and the original video. This method ensures that the audio of the target video matches the beat and is not limited by the types of video templates provided by the video creation platform.

Benefits of technology

It improves the matching of the audio and beats of the target video, enhances the creative freedom of video creators, and improves the efficiency of target video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111275A_ABST
    Figure CN120111275A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, equipment and a medium, and the method comprises the steps: extracting a target video clip matched with the beat information of an original video from a to-be-processed video, obtaining a target video based on the target video clip and the material features of the original video, and generating a video according to the target video. The material features at least comprise audio features of the original video. The method can ensure that the audio of the obtained target video is matched with the beat, achieves the purpose of unpacking the original video for use, and improves the production efficiency of the target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a video generation method, device, equipment and medium. Background Art

[0002] When creating a video on a video creation platform, when a video creator sees a video template of interest, he or she can generate a target video with the same style as the video template by applying the video template of interest to the source video file selected by the video creator with one click.

[0003] In the related art, the application effect of video templates on source video files is not very good, resulting in poor coordination between the audio and picture of the target video. In addition, the video creation platform provides a limited variety of video templates, which cannot provide enough video templates for video creators. Therefore, it is difficult for video creators to freely create target videos that meet their interests. Summary of the invention

[0004] According to one aspect of the present disclosure, a video generation method is provided, the method comprising:

[0005] Extracting a target video segment matching the beat information of the original video from the video to be processed;

[0006] A target video is acquired based on the target video segment and material features of the original video, wherein the material features at least include audio features of the original video.

[0007] According to another aspect of the present disclosure, a method for generating a video template is provided, the method comprising:

[0008] Extracting material information of an original video, wherein the material information of the original video at least includes audio features of the original video;

[0009] Extracting node information of the original video, wherein the node information at least includes beat information of the original video;

[0010] A video template is generated based on the material information of the original video and the node information of the original video.

[0011] According to another aspect of the present disclosure, there is provided a video generating apparatus, comprising:

[0012] An extraction module, used for extracting a target video segment matching the beat information of the original video from the video to be processed;

[0013] The acquisition module is used to acquire the target video based on the target video segment and the material features of the original video, wherein the material features at least include the audio features of the original video.

[0014] According to another aspect of the present disclosure, a device for generating a video template is provided, comprising:

[0015] An extraction module, extracting material information of an original video, extracting node information of the original video, wherein the material information of the original video at least includes audio features of the original video, and the node information at least includes beat information of the original video;

[0016] A generation module is used to generate a video template based on the material information of the original video and the node information of the original video.

[0017] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0018] processor; and,

[0019] A memory for storing programs;

[0020] The program includes instructions, and when the instructions are executed by the processor, the processor executes the method according to the exemplary embodiment of the present disclosure.

[0021] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to the exemplary embodiments of the present disclosure.

[0022] One or more technical solutions provided in the exemplary embodiments of the present disclosure can extract a target video segment that matches the beat information of the original video from the video to be processed, and because the material features of the original video at least include the audio features of the original video, the audio features of the original video match the beat information of the target video segment. On this basis, when obtaining the target video based on the material features of the target video segment and the original video, the audio features of the original video are essentially added to the target video segment to ensure that the audio of the obtained target video matches the beat.

[0023] In addition, the exemplary embodiment of the present disclosure can combine the beat information and material information of the original video with the video to be processed to synthesize a target video when the beat information and material information of the original video are obtained. The target video is not limited by the types of video templates provided by the video creation platform. Therefore, when a video creator of the exemplary embodiment of the present disclosure sees a video of interest but does not have a video template, he or she can directly use the video as the original video, and use the beat information and material information of the original video to create a target video that suits his or her interests, thereby improving the efficiency of target video production. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0025] Figure 1 A schematic diagram showing a flow chart of a video generation method according to an exemplary embodiment of the present disclosure;

[0026] Figure 2 A schematic diagram of a target video clip acquisition process of an exemplary embodiment of the present disclosure is shown;

[0027] Figure 3 A schematic diagram of a process for determining a beat feature of a reference video image according to an exemplary embodiment of the present disclosure is shown;

[0028] Figure 4 A schematic diagram of a transition setting process of an exemplary embodiment of the present disclosure is illustrated;

[0029] Figure 5 A schematic diagram of a method for generating a video template according to an exemplary embodiment of the present disclosure is shown;

[0030] Figure 6 A schematic block diagram of functional modules of a video generating device according to an exemplary embodiment of the present disclosure is shown;

[0031] Figure 7 A schematic block diagram of functional modules of a device for generating a video template according to an exemplary embodiment of the present disclosure is shown;

[0032] Figure 8 A schematic block diagram of a chip according to an exemplary embodiment of the present disclosure is shown;

[0033] Fig. 9 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0034] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0035] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0036] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0037] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0038] When the picture changes more than 24 frames per second, according to the principle of visual persistence, the human eye cannot distinguish a single static picture, so it looks like a smooth and continuous visual effect. Such a continuous picture is called a video. With the development of video technology, the current video form is more abundant, generally composed of multi-modal data composed of multimedia forms such as images, text, and audio. For example: when the video data is multi-modal data, the video data generally contains video frames, subtitles, sound and other information.

[0039] Considering that the multimodal data contained in videos is richer than traditional graphic media, more and more video creators choose videos as media to disseminate information. To improve the efficiency of video production, video creators can apply video templates to source video files to generate target videos with the same style as the video templates.

[0040] Early video templates were usually pre-designed and produced by professional video producers, and the output was generally in the form of video process materials or video templates that support secondary editing. In the era of short videos, video creation platforms can create one-click templates for popular video types. Video creators can select simple video clips, apply one-click templates to the video clips, and directly generate target videos with the same style as the one-click template.

[0041] It can be seen that whether it is to make video templates by professional video producers or to make one-click templates of popular video types on video creation platforms, it requires manual participation in production, and the application effect of video templates is not good, resulting in poor fluency of the target video, such as poor matching of audio and video rhythm of the target video, and the video creation platform provides a limited variety of video templates, which cannot provide enough video templates for video creators. Therefore, it is difficult for video creators to freely create target videos that suit their interests.

[0042] In response to the above problems, an exemplary embodiment of the present disclosure provides a video generation method, which can not only generate a target video through the beat information and material information of the original video to improve the audio and beat matching of the target video, but also get rid of the limitation on the number of video templates of the video creation platform, and directly use the original video to create a target video that suits one's own interests, thereby improving the efficiency of target video production.

[0043] In practical applications, the beat information and material information of the original video of the exemplary embodiment of the present disclosure can be stored in the form of a video template, which can be stored in a video creation platform, etc.

[0044] The video generating method provided by the exemplary embodiment of the present disclosure may be applied to a server or a chip in the server that interacts with a client. The video generating method of the exemplary embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.

[0045] Figure 1 FIG. 2 is a flow chart showing a video generation method according to an exemplary embodiment of the present disclosure. Figure 1 As shown, the video generation method of the exemplary embodiment of the present disclosure may include:

[0046] Step 101: extracting a target video segment that matches the beat information of the original video from the video to be processed. Here, in response to an operation of applying a video template to the video to be processed, the beat information of the original video is obtained from the video template, and then the target video segment that matches the beat information of the original video is extracted from the video to be processed.

[0047] In practical applications, the video duration of the above-mentioned video to be processed may be greater than the video duration of the original video, or may be less than or equal to the video duration of the original video. If the video duration of the video to be processed is greater than the video duration of the original video, a target video segment matching the beat information of the original video may be directly extracted from the video to be processed; if the video duration of the video to be processed is less than or equal to the video duration of the original video, the video to be processed may be expanded first so that the video duration of the video to be processed is greater than the video duration of the original video, and then a target video segment matching the beat information of the original video may be extracted from the video to be processed.

[0048] The beat information of the original video of the exemplary embodiment of the present disclosure can reflect the large-scale action changes or scene switching frequency in the original video. Therefore, the beat information of the original video of the exemplary embodiment of the present disclosure can include the beat characteristics of multiple frames of reference video images. For example, by detecting the image content change rate, a reference video image with a change rate greater than a preset change rate can be obtained, and then the image content change rate of each frame of the reference video image can be defined as the beat characteristic of the reference video image.

[0049] Step 102: Based on the target video clip and the material features of the original video, the target video is obtained, where the material features at least include the audio features of the original video. Here, the material features of the original video can be obtained while the beat information of the original video is obtained from the video template.

[0050] Considering that the material features of the original video at least include the audio features of the original video, and the target video clip matches the beat information of the original video, therefore, the audio features of the original video match the beat information of the target video clip, when obtaining the target video based on the material features of the target video clip and the original video, the audio features of the original video are actually added to the target video clip to ensure that the audio and beat of the obtained target video match. In this case, the application effect of the video template on the video to be processed can be guaranteed.

[0051] In addition, the exemplary embodiment of the present disclosure can combine the beat information and material information of the original video with the video to be processed to synthesize a target video when the beat information and material information of the original video are obtained. The target video is not limited by the types of video templates provided by the video creation platform. Therefore, when a video creator of the exemplary embodiment of the present disclosure sees a video of interest but does not have a video template, he or she can directly use the video as the original video, and use the beat information and material information of the original video to create a target video that suits his or her interests, thereby improving the efficiency of target video production and saving users time and costs.

[0052] As a possible implementation, Figure 2 FIG. 2 shows a schematic diagram of a target video segment acquisition process of an exemplary embodiment of the present disclosure. Figure 2 As shown, the exemplary embodiment of the present disclosure extracts a target video segment that matches the beat information of the original video from the video to be processed, which may include:

[0053] Step 201: Detect the relationship between the video duration of the video to be processed and the video duration of the original video. When the video duration of the video to be processed is greater than the video duration of the original video, it means that a target video segment equal to the video duration of the original video can be obtained from the video to be processed, and therefore, step 202 can be executed. When the video duration of the video to be processed is less than or equal to the video duration of the original video, it means that a target video segment equal to the video duration of the original video cannot be directly obtained from the video to be processed, and therefore, step 204 can be executed.

[0054] Step 202: obtaining a plurality of first candidate video segments from the video to be processed, wherein the video duration of the first candidate video segments is equal to the video duration of the original video.

[0055] In practical applications, a sliding window equal to the video duration of the original video may be defined first, and then the sliding window may be controlled to slide along the time dimension on the video to be processed to obtain a first candidate video segment. Each time the sliding window slides, a first candidate video segment may be captured from the original video.

[0056] Step 203: When the beat information of the first candidate video segment matches the beat information of the original video, the first candidate video segment is determined to be the target video segment. It should be understood that if the correlation between the beat information of the first candidate video segment and the beat information of the original video is greater than the first preset correlation, it can be considered that the beat information of the first candidate video segment matches the beat information of the original video.

[0057] In practical applications, when the original video includes multiple frames of reference video images, the beat information of the original video includes the beat features of the multiple frames of reference video images. If the original video includes multiple consecutive frames of original video images, if the position change rate of the preset key point of a frame of original video image is greater than the preset position change rate, the frame of original video image can be considered as the reference video image of the original video, and the beat features of the reference video image can also be determined based on the position change rate of the preset key point of the frame of original video image.

[0058] Exemplarily, a sliding window can be used to slide in the time dimension of the video to be processed to obtain multiple first candidate video segments, and the beat information of each first candidate video segment can be extracted using a 3D convolutional network. Then, the correlation between the beat information of each first candidate video segment and the original video is calculated using methods such as the Pearson correlation coefficient and the Spearman correlation coefficient.

[0059] When the Pearson correlation coefficient is used to calculate the correlation between the beat information of each first candidate video segment and the original video, the correlation between the beat information of the first candidate video segment and the original video can be expressed by the Pearson correlation coefficient of the beat information of the first candidate video segment and the original video, and accordingly, the first preset correlation can be expressed as the first preset correlation coefficient. In this case, when the Pearson correlation coefficient of the beat information of the first candidate video segment and the original video is greater than the first preset correlation coefficient, it can be considered that the beat information of the first candidate video segment matches the beat information of the original video.

[0060] If the Pearson correlation coefficient between the beat information of a first candidate video segment and the original video is greater than the first preset correlation coefficient, it can be considered that the first candidate video segment meets the requirements of the target video segment for the beat information. Therefore, the candidate video segment can be determined to be the target video segment. Otherwise, it means that the first candidate video segment does not meet the requirements of the target video segment for the beat information, and a new first candidate video segment can be obtained again.

[0061] The above-mentioned first preset correlation coefficient can be the maximum value of the correlation between the beat information of multiple first candidate video segments and the original video, or it can be a custom preset correlation coefficient. For example: when the first preset correlation coefficient is the maximum value of the correlation between the beat characteristics of multiple first candidate video segments and the original video, after obtaining the correlation between the beat characteristics of each first candidate video segment and the original video, it can be considered that the first candidate video with the greatest correlation meets the requirement of the target video segment for the beat information. When the candidate video segment is used as the target video segment, based on the material characteristics of the target video segment and the original video, when the target video is obtained, it can be guaranteed that the target video can restore the beat information of the original video to the greatest extent.

[0062] Step 204: Acquire multiple second candidate video segments from the video to be processed based on a preset video duration, where the preset video duration is less than the video duration of the video to be processed. It should be understood that the method for acquiring each second candidate video segment can refer to the method for acquiring the first candidate video segment above, and will not be repeated here.

[0063] In actual applications, when the exemplary embodiment of the present disclosure obtains the second candidate video segment from the video to be processed, the video duration difference can be determined based on the video duration of the second candidate video segment and the video duration of the original video. If the video duration difference is less than the video duration of the video to be processed, it means that the second video segment can be obtained from the video to be processed again. Therefore, the exemplary embodiment of the present disclosure can determine the preset video duration based on the video duration difference, so that the preset video duration is greater than the video duration difference.

[0064] Step 205: When the beat information of the second candidate video segment matches the beat information of the original video, the second candidate video segment is determined to be the video segment to be spliced.

[0065] In practical applications, the exemplary embodiments of the present disclosure may first refer to the method for obtaining the first candidate video segment, and obtain multiple beat segments of the original video at a preset video length from the beat information of the original video. If the beat information of the second candidate video segment matches the beat segment of the original video at the preset video length, it can be determined that the beat information of the second candidate video segment matches the beat information of the original video. Otherwise, it means that the first candidate video segment cannot match each beat segment of the original video at the preset video length, and the first candidate segment does not match the beat information of the original video. Therefore, the first candidate video segment can be replaced.

[0066] Exemplarily, the exemplary embodiments of the present disclosure can set a sliding window of a preset video length, and then obtain multiple original video segments from the original video by sliding the sliding window on the original video in chronological order. If an original video segment includes one or more frames of reference video images, a beat segment of the original video in the preset video length can be determined based on the beat features of all the reference video images included in the original video segment.

[0067] Exemplarily, for each second candidate video segment, it can be determined whether the beat information of the second candidate video segment is correlated with each beat segment of the original video in the preset video length. When the correlation between the beat information of the second candidate video segment and the beat segments of the original video in the preset video length is greater than the second preset correlation, it can be considered that the beat information of the second candidate video segment matches the beat segments of the original video in the preset video length.

[0068] An exemplary embodiment of the present disclosure can extract the beat information of each second candidate video segment through a 3D convolutional network, and then obtain the correlation between the beat information of the second candidate video segment and the beat segment of the original video at a preset video length through methods such as the Pearson correlation coefficient and the Spearman correlation coefficient.

[0069] When the Pearson correlation coefficient is used to calculate the correlation between the beat information of the second candidate video segment and the beat segments of the original video at the preset video length, the correlation between the beat information of the second candidate video segment and the beat segments of the original video at the preset video length can refer to the Pearson correlation coefficient between the beat information of the second candidate video segment and the beat segments of the original video at the preset video length, and accordingly, the second preset correlation can be expressed as the second preset correlation coefficient. In this case, when the Pearson correlation coefficient between the beat information of the second candidate video segment and the beat segments of the original video at the preset video length is greater than the second preset correlation coefficient, it can be considered that the beat information of the second candidate video segment matches the beat segments of the original video at the preset video length.

[0070] Step 206: updating the video to be processed based on the video segment to be spliced ​​and the video to be processed. Here, the video segment to be spliced ​​and the video to be processed can be spliced ​​to achieve the purpose of updating the video to be processed.

[0071] It can be seen that the exemplary embodiment of the present disclosure can obtain a second candidate video segment that matches the beat segment of the original video at the preset video duration from the video to be processed when the video duration of the video to be processed is less than or equal to the video duration of the original video, and use it as the video segment to be spliced. When splicing it into the video to be processed, the length of the video segment that matches the beat information of the video to be processed and the original video can be increased, thereby ensuring that the rhythm of the obtained target video segment is more matched with the rhythm of the original video.

[0072] When splicing the video segments to be spliced ​​and the video to be processed, the exemplary embodiments of the present disclosure can splice the video segments to be spliced ​​at any position of the video to be processed, or can splice the video segments to be spliced ​​at the splicing position of the video to be processed by setting the splicing position of the video to be processed.

[0073] Exemplarily, when obtaining the second candidate video segment, the area of ​​the video to be processed by the second candidate video segment can be recorded at the same time. If the second candidate video segment matches the beat segment of the original video in a preset video length, then the second candidate video segment is the video segment to be spliced. At the same time, the splicing position of the video to be processed is determined based on the area of ​​the second candidate video segment in the video to be processed. The splicing position of the video to be processed can be adjacent to the area of ​​the video to be processed by the second candidate video segment, or can be separated from the area of ​​the video to be processed by a preset time length. For example, the preset time length can be controlled within 600ms.

[0074] When the video segments to be spliced ​​are spliced ​​at the splicing position of the video to be processed by setting the splicing position of the video to be processed, the process can directly return to step 201 to detect the relationship between the video duration of the updated video to be processed and the video duration of the original video. If the video duration of the updated video to be processed is greater than the video length of the original video, the rhythm matching degree between the obtained target video segment and the original video can be increased, thereby further improving the audio and video synchronization of the obtained target video.

[0075] As a possible implementation manner, the original video of the exemplary embodiment of the present disclosure includes multiple frames of reference video images, and the beat information of the original video includes beat features of the multiple frames of reference video images. Figure 3 FIG. 2 shows a schematic diagram of a process for determining the beat features of a reference video image according to an exemplary embodiment of the present disclosure. Figure 3 As shown, the method of the exemplary embodiment of the present disclosure also includes:

[0076] Step 301: Obtaining positional features of preset key points of multiple frames of original video images included in the original video. The exemplary embodiment of the present disclosure can perform key point detection on the original video image through the target detection model, thereby obtaining positional features of preset key points. The positional features of the preset key points can be key point positional features of a certain target part. The target part can be a certain part of a movable target, or a certain part of a static target.

[0077] When the exemplary embodiment of the present disclosure detects the preset key points of the original video image, the preset key points may be preset feature points, corner points, etc. Taking the corner points as the preset key points as an example, since the corner points may be points with particularly prominent attributes in certain aspects, for example, in the original video image, the corner points may be the connecting points of the contour lines of the object. When the positions of the corner points change, the changes in the position features of the corner points are more obvious. Therefore, when the position features of the preset key points are the corner point position features, the changes in the preset key points between different original video images can be detected more sensitively.

[0078] Step 302: Based on the positional features of the preset key points of each frame of the original video image and the positional features of the preset key points of the next frame of the original video image, determine the positional change rate of the preset key points of each frame of the original video image. Here, the positional change rate of the preset key points of each frame of the original video image of the exemplary embodiment of the present disclosure can track the positional change rate of the preset key points of each frame of the original video image through an optical flow algorithm.

[0079] Step 303: If the position change rate of the preset key point of the original video image is greater than the preset position change rate, determine the beat feature of the reference video image based on the position change rate of the preset key point of the original video image. It should be understood that the key point position feature of the exemplary embodiment of the present disclosure is the position feature of multiple key points. Therefore, when judging whether the position change rate of the key point of the original video image is greater than the preset position change rate, it is actually necessary to judge whether the position change rate of each key point of the original video image is greater than the preset position change rate.

[0080] When the position change rate of each preset key point of the original video image of the current frame is greater than the preset position change rate, it means that the position change rate of the preset key points of the original video image of the current frame is relatively large, and the position difference between the preset key points of the original video image of the current frame and the previous frame is relatively large. The position change of the preset key points of the original video image of the current frame can reflect the beat characteristics of the original video. Therefore, the original video image of the current frame can be determined as the reference video image, and the beat characteristics of the reference video image can be determined based on the position change rate of the preset key points of the original video image of the current frame.

[0081] For example, the key point position change rate of the original video image of the current frame can be directly defined as the beat feature of the reference video image, or the image content change of the reference video image can be determined based on the key point position change rate of the original video image of the current frame, and the image content change of the reference video image can be defined as the beat feature of the reference video image.

[0082] Exemplarily, the exemplary embodiments of the present disclosure may use a corner extraction algorithm to extract corner features of an original video image. The following takes the Shi-Tomasi algorithm as an example to describe the corner feature extraction process of a frame of an original video image.

[0083] For a certain frame of original video image, you can use a fixed window to slide in any direction on the frame of original video image, and compare the degree of change in pixel grayscale of two areas corresponding to the frame of original video image before and after the fixed window slides. If there is a relatively large change in pixel grayscale when sliding in any direction, it can be considered that there is a corner point in the fixed window. The specific steps are as follows:

[0084] First, assuming that the fixed window currently corresponds to a sub-image area I of a certain frame of the original video image, the gradient I of the sub-image area in the x direction can be obtained: x and the gradient I of the sub-image area in the y direction y .

[0085]

[0086] Secondly, the gradient product of the sub-image area in two directions is calculated. The gradient product of the two directions can be the product of two gradients in the same direction or the product of two gradients in different directions. The product of the gradients in the two y directions of the sub-image area The product of the gradient in the x direction and the gradient in the y direction of the sub-image area is I xy =I x I y .

[0087] Again, the Gaussian function g(·) is used to multiply the gradient product of the two x directions of the sub-image area The product of the gradients in the two y directions of the sub-image area And the product of the gradient in the x direction and the gradient in the y direction of the sub-image area I xy Gaussian weighting is performed to generate element A, element C, and element B of the matrix M.

[0088] Among them, w represents the weight.

[0089] Next, the Harris response value of each pixel in the sub-image area I is calculated, and the Harris response value r that is less than the response value threshold t is set to zero, thereby obtaining a Harris response value set R of the pixels in the sub-image area.

[0090] Where, R = {R: detM-α(trM)2 >t}, where tr represents the trace of the matrix M, det represents the determinant of the matrix M, and α represents an empirical constant that can be in the range of (0.04, 0.06).

[0091] Finally, the Harris response value of the sub-image area is non-maximum suppressed in a window neighborhood of 3×3 or 5×5 scale, so as to obtain the pixel point corresponding to the maximum Harris response value in the sub-image area as the corner point.

[0092] As a possible implementation, the exemplary embodiment of the present disclosure may further perform transition settings on the target video segment before acquiring the target video based on the material features of the target video segment and the original video. Figure 4 The following is a schematic diagram of the transition setting process of the exemplary embodiment of the present disclosure. Figure 4 As shown, the method of the exemplary embodiment of the present disclosure may also include:

[0093] Step 401: Acquire a target transition feature, which matches the target video images before and after the transition of the target video segment.

[0094] The target video images before and after the transition of the exemplary embodiment of the present disclosure may include the pre-transition video image and the post-transition video image included in the target video segment. The pre-transition video image may be one frame or multiple frames, and the post-transition video image may be one frame or multiple frames.

[0095] In practical applications, the grayscale difference of each pixel between two adjacent target video images included in the target video segment can be calculated, and its absolute value can be taken to obtain a differential image of the two adjacent target video images. If the minimum value of the grayscale difference of the differential image exceeds the preset difference, it can be considered that there is a transition between the two adjacent target video images. At this time, the two adjacent target video images can be used as the target video images before and after the transition, or the video segment containing the two adjacent target video images can be used as the target video images before and after the transition.

[0096] On this basis, the matching degree between the target video image before and after the transition and different preset transition features can be compared. If the matching degree between a target video image before and after the transition and a preset transition feature is greater than the preset matching degree, the preset transition feature can be considered as the target transition feature. Here, the preset transition feature can be determined based on the pre-stored transition feature, or it can be determined based on the transition feature of the original video.

[0097] When determining the preset transition feature based on the transition feature of the original video, if the matching degree between a target video image before and after the transition and a preset transition feature is greater than the preset matching degree, the preset transition feature can be used as part of the video template, and when the video template is applied to the target video clip, the preset transition feature acquisition operation can be completed for the video template at the same time. The preset transition feature can be various transition types, such as flash to white, flash to black, fade in and fade out, wipe transition, superimpose, use of empty shots, etc., but is not limited thereto.

[0098] Step 402: Setting the target video images before and after the transition based on the target transition feature. The setting of the transition here can be processing the target video images before and after the transition based on the target transition feature, or inserting a new transition segment to realize the target transition feature.

[0099] When the target transition feature matches the target video images before and after the transition of the target video clip, the minimum grayscale difference of the differential image of the two adjacent frames of the target video image included in the target video image before and after the transition exceeds the preset difference. Therefore, the target transition feature is also related to the beat information of the target video clip. Therefore, by setting the transition of the target video images before and after the transition based on the target transition feature, it can also be ensured that the transition effect of the obtained target video matches the video beat, thereby further improving the target video effect.

[0100] As a possible implementation, the material features of the exemplary embodiment of the present disclosure may include multiple groups of visual material features, each group of visual material features includes basic information of the visual material and the retention time of the basic information of the visual material. Here, in terms of material type, it can also include text information of the original video, and even visual materials such as special effect information and sticker information. If the material type of the material feature is not distinguished, according to the feature type, the material features of the exemplary embodiment of the present disclosure may include material content, material display location, and material style, etc.

[0101] In practical applications, the method of the exemplary embodiment of the present disclosure may also include: if the basic information of the visualization material of multiple consecutive frames of original video images is the same, aggregating the basic information of the visualization material of the multiple consecutive frames of original video images to obtain the aggregation result of the basic information of the visualization material, determining the retention time of the basic information of the visualization material based on the timestamps of the multiple consecutive frames of original video images, and obtaining a set of visualization material features based on the aggregation result of the basic information of the visualization material and the retention time of the original video images.

[0102] Exemplarily, when the basic information of the visualization material of two adjacent frames of original video images is the same, the material content, material display position and material style may be the same. When aggregating the basic information of the visualization material of two adjacent frames of original video images, the frame numbers of the two adjacent frames of original video images may be obtained as the timestamps of the two adjacent frames of original video images, which are used to determine the retention time of the basic information of the visualization material.

[0103] When obtaining a target video based on the material features of the target video segment and the original video, the visual material features can be migrated to the target video segment based on the basic information of the visual material included in each group of visual material features and the retention time of the basic information of the visual material. For example, when the visual material is the subtitles of the original video, the subtitles of the original video can be migrated to the target video segment according to the content of the subtitles, the layout area of ​​the subtitles, the style of the subtitles, etc., so as to obtain the target video.

[0104] The exemplary embodiments of the present disclosure also provide a method for generating a video template, which can automatically and quickly generate a video template, and the application effect of the video template is relatively good, which can ensure that the generated target video has a relatively good audio and video synchronization effect.

[0105] Figure 5 FIG. 2 is a flow chart showing a method for generating a video template according to an exemplary embodiment of the present disclosure. Figure 5 As shown, the method for generating a video template of an exemplary embodiment of the present disclosure may include:

[0106] Step 501: extracting the material information of the original video, the material information of the original video at least includes the audio features of the original video. The audio features of the original video may include the audio features of at least one sound source, for example, the audio of these sound sources may include background music, human voice, environmental sound, and even equipment noise.

[0107] In practical applications, the original audio can be denoised, and then the original video can be used for audio recognition and separation based on a deep neural network-based music source feature extraction and separation algorithm (for example: SA-CEDN-4FEM) to obtain background music and vocals, etc.

[0108] Step 502: Extract node information of the original video, where the node information includes at least the beat information of the original video. The beat information of the original video here can refer to the previous text and will not be repeated here. In order to increase the beat information extraction effect of the original video, the exemplary embodiment of the present disclosure can also perform noise reduction and correction processing on the original video to ensure that the original video is not distorted.

[0109] Step 503: Generate a video template based on the material information of the original video and the node information of the original video. When the video template is used to generate a target video, the above-mentioned video generation method can be referred to, which will not be described in detail here.

[0110] The exemplary embodiment of the present disclosure obtains a video template for users to quickly edit by intelligently disassembling and converting the multimodal elements of the original video, which can not only improve the generation speed of the video template and greatly shorten the creation speed of the video template, but also directly produce a target video with the same style as the original video when the video creator does not have professional video editing technology. Moreover, the video creator can obtain a video template that can be used out of the box for any video through the method of the exemplary embodiment of the present disclosure, which not only improves the efficiency and quality of video production, but also saves time and cost for users.

[0111] The exemplary embodiments of the present disclosure can also upload the video template to the video creation platform after generating the video template, if authorized, for use by video creators, thereby increasing the number and types of video templates on the video creation platform and improving the content ecology richness of the video creation platform;

[0112] As a possible implementation, the material features of the exemplary embodiment of the present disclosure further include multiple groups of visual material features of the original video, each group of visual material features including basic information of the visual material and retention time of the basic information of the visual material. Here, the method of the exemplary embodiment of the present disclosure may further include:

[0113] Extracting basic information of visualization material of each frame of the original video image included in the original video; if the basic information of visualization material of multiple consecutive frames of original video images is the same, aggregating the basic information of visualization material of the multiple consecutive frames of original video images to obtain an aggregation result of the basic information of visualization material; determining the retention time of the basic information of visualization material based on the timestamps of the multiple consecutive frames of original video images; and obtaining a set of visualization material features based on the aggregation result of the basic information of visualization material and the retention time of the basic information of visualization material.

[0114] In terms of material types, it can also include text information of the original video, and even visual materials such as special effect information and sticker information. If the material type of material features is not distinguished, the material features of the exemplary embodiment of the present disclosure can include material content, material display location, and material style, etc., according to the feature type. The following takes the text information of the original video as an example to describe the method for extracting text features of the original video of the exemplary embodiment of the present disclosure.

[0115] First, preprocess each frame of the original video image included in the original video to obtain the text image corresponding to each frame of the original video image. For example, grayscale, binarization, noise reduction, tilt correction and text segmentation can be performed on each frame of the original video image to obtain the corresponding text image.

[0116] Secondly, style and position detection and character recognition are performed on each character of the text image to obtain the text style, text display position and text content.

[0117] For example, the text image can be extracted through a convolutional network to obtain the image features of the text image. If the text image contains Chinese characters or other characters with a large number of characters, the image features of the text image can also be reduced in dimension to reduce the amount of subsequent data processing. Image features can be classified through a classifier to obtain the display position and character style of each character. The character style here can be character color, character font, character special effects, character direction, character length, character overlap, character density, etc. At the same time, natural language models such as RNN, CRNN, transformer and other natural language models can also be used to identify the image features of text images to obtain the character content included in the text.

[0118] Next, considering that the time interval between adjacent frames of original video images is relatively short, it is possible that the original video images of consecutive frames contain the same basic text information. Therefore, all character features of two adjacent frames of text images can be compared. If all character features are the same, all character features of the two adjacent frames can be merged. At the same time, the retention time of each character can be updated to obtain a set of character features.

[0119] Finally, a text feature that meets the format requirements may be generated according to the saving format of the text feature, and the saving of the text feature may include (character content, character position, character style, and character retention time).

[0120] As a possible implementation manner, the original video of the exemplary embodiment of the present disclosure includes multiple frames of reference video images, the beat information of the original video includes beat features of the multiple frames of reference video images, and the node information of the original video is extracted, including:

[0121] The position features of preset key points of multiple frames of original video images included in the original video are obtained, and the position change rate of the preset key points of each original video image is determined based on the position features of the preset key points of each frame of the original video image and the position features of the preset key points of the next frame of the original video image. If the position change rate of the preset key points of the original video image is greater than the preset position change rate, the beat features of the reference video image are determined based on the position changes of the preset key points of the original video image.

[0122] When the node information of the original video of the exemplary embodiment of the present disclosure also includes the transition features of the original video, extracting the node information of the original video may also include: obtaining the grayscale change of two adjacent frames of original video images included in the original video, if the grayscale change of two adjacent frames of original video images is greater than a preset grayscale change, determining the transition features of the original video based on the two adjacent frames of the original video images.

[0123] One or more technical solutions provided in the exemplary embodiments of the present disclosure can extract a target video segment that matches the beat information of the original video from the video to be processed, and because the material features of the original video at least include the audio features of the original video, the audio features of the original video match the beat information of the target video segment. On this basis, when obtaining the target video based on the material features of the target video segment and the original video, the audio features of the original video are essentially added to the target video segment to ensure that the audio of the obtained target video matches the beat.

[0124] In addition, the exemplary embodiment of the present disclosure can combine the beat information and material information of the original video with the video to be processed to synthesize a target video when the beat information and material information of the original video are obtained. The target video is not limited by the types of video templates provided by the video creation platform. Therefore, when a video creator of the exemplary embodiment of the present disclosure sees a video of interest but does not have a video template, he or she can directly use the video as the original video, and use the beat information and material information of the original video to create a target video that suits his or her interests, thereby improving the efficiency of target video production.

[0125] In summary, the exemplary embodiments of the present disclosure can parse the original video for multimodal elements of text, audio, and images, and convert them into video templates that can be edited by users, so as to achieve the purpose of automatically and quickly generating video templates, which can greatly shorten the production process of video templates. For video creators, video creators save video templates on the video creation platform. Under authorized use, the number and types of templates on the video creation platform can be increased, and the richness of the platform's content ecology can be improved; for video creators, video creators can use out-of-the-box video templates for any video, which improves the efficiency and quality of video production, and can also save time and costs for video creators.

[0126] Moreover, the video template generation method of the exemplary embodiment of the present invention can be connected with the video editing link to accurately insert resources such as text features, audio features, transition features, etc. into the time nodes of the corresponding target video clips to generate a target video, which supports further editing, resource replacement and key frame fine-tuning.

[0127] The above mainly introduces the solution provided by the embodiment of the present disclosure from the perspective of the server. It is understandable that, in order to implement the above functions, the server includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0128] The embodiments of the present disclosure may divide the server into functional units according to the above method examples. For example, each functional module may be divided corresponding to each function, or two or more functions may be integrated into one processing module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present disclosure is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0129] In the case of dividing each functional module according to each function, an exemplary embodiment of the present disclosure provides a video generating device, which may be a server or a chip applied to a server. Figure 6 FIG. 2 shows a schematic block diagram of functional modules of a video generation device according to an exemplary embodiment of the present disclosure. Figure 6 As shown, the video generating device 600 includes:

[0130] An extraction module 601 is used to extract a target video segment that matches the beat information of the original video from the video to be processed;

[0131] The acquisition module 602 is used to acquire the target video based on the target video segment and the material features of the original video, where the material features at least include the audio features of the original video.

[0132] In one possible implementation, the extraction module 601 is used to obtain multiple first candidate video segments from the video to be processed when the video length of the video to be processed is greater than the video length of the original video, and if the beat information of the first candidate video segment matches the beat information of the original video, determine that the first candidate video segment is the target video segment, and the video length of the first candidate video segment is equal to the video length of the original video.

[0133] In a possible implementation, if the correlation between the beat information of the first candidate video segment and the beat information of the original video is greater than a first preset correlation, the beat information of the first candidate video segment matches the beat information of the original video.

[0134] In a possible implementation, the extraction module 601 is also used to obtain multiple second candidate video segments from the video to be processed based on a preset video length when the video length of the video to be processed is less than or equal to the video length of the original video, and the preset video length is less than the video length of the video to be processed; if the beat information of the second candidate video segment matches the beat information of the original video, determine that the second candidate video segment is the video segment to be spliced, and update the video to be processed based on the video segment to be spliced ​​and the video to be processed.

[0135] In one possible implementation, the extraction module 601 is used to determine a video duration difference based on the video duration of the second candidate video segment and the video duration of the original video. If the video duration difference is less than the video duration of the video to be processed, the preset video duration is determined based on the video duration difference, and the preset video duration is greater than the video duration difference.

[0136] In a possible implementation, the extraction module 601 is used to obtain multiple beat segments of the original video in the preset video length from the beat information of the original video. If the beat information of the second candidate video segment matches each beat segment of the original video in the preset video length, it is determined that the beat information of the second candidate video segment matches the beat information of the original video.

[0137] In a possible implementation, when the correlation between the beat information of the second candidate video segment and the beat segment of the original video at the preset video length is greater than a second preset correlation, the beat information of the second candidate video segment matches the beat segment of the original video at the preset video length.

[0138] In a possible implementation, the original video includes multiple frames of reference video images, and the beat information of the original video includes beat features of the multiple frames of reference video images;

[0139] The extraction module 601 is also used to obtain the position features of preset key points of multiple frames of original video images included in the original video, and determine the position change rate of the preset key points of each frame of the original video image based on the position features of the preset key points of each frame of the original video image and the position features of the preset key points of the next frame of the original video image. If the position change rate of the preset key points of the original video image is greater than the preset position change rate, the beat features of the reference video image are determined based on the position change rate of the key points of the original video image.

[0140] In one possible implementation, the acquisition module 602 is used to acquire target transition features, and the target transition features match the target video images before and after the transition of the target video clip. The device also includes a setting module 603, and the setting module 603 is used to perform transition settings for the target video images before and after the transition based on the target transition features.

[0141] In a possible implementation, the material features also include multiple groups of visual material features of the original video, each group of visual material features includes basic information of the visual material and the retention time of the basic information of the visual material, and the extraction module 601 is also used to extract the basic information of the visual material of each frame of the original video image included in the original video; if the basic information of the visual material of multiple consecutive frames of the original video image is the same, the basic information of the visual material of the multiple consecutive frames of the original video image is aggregated to obtain the aggregation result of the basic information of the visual material; based on the timestamps of the multiple consecutive frames of the original video image, the retention time of the basic information of the visual material is determined; based on the aggregation result of the basic information of the visual material and the retention time of the basic information of the visual material, a group of visual material features is obtained.

[0142] In the case of dividing each functional module according to each function, an exemplary embodiment of the present disclosure provides a device for generating a video template, and the device for generating a video template may be a server or a chip applied to a server. Figure 7 FIG. 1 shows a schematic block diagram of functional modules of a device for generating a video template according to an exemplary embodiment of the present disclosure. Figure 7 As shown, the video template generation device 700 includes:

[0143] An extraction module 701 extracts material information of an original video and node information of the original video, wherein the material information of the original video at least includes audio features of the original video, and the node information at least includes beat information of the original video;

[0144] The generating module 702 is used to generate a video template based on the material information of the original video and the node information of the original video.

[0145] In a possible implementation, the material features also include multiple groups of visual material features of the original video, each group of the visual material features includes basic information of the visual material and retention time of the basic information of the visual material, the extraction module 701 is used to extract the basic information of the visual material of each frame of the original video image included in the original video, if the basic information of the visual material of multiple consecutive frames of the original video image is the same, aggregate the basic information of the visual material of the multiple consecutive frames of the original video image to obtain the aggregation result of the basic information of the visual material, and determine the retention time of the basic information of the visual material based on the timestamps of the multiple consecutive frames of the original video image;

[0146] In a possible implementation, the original video includes multiple frames of reference video images, and the beat information of the original video includes beat features of the multiple frames of reference video images. The extraction module 701 is used to obtain the position features of preset key points of the multiple frames of original video images included in the original video, and determine the position change rate of the preset key points of each original video image based on the position features of the preset key points of each frame of the original video image and the position features of the preset key points of the next frame of the original video image. If the position change rate of the preset key points of the original video image is greater than the preset position change rate, the beat features of the reference video image are determined based on the position change of the preset key points of the original video image.

[0147] In a possible implementation, the node information of the original video also includes the transition features of the original video, and the extraction module 701 is also used to obtain the grayscale change of two adjacent frames of original video images included in the original video. If the grayscale change of the two adjacent frames of the original video images is greater than a preset grayscale change, the transition features of the original video are determined based on the two adjacent frames of the original video images.

[0148] Figure 8 Schematic block diagram of a chip according to an exemplary embodiment of the present disclosure is shown. Figure 8 As shown, the chip 800 includes one or more (including two) processors 801 and a communication interface 802. The communication interface 802 can support the server to perform the data sending and receiving steps in the above image processing method, and the processor 801 can support the server to perform the data processing steps in the above image processing method.

[0149] Optional, such as Figure 8 As shown, the chip 800 also includes a memory 803, which may include a read-only memory and a random access memory, and provides operation instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory (NVRAM).

[0150] In some embodiments, Figure 8 As shown, the processor 801 performs corresponding operations by calling the operation instructions stored in the memory (the operation instructions may be stored in the operating system). The processor 801 controls the processing operations of any one of the terminal devices, and the processor may also be called a central processing unit (CPU). The memory 803 may include a read-only memory and a random access memory, and provides instructions and data to the processor 801. A portion of the memory 803 may also include NVRAM. For example, in an application, the memory, the communication interface, and the memory are coupled together through a bus system, wherein the bus system may include a power bus, a control bus, and a status signal bus in addition to a data bus. However, for the sake of clarity, in Figure 8 Various buses are labeled as bus system 804 .

[0151] The method disclosed in the above-mentioned embodiment of the present disclosure can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (digital signal processing, DSP), an ASIC, a field-programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiment of the present disclosure can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware.

[0152] The exemplary embodiment of the present disclosure also provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to cause the electronic device to perform the method according to the embodiment of the present disclosure when executed by the at least one processor.

[0153] The exemplary embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform the method according to the embodiments of the present disclosure.

[0154] The exemplary embodiments of the present disclosure further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, the computer is used to enable the computer to perform the method according to the embodiments of the present disclosure.

[0155] refer to Fig. 9 , a block diagram of an electronic device 900 that can be used as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0156] like Fig. 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0157] like Fig. 9As shown, multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. The input unit 906 can be any type of device that can input information to the electronic device 900, and the input unit 906 can receive input digital or character information, and generate key signal input related to the user settings and / or function control of the electronic device. The output unit 907 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 908 can include but is not limited to a disk, an optical disk. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth TM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0158] like Fig. 9 As shown, the computing unit 901 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above. For example, in some embodiments, the method of the exemplary embodiment of the present disclosure may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via ROM 902 and / or communication unit 909. In some embodiments, the computing unit 901 can be configured to perform the method of the exemplary embodiment of the present disclosure in any other appropriate manner (e.g., by means of firmware).

[0159] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0160] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0161] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0162] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0163] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0164] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.

[0165] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the process or function described in the embodiment of the present disclosure is executed in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user device or other programmable device. The computer program or instruction may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instruction may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired or wireless means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server, data center, etc. that integrates one or more available media. The available medium may be a magnetic medium, for example, a floppy disk, a hard disk, a tape; it may also be an optical medium, for example, a digital video disc (DVD); it may also be a semiconductor medium, for example, a solid state drive (SSD).

[0166] Although the present disclosure has been described in conjunction with specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present disclosure. Accordingly, this specification and the drawings are merely exemplary illustrations of the present disclosure as defined by the appended claims, and are deemed to have covered any and all modifications, variations, combinations or equivalents within the scope of the present disclosure. Obviously, those skilled in the art may make various modifications and variations to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is also intended to include these modifications and variations.

Claims

1. A video generation method, It is characterized in that The method comprises: Extracting a target video segment matching the beat information of the original video from the video to be processed; A target video is acquired based on the target video segment and material features of the original video, wherein the material features at least include audio features of the original video.

2. The method according to claim 1, It is characterized in that The step of extracting a target video segment matching the beat information of the original video from the video to be processed includes: When the video duration of the video to be processed is greater than the video duration of the original video, obtaining a plurality of first candidate video segments from the video to be processed, wherein the video duration of the first candidate video segments is equal to the video duration of the original video; If the beat information of the first candidate video segment matches the beat information of the original video, the first candidate video segment is determined to be the target video segment.

3. The method according to claim 2, It is characterized in that If the correlation between the beat information of the first candidate video segment and the beat information of the original video is greater than a first preset correlation, the beat information of the first candidate video segment matches the beat information of the original video.

4. The method according to claim 2, It is characterized in that The step of extracting a target video segment matching the beat information of the original video from the video to be processed further includes: When the video duration of the video to be processed is less than or equal to the video duration of the original video, obtaining a plurality of second candidate video segments from the video to be processed based on a preset video duration, wherein the preset video duration is less than the video duration of the video to be processed; If the beat information of the second candidate video segment matches the beat information of the original video, determining that the second candidate video segment is the video segment to be spliced; The video to be processed is updated based on the video segments to be spliced ​​and the video to be processed.

5. The method according to claim 4, It is characterized in that The step of extracting a target video segment matching the beat information of the original video from the video to be processed further includes: Determine a video duration difference based on the video duration of the second candidate video segment and the video duration of the original video; If the video duration difference is smaller than the video duration of the to-be-processed video, the preset video duration is determined based on the video duration difference, and the preset video duration is greater than the video duration difference.

6. The method according to claim 4, It is characterized in that The step of extracting a target video segment matching the beat information of the original video from the video to be processed further includes: Acquire multiple beat segments of the original video at the preset video duration from the beat information of the original video; If the beat information of the second candidate video segment matches each beat segment of the original video in the preset video duration, it is determined that the beat information of the second candidate video segment matches the beat information of the original video.

7. The method according to claim 6, It is characterized in that When the correlation between the beat information of the second candidate video segment and the beat segment of the original video at the preset video duration is greater than the second preset correlation, the beat information of the second candidate video segment matches the beat segment of the original video at the preset video duration.

8. The method according to any one of claims 1 to 7, It is characterized in that The original video includes multiple frames of reference video images, the beat information of the original video includes beat features of the multiple frames of reference video images, and the method further includes: Acquire position features of preset key points of multiple frames of original video images included in the original video; Determine the position change rate of the preset key points of each frame of the original video image based on the position features of the preset key points of each frame of the original video image and the position features of the preset key points of the next frame of the original video image; If the position change rate of the preset key point of the original video image is greater than the preset position change rate, the beat feature of the reference video image is determined based on the position change rate of the preset key point of the original video image.

9. The method according to any one of claims 1 to 7, It is characterized in that The method further comprises: Acquire a target transition feature, wherein the target transition feature matches the target video images before and after the transition of the target video segment; The target video images before and after the transition are set for transition based on the target transition feature.

10. The method according to any one of claims 1 to 7, It is characterized in that The material features further include multiple groups of visual material features of the original video, each group of visual material features includes basic information of the visual material and retention time of the basic information of the visual material, and the method further includes: Extracting basic information of visual materials of each frame of original video image included in the original video; If the basic information of the visualization material of the consecutive multiple frames of the original video image is the same, aggregating the basic information of the visualization material of the consecutive multiple frames of the original video image to obtain an aggregation result of the basic information of the visualization material; Determining the retention time of the basic information of the visualization material based on the timestamps of the continuous multiple frames of the original video image; Based on the aggregation result of the basic information of the visualization material and the retention time of the basic information of the visualization material, a group of visualization material features is obtained.

11. A video generating device, It is characterized in that include: An extraction module, used for extracting a target video segment matching the beat information of the original video from the video to be processed; The acquisition module is used to obtain the target video based on the target video segment and the material features of the original video, wherein the material features at least include the audio features of the original video.

12. An electronic device, It is characterized in that include: processor; as well as, A memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 10.

13. A non-transitory computer-readable storage medium, It is characterized in that The non-transitory computer-readable storage medium stores computer instructions for causing the computer to execute the method according to any one of claims 1 to 10.