Video generation method and apparatus, device, and medium
By extracting the target video clips that match the original video beat information and generating a video template based on the material characteristics, the problem of poor coordination between the target video audio and the picture in the prior art is solved, and a more efficient video creation process is achieved.
Patent Information
- Application Number
- PCT/CN2024/136507
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-04
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-12
AI Technical Summary
When the existing video creation platform applies video templates, the audio and picture matching of the target video is poor, and the types of video templates are limited, which limits the creative freedom of video creators.
The target video clips matching the original video beat information from the to-process video and obtain the target video based on the material characteristics of the target video clip and the original video. The method includes extracting material information and node information of the original video and generating a video template to achieve matching audio and beats.
It improves the matching of the audio and beats of the target video, liberates the creative limitations of video creators, and allows them to create target videos that meet their interests more freely, improving the efficiency of video production.
Smart Images

Figure CN2024136507_12062025_PF_FP_ABST
Abstract
Description
Video generation method, device, equipment and medium
[0001] This application claims priority to Chinese Patent Application No. 202311651138.3 filed on December 4, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The present disclosure relates to a video generation method, apparatus, device and medium. Background Art
[0003] When creating videos on a video creation platform, when a video creator sees a video template of interest, he or she can apply the video template of interest to the source video file selected by the video creator with one click to generate a target video with the same style as the video template.
[0004] In related technologies, the application of video templates to source video files is not very effective, resulting in poor audio and visual coordination in the resulting target video. Furthermore, video creation platforms offer a limited variety of video templates, which cannot provide sufficient video templates for video creators. Therefore, video creators find it difficult to freely create target videos that suit their interests. Summary of the Invention
[0005] According to one aspect of the present disclosure, a video generation method is provided, the method comprising:
[0006] Extracting a target video segment that matches the beat information of the original video from the video to be processed;
[0007] A target video is acquired based on the target video clip and material features of the original video, where the material features at least include audio features of the original video.
[0008] According to another aspect of the present disclosure, a method for generating a video template is provided, the method comprising:
[0009] Extracting material information of the original video, where the material information of the original video at least includes audio features of the original video;
[0010] Extracting node information of the original video, where the node information at least includes beat information of the original video;
[0011] A video template is generated based on the material information of the original video and the node information of the original video.
[0012] According to another aspect of the present disclosure, there is provided a video generating apparatus, comprising:
[0013] An extraction module is used to extract a target video segment that matches the beat information of the original video from the video to be processed;
[0014] The acquisition module is used to acquire the target video based on the target video clip and the material features of the original video, where the material features at least include the audio features of the original video.
[0015] According to another aspect of the present disclosure, a device for generating a video template is provided, comprising:
[0016] An extraction module extracts material information of an original video and node information of the original video, wherein the material information of the original video at least includes audio features of the original video, and the node information at least includes beat information of the original video;
[0017] A generation module is used to generate a video template based on the material information of the original video and the node information of the original video.
[0018] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0019] processor; and,
[0020] Memory for storing programs;
[0021] The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to the exemplary embodiment of the present disclosure.
[0022] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to the exemplary embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0024] FIG1 is a schematic flow chart showing a video generation method according to an exemplary embodiment of the present disclosure;
[0025] FIG2 shows a schematic diagram of a target video segment acquisition process according to an exemplary embodiment of the present disclosure;
[0026] FIG3 shows a schematic diagram of a process for determining beat features of a reference video image according to an exemplary embodiment of the present disclosure;
[0027] FIG4 illustrates a schematic diagram of a transition setting process according to an exemplary embodiment of the present disclosure;
[0028] FIG5 shows a schematic flow chart of a method for generating a video template according to an exemplary embodiment of the present disclosure;
[0029] FIG6 shows a schematic block diagram of functional modules of a video generating device according to an exemplary embodiment of the present disclosure;
[0030] FIG7 shows a schematic block diagram of functional modules of a device for generating a video template according to an exemplary embodiment of the present disclosure;
[0031] FIG8 shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure; and
[0032] FIG9 shows a block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION
[0033] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0034] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0035] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0036] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0037] When the image changes at a rate of more than 24 frames per second, the human eye cannot distinguish a single static image due to the persistence of vision principle, resulting in a smooth, continuous visual effect. This continuous image is called a video. With the development of video technology, video formats have become more diverse, and it is generally multimodal data composed of multimedia formats such as images, text, and audio. For example, when video data is multimodal, it generally contains video frames, subtitles, and sound information.
[0038] Considering that the multimodal data contained in videos is richer than traditional graphic media, more and more video creators choose videos as a medium for information dissemination. To improve the efficiency of video production, video creators can apply video templates to source video files to generate target videos with the same style as the video templates.
[0039] Early video templates were typically pre-designed and produced by professional video producers, typically in the form of video process materials or video templates that could be re-edited. In the era of short videos, video creation platforms can now create one-click templates for popular video types. Video creators can select a simple video clip, apply the one-click template to it, and directly generate a target video with the same style as the one-click template.
[0040] It can be seen that whether video templates are made by professional video producers or one-click templates for popular video types can be made on video creation platforms, manual participation in production is required, and the application effect of video templates is not good, resulting in poor fluency of the target video obtained, such as poor matching of audio and video rhythm of the target video, and the video creation platform provides a limited variety of video templates, which cannot provide enough video templates for video creators. Therefore, it is difficult for video creators to freely create target videos that suit their interests.
[0041] In response to the above problems, an exemplary embodiment of the present disclosure provides a video generation method, which can not only generate a target video through the beat information and material information of the original video to improve the audio and beat matching of the target video, but also get rid of the limitation on the number of video templates of the video creation platform, and directly use the original video to create a target video that suits one's own interests, thereby improving the efficiency of target video production.
[0042] In practical applications, the beat information and material information of the original video of the exemplary embodiment of the present disclosure can be stored in the form of a video template, which can be stored in a video creation platform, etc.
[0043] The video generation method provided by the exemplary embodiment of the present disclosure can be applied to a server or a chip in the server that interacts with a client. The video generation method of the exemplary embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.
[0044] FIG1 shows a flow chart of a video generation method according to an exemplary embodiment of the present disclosure. As shown in FIG1 , the video generation method according to an exemplary embodiment of the present disclosure may include:
[0045] Step 101: Extracting a target video segment that matches the tempo information of the original video from the video to be processed. This may be in response to applying a video template to the video to be processed, obtaining the tempo information of the original video from the video template, and then extracting a target video segment that matches the tempo information of the original video from the video to be processed.
[0046] In practical applications, the video duration of the video to be processed may be greater than the video duration of the original video, or may be less than or equal to the video duration of the original video. If the video duration of the video to be processed is greater than the video duration of the original video, a target video segment that matches the tempo information of the original video may be directly extracted from the video to be processed. If the video duration of the video to be processed is less than or equal to the video duration of the original video, the video to be processed may be first expanded so that the video duration of the video to be processed is greater than the video duration of the original video, and then a target video segment that matches the tempo information of the original video may be extracted from the video to be processed.
[0047] The tempo information of the original video of the exemplary embodiment of the present disclosure can reflect the frequency of large-scale action changes or scene switching in the original video. Therefore, the tempo information of the original video of the exemplary embodiment of the present disclosure can include the tempo characteristics of multiple frames of reference video images. For example, by detecting the image content change rate, a reference video image with a change rate greater than a preset change rate can be obtained, and then the image content change rate of each frame of the reference video image can be defined as the tempo characteristic of the reference video image.
[0048] Step 102: Based on the target video clip and the material features of the original video, the target video is obtained, where the material features include at least the audio features of the original video. Here, the material features of the original video may be obtained while simultaneously obtaining the beat information of the original video from the video template.
[0049] Considering that the material features of the original video include at least its audio features, and the target video clip matches the tempo information of the original video, the audio features of the original video match the tempo information of the target video clip. When obtaining the target video based on the material features of the target video clip and the original video, the audio features of the original video are essentially added to the target video clip to ensure that the audio and tempo of the obtained target video match. In this case, the application effect of the video template on the video to be processed can be guaranteed.
[0050] In addition, the exemplary embodiment of the present disclosure can combine the beat information and material information of the original video with the video to be processed to synthesize a target video when the beat information and material information of the original video are obtained. The target video is not limited by the types of video templates provided by the video creation platform. Therefore, when a video creator of the exemplary embodiment of the present disclosure sees a video of interest but does not have a video template, he or she can directly use the video as the original video, and use the beat information and material information of the original video to create a target video that suits his or her interests, thereby improving the efficiency of target video production and saving users time and costs.
[0051] As a possible implementation, FIG2 shows a schematic diagram of a target video segment acquisition process of an exemplary embodiment of the present disclosure. As shown in FIG2, the exemplary embodiment of the present disclosure extracts a target video segment that matches the beat information of the original video from the video to be processed, which may include:
[0052] Step 201: Detect the relationship between the video duration of the video to be processed and the video duration of the original video. If the video duration of the video to be processed is greater than the video duration of the original video, it indicates that a target video segment with the same video duration as the original video can be obtained from the video to be processed, and therefore, step 202 can be executed. If the video duration of the video to be processed is less than or equal to the video duration of the original video, it indicates that a target video segment with the same video duration as the original video cannot be directly obtained from the video to be processed, and therefore, step 204 can be executed.
[0053] Step 202: Acquire multiple first candidate video segments from the video to be processed, where the video duration of the first candidate video segments is equal to the video duration of the original video.
[0054] In practical applications, a sliding window equal to the video duration of the original video can be defined. The window is then controlled to slide along the time dimension of the processed video to obtain the first candidate video segment. Each time the window slides, a first candidate video segment is captured from the original video.
[0055] Step 203: When the tempo information of the first candidate video segment matches the tempo information of the original video, the first candidate video segment is determined to be the target video segment. It should be understood that if the correlation between the tempo information of the first candidate video segment and the tempo information of the original video is greater than a first preset correlation, then the tempo information of the first candidate video segment can be considered to match the tempo information of the original video.
[0056] In practical applications, when an original video includes multiple frames of reference video images, the tempo information of the original video includes the tempo features of the multiple frames of reference video images. Specifically, if the original video includes multiple consecutive frames of original video images, and the position change rate of a preset key point in a frame of original video image is greater than the preset position change rate, the frame of original video image can be considered as the reference video image of the original video. The tempo features of the reference video image can also be determined based on the position change rate of the preset key point in the frame of original video image.
[0057] For example, a sliding window can be used to slide in the time dimension of the video to be processed to obtain multiple first candidate video segments, and the beat information of each first candidate video segment can be extracted using a 3D convolutional network. Then, the correlation between the beat information of each first candidate video segment and the original video is calculated using methods such as the Pearson correlation coefficient and the Spearman correlation coefficient.
[0058] When the Pearson correlation coefficient is used to calculate the correlation between the beat information of each first candidate video segment and the original video, the correlation between the beat information of the first candidate video segment and the original video can be expressed by the Pearson correlation coefficient of the beat information of the first candidate video segment and the original video. Accordingly, the first preset correlation can be expressed as the first preset correlation coefficient. In this case, when the Pearson correlation coefficient of the beat information of the first candidate video segment and the original video is greater than the first preset correlation coefficient, it can be considered that the beat information of the first candidate video segment matches the beat information of the original video.
[0059] If the Pearson correlation coefficient between the beat information of a first candidate video clip and the original video is greater than the first preset correlation coefficient, it can be considered that the first candidate video clip meets the requirements of the target video clip for the beat information. Therefore, the candidate video clip can be determined to be the target video clip. Otherwise, it means that the first candidate video clip does not meet the requirements of the target video clip for the beat information, and a new first candidate video clip can be obtained again.
[0060] The above-mentioned first preset correlation coefficient can be the maximum value of the correlation between the beat information of multiple first candidate video clips and the original video, or it can be a custom preset correlation coefficient. For example: when the first preset correlation coefficient is the maximum value of the correlation between the beat characteristics of multiple first candidate video clips and the original video, after obtaining the correlation between the beat characteristics of each first candidate video clip and the original video, it can be considered that the first candidate video with the greatest correlation meets the requirements of the target video clip for the beat information. When the candidate video clip is used as the target video clip, based on the material characteristics of the target video clip and the original video, when the target video is obtained, it can be guaranteed that the target video can restore the beat information of the original video to the greatest extent.
[0061] Step 204: Obtain multiple second candidate video segments from the video to be processed based on a preset video duration, where the preset video duration is less than the video duration of the video to be processed. It should be understood that the method for obtaining each second candidate video segment can refer to the method for obtaining the first candidate video segment described above, and will not be further described here.
[0062] In actual applications, when the exemplary embodiment of the present disclosure obtains the second candidate video segment from the video to be processed, the video duration difference can be determined based on the video duration of the second candidate video segment and the video duration of the original video. If the video duration difference is less than the video duration of the video to be processed, it means that the second video segment can be obtained from the video to be processed again. Therefore, the exemplary embodiment of the present disclosure can determine the preset video duration based on the video duration difference, so that the preset video duration is greater than the video duration difference.
[0063] Step 205: When the beat information of the second candidate video segment matches the beat information of the original video, the second candidate video segment is determined to be the video segment to be spliced.
[0064] In actual application, the exemplary embodiment of the present disclosure may first refer to the method for obtaining the first candidate video segment, and obtain multiple beat segments of the original video at a preset video length from the beat information of the original video. If the beat information of the second candidate video segment matches the beat segment of the original video at the preset video length, it can be determined that the beat information of the second candidate video segment matches the beat information of the original video. Otherwise, it means that the first candidate video segment cannot match each beat segment of the original video at the preset video length, and the first candidate segment does not match the beat information of the original video. Therefore, the first candidate video segment can be replaced.
[0065] Exemplarily, the exemplary embodiment of the present disclosure can set a sliding window of a preset video length, and then obtain multiple original video segments from the original video by sliding the sliding window on the original video in a chronological order. If an original video segment includes one or more frames of reference video images, a beat segment of the original video at the preset video length can be determined based on the beat features of all the reference video images included in the original video segment.
[0066] Exemplarily, for each second candidate video segment, it can be determined whether the beat information of the second candidate video segment is correlated with each beat segment of the original video in the preset video length. When the correlation between the beat information of the second candidate video segment and the beat segments of the original video in the preset video length is greater than the second preset correlation, it can be considered that the beat information of the second candidate video segment matches the beat segments of the original video in the preset video length.
[0067] An exemplary embodiment of the present disclosure can extract the beat information of each second candidate video segment through a 3D convolutional network, and then obtain the correlation between the beat information of the second candidate video segment and the beat segment of the original video at a preset video length through methods such as the Pearson correlation coefficient and the Spearman correlation coefficient.
[0068] When the Pearson correlation coefficient is used to calculate the correlation between the beat information of the second candidate video segment and the beat segments of the original video at the preset video length, the correlation between the beat information of the second candidate video segment and the beat segments of the original video at the preset video length can refer to the Pearson correlation coefficient between the beat information of the second candidate video segment and the beat segments of the original video at the preset video length. Accordingly, the second preset correlation can be expressed as a second preset correlation coefficient. In this case, when the Pearson correlation coefficient between the beat information of the second candidate video segment and the beat segments of the original video at the preset video length is greater than the second preset correlation coefficient, it can be considered that the beat information of the second candidate video segment matches the beat segments of the original video at the preset video length.
[0069] Step 206: Updating the video to be processed based on the video segments to be spliced and the video to be processed. Here, the video segments to be spliced and the video to be processed can be spliced together to achieve the purpose of updating the video to be processed.
[0070] It can be seen that the exemplary embodiment of the present disclosure can obtain a second candidate video segment that matches the beat segment of the original video at the preset video length from the video to be processed when the video length of the video to be processed is less than or equal to the video length of the original video, and use it as the video segment to be spliced. When splicing it into the video to be processed, the length of the video segment that matches the beat information of the video to be processed and the original video can be increased, thereby ensuring that the rhythm of the obtained target video segment is more matched with the rhythm of the original video.
[0071] When splicing the video clip to be spliced and the video to be processed, the exemplary embodiment of the present disclosure can splice the video clip to be spliced at any position of the video to be processed, or can splice the video clip to be spliced at the splicing position of the video to be processed by setting the splicing position of the video to be processed.
[0072] Exemplarily, when obtaining the second candidate video segment, the area of the video to be processed can be recorded at the same time. If the second candidate video segment matches the beat segment of the original video in the preset video length, then the second candidate video segment is the video segment to be spliced. At the same time, the splicing position of the video to be processed is determined based on the area of the video to be processed by the second candidate video segment. The splicing position of the video to be processed can be adjacent to the area of the video to be processed by the second candidate video segment, or it can be separated from the area of the video to be processed by a preset time length. For example, the preset time length can be controlled within 600ms.
[0073] When the video clip to be spliced is spliced at the splicing position of the video to be processed by setting the splicing position of the video to be processed, the process can directly return to step 201 to detect the relationship between the video length of the updated video to be processed and the video length of the original video. If the video length of the updated video to be processed is greater than the video length of the original video, the rhythm matching degree between the obtained target video clip and the original video can be increased, thereby further improving the audio and video synchronization of the obtained target video.
[0074] As a possible implementation, the original video of the exemplary embodiment of the present disclosure includes multiple frames of reference video images, and the beat information of the original video includes beat features of the multiple frames of reference video images. FIG3 shows a schematic diagram of the process of determining the beat features of the reference video images of the exemplary embodiment of the present disclosure. As shown in FIG3, the method of the exemplary embodiment of the present disclosure further includes:
[0075] Step 301: Obtain positional features of preset key points in multiple frames of raw video images included in the original video. In exemplary embodiments of the present disclosure, key point detection can be performed on the original video images using an object detection model to obtain positional features of the preset key points. The positional features of the preset key points can be key point positional features of a target part. The target part can be a part of a movable target or a part of a static target.
[0076] When the exemplary embodiment of the present disclosure detects the preset key points of the original video image, the preset key points may be preset feature points, corner points, etc. Taking the corner points as the preset key points as an example, since the corner points may be points with particularly prominent attributes in certain aspects, for example: in the original video image, the corner points may be the connecting points of the object contour lines. When the positions of the corner points change, the changes in the corner point position features are more obvious. Therefore, when the position features of the preset key points are the corner point position features, the changes in the preset key points between different original video images can be detected more sensitively.
[0077] Step 302: Determine the position change rate of the preset key points in each frame of the original video image based on the positional features of the preset key points in each frame of the original video image and the positional features of the preset key points in the next frame of the original video image. In an exemplary embodiment of the present disclosure, the position change rate of the preset key points in each frame of the original video image can be tracked using an optical flow algorithm.
[0078] Step 303: If the position change rate of the preset key points in the original video image is greater than the preset position change rate, determine the beat characteristics of the reference video image based on the position change rate of the preset key points in the original video image. It should be understood that the key point position characteristics of the exemplary embodiment of the present disclosure are position characteristics of multiple key points. Therefore, when determining whether the position change rate of the key points in the original video image is greater than the preset position change rate, it is actually necessary to determine whether the position change rate of each key point in the original video image is greater than the preset position change rate.
[0079] When the position change rate of each preset key point of the current frame original video image is greater than the preset position change rate, it means that the position change rate of the preset key point of the current frame original video image is relatively large, and the position difference between the preset key points of the current frame original video image and the previous frame original video image is relatively large. The position change of the preset key point of the current frame original video image can reflect the beat characteristics of the original video. Therefore, the current frame original video image can be determined as the reference video image, and the beat characteristics of the reference video image can be determined based on the position change rate of the preset key point of the current frame original video image.
[0080] For example, the key point position change rate of the current frame original video image can be directly defined as the beat feature of the reference video image, or the image content change of the reference video image can be determined based on the key point position change rate of the current frame original video image, and the image content change of the reference video image can be defined as the beat feature of the reference video image.
[0081] For example, the exemplary embodiments of the present disclosure may use a corner extraction algorithm to extract corner features of an original video image. The following describes the corner feature extraction process of a frame of an original video image using the Shi-Tomasi algorithm as an example.
[0082] For a given frame of raw video, you can slide a fixed window in any direction across the frame and compare the grayscale changes in two areas of the frame before and after the window is slid. If significant grayscale changes occur in either direction, it can be considered that a corner exists within the fixed window. The specific steps are as follows:
[0083] First, assuming that the fixed window currently corresponds to a sub-image area I of a certain frame of the original video image, the gradient I of the sub-image area in the x direction can be obtained. x and the gradient I of the sub-image area in the y direction y .
[0084] Secondly, the gradient product of the two directions of the sub-image area is calculated. The gradient product of the two directions can be the product of two gradients in the same direction or the product of two gradients in different directions. The product of the gradients in the two y directions of the sub-image area The product of the gradient in the x direction and the gradient in the y direction of the sub-image area is I xy =I x I y .
[0085] Again, use the Gaussian function g(·) to multiply the gradient product of the two x directions of the sub-image area The product of the gradients in the two y directions of the sub-image area And the product of the gradient in the x direction and the gradient in the y direction of the sub-image area I xy Perform Gaussian weighting to generate element A, element C, and element B of the matrix M.
[0086] Among them, w represents the weight.
[0087] Next, the Harris response value of each pixel in the sub-image area I is calculated, and the Harris response value r that is less than the response value threshold t is set to zero, thereby obtaining the Harris response value set R of the pixels in the sub-image area.
[0088] Where, R = {R: detM - α(trM) 2 >t}, where tr represents the trace of the matrix M, det represents the determinant of the matrix M, and α represents an empirical constant that can be in the range of (0.04, 0.06).
[0089] Finally, non-maximum suppression is performed on the Harris response value of the sub-image area within a 3×3 or 5×5 scale window neighborhood to obtain the pixel point corresponding to the maximum Harris response value in the sub-image area as the corner point.
[0090] As a possible implementation, the exemplary embodiment of the present disclosure may further perform transition settings on the target video segment before acquiring the target video based on the material features of the target video segment and the original video. FIG4 illustrates a schematic diagram of the transition setting process of the exemplary embodiment of the present disclosure. As shown in FIG4 , the method of the exemplary embodiment of the present disclosure may further include:
[0091] Step 401: Acquire a target transition feature, which matches the target video images before and after the transition of the target video clip.
[0092] The target video images before and after the transition of the exemplary embodiment of the present disclosure may include the pre-transition video image and the post-transition video image included in the target video clip. The pre-transition video image may be one frame or multiple frames, and the post-transition video image may be one frame or multiple frames.
[0093] In practical applications, the grayscale difference between each pixel in the target video clip can be calculated and its absolute value taken to obtain a differential image of the two adjacent target video frames. If the minimum grayscale difference in the differential image exceeds a preset difference, it can be determined that a transition exists between the two adjacent target video frames. In this case, the two adjacent target video frames can be used as the target video images before and after the transition, or a video clip containing the two adjacent target video frames can be used as the target video images before and after the transition.
[0094] On this basis, the matching degree between the target video image before and after the transition and different preset transition features can be compared. If the matching degree between a target video image before and after the transition and a preset transition feature is greater than the preset matching degree, the preset transition feature can be considered as the target transition feature. The preset transition feature can be determined based on pre-stored transition features or based on the transition features of the original video.
[0095] When determining a preset transition feature based on the transition features of the original video, if the degree of match between a target video image before and after the transition and a preset transition feature is greater than the preset match degree, the preset transition feature can be used as part of the video template. When the video template is applied to the target video clip, the preset transition feature can be obtained from the video template. The preset transition feature can be various transition types, such as, but not limited to, flash to white, flash to black, fade in and fade out, wipe transition, dissolve, and the use of an empty shot.
[0096] Step 402: Based on the target transition feature, the target video image before and after the transition is set to transition. The transition setting performed here can be based on the target transition feature to process the target video image before and after the transition, or insert a new transition paragraph to achieve the target transition feature.
[0097] When the target transition feature matches the target video images before and after the transition of the target video clip, the minimum grayscale difference of the differential image of the two adjacent frames of target video images included in the target video images before and after the transition exceeds the preset difference. Therefore, the target transition feature is also related to the beat information of the target video clip. Therefore, by setting the transition of the target video images before and after the transition based on the target transition feature, it can also be ensured that the transition effect of the obtained target video matches the video beat, thereby further improving the target video effect.
[0098] As a possible implementation, the material features of the exemplary embodiment of the present disclosure may include multiple sets of visual material features, each set of visual material features including basic visual material information and the retention duration of the basic visual material information. Here, in terms of material type, it can also include text information of the original video, and even visual materials such as special effects information and sticker information. If the material type of the material features is not distinguished, the material features of the exemplary embodiment of the present disclosure can include material content, material display location, and material style, etc., according to the feature type.
[0099] In practical applications, the method of the exemplary embodiment of the present disclosure may further include: if the basic information of the visualization material of multiple consecutive frames of original video images is the same, aggregating the basic information of the visualization material of the multiple consecutive frames of original video images to obtain the aggregation result of the basic information of the visualization material, determining the retention time of the basic information of the visualization material based on the timestamps of the multiple consecutive frames of original video images, and obtaining a set of visualization material features based on the aggregation result of the basic information of the visualization material and the retention time of the original video images.
[0100] For example, when the basic information of the visualization materials of two adjacent frames of original video images is the same, the material content, material display location, and material style may be the same. When aggregating the basic information of the visualization materials of two adjacent frames of original video images, the frame numbers of the two adjacent frames of original video images can be obtained as the timestamps of the two adjacent frames of original video images, which are used to determine the retention period of the basic information of the visualization materials.
[0101] When acquiring a target video based on the material features of the target video clip and the original video, the visual material features can be migrated to the target video clip based on the basic visual material information and retention duration of each set of visual material features. For example, if the visual material is subtitles from the original video, the subtitles from the original video can be migrated to the target video clip based on the subtitle content, subtitle layout area, subtitle style, etc., to obtain the target video.
[0102] The exemplary embodiments of the present disclosure also provide a method for generating a video template, which can automatically and quickly generate a video template, and the application effect of the video template is relatively good, which can ensure that the generated target video has a relatively good audio and video synchronization effect.
[0103] FIG5 shows a schematic flow chart of a method for generating a video template according to an exemplary embodiment of the present disclosure. As shown in FIG5 , the method for generating a video template according to an exemplary embodiment of the present disclosure may include:
[0104] Step 501: Extracting original video material information, which includes at least audio features of the original video. The audio features of the original video may include audio features of at least one sound source, such as background music, human voice, ambient sound, or even device noise.
[0105] In practical applications, the original audio can be subjected to noise reduction processing, and then the original video can be subjected to audio recognition and separation based on a music source feature extraction and separation algorithm based on a deep neural network (for example: SA-CEDN-4FEM) to obtain background music and vocals, etc.
[0106] Step 502: Extract node information from the original video, which includes at least the original video's tempo information. The tempo information of the original video can be found in the previous section and will not be further elaborated here. To enhance the effectiveness of extracting tempo information from the original video, exemplary embodiments of the present disclosure may also perform noise reduction and de-rating on the original video to ensure that the original video is not distorted.
[0107] Step 503: Generate a video template based on the material information of the original video and the node information of the original video. When using the video template to generate the target video, you can refer to the above video generation method, which will not be described here.
[0108] The exemplary embodiments of the present disclosure intelligently decompose and transform the multimodal elements of an original video to create video templates for users to quickly edit. This not only speeds up the generation of video templates and significantly reduces the creation time, but also allows video creators to directly create target videos with the same style as the original video, even without professional video editing skills. Furthermore, video creators can use the methods of the exemplary embodiments of the present disclosure to obtain ready-to-use video templates for any video, which not only improves the efficiency and quality of video production but also saves time and costs for users.
[0109] The exemplary embodiments of the present disclosure can also upload the video template to the video creation platform after generating it, with authorization, for use by video creators, thereby increasing the number and types of video templates on the video creation platform and improving the content ecosystem richness of the video creation platform;
[0110] As a possible implementation, the material features of the exemplary embodiment of the present disclosure further include multiple sets of visual material features of the original video, each set of visual material features including basic visual material information and the retention duration of the basic visual material information. Here, the method of the exemplary embodiment of the present disclosure may further include:
[0111] Extract basic information of the visualization material of each frame of the original video image included in the original video; if the basic information of the visualization material of multiple consecutive frames of the original video image is the same, aggregate the basic information of the visualization material of the multiple consecutive frames of the original video image to obtain an aggregation result of the basic information of the visualization material; determine the retention time of the basic information of the visualization material based on the timestamps of the multiple consecutive frames of the original video image; and obtain a set of visualization material features based on the aggregation result of the basic information of the visualization material and the retention time of the basic information of the visualization material.
[0112] Material types can also include textual information from the original video, and even visual materials such as special effects and stickers. Without distinguishing between material types, material features can be classified by feature type. The material features of the exemplary embodiments of this disclosure can include material content, material display location, and material style. The following uses textual information from an original video as an example to describe the method for extracting textual features from an original video in the exemplary embodiments of this disclosure.
[0113] First, each frame of the original video image is preprocessed to obtain the text image corresponding to each frame of the original video image. For example, each frame of the original video image can be grayscaled, binarized, denoised, tilted, and segmented to obtain the corresponding text image.
[0114] Secondly, style and position detection and character recognition are performed on each character in the text image to obtain the text style, text display position and text content.
[0115] For example, a convolutional network can be used to extract features from text images to obtain image features of the text images. If the text image contains Chinese characters or other characters with a relatively large number of characters, the image features of the text image can also be subjected to dimensionality reduction processing to reduce the amount of subsequent data processing. Image features can be classified by a classifier to obtain the display position and character style of each character. The character style here can be character color, character font, character special effects, character direction, character length, text overlap, text density, etc. At the same time, natural language models such as RNN, CRNN, transformer, etc. can also be used to identify the image features of text images to obtain the character content included in the text.
[0116] Next, considering that the time interval between adjacent frames of original video images is relatively short, it is possible that the original video images of consecutive frames contain the same basic text information. Therefore, all character features of two adjacent frames of text images can be compared. If all character features are the same, all character features of the two adjacent frames can be merged. At the same time, the retention time of each character can be updated to obtain a set of character features.
[0117] Finally, a text feature that meets the format requirements can be generated according to the storage format of the text feature. The storage of the text feature can include (character content, character position, character style, and character retention time).
[0118] As a possible implementation, the original video of the exemplary embodiment of the present disclosure includes multiple frames of reference video images, the beat information of the original video includes beat features of the multiple frames of reference video images, and extracting the node information of the original video includes:
[0119] The position features of the preset key points of multiple frames of original video images included in the original video are obtained, and the position change rate of the preset key points of each original video image is determined based on the position features of the preset key points of each frame of the original video image and the position features of the preset key points of the next frame of the original video image. If the position change rate of the preset key points of the original video image is greater than the preset position change rate, the beat features of the reference video image are determined based on the position change of the preset key points of the original video image.
[0120] When the node information of the original video of the exemplary embodiment of the present disclosure also includes the transition features of the original video, extracting the node information of the original video may further include: obtaining the grayscale change of two adjacent frames of original video images included in the original video, if the grayscale change of the two adjacent frames of original video images is greater than the preset grayscale change, determining the transition features of the original video based on the two adjacent frames of the original video images.
[0121] One or more technical solutions provided in exemplary embodiments of the present disclosure can extract a target video segment from a video to be processed that matches the tempo information of the original video. Because the material features of the original video include at least its audio features, the audio features of the original video match the tempo information of the target video segment. On this basis, when a target video is obtained based on the material features of the target video segment and the original video, the audio features of the original video are essentially added to the target video segment, ensuring that the audio and tempo of the obtained target video match.
[0122] In addition, the exemplary embodiment of the present disclosure can combine the beat information and material information of the original video with the video to be processed to synthesize the target video when the beat information and material information of the original video are obtained. The target video is not limited by the types of video templates provided by the video creation platform. Therefore, when the video creator of the exemplary embodiment of the present disclosure sees a video of interest but does not have a video template, he or she can directly use the video as the original video, and use the beat information and material information of the original video to create a target video that suits his or her interests, thereby improving the efficiency of target video production.
[0123] In summary, the exemplary embodiments of the present disclosure can parse the original video into multimodal elements of text, audio, and images, and convert them into video templates that can be edited by users, so as to achieve the purpose of automatically and quickly generating video templates, which can greatly shorten the video template production process. For video creators, video creators save video templates on the video creation platform. Under authorized use, the number and types of templates on the video creation platform can be increased, and the richness of the platform's content ecology can be improved; for video creators, video creators can use video templates that are ready to use for any video, which improves the efficiency and quality of video production and can also save time and costs for video creators.
[0124] Moreover, the video template generation method of the exemplary embodiment of the present disclosure can be connected with the video editing link to accurately insert resources such as text features, audio features, transition features, etc. into the time nodes of the corresponding target video clips to generate a target video, which supports further editing, resource replacement and key frame fine-tuning.
[0125] The above mainly introduces the solution provided by the embodiment of the present disclosure from the perspective of the server. It can be understood that in order to realize the above functions, the server includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0126] The embodiments of the present disclosure can divide the server into functional units according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present disclosure is schematic and is only a logical functional division. In actual implementation, there may be other division methods.
[0127] In the case of dividing each functional module according to each function, the exemplary embodiment of the present disclosure provides a video generation device, which can be a server or a chip applied to a server. Figure 6 shows a schematic block diagram of the functional modules of the video generation device according to the exemplary embodiment of the present disclosure. As shown in Figure 6, the video generation device 600 includes:
[0128] Extraction module 601, used to extract target video segments that match the beat information of the original video from the video to be processed;
[0129] The acquisition module 602 is configured to acquire a target video based on the target video segment and material features of the original video, where the material features at least include audio features of the original video.
[0130] In one possible implementation, the extraction module 601 is used to obtain multiple first candidate video segments from the video to be processed when the video length of the video to be processed is greater than the video length of the original video. If the beat information of the first candidate video segment matches the beat information of the original video, the first candidate video segment is determined to be the target video segment, and the video length of the first candidate video segment is equal to the video length of the original video.
[0131] In a possible implementation, if the correlation between the beat information of the first candidate video segment and the beat information of the original video is greater than a first preset correlation, the beat information of the first candidate video segment matches the beat information of the original video.
[0132] In one possible implementation, the extraction module 601 is also used to obtain multiple second candidate video segments from the video to be processed based on a preset video length when the video length of the video to be processed is less than or equal to the video length of the original video, and the preset video length is less than the video length of the video to be processed. If the beat information of the second candidate video segment matches the beat information of the original video, the second candidate video segment is determined to be the video segment to be spliced, and the video to be processed is updated based on the video segment to be spliced and the video to be processed.
[0133] In one possible implementation, the extraction module 601 is used to determine a video duration difference based on the video duration of the second candidate video segment and the video duration of the original video. If the video duration difference is less than the video duration of the video to be processed, the preset video duration is determined based on the video duration difference, and the preset video duration is greater than the video duration difference.
[0134] In one possible implementation, the extraction module 601 is used to obtain multiple beat segments of the original video in the preset video length from the beat information of the original video. If the beat information of the second candidate video segment matches each beat segment of the original video in the preset video length, it is determined that the beat information of the second candidate video segment matches the beat information of the original video.
[0135] In one possible implementation, when the correlation between the beat information of the second candidate video segment and the beat segment of the original video at the preset video length is greater than a second preset correlation, the beat information of the second candidate video segment matches the beat segment of the original video at the preset video length.
[0136] In a possible implementation, the original video includes multiple frames of reference video images, and the beat information of the original video includes beat features of the multiple frames of reference video images;
[0137] The extraction module 601 is also used to obtain the position features of the preset key points of the multiple frames of original video images included in the original video, and determine the position change rate of the preset key points of each frame of the original video image based on the position features of the preset key points of the original video image in each frame and the position features of the preset key points of the original video image in the next frame. If the position change rate of the preset key points of the original video image is greater than the preset position change rate, the beat features of the reference video image are determined based on the position change rate of the key points of the original video image.
[0138] In one possible implementation, the acquisition module 602 is used to obtain target transition features, and the target transition features match the target video images before and after the transition of the target video clip. The device also includes a setting module 603, and the setting module 603 is used to perform transition settings on the target video images before and after the transition based on the target transition features.
[0139] In a possible implementation, the material features also include multiple groups of visual material features of the original video, each group of visual material features includes basic information of the visual material and the retention time of the basic information of the visual material, and the extraction module 601 is also used to extract the basic information of the visual material of each frame of the original video image included in the original video; if the basic information of the visual material of multiple consecutive frames of the original video image is the same, the basic information of the visual material of the multiple consecutive frames of the original video image is aggregated to obtain the aggregation result of the basic information of the visual material; based on the timestamps of the multiple consecutive frames of the original video image, the retention time of the basic information of the visual material is determined; based on the aggregation result of the basic information of the visual material and the retention time of the basic information of the visual material, a group of visual material features is obtained.
[0140] In the case of dividing each functional module according to each function, the exemplary embodiment of the present disclosure provides a device for generating a video template. The device for generating a video template can be a server or a chip applied to a server. Figure 7 shows a schematic block diagram of the functional modules of the device for generating a video template according to an exemplary embodiment of the present disclosure. As shown in Figure 7, the device for generating a video template 700 includes:
[0141] Extraction module 701, extracting material information of the original video, extracting node information of the original video, wherein the material information of the original video at least includes audio features of the original video, and the node information at least includes beat information of the original video;
[0142] The generating module 702 is configured to generate a video template based on the material information of the original video and the node information of the original video.
[0143] In one possible implementation, the material features further include multiple groups of visual material features of the original video, each group of visual material features including basic visual material information and a retention duration of the basic visual material information. The extraction module 701 is configured to extract the basic visual material information of each frame of the original video image included in the original video; if the basic visual material information of multiple consecutive frames of the original video image is the same, the basic visual material information of the consecutive multiple frames of the original video image is aggregated to obtain an aggregation result of the basic visual material information; and the retention duration of the basic visual material information is determined based on the timestamps of the consecutive multiple frames of the original video image.
[0144] In one possible implementation, the original video includes multiple frames of reference video images, and the beat information of the original video includes beat features of the multiple frames of reference video images. The extraction module 701 is used to obtain the position features of the preset key points of the multiple frames of original video images included in the original video, and determine the position change rate of the preset key points of each original video image based on the position features of the preset key points of each frame of the original video image and the position features of the preset key points of the next frame of the original video image. If the position change rate of the preset key points of the original video image is greater than the preset position change rate, the beat features of the reference video image are determined based on the position change of the preset key points of the original video image.
[0145] In one possible implementation, the node information of the original video also includes the transition features of the original video, and the extraction module 701 is further used to obtain the grayscale changes of two adjacent frames of original video images included in the original video. If the grayscale changes of the two adjacent frames of the original video images are greater than the preset grayscale changes, the transition features of the original video are determined based on the two adjacent frames of the original video images.
[0146] Figure 8 shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure. As shown in Figure 8, chip 800 includes one or more (including two) processors 801 and a communication interface 802. Communication interface 802 can support the server in executing the data transmission and reception steps in the above-mentioned image processing method, and processor 801 can support the server in executing the data processing steps in the above-mentioned image processing method.
[0147] Optionally, as shown in FIG8 , the chip 800 further includes a memory 803 , which may include a read-only memory and a random access memory, and provides operating instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory (NVRAM).
[0148] In some embodiments, as shown in FIG8 , the processor 801 performs corresponding operations by calling operation instructions stored in the memory (the operation instructions may be stored in the operating system). The processor 801 controls the processing operations of any one of the terminal devices, and the processor may also be referred to as a central processing unit (CPU). The memory 803 may include a read-only memory and a random access memory, and provides instructions and data to the processor 801. A portion of the memory 803 may also include NVRAM. For example, in an application, the memory, the communication interface, and the memory are coupled together through a bus system, wherein the bus system may include, in addition to the data bus, a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are labeled as bus system 804 in FIG8 .
[0149] The methods disclosed in the above embodiments of the present disclosure can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0150] The exemplary embodiments of the present disclosure further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being configured to cause the electronic device to perform a method according to an exemplary embodiment of the present disclosure when executed by the at least one processor.
[0151] Exemplary embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present disclosure.
[0152] Exemplary embodiments of the present disclosure further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to perform the method according to the embodiment of the present disclosure.
[0153] With reference to Figure 9, a block diagram of an electronic device 900 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0154] As shown in Figure 9, electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In RAM 903, various programs and data required for the operation of device 900 can also be stored. Computing unit 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.
[0155] As shown in Figure 9, multiple components within electronic device 900 are connected to an I / O interface 905, including an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. Input unit 906 can be any type of device capable of inputting information into electronic device 900. Input unit 906 can receive input digital or character information and generate key input signals related to user settings and / or function control of the electronic device. Output unit 907 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 908 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 909 allows electronic device 900 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0156] As shown in Figure 9, the computing unit 901 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above. For example, in some embodiments, the method of the exemplary embodiments of the present disclosure may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. In some embodiments, the computing unit 901 can be configured to perform the method of the exemplary embodiments of the present disclosure in any other appropriate manner (e.g., by means of firmware).
[0157] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0158] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0159] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0161] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0162] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0163] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user device, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).
[0164] Although the present disclosure has been described with reference to specific features and embodiments thereof, it will be apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present disclosure. Accordingly, this specification and the drawings are merely illustrative of the present disclosure as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present disclosure. Obviously, those skilled in the art may make various modifications and variations to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, the present disclosure is intended to include such modifications and variations if they fall within the scope of the claims of the present disclosure and their equivalents.
Claims
1. A video generation method, comprising: Extracting a target video segment matching the beat information of the original video from the video to be processed; A target video is acquired based on the target video segment and material features of the original video, wherein the material features at least include audio features of the original video.
2. The method according to claim 1, wherein: The step of extracting a target video segment matching the beat information of the original video from the video to be processed includes: When the video duration of the video to be processed is greater than the video duration of the original video, obtaining a plurality of first candidate video segments from the video to be processed, wherein the video duration of the first candidate video segments is equal to the video duration of the original video; If the beat information of the first candidate video segment matches the beat information of the original video, the first candidate video segment is determined to be the target video segment.
3. The method according to claim 2, wherein: If the correlation between the beat information of the first candidate video segment and the beat information of the original video is greater than a first preset correlation, the beat information of the first candidate video segment matches the beat information of the original video.
4. The method according to claim 2, wherein: The step of extracting a target video segment matching the beat information of the original video from the video to be processed further includes: When the video duration of the video to be processed is less than or equal to the video duration of the original video, obtaining a plurality of second candidate video segments from the video to be processed based on a preset video duration, wherein the preset video duration is less than the video duration of the video to be processed; If the beat information of the second candidate video segment matches the beat information of the original video, determining that the second candidate video segment is the video segment to be spliced; The video to be processed is updated based on the video segments to be spliced and the video to be processed.
5. The method according to claim 4, wherein: The step of extracting a target video segment matching the beat information of the original video from the video to be processed further includes: Determine a video duration difference based on the video duration of the second candidate video segment and the video duration of the original video; If the video duration difference is smaller than the video duration of the to-be-processed video, the preset video duration is determined based on the video duration difference, and the preset video duration is greater than the video duration difference.
6. The method according to claim 4, wherein: The step of extracting a target video segment matching the beat information of the original video from the video to be processed further includes: Acquire multiple beat segments of the original video at the preset video duration from the beat information of the original video; If the beat information of the second candidate video segment matches each beat segment of the original video in the preset video duration, it is determined that the beat information of the second candidate video segment matches the beat information of the original video.
7. The method according to claim 6, wherein: When the correlation between the beat information of the second candidate video segment and the beat segment of the original video at the preset video duration is greater than the second preset correlation, the beat information of the second candidate video segment matches the beat segment of the original video at the preset video duration.
8. The method according to any one of claims 1 to 7, wherein: The original video includes multiple frames of reference video images, the beat information of the original video includes beat features of the multiple frames of reference video images, and the method further includes: Acquire position features of preset key points of multiple frames of original video images included in the original video; Determine the position change rate of the preset key points of each frame of the original video image based on the position features of the preset key points of each frame of the original video image and the position features of the preset key points of the next frame of the original video image; If the position change rate of the preset key point of the original video image is greater than the preset position change rate, the beat feature of the reference video image is determined based on the position change rate of the preset key point of the original video image.
9. The method according to any one of claims 1 to 8, further comprising: Acquire a target transition feature, wherein the target transition feature matches the target video images before and after the transition of the target video segment; The target video images before and after the transition are set for transition based on the target transition feature.
10. The method according to any one of claims 1 to 9, wherein: The material features further include multiple groups of visual material features of the original video, each group of visual material features includes basic information of the visual material and retention time of the basic information of the visual material, and the method further includes: Extracting basic information of visual materials of each frame of original video image included in the original video; If the basic information of the visualization material of the consecutive multiple frames of the original video image is the same, aggregating the basic information of the visualization material of the consecutive multiple frames of the original video image to obtain an aggregation result of the basic information of the visualization material; Determining the retention time of the basic information of the visualization material based on the timestamps of the continuous multiple frames of the original video image; Based on the aggregation result of the basic information of the visualization material and the retention time of the basic information of the visualization material, a group of visualization material features is obtained.
11. A video generating device, comprising: An extraction module is configured to extract a target video segment matching the beat information of the original video from the video to be processed; The acquisition module is configured to obtain the target video based on the target video segment and the material features of the original video, wherein the material features at least include the audio features of the original video.
12. An electronic device comprising: processor; as well as, A memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Music matching method and device, terminal and storage medium
CN110839173A
Video processing method and device, electronic equipment and storage medium
CN113727038A
Cross-modal retrieval method and device, network training method and device, equipment and medium
CN114817655A
Video editing method and device, equipment and storage medium
CN115412764A
Method and apparatus for processing video, electronic device and storage medium
EP4125089A1
Cited By
Material processing method and related device
CN121418605A