Video generation method and device, equipment, medium and product

By obtaining the keyframe feature sequences of the video to be edited and the reference video, and using the pre-trained clip sequence prediction model to generate edited videos that meet the reference video style, the problems of personalized needs and editing error accumulation in the prior art are solved, and high-quality personalized video editing is achieved.

CN120281965APending Publication Date: 2025-07-08UNIVERSITY OF INTERNATIONAL BUSINESS & ECONOMICS (UIBE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510433221.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing intelligent video editing technology cannot meet personalized needs, the editing works are the same, and there are quality problems such as video incoherence and audio and video out of synchronization due to cumulative editing errors.

Method used

By obtaining the keyframe feature sequences of the video to be edited and the reference video, the video sequence is predicted using the pre-trained clip sequence prediction model, and stitching is performed to generate a clip video that conforms to the reference video style.

Benefits of technology

Personalized video editing is realized, reducing the accumulation of editing errors, and the generated video conforms to the reference video style, improving the targetedness and quality of the editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281965A_ABST
    Figure CN120281965A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method and device, equipment, a medium and a product. The method comprises the following steps: acquiring a to-be-edited video and a reference video corresponding to the to-be-edited video; determining a key frame feature sequence of the to-be-edited video and a key frame feature sequence of the reference video; through a pre-trained editing sequence prediction model, based on the key frame feature sequence of the to-be-edited video and the key frame feature sequence of the reference video, predicting to obtain an edited video sequence; and splicing the edited video sequences to obtain an edited video conforming to the reference video style. According to the technical scheme, the to-be-edited video is automatically edited on the basis of the reference video, the edited video conforming to the reference video style is obtained, edition is more targeted, and therefore the personalized requirement of a user for video edition is met, in addition, the edited video sequence is directly output by adopting the edition sequence prediction model, the user experience is improved, and the user experience is improved. And the editing process is simplified, so that the editing error accumulation is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of video processing, and in particular, to a video generation method, apparatus, device, medium, and product. Background Art

[0002] With the development of technology and the improvement of storage, a large number of videos have been generated in various industries. Due to the large volume of videos and the limited attention and energy of the audience, video editing has received extensive attention. In order to save costs and lower the threshold, intelligent video editing methods have gradually become popular.

[0003] In current intelligent video editing technologies, video thumbnails are often screened by intelligent methods and highlight segments are extracted by intelligent means. Although existing methods can obtain wonderful segments in the video and perform simple combinations, there is still a certain distance from practicality. First, the works edited by existing intelligent video editing methods are all the same and cannot meet personalized needs, and feature collapse is likely to occur. Second, there are too many intermediate steps in the editing of existing intelligent video editing methods, and editing error accumulation is likely to occur, resulting in quality problems such as discontinuous edited videos and out-of-sync audio and video. Summary of the Invention

[0004] The present disclosure provides a video generation method, apparatus, device, medium, and product, which meet the personalized needs of users for video editing and reduce the accumulation of editing errors.

[0005] According to one aspect of the present disclosure, there is provided a video generation method, including:

[0006] Obtaining a video to be edited and a reference video corresponding to the video to be edited;

[0007] Determining a key frame feature sequence of the video to be edited and a key frame feature sequence of the reference video;

[0008] Predicting a completed edited video sequence corresponding to the video to be edited based on the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video by a pre-trained editing sequence prediction model;

[0009] Splicing the completed edited video sequence corresponding to the video to be edited to obtain an edited video conforming to the style of the reference video.

[0010] According to another aspect of the present disclosure, there is provided a video generation apparatus, including:

[0011] A video acquisition module, configured to obtain a video to be edited and a reference video corresponding to the video to be edited;

[0012] A key frame feature sequence determination module, configured to determine the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video;

[0013] A clipped video sequence prediction module, configured to predict a completed clipped video sequence corresponding to the video to be clipped based on the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video through a pre-trained clipped sequence prediction model;

[0014] A clipped video sequence splicing module, configured to splice the completed clipped video sequence corresponding to the video to be clipped to obtain a clipped video conforming to the style of the reference video.

[0015] According to another aspect of the present disclosure, there is provided an electronic device, the electronic device includes:

[0016] At least one processor;

[0017] And a memory communicatively connected to the at least one processor;

[0018] Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the video generation method according to any embodiment of the present disclosure.

[0019] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the video generation method according to any embodiment of the present disclosure when executed by a processor.

[0020] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, and the computer program implements the video generation method according to any one of the embodiments of the present disclosure when executed by a processor.

[0021] The technical solution of the embodiment of the present disclosure includes obtaining a video to be clipped and a reference video corresponding to the video to be clipped; determining the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video; predicting, by a pre-trained clip sequence prediction model based on the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video, a clipped video sequence corresponding to the video to be clipped; and splicing the clipped video sequence corresponding to the video to be clipped to obtain a clipped video that conforms to the style of the reference video. In the above technical solution, the video to be clipped is automatically clipped based on the reference video to obtain a clipped video that conforms to the style of the reference video, making the clip more targeted, thereby meeting the personalized needs of users for video clipping. In addition, the present disclosure directly outputs the clipped video sequence by using the clip sequence prediction model, simplifies the clip process, and thus reduces the accumulation of clip errors.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0024] Figure 1 is a flowchart of a video generation method provided according to an embodiment of the present disclosure;

[0025] Figure 2 is a flowchart of another video generation method provided according to an embodiment of the present disclosure;

[0026] Figure 3 is a flowchart of another video generation method provided according to an embodiment of the present disclosure;

[0027] Figure 4 is a flowchart of another video generation method provided according to an embodiment of the present disclosure;

[0028] Figure 5 is a flowchart of another video generation method provided according to an embodiment of the present disclosure;

[0029] Figure 6 is a schematic diagram of a video clipping method provided according to an embodiment of the present disclosure;

[0030] Figure 7It is a schematic structural diagram of a personalized video automatic editing software system provided according to an embodiment of the present disclosure;

[0031] Figure 8 It is a schematic structural diagram of a video generation device provided according to an embodiment of the present disclosure;

[0032] Figure 9 It is a schematic structural diagram of an electronic device for implementing the video generation method according to an embodiment of the present disclosure. Specific embodiments

[0033] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the technical solutions of the present disclosure all comply with the relevant regulations of national laws and regulations.

[0035] The technical field and related terms of the embodiments of the present disclosure will be briefly described below.

[0036] Video automatic editing: Video automatic editing refers to using computer technology and artificial intelligence algorithms to perform intelligent analysis, recognition, and processing on original video materials, and automatically completing a series of workflow processes such as video content selection, editing, combination, and optimization. This technology can efficiently generate logical and ornamental video clips according to user needs or preset rules, without manual editing one by one, greatly improving the efficiency and quality of video production.

[0037] Prompt Learning: Prompt learning refers to a machine learning technique in which the model aids the learning process by receiving externally provided prompt information. These prompt information can be prior knowledge, guiding rules, or specific problem-solving solutions. In this learning mode, the training of the model not only depends on the data itself, but through the guidance of prompt information, it accelerates the learning process, improves learning efficiency and accuracy, and thus achieves better learning results in a shorter time. Prompt learning is particularly suitable for situations where data is sparse or difficult to learn directly from the data, and it provides a new paradigm for machine learning. Currently, there are two types of prompt learning in the field of large language models, one is discrete prompt learning, and the other is continuous prompt learning. Discrete prompt learning is to add text information to the training samples to specify the model to complete a specific task. These prompts need to be manually edited text, and the effect is unstable. The prompt of continuous prompt learning is no longer text, but a continuous feature vector, and this prompt is finally determined through training, rather than artificial design.

[0038] Multi-modal Continuous Prompt: The continuous prompt is not text, but a feature vector generated through training, and usually it is trained based on a language model. If the continuous prompt is jointly learned according to visual information and text information, it is called a multi-modal continuous prompt.

[0039] Self-attention Model: The attention model was proposed by Turing Award winner Yoshua Bengio in 2014 and has been widely applied in various prediction fields in recent years. In 2017, the self-attention model was proposed by Google. The self-attention model essentially models the correlation between different parts of the entire input, which not only avoids the influence of the cumulative error of the first t - 1 time slices on the prediction of the t-th time slice in the existing deep learning-based models, but also avoids the model being limited by the square or rectangular receptive field. It can determine the shape and type of its own receptive field, improving the accuracy and efficiency of the model.

[0040] In the existing video editing technology, directly extracting relevant features from the video to be edited and generating an editing strategy is likely to cause feature collapse, resulting in the same style of edited videos. In addition, the existing video editing needs to first generate a scene description from the video, then generate an editing strategy based on the scene description, and finally generate the edit according to the editing strategy, which requires multiple processes and is prone to error accumulation, seriously reducing the quality of the edited video. For this reason, the embodiments of the present disclosure provide a video generation method, device, equipment, medium, and product, which can effectively solve this problem. The following will further describe in detail the video generation method, device, equipment, medium, and product provided by the embodiments of the present disclosure.

[0041] Figure 1 The following is a flowchart of a video generation method provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of automatic video editing. The method can be executed by a video generation device, which can be implemented in the form of hardware and / or software, and the video generation device can be configured in electronic devices such as terminals and servers. As Figure 1 shown, the method includes:

[0042] S110. Obtain the video to be edited and the reference video corresponding to the video to be edited.

[0043] Among them, the video to be edited refers to the video to be edited, which can be obtained by splicing multiple videos or can be a long video. The reference video is a video provided or given by the user, which can be a sample video or a video that the user has edited before; the reference video can provide a reference for the editing of the video to be edited, so that the finally obtained edited video conforms to the style of the reference video.

[0044] Exemplarily, the video to be edited and the reference video corresponding to the video to be edited can be read from a preset storage path of the electronic device, and the video to be edited and the reference video corresponding to the video to be edited can also be obtained from other devices communicatively connected to the electronic device.

[0045] S120. Determine the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video.

[0046] Among them, the key frame feature sequence refers to a sequence composed of the features of multiple key frames in the video, which can include but is not limited to features such as the brightness, chroma, noise, scene, camera movement, and emotion of the video.

[0047] Specifically, key frames of the video to be edited and the reference video are extracted respectively to obtain the key frames of the video to be edited and the reference video, and then feature extraction is performed on the key frames of the video to be edited and the reference video respectively to obtain the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video.

[0048] S130. Based on the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video, predict the edited video sequence corresponding to the video to be edited through a pre-trained editing sequence prediction model.

[0049] Among them, the clip sequence prediction model is a pre-trained deep learning model, which can be a self-attention model or a variant of the self-attention model, etc., and will not be specifically limited here. The clipped video sequence refers to the video sequence clipped from the video to be clipped, which can include the highlight moments or key moment video frames of the video to be clipped. For example, the highlight moment or key moment video frames can be the video frames corresponding to shooting actions, jumping actions, or hitting actions, etc.

[0050] Exemplarily, one or more of the video summary public dataset, the film and television promotional video public dataset, and the camera movement dataset can be used as the sample video to be clipped, and the clip sequence prediction model can be trained using a cross-entropy or maximum likelihood estimation loss function, etc. During the training process of the clip sequence prediction model, the input of the clip sequence prediction model can be the key frame feature sequence of the sample video to be clipped, and the output is the video sequence of the highlight shots or key shots in the sample video to be clipped.

[0051] S140. Concatenate the clipped video sequences corresponding to the video to be clipped to obtain a clipped video that conforms to the reference video style.

[0052] Among them, the clipped video that conforms to the reference video style refers to a video with the same photography or editing style as the reference video. The video style can be montage, narrative, rhythm-driven, documentary, or an undefined style, etc. It can be understood that if the style of the reference video is the montage style, the obtained clipped video is the montage style; if the style of the reference video is the rhythm-driven style, the obtained clipped video is the rhythm-driven style, and the same applies to other styles, which will not be elaborated here.

[0053] Specifically, the clipped video sequences corresponding to the video to be clipped can be concatenated in chronological order to obtain a clipped video that conforms to the reference video style.

[0054] The technical solution of the embodiments of the present disclosure includes obtaining the video to be clipped and the reference video corresponding to the video to be clipped; determining the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video; predicting the clipped video sequence corresponding to the video to be clipped based on the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video through a pre-trained clip sequence prediction model; and concatenating the clipped video sequences corresponding to the video to be clipped to obtain a clipped video that conforms to the reference video style. In the above technical solution, the video to be clipped is automatically clipped based on the reference video to obtain a clipped video that conforms to the reference video style, making the clip more targeted, thereby meeting the personalized needs of users for video clipping. In addition, the present disclosure directly outputs the clipped video sequence using the clip sequence prediction model, simplifying the clip process and thus reducing the accumulation of clip errors.

[0055] Figure 2 The flowchart of another video generation method provided by an embodiment of the present disclosure. The various alternative solutions in the video generation method provided in this embodiment can be combined with those in the above-mentioned embodiments. Based on the above embodiments, this embodiment further refines the steps of determining the key frame feature sequence.

[0056] As Figure 2 shown, the method includes:

[0057] S210. Obtain the video to be clipped and the reference video corresponding to the video to be clipped.

[0058] S220. Extract key frames from the video to be clipped to obtain the key frames of the video to be clipped.

[0059] S230. Extract key frames from the reference video corresponding to the video to be clipped to obtain the key frames of the reference video.

[0060] S240. Perform attribute analysis on the key frames of the video to be clipped and the key frames of the reference video respectively to obtain the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video.

[0061] S250. Based on the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video, predict the completed clipped video sequence corresponding to the video to be clipped through a pre-trained clip sequence prediction model.

[0062] S260. Splice the completed clipped video sequence corresponding to the video to be clipped to obtain a clipped video that conforms to the style of the reference video.

[0063] In the embodiment of the present disclosure, key frames of a video can be extracted by a scene change detection method, a motion energy analysis method, or a neural network model.

[0064] Exemplarily, a reinforcement learning network can be used to extract key frames from the video to be clipped and the reference video.

[0065] In the embodiment of the present disclosure, structured attribute analysis and unstructured attribute analysis can be performed on key frames to obtain a key frame feature sequence. Among them, structured attributes can include brightness, chroma, noise, scene type, and camera movement, and unstructured attributes include emotion, and emotion can be anger, disgust, fear, happiness, sadness, or surprise.

[0066] Based on the above embodiments, optionally, key frames of the reference video corresponding to the video to be clipped are extracted to obtain key frames of the reference video, including: obtaining initial key frames, and determining the mean value of the feature vectors corresponding to the initial key frames; determining the key frames of the reference video according to the mean value of the feature vectors corresponding to the initial key frames through a trained key frame extraction model; wherein, the key frame extraction model is a reinforcement learning model trained by a preset hierarchical reward function for state information and action information, the preset hierarchical reward function includes a root reward function and a leaf reward function, the state information includes the mean value of the feature vectors of the key frames, and the action information includes actions executed based on the root policy and the leaf policy.

[0067] Among them, the initial key frame refers to the key frame obtained through initialization operations.

[0068] Exemplarily, the steps of obtaining the initial key frames may include: uniformly sampling the reference video at a step size of 2 to obtain N video frames, calculating the histograms of hue H, saturation S, and brightness V for the sampled video frames, and then combining them in a 9:3:1 manner to obtain k represents the k-th video frame sampled; calculating The formula for may be:

[0069]

[0070] Among them, represents the histogram of hue H, represents the histogram of saturation S, represents the histogram of brightness V. Further, perform differencing on the sequence to obtain Furthermore, select The largest K video frames in are used as the initial key frames.

[0071] In the embodiments of the present disclosure, the mean value of the feature vectors refers to the mean value of the feature vectors of multiple key frames.

[0072] Specifically, the steps of determining the mean value of the feature vectors corresponding to the key frames include: inputting the K key frames one by one into a pre-trained Resnet50 model, the Resnet50 model outputs the feature vectors corresponding to the K key frames respectively, and then calculates the mean value of the feature vectors corresponding to the K key frames respectively to obtain the mean value of the feature vectors corresponding to the key frames.

[0073] In the embodiments of the present disclosure, the key frame extraction model is a reinforcement learning model trained by a preset hierarchical reward function for state information and action information, and the reinforcement learning model may be an A2C (Advantage Actor-Critic) reinforcement learning network or a reinforcement learning model with other network architectures.

[0074] Exemplarily, the key frame extraction model selects a target key frame from K key frames through the root policy, and the leaf policy performs one of the following operations on the selected target key frame: shifting left by a preset number of steps or shifting right by a preset number of steps. The present disclosure proposes a hierarchical reward function for the root policy and the leaf policy. The hierarchical reward function includes a root reward function and a leaf reward function. Among them, the root reward function r h is shown in the following formula:

[0075] r h =(I t -I t-1 )+(A t -A t-1 );

[0076]

[0077] A t =AesNet(s t );

[0078] Among them, I t represents the amount of information at the t-th iteration, which is represented by the product of the conditional probabilities of selecting the i-th key frame by the state s t of the key frame extraction model at the t-th iteration. s t is the mean of the feature vectors extracted by the pre-trained Resnet50 model, which refers to the root policy selecting a certain key frame. AesNet represents an image aesthetic evaluation network, and the output is the mean of the aesthetic scores of the key frames. r h represents the change values of the amount of information and the aesthetic score at the current iteration and the previous iteration.

[0079] The leaf reward function r v is shown in the following formula:

[0080]

[0081] Among them, cosine(·,·) is to calculate the cosine similarity, represents the feature vector extracted by the pre-trained Resnet50 model for the i-th key frame, represents the feature vector extracted by the pre-trained Resnet50 model for the j-th key frame.

[0082] Training the A2C reinforcement learning network through the above hierarchical reward function can effectively improve the accuracy of the finally output K key frames.

[0083] It should be noted that the principle of the key frame extraction step of the video to be clipped is the same as that of the reference video, and will not be elaborated here.

[0084] Based on the above embodiments, optionally, the actions performed based on the root policy and the leaf policy include the actions performed based on the root policy and the actions performed based on the leaf policy. The actions performed based on the root policy include selecting a target key frame from multiple key frames. The actions performed based on the leaf policy include one of the following operations: shifting the target key frame 5 frames to the left, shifting the target key frame 10 frames to the left, shifting the target key frame 5 frames to the right, and shifting the target key frame 10 frames to the right.

[0085] It should be noted that shifting the target key frame 5 frames to the left, shifting the target key frame 10 frames to the left, shifting the target key frame 5 frames to the right, and shifting the target key frame 10 frames to the right are set according to multiple trials. The advantage of such a setting is that during the process of selecting key frames, the video can be quickly browsed (achieved by a large number of frames of 10 frames), and at the same time, precise fine-tuning can be performed to select the most suitable key frame (achieved by a small number of frames of 5 frames).

[0086] Based on the above embodiments, optionally, the key frame feature sequence includes one or more of: a brightness histogram sequence, a hue histogram sequence, a saturation histogram sequence, a noise histogram sequence, a saliency map sequence, a camera motion sequence, and an emotional similarity sequence.

[0087] Among them, the brightness histogram sequence, the hue histogram sequence, the saturation histogram sequence, the noise histogram sequence, the saliency map sequence, and the camera motion sequence are structured attributes, and the emotional similarity sequence is an unstructured attribute.

[0088] It should be noted that the brightness histogram sequence, the hue histogram sequence, the saturation histogram sequence, the noise histogram sequence, the saliency map sequence, the camera motion sequence, and the emotional similarity sequence can comprehensively characterize the features of the video, providing an accurate and reliable data basis for subsequent prediction of the completed video sequence.

[0089] Exemplarily, the structured attribute analysis step may include: decomposing each key frame into a luminance map and a reflectance map through a decoupling network (such as Decom-Net). For the key frame sequence a luminance map sequence and a reflectance map sequence can be obtained. Calculating the saliency map of each key frame through a saliency calculation network (such as BASNet). For the key frame sequence a saliency map sequence can be obtained. The noise map of each key frame can be obtained through a noise estimation network. For the key frame sequence a noise map sequence can be obtained, where the calculation formula of the noise estimation network is:

[0090]

[0091] where N iThe noise map representing the i-th key frame, The gradient map of the i-th key frame in the x direction, The gradient map of the i-th key frame in the y direction.

[0092] Furthermore, for the brightness attribute, the brightness histogram of each key frame can be calculated through the brightness map sequence to obtain the brightness histogram sequence. For the chromaticity attribute, the reflection map sequence can be projected onto the HSV space to obtain the hue histogram sequence and the saturation histogram sequence. For the noise attribute, the noise histogram sequence can be obtained through the noise map sequence. For the camera movement attribute, the motion vectors of the significant objects between every two frames in the saliency map sequence are calculated to form the camera movement sequence.

[0093] The unstructured attribute analysis steps may include: For high-quality clips, emotions need to be considered, but it is difficult to directly establish a connection between emotions and visual features. Based on the emotion system of American psychologist Paul Ekman, this disclosure classifies emotions into anger, disgust, fear, happiness, sadness, and surprise. For the six emotions, this disclosure designs a prompt template composed of continuous vectors for each emotion, as shown in the following formula:

[0094]

[0095] Among them, F(C k ) represents the prompt template selected when training the C k emotion type, represents the E p th random vector, E p represents the length of the continuous vector, [C k represents the vector representation corresponding to C k , and C k can be one of Anger, surprise, disgust, enjoyment, fear, and sadness.

[0096] Through the above prompt template, this disclosure sets the corresponding prompt vectors for each emotion type. Through the image-emotion label pairs in the image emotion classification dataset, multi-modal continuous prompt vectors can be trained.

[0097] Specifically, this disclosure uses the CLIP model as the feature extractor, and then fine-tunes the multi-modal continuous prompt vectors through the similarity between the image features and the text features, as shown in the following formula:

[0098]

[0099] Among them, E T represents the text feature extractor, E ILet \(F\) represent the image feature extractor, \(I\) represent the image corresponding to the key frame; \(\tau\) represents the temperature coefficient, which is used to adjust the distribution of similarity scores. The loss function used in the training process can be the cross-entropy loss function, as shown in the following formula:

[0100]

[0101] Among them, represents the true value of the emotion type, and \(i\) represents the emotion type. After training, during the video feature extraction process, the similarity between \(F(C k ) and the key frame can be calculated, thereby forming an emotion similarity sequence.

[0102] Exemplarily, for the reference video \(V ref \), \(K\) key frames can be extracted. The key frame feature sequence ref of the reference video \(V is obtained by splicing the brightness histogram sequence hue histogram sequence saturation histogram sequence noise histogram sequence salience map sequence camera movement sequence and the emotion similarity sequence For the video to be edited \(V un \), key frames can be extracted, and the value of can be 1 / 50 of the total length of the video or other ratios. The key frame feature sequence un of the video to be edited \(V is obtained by splicing the brightness histogram sequence of key frames hue histogram sequence saturation histogram sequence noise histogram sequence salience map sequence camera movement sequence and the emotion similarity sequence spliced together.

[0103] The technical solution of the embodiments of the present disclosure obtains a key frame feature sequence that can comprehensively represent video features through key frame extraction and attribute analysis, providing an accurate and reliable data basis for subsequent prediction of the edited video sequence.

[0104] Figure 3 FIG. is a flowchart of another video generation method provided by the embodiments of the present disclosure. The method of this embodiment can be combined with each optional solution in the video generation method provided in the above embodiments. On the basis of the above embodiments, this embodiment adds steps for music matching and music addition.

[0105] As shown in Figure 3 the figure, the method includes:

[0106] S310. Obtain the video to be clipped and the reference video corresponding to the video to be clipped.

[0107] S320. Determine the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video.

[0108] S330. Based on the camera movement sequence in the key frame feature sequence of the reference video, match a candidate music set from the music library.

[0109] S340. Based on the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video, use a pre-trained clip sequence prediction model to predict the completed clip video sequence corresponding to the video to be clipped.

[0110] S350. Stitch the completed clip video sequence corresponding to the video to be clipped to obtain a clipped video that conforms to the style of the reference video.

[0111] S360. Based on the clipped video that conforms to the style of the reference video, sample the camera movement sequence in the key frame feature sequence of the video to be clipped to obtain the camera movement sequence corresponding to the clipped video that conforms to the style of the reference video.

[0112] S370. Based on the camera movement sequence corresponding to the clipped video that conforms to the style of the reference video, match the target background music from the candidate music set.

[0113] S380. Add the target background music to the clipped video that conforms to the style of the reference video to obtain the target clipped video.

[0114] In the embodiments of the present disclosure, the music library may include multiple pieces of music, and each piece of music is labeled with camera movement timing points, so that the camera movement sequence of the reference video can be used to find a preset number of pieces of music as the candidate music set from the music library by the Hungarian matching algorithm or other matching algorithms according to the time points. Further, based on the timing of the clipped video that conforms to the style of the reference video, the camera movement sequence in the key frame feature sequence of the video to be clipped is sampled to obtain the camera movement sequence corresponding to the clipped video that conforms to the style of the reference video and then the most suitable target background music is matched from the candidate music set according to and the target background music is added to the clipped video that conforms to the style of the reference video to obtain the target clipped video.

[0115] Exemplarily,Figure 4 The flowchart of another video generation method provided by an embodiment of the present disclosure. As Figure 4 shown, first, load the weights of multiple algorithm models, including a key frame extraction model, a decoupling network, a saliency calculation network, a noise estimation network, and a clip sequence prediction model. Second, load the video to be clipped. Third, determine whether the video to be clipped is successfully loaded. If not, prompt the user to load the video to be clipped. Fourth, in the case of successful loading, preprocess the video to be clipped, including key frame extraction operations, structured attribute analysis operations, unstructured attribute analysis operations, and feature splicing operations, etc. Fifth, load the reference video. Sixth, determine whether the reference video is successfully loaded. If not, prompt the user to load the reference video. Seventh, in the case of successful loading, preprocess the reference video, including key frame extraction operations, structured attribute analysis operations, unstructured attribute analysis operations, and feature splicing operations, etc. Eighth, according to the camera movement sequence of the reference video find the Top-5 music from the music library by the Hungarian matching algorithm at the time point. Ninth, perform a clip sequence prediction operation to obtain the completed clipped video sequence, and then splice the completed clipped video sequence to obtain a clipped video that conforms to the style of the reference video. Tenth, based on the time sequence of the clipped video that conforms to the style of the reference video, sample the camera movement sequence in the key frame feature sequence of the video to be clipped to obtain the camera movement sequence corresponding to the clipped video that conforms to the style of the reference video and then according to match the most suitable Top-1 music from the candidate music set by the Hungarian matching algorithm at the time point, and then add the Top-1 music to the clipped video that conforms to the style of the reference video to obtain the target clipped video.

[0116] The technical solution of the embodiment of the present disclosure effectively improves the accuracy of the background music by matching music with the camera movement sequence in two stages, thereby improving the quality of the clipped video.

[0117] Figure 5 The flowchart of a video generation method provided by an embodiment of the present disclosure. The method of this embodiment can be combined with each optional solution in the video generation method provided in the above embodiment. On the basis of the above embodiments, this embodiment further refines the prediction of the completed clipped video sequence corresponding to the video to be clipped based on the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video by the pre-trained clip sequence prediction model.

[0118] As Figure 5 shown, the method includes:

[0119] S410. Obtain the video to be clipped and the reference video corresponding to the video to be clipped.

[0120] S420. Determine the key-frame feature sequence of the video to be clipped and the key-frame feature sequence of the reference video.

[0121] S430. Input the key-frame feature sequence of the video to be clipped into the self-attention encoding layer to obtain the encoded feature sequence corresponding to the video to be clipped.

[0122] S440. Input the key-frame feature sequence of the reference video and the encoded feature sequence corresponding to the video to be clipped into the cross-attention layer to obtain a first cross feature and a second cross feature.

[0123] S450. Concatenate the first cross feature and the second cross feature to obtain a target cross feature.

[0124] S460. Perform convolution processing on the target cross feature to obtain the predicted feature corresponding to the target cross feature.

[0125] S470. Based on the predicted feature corresponding to the target cross feature, extract the completed clipped video sequence corresponding to the video to be clipped from the video to be clipped.

[0126] S480. Concatenate the completed clipped video sequences corresponding to the video to be clipped to obtain a clipped video that conforms to the style of the reference video.

[0127] In the embodiments of the present disclosure, the clip sequence prediction model includes a self-attention encoding layer and a cross-attention layer. The self-attention encoding layer is used to encode the key-frame feature sequence of the video to be clipped. The cross-attention layer is used to perform interactive learning on the key-frame feature sequence of the reference video and the encoded feature sequence corresponding to the video to be clipped, so as to achieve feature enhancement.

[0128] Exemplarily, on the basis of the self-attention mechanism, the self-attention encoder of the present disclosure adds multi-head attention (MHA) and a feed-forward neural network (FFN). The encoding process is shown by the following formula:

[0129]

[0130]

[0131] Among them, f0 represents the data input at the beginning of the self-attention encoder. f0 is the key-frame feature sequence X of the video to be clipped un split into lengths of and each element feature dimension is sequence. The self-attention encoder can be obtained by cascading k = 6 self-attention encoding layers, and i represents the i-th self-attention encoding layer. Q i , K i , V i are respectively the query, key, and value vectors obtained by weighting the output of the (i - 1)-th self-attention encoding layer with parameter matrices. and are learnable parameter matrices, N H represents the number of heads in the multi-head attention mechanism. The calculation formula of the multi-head attention mechanism is:

[0132] MHA(Q, K, V) = Concat[Attention1(Q, K, V), …, Attention h (Q, K, V)];

[0133]

[0134] Among them, Q, K, and V respectively represent the query vector (Query), key vector (Key), and value vector (Value). By calculating the similarity between the query vector and the key vector, the weights of the attention can be obtained. Then, through the weights of the attention, the attention values of the attention heads are obtained on the value vector. Furthermore, the attention values of each attention head are concatenated together through a concatenation operation (Concat) to obtain MHA(Q, K, V).

[0135] Furthermore, the temporal relationship of each video segment (one or more key frames) in i-1 can be modeled through the multi-head attention mechanism and residual, and the information in i-1 is retained, avoiding model degradation during training. Furthermore, through the feed-forward neural network and residual dimension adjustment, the output f i of the i-th self-attention encoding layer can be obtained, so it can be known that the output of the k-th self-attention encoding layer is which is also the final output of the self-attention encoder. The self-attention encoder enables the features of each input video segment to contain the information of the position where the segment is located and the temporal context information.

[0136] Furthermore, the present disclosure uses the key frame feature sequence X Ref of the reference video to decode the encoded feature sequence through cross-attention modeling, as shown in the following formula:

[0137]

[0138] Continue to calculate and fk The cross multi-head attention is as shown in the following formula:

[0139]

[0140] Wherein, and are learnable parameter matrices, After passing through k = 6 layers of cross-attention layers, the output features are the first cross feature and the second cross feature Concatenate and to obtain the target cross feature

[0141] Furthermore, input f o into the recognition module. The recognition module obtains the predicted feature corresponding to the target cross feature through a convolution as shown in the following formula:

[0142] p p = Softmax(Conv(f o ));

[0143] For the key frame sequence of the video to be clipped, predictions conforming to the key frame sequence of the reference video can be obtained. For each key frame in the key frame sequence of the reference video, the start and end timestamps (t s , t e ) in the video to be clipped can be matched, as shown in the following formula:

[0144]

[0145] According to the above formula, the clipped video sequence corresponding to the video to be clipped can be extracted from the video to be clipped

[0146] Figure 6 is a schematic diagram of a video clipping method provided according to an embodiment of the present disclosure. The video clipping method includes a reference video analysis part and a video automatic clipping part. Among them, the reference video analysis part includes operations such as key frame extraction operation, structured attribute analysis operation, and unstructured attribute analysis operation, etc.; the video automatic clipping part includes operations such as clipping sequence prediction operation and music matching operation, etc.

[0147] Figure 7It is a schematic structural diagram of a personalized video automatic editing software system provided according to an embodiment of the present disclosure. The personalized video automatic editing software system includes an initialization module, a video preprocessing module, and a video editing module. Among them, the initialization module is used to load the weights of multiple algorithm models, and the algorithm models include a key frame extraction model, a decoupling network, a saliency calculation network, a noise estimation network, and a clip sequence prediction model. The video preprocessing module includes key frame extraction operations, structured attribute analysis operations, unstructured attribute analysis operations, and feature splicing operations. The video editing module includes clip sequence prediction operations and music matching operations.

[0148] Figure 8 It is a schematic structural diagram of a video generation device provided according to an embodiment of the present disclosure. As Figure 8 shown, the device includes:

[0149] A video acquisition module 510, configured to acquire a video to be edited and a reference video corresponding to the video to be edited;

[0150] A key frame feature sequence determination module 520, configured to determine the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video;

[0151] A clipped video sequence prediction module 530, configured to predict, based on the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video through a pre-trained clip sequence prediction model, a clipped completed video sequence corresponding to the video to be edited;

[0152] A clipped video sequence splicing module 540, configured to splice the clipped completed video sequence corresponding to the video to be edited to obtain a clipped video that conforms to the style of the reference video.

[0153] The technical solution of the embodiment of the present disclosure is to acquire a video to be edited and a reference video corresponding to the video to be edited; determine the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video; predict, based on the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video through a pre-trained clip sequence prediction model, a clipped completed video sequence corresponding to the video to be edited; and splice the clipped completed video sequence corresponding to the video to be edited to obtain a clipped video that conforms to the style of the reference video. In the above technical solution, the video to be edited is automatically edited based on the reference video to obtain a clipped video that conforms to the style of the reference video, making the editing more targeted, thereby meeting the personalized needs of users for video editing. In addition, the present disclosure directly outputs the clipped completed video sequence by using the clip sequence prediction model, simplifies the editing process, and thus reduces the accumulation of editing errors.

[0154] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the key frame feature sequence determination module 520 includes:

[0155] A key frame extraction unit for the video to be clipped, configured to extract key frames from the video to be clipped to obtain the key frames of the video to be clipped;

[0156] A key frame extraction unit for the reference video, configured to extract key frames from the reference video corresponding to the video to be clipped to obtain the key frames of the reference video;

[0157] An attribute analysis unit, configured to perform attribute analysis on the key frames of the video to be clipped and the key frames of the reference video respectively, to obtain the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video.

[0158] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the key frame extraction unit for the reference video may specifically be configured to:

[0159] Obtain an initial key frame, and determine the mean value of the feature vectors corresponding to the initial key frame;

[0160] Determine the key frames of the reference video according to the mean value of the feature vectors corresponding to the initial key frame through a trained key frame extraction model;

[0161] Wherein, the key frame extraction model is a reinforcement learning model trained by a preset hierarchical reward function for state information and action information, the preset hierarchical reward function includes a root reward function and a leaf reward function, the state information includes the mean value of the feature vectors of the key frames, and the action information includes actions performed based on a root policy and a leaf policy.

[0162] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the actions performed based on the root policy and the leaf policy include actions performed based on the root policy and actions performed based on the leaf policy. The actions performed based on the root policy include selecting a target key frame from multiple key frames, and the actions performed based on the leaf policy include one of the following operations: shifting the target key frame 5 frames to the left, shifting the target key frame 10 frames to the left, shifting the target key frame 5 frames to the right, and shifting the target key frame 10 frames to the right.

[0163] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the key frame feature sequence includes one or more of: a brightness histogram sequence, a hue histogram sequence, a saturation histogram sequence, a noise histogram sequence, a saliency map sequence, a camera movement sequence, and an emotional similarity sequence.

[0164] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the video generation device further includes:

[0165] A candidate music set determination module, configured to match a candidate music set from a music library based on the camera movement sequence in the key frame feature sequence of the reference video;

[0166] Correspondingly, the video generation device further includes:

[0167] A target background music determination module, configured to sample the camera movement sequence in the key frame feature sequence of the video to be clipped based on the clipped video that conforms to the style of the reference video, so as to obtain the camera movement sequence corresponding to the clipped video that conforms to the style of the reference video; and match a target background music from the candidate music set based on the camera movement sequence corresponding to the clipped video that conforms to the style of the reference video;

[0168] A target clipped video determination module, configured to add the target background music to the clipped video that conforms to the style of the reference video to obtain a target clipped video.

[0169] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the clipped sequence prediction model includes a self-attention encoding layer and a cross-attention layer;

[0170] Correspondingly, the clipped video sequence prediction module 530 may specifically be configured to:

[0171] Input the key frame feature sequence of the video to be clipped into the self-attention encoding layer to obtain an encoded feature sequence corresponding to the video to be clipped;

[0172] Input the key frame feature sequence of the reference video and the encoded feature sequence corresponding to the video to be clipped into the cross-attention layer to obtain a first cross feature and a second cross feature;

[0173] Concatenate the first cross feature and the second cross feature to obtain a target cross feature;

[0174] Perform convolution processing on the target cross feature to obtain a prediction feature corresponding to the target cross feature;

[0175] Extract the clipped completed video sequence corresponding to the video to be clipped from the video to be clipped based on the prediction feature corresponding to the target cross feature.

[0176] The video generation device provided in the embodiments of the present disclosure may execute the video generation method provided in any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0177] Figure 9The structural schematic diagram of the electronic device 10 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0178] As Figure 9 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The I / O interface 15 is also connected to the bus 14.

[0179] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0180] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the video generation method, which includes:

[0181] Obtaining the video to be clipped and the reference video corresponding to the video to be clipped;

[0182] Determine the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video;

[0183] Based on the key frame feature sequence of the video to be clipped and the key frame feature sequence of the reference video, predict the completed clipped video sequence corresponding to the video to be clipped through a pre-trained clip sequence prediction model;

[0184] Stitch the completed clipped video sequence corresponding to the video to be clipped to obtain a clipped video that conforms to the style of the reference video.

[0185] In some embodiments, the video generation method can be implemented as a computer program, which is tangibly included in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by the processor 11, one or more steps of the video generation method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the video generation method by any other suitable means (e.g., by means of firmware).

[0186] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a dedicated or general-purpose programmable processor, which can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0187] The computer programs for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0188] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0189] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0190] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0191] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0192] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this disclosure can be achieved, and no limitation is imposed herein.

[0193] The embodiments of the present disclosure also provide a computer program product, including a computer program, which when executed by a processor, implements the video generation method provided in any embodiment of the present disclosure.

[0194] In the process of implementing the computer program product, the computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0195] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A video generation method, characterized in that, Including: Obtain the video to be edited and the reference video corresponding to the video to be edited; Determine the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video; Based on the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video, use the pre-trained editing sequence prediction model to predict the completed video sequence corresponding to the video to be edited; Stitch the completed video sequence corresponding to the video to be edited to obtain an edited video that conforms to the style of the reference video.

2. The method according to claim 1, wherein The determining the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video includes: Extract key frames from the video to be edited to obtain the key frames of the video to be edited; Extract key frames from the reference video corresponding to the video to be edited to obtain the key frames of the reference video; Perform attribute analysis on the key frames of the video to be edited and the key frames of the reference video respectively to obtain the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video.

3. The method according to claim 2, characterized in that, The extracting key frames from the reference video corresponding to the video to be edited to obtain the key frames of the reference video includes: Obtain the initial key frame and determine the mean value of the feature vectors corresponding to the initial key frame; Use the trained key frame extraction model to determine the key frames of the reference video according to the mean value of the feature vectors corresponding to the initial key frame; Wherein, the key frame extraction model is a reinforcement learning model trained by a preset hierarchical reward function for state information and action information. The preset hierarchical reward function includes a root reward function and a leaf reward function. The state information includes the mean value of the feature vectors of the key frames, and the action information includes the actions performed based on the root policy and the leaf policy.

4. The method according to claim 3, characterized in that The actions performed based on the root policy and the leaf policy include the actions performed based on the root policy and the actions performed based on the leaf policy. The actions performed based on the root policy include selecting a target key frame from multiple key frames. The actions performed based on the leaf policy include one of the following operations: shifting the target key frame 5 frames to the left, shifting the target key frame 10 frames to the left, shifting the target key frame 5 frames to the right, and shifting the target key frame 10 frames to the right.

5. The method according to claim 1, characterized in that The key frame feature sequence includes one or more of the following: brightness histogram sequence, hue histogram sequence, saturation histogram sequence, noise histogram sequence, saliency map sequence, camera movement sequence, and emotional similarity sequence.

6. The method according to claim 5, characterized in that After determining the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video, it further includes: Based on the camera movement sequence in the key frame feature sequence of the reference video, match a candidate music set from the music library; Correspondingly, after stitching the completed video sequence corresponding to the video to be edited to obtain an edited video that conforms to the style of the reference video, it further includes: Based on the edited video that conforms to the style of the reference video, sample the camera movement sequence in the key frame feature sequence of the video to be edited to obtain the camera movement sequence corresponding to the edited video that conforms to the style of the reference video. Based on the camera movement sequence corresponding to the edited video that conforms to the reference video style, a target background music is matched from the candidate music set; The target background music is added to the edited video that conforms to the reference video style to obtain a target edited video.

7. The method according to claim 1, characterized in that, The edited sequence prediction model includes a self-attention encoding layer and a cross-attention layer; Correspondingly, the edited sequence prediction model that has been pre-trained predicts the edited video sequence corresponding to the video to be edited based on the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video, including: The key frame feature sequence of the video to be edited is input into the self-attention encoding layer to obtain the encoded feature sequence corresponding to the video to be edited; The key frame feature sequence of the reference video and the encoded feature sequence corresponding to the video to be edited are input into the cross-attention layer to obtain a first cross feature and a second cross feature; The first cross feature and the second cross feature are concatenated to obtain a target cross feature; The target cross feature is subjected to convolution processing to obtain the prediction feature corresponding to the target cross feature; Based on the prediction feature corresponding to the target cross feature, the edited video sequence corresponding to the video to be edited is extracted from the video to be edited.

8. A video generation device, characterized in that, Including: A video acquisition module for acquiring the video to be edited and the reference video corresponding to the video to be edited; A key frame feature sequence determination module for determining the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video; An edited video sequence prediction module for predicting the edited video sequence corresponding to the video to be edited through a pre-trained edited sequence prediction model based on the key frame feature sequence of the video to be edited and the key frame feature sequence of the reference video; An edited video sequence splicing module for splicing the edited video sequence corresponding to the video to be edited to obtain an edited video that conforms to the reference video style.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the video generation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the video generation method according to any one of claims 1-7 when executed by a processor.

11. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program implements the video generation method according to any one of claims 1-7 when executed by a processor.