Method and apparatus for synthesizing video, electronic device, computer storage medium
By extracting features from reference videos and candidate materials using a large model, and automatically selecting and combining materials using a multimodal large language model, the problem of video compositing that relies on templates and manual editing in existing technologies is solved, and an automated and efficient video compositing process is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2024-12-20
- Publication Date
- 2026-06-23
AI Technical Summary
Existing video synthesis methods rely on specific templates or manual editing, and cannot automatically select and synthesize similar videos based on any reference video. The process is cumbersome and consumes a lot of computing resources.
The system uses a large model to extract features from reference videos and candidate materials, and automatically selects target materials and combines them into a target video through a multimodal large language model, reducing user operations and computational resource consumption.
It enables automatic selection and synthesis of videos based on any reference video, reducing user operations, improving user experience, and freeing up computing resources.
Smart Images

Figure CN122269075A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image / video processing technology, and more particularly to the fields of large-scale modeling and intelligent video. Specifically, this disclosure relates to a method and apparatus for synthesizing video, an electronic device, and a computer-readable storage medium. Background Technology
[0002] With the rise of short video platforms, more and more people are combining their photos and videos into a richly textured beat-synced video or template video, and then sharing it on social media platforms.
[0003] At the same time, more and more people are choosing to store their photos or videos in the cloud to save space on their phones. Summary of the Invention
[0004] This disclosure provides a method and apparatus for synthesizing video, an electronic device, and a computer-readable storage medium.
[0005] According to a first aspect of this disclosure, a method for synthesizing video is provided, the method comprising:
[0006] Extract reference video features from user-uploaded reference videos;
[0007] Extract the material features of multiple candidate materials uploaded by the user, wherein the candidate materials are image materials or video materials;
[0008] At least one or more target materials are selected from a plurality of candidate materials based on the reference video features and the material features;
[0009] Based on the reference video, the target materials are combined to synthesize the target video.
[0010] According to a second aspect of this disclosure, an apparatus for synthesizing video is provided, the apparatus comprising:
[0011] The first feature module is used to extract reference video features from user-uploaded reference videos;
[0012] The second feature module is used to extract the candidate material features of multiple candidate materials uploaded by the user, wherein the candidate materials are image materials or video materials;
[0013] The feature comparison module is used to select one or more target materials from a plurality of candidate materials based at least on the reference video features and the candidate material features;
[0014] The video module is used to combine the target materials based on the reference video to synthesize a target video.
[0015] According to a third aspect of this disclosure, an electronic device is provided, the electronic device comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to at least one of the aforementioned processors; wherein,
[0018] The memory stores instructions that can be executed by at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method for synthesizing video.
[0019] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described method of synthesizing video.
[0020] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method for synthesizing video.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0022] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0023] Figure 1 This is a schematic flowchart of a method for synthesizing video provided in an embodiment of this disclosure;
[0024] Figure 2 This is a flowchart illustrating some steps of another method for synthesizing video provided in this embodiment of the present disclosure;
[0025] Figure 3 This is a flowchart illustrating some steps of another method for synthesizing video provided in this embodiment of the present disclosure;
[0026] Figure 4 This is a flowchart illustrating some steps of another method for synthesizing video provided in this embodiment of the present disclosure;
[0027] Figure 5 This is a flowchart illustrating some steps of another method for synthesizing video provided in this embodiment of the present disclosure;
[0028] Figure 6 This is a flowchart illustrating some steps of another method for synthesizing video provided in this embodiment of the present disclosure;
[0029] Figure 7 It is an interactive interface for obtaining reference videos and user prompts;
[0030] Figure 8 This is a schematic diagram of a specific embodiment of a method for synthesizing video provided in this disclosure;
[0031] Figure 9 This is a schematic diagram of the structure of a device for synthesizing video provided in an embodiment of this disclosure;
[0032] Figure 10 This is a block diagram of an electronic device used to implement the method for synthesizing video according to embodiments of the present disclosure. Detailed Implementation
[0033] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0034] Some related technologies can provide users with editors and preset transition effects, allowing users to manually edit and generate videos using their own photos or videos.
[0035] In some related technologies, templates are provided so that users can select photos and videos and then automatically synthesize videos.
[0036] However, the above methods rely on templates or editors provided by certain software to create videos. They cannot generate similar videos based on any video, nor can they automatically select photos or videos taken by the user based on reference videos. Users need to make manual selections, making the process relatively cumbersome.
[0037] The methods, apparatuses, electronic devices, and computer-readable storage media for synthesizing videos provided in this disclosure are intended to solve at least one of the above-mentioned technical problems of the prior art.
[0038] The method for synthesizing video provided in this disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in memory. Alternatively, the method can be executed by a server.
[0039] Figure 1 A flowchart illustrating a method for synthesizing video according to an embodiment of this disclosure is shown. Figure 1 As shown in the figure, the method for synthesizing video provided in this embodiment of the present disclosure includes steps S110, S120, S130 and S140.
[0040] In step S110, the reference video features of the user-uploaded reference video are extracted;
[0041] In step S120, the candidate material features of multiple candidate materials uploaded by the user are extracted respectively. The candidate materials are image materials or video materials.
[0042] In step S130, at least one or more target materials are selected from multiple candidate materials based on the features of the reference video and the features of the candidate materials;
[0043] In step S140, the target materials are combined based on the reference video to synthesize the target video.
[0044] For example, in step S110, the reference video uploaded by the user can be a video uploaded by the user through an interactive device or interactive page.
[0045] In some possible implementations, a pre-trained large model can be used to extract reference video features from a reference video. Specifically, the reference video and prompt text can be input into the large model to prompt it to extract reference video features from the reference video.
[0046] The prompt text is the input text used to guide the model to perform a specific task.
[0047] The characteristics of a reference video can be key information about the reference video, such as the font, decorations, resolution, and duration of the text used in the reference video.
[0048] In some possible implementations, in step S120, the multiple candidate materials uploaded by the user can be images or videos uploaded by the user to the cloud platform.
[0049] Since the purpose of extracting candidate material features is to filter out target materials from the candidate materials that can form videos similar to the reference video, the extracted candidate material features and the extracted reference video features should be of the same type. Therefore, a large model that extracts reference video features can be used to extract candidate material features from the candidate materials.
[0050] Specifically, candidate materials and prompt text can be input into the large model to prompt the large model to extract candidate material features.
[0051] In some possible implementations, step S120 can be performed in advance, that is, step S120 is executed immediately after the user uploads the image or video to the cloud platform.
[0052] In some possible implementations, in step S130, candidate material features that match the reference video features can be matched, and the candidate material features that match the reference video features with a high degree of matching can be selected as the target material.
[0053] In some possible implementations, the reference video may also contain aspects that users are not satisfied with. Users can input their improvement requirements for the reference video through interactive devices or interfaces. The user's improvement requirements can also participate in the selection of target material from candidate materials. That is, the target material is selected from candidate materials based on the matching degree between the characteristics of candidate materials, the characteristics of the reference video, and the user's improvement requirements.
[0054] In some possible implementations, in step S140, the target materials are combined in sequence based on the reference video to generate the target video.
[0055] In some specific implementations, the target material, reference video, and user input improvement requirements can be used to generate a prompt template for the input multimodal large language model, generate a prompt multimodal large language model reference video and improvement requirements, and synthesize a video prompt based on the target material.
[0056] Multimodal Large Language Models (MLLM) are built upon Large Language Models (LLM) and Large Vision Models (LVM). They can handle various media data types, including text, images, audio, and video. Through joint training, they learn the relationships between data from different modalities, improving the model's performance and generalization ability. Their core lies in cross-modal information fusion and understanding, enabling the model to grasp the deeper meaning behind the data more comprehensively and accurately.
[0057] The prompt template may include: prompting a multimodal large model to obtain information from reference videos and improvement requirements, combine target materials, and generate a target video.
[0058] For example, the prompt template could be: "You are a professional video generator. Based on the current video {}, and using {}, generate a video that meets {}." The generated prompt could be: "You are a professional video generator. Based on the current video {reference video}, and using {target material}, generate a video that meets {improvement requirements}."
[0059] In the video synthesis method provided in this disclosure, any reference video uploaded by the user can be referenced, and images or videos can be automatically selected to synthesize a shareable target video. The process of generating the target video can refer to any video without relying on a specific template, and the automatic selection of materials and automatic video synthesis reduces user operations, improves user experience, and also releases the computing resources required for user interaction.
[0060] The method for synthesizing videos provided in this disclosure will now be described in detail.
[0061] In some possible implementations, the reference video features may include video frame features and video audio features.
[0062] In some possible implementations, the video frame and audio of the reference video can be separated, and a visual multimodal model, namely VLM (Vision Language Models), can be used to extract features from the video frame, while an audio multimodal model can be used to extract features from the audio.
[0063] In some possible implementations, the same multimodal large model can also be used to extract features from video footage and video audio.
[0064] Figure 2 The diagram illustrates a flowchart of one approach for feature extraction from video footage and audio using a multimodal large model. Figure 2 As shown, using a multimodal large model to extract features from video footage and video audio may include step S210.
[0065] In step S210, video frame features of the reference video and video speech features of the reference video are extracted based on the pre-trained multimodal large model.
[0066] Among them, video image features include the video image style of the reference video, the number of corresponding frames in the reference video, and the frame parameters; video audio features include the music style and rhythm of the video audio.
[0067] In some possible implementations, in step S210, the reference video can be directly input into a pre-trained multimodal large model to obtain video image features and video speech features.
[0068] Specifically, a reference video can be input into a pre-generated prompt template to generate a multimodal large language model that obtains the video frame segmentation, the number of frames and frame parameters of the reference video, as well as the music style and rhythm of the video audio.
[0069] Alternatively, the video footage used in the reference video can be determined based on the video transitions in the reference video, and the number of video footage frames can be obtained by counting.
[0070] For each video clip, a multimodal large model feature extraction module is used to obtain the image parameters of each video clip, such as resolution, duration, animation style, and color tone. Simultaneously, the video clips are identified, including the people appearing in the clips and their corresponding information (such as gender, age, and what they are doing), as well as the scenes in the clips and their corresponding scene information, such as scene type.
[0071] The style of a reference video can be determined by the color tone of the video footage, the characters appearing in the footage and their corresponding information, and the scenes in the footage and their corresponding information.
[0072] The feature extraction module of the multimodal large model is used to identify music appearing in video speech. It can identify the music name based on the music content, obtain the music style and rhythm based on the music name, and determine the music style and rhythm by identifying the music's pitch, loudness and timbre.
[0073] Figure 3 This diagram illustrates a flowchart of an implementation method for obtaining target material based on the video frame features and audio-visual features of a reference video, as shown below. Figure 3 As shown, obtaining target material based on the video frame features and video audio features of the reference video may include steps S310 and S320.
[0074] In step S310, based on the pre-trained large model, the matching degree between the candidate material features and the video scene style is obtained;
[0075] In step S320, a target material is selected from multiple candidate materials based on the image matching degree and image parameters, with the number of images being equal to the number of images.
[0076] In some possible implementations, in step S310, a multimodal large model as described in step S210 can be used to obtain candidate material features.
[0077] This involves using a multimodal large-scale model feature extraction module to obtain the image parameters of each candidate material, such as resolution, duration, animation style, and color tone. Simultaneously, the candidate materials are identified, including the people appearing in the materials and their corresponding information (such as gender, age, and what they are doing), as well as the scenes and their corresponding scene information, such as scene type.
[0078] The visual style of candidate materials can be determined based on their color tone, the characters appearing in the materials and their corresponding character information, and the scenes in the materials and their corresponding scene information.
[0079] The similarity between the visual style of the candidate material and the visual style of the video is used as the matching degree between the candidate material features and the video visual style, i.e., the visual matching degree.
[0080] In some possible implementations, in step S320, multiple target materials are selected according to the matching degree between each candidate material and the video scene style from high to low. At this time, the number of target materials can be greater than the number of scenes.
[0081] Based on the image parameters, multiple target materials are filtered, and target materials that do not meet the image parameters are deleted, resulting in a number of target materials equal to the number of images.
[0082] In some possible implementations, candidate materials can also be selected based on the degree of matching between the candidate materials and the video and audio features.
[0083] Figure 4 The diagram illustrates a process for filtering candidate materials based on the degree of matching between candidate materials and video / audio features. Figure 4 As shown, filtering candidate materials based on the degree of matching between candidate materials and video / audio features may include steps S410 and S420.
[0084] In step S410, based on the pre-trained large model, the music matching degree between the candidate material features and the music style of the reference video is obtained;
[0085] In step S420, a target material is selected from multiple candidate materials based on the image matching degree, music matching degree, and image parameters, with the number of images equal to the number of images.
[0086] In some possible implementations, in step S410, the visual style of the candidate material can be obtained based on the method described in step S310.
[0087] The similarity between the visual style of the candidate material and the music style of the reference video is used as the matching degree between the candidate material features and the music style of the reference video, i.e., the music matching degree.
[0088] In some possible implementations, in step S420, the matching degree between the candidate material and the reference video can be obtained by weighted summing of the image matching degree and music matching degree corresponding to the candidate material features.
[0089] Select multiple target materials based on the matching degree between each candidate material and the reference video, from high to low. At this time, the number of target materials can be greater than the number of frames.
[0090] Based on the image parameters, multiple target materials are filtered, and target materials that do not meet the image parameters are deleted, resulting in a number of target materials equal to the number of images.
[0091] In some possible implementations, the reference video may also contain aspects that users are not satisfied with. Users can input their improvement requirements for the reference video through interactive devices or interfaces. The user's improvement requirements can also participate in the selection of target material from candidate materials. That is, the target material is selected from candidate materials based on the matching degree between the characteristics of candidate materials, the characteristics of the reference video, and the user's improvement requirements.
[0092] Figure 5 The diagram illustrates a flowchart of an implementation method for selecting target material from candidate materials based on the degree of matching between candidate material features, reference video features, and user improvement requirements. Figure 5 As shown, selecting target material from candidate materials based on the matching degree of candidate material features, reference video features, and user improvement requirements may include steps S510, S520, and S530.
[0093] In step S510, user prompts uploaded by the user are obtained. The user prompts include the user's requirements for the video visuals of the target video and / or the user's requirements for the video audio of the target video.
[0094] In step S520, the video frame features are modified according to the video frame requirements to generate frame reference features; the video audio features are modified according to the video audio requirements to generate audio reference features.
[0095] In step S530, one or more target materials are selected from multiple candidate materials based on the image reference features, the voice reference features, and the candidate material features.
[0096] In some possible implementations, in step S510, the user-uploaded user prompt is the user's request for improvement of the reference video, which can be input by the user through an interactive device or interface.
[0097] It can include the user's requirements for the video image of the target video, such as requirements for video image parameters (e.g., resolution, duration, animation style, color tone, font used, whether there is a watermark, etc.) or requirements for video image style.
[0098] It can also include the user's requirements for the audio and video of the target video, such as the style of the background music, or only the music with the human voice removed.
[0099] In some possible implementations, in step S520, based on the video frame requirements, if there is a conflict between the video frame requirements and the video frame features, the video frame requirements are used as the frame reference features. For example, if the resolution of the video frame requirements is 1080, but the resolution of the video frame features is 720, then the resolution in the final frame reference features is 1080.
[0100] Similarly, based on the video-to-speech requirements, in cases where the video-to-speech requirements conflict with video-to-speech features, the video-to-speech requirements are used as the speech reference features.
[0101] In some possible implementations, in step S530, the degree of matching between the picture style of the candidate material and the video picture style in the picture reference features is taken as the matching degree between the candidate material features and the picture reference features, and the similarity between the picture style of the candidate material and the music style in the speech reference features is taken as the matching degree between the candidate material features and the speech reference features.
[0102] The matching degree of the candidate material is obtained by weighted summing the matching degree between the candidate material features and the image reference features and the matching degree between the candidate material features and the speech reference features.
[0103] Select multiple target materials according to the matching degree of each candidate material from high to low. At this time, the number of target materials can be greater than the number of images in the image reference features.
[0104] Based on the image parameters in the image reference features, multiple target materials are filtered out. Target materials that do not meet the image parameters in the image reference features are deleted, and the number of target materials obtained is the number of images in the image reference features to obtain the target materials required by the foot massage user.
[0105] After obtaining the target materials, the target materials can be combined with reference videos to create the target video.
[0106] In some possible implementations, background music can be selected from a preset music library based on the music style and rhythm of the reference video; the target material can be combined into the video frame of the target video, and the background music can be used as the video audio of the target video.
[0107] In other words, music with the same style and rhythm as the background music in the reference video is selected as the background music for the target video.
[0108] Since the target material is obtained based on the music style and rhythm of the reference video, the style of the target material matches the music style and rhythm of the reference video. Therefore, the music style and rhythm of the music selected based on the music style and rhythm of the reference video match the video footage composed of the target material.
[0109] In this way, the video footage and background music of the target video will match.
[0110] In some possible implementations, the reference video may contain beats, where the video frame transitions and music rhythms overlap. The beat features of the reference video can be obtained, and a target video with similar beats can be generated by referencing these beat features.
[0111] Figure 6 The diagram illustrates a flowchart of an implementation method for obtaining the beat features of a reference video and generating a target video with similar beat features, as shown below. Figure 6 As shown, obtaining the beat features of the reference video and generating a target video with similar beat features may include steps S610, S620, and S630.
[0112] In step S610, the joint features of the video frame and the video speech of the reference video are extracted based on the pre-trained large model;
[0113] In step S620, the overlap time point between the video frame switching and the music rhythm is determined based on the joint features, and the timing features of the reference video are determined based on the overlap time point.
[0114] In step S630, the target material is switched at the overlapping time point of the target video.
[0115] In some possible implementations, in steps S610 and S620, the switching time points of the reference video frames and the music drum beats of the reference video audio can be obtained based on the large model. If the switching time point of the reference video frames is consistent with the time point corresponding to the music drum beats, then the time point is marked as the overlapping time point.
[0116] All overlapping time points of the acquired reference video are used as the timing features of the reference video.
[0117] In some possible implementations, in step S630, since the background music of the target video is selected based on the music style and rhythm of the reference video, the background music of the target video is consistent with the background music of the reference video in terms of music style and rhythm. The drum beats of the background music of the reference video also basically overlap with the drum beats of the background music of the target video. Therefore, the overlap time point of the reference video can be used as the overlap time point of the target video.
[0118] In some possible implementations, the target material can be switched at the overlapping time point of the target video, i.e., the screen can be switched. This can make the time point of the lake surface switch coincide with the drum beat of the background music of the target video, thus achieving the timing of the target video.
[0119] In some possible implementations, the target material is determined based on the number of frames, meaning the number of target materials is the same as the number of frames. However, since different materials correspond to different durations, the duration of the video frame after combining the target materials may not be consistent with the duration of the reference video. Therefore, it is necessary to process the video frame after combining the target materials to make the duration of the video frame consistent with the duration of the reference video.
[0120] Specifically, when the sum of the time lengths of all target materials is less than the time length of the reference video, a pre-trained large model is used to dynamically process the image materials in the target materials, generating dynamic materials corresponding to the image materials; the dynamic materials are then combined with the target materials to synthesize the target video.
[0121] If the sum of the durations of all target materials is greater than the duration of the reference video, a pre-trained large model is used to edit the video materials in the target materials, and the target video is synthesized based on the edited target materials.
[0122] The sum of the durations of all target materials is less than the duration of the reference video. In other words, the duration of the video frame after combining the target materials is less than the duration of the reference video. Therefore, the video frame length of the target video can be increased by making the image materials dynamic.
[0123] Among them, the dynamicization of image materials can be achieved based on a large model, that is, the image materials are input into the large model, the large model analyzes the image materials, generates images similar to the image materials, and combines them with the image materials to form a video.
[0124] The generated video is used to replace the original image materials and combined with other target materials to generate the target video.
[0125] The sum of the durations of all target materials is greater than the duration of the reference video. In other words, the duration of the video frame after combining the target materials is greater than the duration of the reference video. Therefore, the length of the target video frame can be reduced by editing the video materials.
[0126] Editing video footage can also be achieved using a large model. This involves inputting the video footage into the large model, which then analyzes and extracts features from the footage to determine the amount of information contained in each frame. Frames with less information are then cut out, while frames with more information are retained.
[0127] The target video is generated by combining the edited video footage with other target footage to replace the original video footage.
[0128] The process of combining target materials to generate a target video can be achieved based on a large model. The reference video and target materials are input into the large model, which determines the combination order of the target materials based on the similarity between the target materials and the image frames of the reference video at different time points. The target materials are then combined according to the combination order to generate the target video.
[0129] The method for synthesizing video provided in this disclosure will be described in detail below with reference to a specific embodiment.
[0130] Figure 7 The interactive interface for obtaining reference videos and user prompts is shown; such as Figure 7 As shown, users can upload a reference video by clicking the "+" sign below "Select a reference video" on the interactive interface, and enter a description of the desired video in the box below "Do you have any special requirements for the generated video?", such as... Figure 7 The phrase "choose photos with your child whenever possible, and the final video should be 1080 resolution" is the user's description of the desired video, i.e., the user prompt. Clicking the "Start Generation" button initiates the method for synthesizing video provided in this embodiment of the present disclosure to obtain the target video.
[0131] Figure 8 The diagram illustrates a process schematic of a specific embodiment of a method for synthesizing video provided in this disclosure, as shown below. Figure 8As shown, the user sees a decent travel video but doesn't know who made it or which application they used. However, if they want to create a similar video, they can click the "+" sign below "Select a video to reference" on the interactive interface to upload the travel video or its link. In the box below "What special requirements do you have for the generated video?" on the interactive interface, they can enter a description of the desired video: "Select photos from my trip to Jiuzhaigou last year and create an identical video that features my off-road vehicle, without any logos or watermarks, and in 1080p HD resolution."
[0132] Upon receiving a user-uploaded reference video or link to a reference video, along with a description of the desired video, feature extraction is performed on the reference video. The analysis reveals that the reference video is 18 seconds long, contains 9 key moments, and includes 10 pieces of material, including 4 photos and 6 videos. The theme is travel, and it incorporates various elements such as highways, mountains, cars, and sunny days. Each segment has dynamic text, which are all characteristics of the reference video.
[0133] You can remove the videos and photos and use them as a template, then fill in the target materials you acquire later.
[0134] With user authorization, the cloud platform performs AI (artificial intelligence) analysis on the user-uploaded photos and videos, extracting features. Based on the features of the reference video and the extracted features, it selects materials from the user-uploaded photos and videos to obtain target materials. Based on a large model, it fills the target materials into the template corresponding to the reference video according to the features of the reference video, synthesizing the target video. After synthesis, the target video is sent back to the user.
[0135] Based on and Figure 1 The method shown follows the same principle. Figure 9 A schematic diagram of the structure of an apparatus for synthesizing video according to an embodiment of this disclosure is shown, as follows: Figure 9 As shown, the device 90 for synthesizing video may include:
[0136] The first feature module 910 is used to extract reference video features of the user-uploaded reference video;
[0137] The second feature module 920 is used to extract the candidate material features of multiple candidate materials uploaded by the user, where the candidate materials are image materials or video materials;
[0138] The feature comparison module 930 is used to select one or more target materials from multiple candidate materials based on at least the features of the reference video and the features of the candidate materials.
[0139] Video module 940 is used to combine target materials based on reference video to synthesize a target video.
[0140] In the video synthesis apparatus provided in this embodiment, any reference video uploaded by the user can be referenced to automatically select images or videos to synthesize a shareable target video. The process of generating the target video can refer to any video without relying on a specific template, and the automatic selection of materials and automatic video synthesis reduces user operations, improves user experience, and also frees up computing resources required for user interaction.
[0141] It is understood that the above-described modules of the video synthesis apparatus in the embodiments of this disclosure have the ability to implement... Figure 1 The embodiments shown illustrate the functionality of corresponding steps in the method for synthesizing video. This functionality can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described functions. These modules can be software and / or hardware, and each module can be implemented individually or integrated from multiple modules. For a detailed description of the functions of each module in the above-described video synthesis apparatus, please refer to [link to relevant documentation]. Figure 1 The corresponding description of the video synthesis method in the illustrated embodiments will not be repeated here.
[0142] In some possible implementations, the reference video features include video frame features and video speech features; the first feature module is used to: extract video frame features of the reference video and video speech features of the reference video based on a pre-trained multimodal large model; wherein, the video frame features include the video frame style of the reference video, the number of frames corresponding to the reference video, and the frame parameters; the video speech features include the music style and music rhythm of the video speech.
[0143] In some possible implementations, candidate material features include information about the people appearing in the candidate material and scene information about the scene corresponding to the candidate material; the feature comparison module is used to: obtain the matching degree between the candidate material features and the video scene style based on a pre-trained large model; and select target material from multiple candidate materials in a number equal to the number of scenes based on the matching degree and scene parameters.
[0144] In some possible implementations, the feature comparison module is used to: obtain the music matching degree between the features of candidate materials and the music style of the reference video based on a pre-trained large model; and select target materials from multiple candidate materials in a number equal to the number of frames based on the image matching degree, music matching degree, and image parameters.
[0145] In some possible implementations, the video module is used to: select background music from a preset music library based on the music style and rhythm of the reference video; combine the target material into video frames of the target video, and use the background music as the video audio of the target video.
[0146] In some possible implementations, the reference video features include timing features; the second feature module is used to: extract joint features of the video frames and audio of the reference video based on a pre-trained large model; determine the overlap time points of video frame transitions and music rhythms based on the joint features; and determine the timing features of the reference video based on the overlap time points.
[0147] In some possible implementations, the video module is used to: set the switching of target footage at the overlapping time points of the target video.
[0148] In some possible implementations, the feature comparison module is used to: obtain user-uploaded prompts, which include the user's requirements for the video frame and / or the user's requirements for the video audio of the target video; modify the video frame features according to the video frame requirements to generate frame reference features; modify the video audio features according to the video audio requirements to generate audio reference features; and select one or more target materials from multiple candidate materials based on the frame reference features, audio reference features, and candidate material features.
[0149] In some possible implementations, the video module is used to: use a pre-trained large model to dynamically process the image materials in the target materials, generating dynamic materials corresponding to the image materials, provided that the sum of the time lengths corresponding to all target materials is less than the time length of the reference video; and combine the dynamic materials with the target materials to synthesize the target video.
[0150] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0151] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0152] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0153] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method for synthesizing video as provided in the embodiments of this disclosure.
[0154] Compared to existing technologies, this electronic device can automatically select images or videos from any user-uploaded reference video to synthesize a shareable target video. The process of generating the target video can reference any video without relying on a specific template, and it automatically selects and synthesizes materials, reducing user operations, improving the user experience, and freeing up computing resources previously required for user interaction.
[0155] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform a method for synthesizing video as provided in the embodiments of this disclosure.
[0156] Compared to existing technologies, this readable storage medium can reference any user-uploaded video, automatically select images or videos, and synthesize a shareable target video. The process of generating the target video can reference any video without relying on a specific template, and automatically selects and synthesizes materials, reducing user operations, improving the user experience, and freeing up computing resources previously required for user interaction.
[0157] The computer program product includes a computer program that, when executed by a processor, implements a method for synthesizing video as provided in embodiments of this disclosure.
[0158] Compared to existing technologies, this computer program can automatically select images or videos from any user-uploaded reference video to synthesize a shareable target video. The process of generating the target video can reference any video without relying on a specific template, and it automatically selects and synthesizes materials, reducing user operations, improving the user experience, and freeing up computing resources previously required for user interaction.
[0159] Figure 10 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0160] like Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0161] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0162] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the method of synthesizing video. For example, in some embodiments, the method of synthesizing video may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the method of synthesizing video described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform the method of synthesizing video by any other suitable means (e.g., by means of firmware).
[0163] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0164] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0165] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0167] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0168] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0169] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0170] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for synthesizing video, comprising: Extract reference video features from user-uploaded reference videos; Extract the candidate material features from multiple candidate materials uploaded by the user, wherein the candidate materials are image materials or video materials; At least one or more target materials are selected from a plurality of candidate materials based on the reference video features and the candidate material features; Based on the reference video, the target materials are combined to synthesize the target video.
2. The method according to claim 1, wherein, The reference video features include video frame features and video audio features; the extraction of reference video features from user-uploaded reference videos includes: Based on a pre-trained multimodal large model, video frame features and video speech features of the reference video are extracted from the video frame of the reference video. The video image features include the video image style of the reference video, the number of frames corresponding to the reference video, and the image parameters; the video audio features include the music style and music rhythm of the video audio.
3. The method according to claim 2, wherein, The candidate material features include the character information of the people appearing in the candidate material and the scene information of the scene corresponding to the candidate material; The step of selecting one or more target materials from a plurality of candidate materials based at least on the reference video features and the candidate material features includes: Based on a pre-trained large model, the matching degree between the candidate material features and the video scene style is obtained; Based on the image matching degree and the image parameters, target material is selected from a plurality of candidate materials in a number equal to the number of images.
4. The method according to claim 3, wherein, The step of selecting one or more target materials from a plurality of candidate materials based at least on the reference video features and the candidate material features includes: Based on a pre-trained large model, the music matching degree between the candidate material features and the music style of the reference video is obtained; Based on the image matching degree, the music matching degree, and the image parameters, target material is selected from a plurality of candidate materials in a number equal to the number of images.
5. The method according to claim 2, wherein, The step of combining the target materials based on the reference video to synthesize the target video includes: Background music is selected from a preset music library based on the music style and rhythm of the reference video; The target materials are combined to form the video frame of the target video, and the background music is used as the video audio of the target video.
6. The method according to claim 2, wherein, The reference video features include timing features; The reference video features include video frame features and video audio features; The extraction of reference video features from user-uploaded reference videos includes: Based on a pre-trained large model, the joint features of the video frames and the video audio of the reference video are extracted; Based on the joint features, the overlap time points of video screen switching and music rhythm are determined, and the timing features of the reference video are determined according to the overlap time points.
7. The method according to claim 6, wherein, The step of combining the target materials based on the reference video to synthesize the target video includes: The target material is switched at the overlapping time points of the target video.
8. The method according to claim 2, wherein, The step of selecting one or more target materials from a plurality of candidate materials based at least on the reference video features and the candidate material features includes: Obtain user-uploaded prompts, which include the user's requirements for the video visuals of the target video and / or the user's requirements for the video audio of the target video. The video frame features are modified according to the video frame requirements to generate frame reference features; the video audio features are modified according to the video audio requirements to generate audio reference features. One or more target materials are selected from a plurality of candidate materials based on the image reference features, the voice reference features, and the candidate material features.
9. The method according to claim 1, wherein, The step of combining the target materials based on the reference video to synthesize the target video includes: If the sum of the time lengths corresponding to all the target materials is less than the time length of the reference video, a pre-trained large model is used to dynamically process the image materials in the target materials to generate dynamic materials corresponding to the image materials. The dynamic material is combined with the target material to synthesize the target video.
10. An apparatus for synthesizing video, comprising: The first feature module is used to extract reference video features from user-uploaded reference videos; The second feature module is used to extract the candidate material features of multiple candidate materials uploaded by the user, wherein the candidate materials are image materials or video materials; The feature comparison module is used to select one or more target materials from a plurality of candidate materials based at least on the reference video features and the candidate material features; The video module is used to combine the target materials based on the reference video to synthesize a target video.
11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-9.
13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.