Video stitching method and device based on text description, storage medium and equipment
By adjusting the resolution and frame rate of the video clips in video stitching and using the fine-tuned video frame interpolation model, the video quality problems caused by different resolution and frame rates in video stitching are solved, achieving high-quality and smooth video stitching effects.
Patent Information
- Application Number
- CN202510195674.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the video stitching process, due to the different resolutions and frame rates of different video clips, the video quality after stitching is poor.
By finding multiple matching video clips in the video library based on the text description of the spliced video, and adjusting them to the same resolution and aspect ratio, inserting interpolated video frames using the fine-tuned video frame interpolation model to make the frame rate consistent, and finally splicing multiple video clips.
Ensure that the stitched video is visually consistent with the playback fluency, and generate customized stitched videos that meet user needs, avoiding the problem of video quality degradation.
Smart Images

Figure CN120091175A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to a video splicing method, device, storage medium, and equipment based on text description. Background Art
[0002] Video splicing technology refers to selecting appropriate video clips from a video clip library and splicing them into a final video according to requirements.
[0003] During the video splicing process, text encoding technology is usually used to convert customer requirements from text description into text vectors, and then the text vectors are matched with the video clip tags in the video library for similarity, and the most suitable video clips are selected from them for splicing.
[0004] However, the resolution and frame rate of each video clip are different, resulting in poor quality of the spliced video. Summary of the Invention
[0005] This application provides a video splicing method, device, storage medium, and equipment based on text description, which is used to solve the problem of poor quality of the spliced video caused by splicing video clips with different resolutions and frame rates. The technical solution is as follows:
[0006] According to the first aspect of this application, a video splicing method based on text description is provided, and the method includes:
[0007] Search for multiple matching video clips in the video library according to the text description of the spliced video;
[0008] Adjust the multiple video clips to the same resolution and aspect ratio;
[0009] For each video clip, insert interpolated video frames between the original video frames in the video clip by using the fine-tuned video frame interpolation model, so that the frame rates of the multiple video clips are the same. The fine-tuned video frame interpolation model is obtained by jointly fine-tuning the original video frame interpolation model using a public dataset and the dataset in the video library;
[0010] Splice the multiple video clips to obtain the spliced video.
[0011] In a possible implementation, the adjusting the multiple video clips to the same resolution and aspect ratio includes:
[0012] Select a target video clip with the highest resolution from the multiple video clips, and obtain the target resolution and target aspect ratio of the target video clip;
[0013] Pad other video segments with black borders so that the aspect ratio of each padded video segment is equal to the target aspect ratio;
[0014] Use interpolation upsampling or interpolation downsampling to adjust the resolution of each video segment so that the resolution of each adjusted video segment is equal to the target resolution.
[0015] In a possible implementation, the using interpolation upsampling or interpolation downsampling to adjust the resolution of each video segment includes:
[0016] Calculate the target resolution ratio according to the target resolution;
[0017] For each video segment, use black border padding to adjust the resolution of the video segment to an intermediate resolution, and the resolution ratio of the intermediate resolution is equal to the target resolution ratio;
[0018] If the intermediate resolution is lower than the target resolution, use interpolation upsampling to proportionally adjust the intermediate resolution to the target resolution;
[0019] If the intermediate resolution is higher than the target resolution, use interpolation downsampling to proportionally adjust the intermediate resolution to the target resolution.
[0020] In a possible implementation, the using the fine-tuned video frame interpolation model to insert interpolated video frames between the original video frames in the video segment includes:
[0021] Select the video segment with the highest frame rate from the multiple video segments and obtain the target frame rate of the video segment;
[0022] For each video segment, if the target frame rate is an integer multiple of the frame rate of the video segment, use the fine-tuned video frame interpolation model to insert interpolated video frames between the original video frames in the video segment so that the frame rate of the video segment is equal to the target frame rate;
[0023] If the target frame rate is not an integer multiple of the frame rate of the video segment, calculate the integer multiple frame rate that is greater than and closest to the target frame rate according to the frame rate of the video segment; use the fine-tuned video frame interpolation model to insert interpolated video frames between the original video frames in the video segment so that the frame rate of the video segment is equal to the integer multiple frame rate; calculate the difference N between the integer multiple frame rate and the target frame rate, and delete the last N video frames played per second in the video segment so that the frame rate of the video segment is equal to the target frame rate.
[0024] In one possible implementation, the step of searching for multiple matching video segments in the video library according to the text description of the spliced video includes:
[0025] Obtain the text description of the spliced video, and use a text encoder to encode the text description to obtain a text vector;
[0026] Obtain the label vectors of each video segment in the video library;
[0027] Calculate the cosine similarity between the text vector and each label vector;
[0028] Select multiple video segments that match the text description according to the cosine similarity.
[0029] In one possible implementation, the step of obtaining the label vectors of each video segment in the video library includes:
[0030] Use the MiniCPM model to process each video segment in the video library to obtain video labels;
[0031] Use a text encoder to encode each video label to obtain the label vector of each video segment.
[0032] In one possible implementation, the method further includes:
[0033] Randomly select multiple video segments from the video library;
[0034] For each video segment, take every three consecutive video frames as a data group, and use the middle frame in the data group as the label image, and the front and back two frames as the input images to form the first training data;
[0035] Extract multiple groups of second training data from the public dataset, each group of second training data includes three consecutive video frames, and the middle frame in the three consecutive video frames is the label image, and the front and back two frames are the input images;
[0036] Use the first training data, the second training data and a preset loss function to fine-tune the original video frame interpolation model.
[0037] According to the second aspect of the present application, there is provided a video splicing device based on text description, the device includes:
[0038] A search module, configured to search for multiple matching video segments in the video library according to the text description of the spliced video;
[0039] An adjustment module, configured to adjust the multiple video segments to the same resolution and aspect ratio;
[0040] An interpolation frame module, configured to, for each video segment, insert interpolated video frames between the original video frames in the video segment by using a fine-tuned video frame interpolation model, so that the frame rates of the multiple video segments are the same. The fine-tuned video frame interpolation model is obtained by jointly fine-tuning the original video frame interpolation model by using a public dataset and a dataset in the video library;
[0041] A splicing module, configured to splice the multiple video segments to obtain the spliced video.
[0042] According to a third aspect of the present application, there is provided a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned video splicing method based on text description.
[0043] According to a fourth aspect of the present application, there is provided a computer device, which includes the above-mentioned video splicing device based on text description.
[0044] The beneficial effects of the technical solution provided by the present application at least include:
[0045] Multiple video segments that match can be found in the video library through text description, which can accurately match the user's needs and generate a customized spliced video that meets the user's needs; by adjusting the resolution, aspect ratio, and frame rate of each video segment to be consistent, it can ensure that the spliced video is visually and smoothly played; by jointly fine-tuning the video frame interpolation model by using a public dataset and a dataset in the video library, it can not only improve its performance in specific video segments, but also prevent the problem of catastrophic forgetting of the model through joint fine-tuning, ensuring the video smoothness of the generated spliced video. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 is a flowchart of a video splicing method based on text description provided by an embodiment of the present application;
[0048] Figure 2 is a flowchart of a video splicing method based on text description provided by an embodiment of the present application;
[0049] Figure 3 is a schematic diagram of splicing a waterfall video provided by an embodiment of the present application;
[0050] Figure 4 It is a schematic diagram for training a video frame interpolation model provided by an embodiment of the present application;
[0051] Figure 5 It is a structural block diagram of a video splicing device based on text description provided by an embodiment of the present application. Detailed implementation manners
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0053] As Figure 1 shown, it shows a method flow chart of a video splicing method based on text description provided by an embodiment of the present application. This video splicing method based on text description can be applied to a computer device. The video splicing method based on text description may include:
[0054] Step 101: Search for multiple matching video clips in the video library according to the text description of the spliced video.
[0055] The text description is a text for describing the semantic information of the spliced video. Among them, the more detailed the text description, the more the matching video clips meet the user's needs.
[0056] The video library includes video clips of multiple themes, and parameters such as the resolution, aspect ratio, and frame rate of different video clips may be different.
[0057] In this embodiment, the computer extracts the semantic information of each video clip in the video library, matches the extracted semantic information with the received text description, and screens multiple video clips that match the text description according to the similarity.
[0058] Step 102: Adjust the multiple video clips to the same resolution and aspect ratio.
[0059] The computer device can first determine the target resolution and target aspect ratio, then use black border filling to adjust each video clip to the target aspect ratio, and finally use interpolation upsampling or interpolation downsampling to adjust each video clip to the target resolution. For the specific adjustment process, please refer to the following description and will not be elaborated here.
[0060] Among them, the target resolution and target aspect ratio can be set values or values determined according to the selected multiple video clips, which are not limited in this embodiment.
[0061] Step 103: For each video segment, use the fine-tuned video frame interpolation model to insert interpolated video frames between the original video frames in the video segment, so that the frame rates of multiple video segments are the same. The fine-tuned video frame interpolation model is obtained by jointly fine-tuning the original video frame interpolation model using the public dataset and the dataset in the video library.
[0062] The original video frame interpolation model is used to insert video frames in the video segment to make the frame rates of each video segment the same.
[0063] In the related art, it is necessary to use a public dataset to train the video frame interpolation model, and then use the video library to fine-tune the video frame interpolation model. However, when fine-tuning the video frame interpolation model, the video frame interpolation model is prone to catastrophic forgetting, that is, while the video frame interpolation model adapts to the new task, it forgets its ability on the original task. To avoid this phenomenon, during the process of fine-tuning the video frame interpolation model, it is necessary to jointly fine-tune the video frame interpolation model using the public dataset and the dataset in the video library to balance between the customized task and the general task and prevent the video frame interpolation model from overfitting to the new data.
[0064] For a certain video segment, the video frame interpolation model can select two consecutive original video frames, generate an interpolated video frame based on these two original video frames, and insert the interpolated video frame between these two original video frames to increase the frame rate of the video segment.
[0065] Step 104: Concatenate multiple video segments to obtain a concatenated video.
[0066] After obtaining each video segment with the same resolution, aspect ratio, and frame rate, the computer device can concatenate these video segments in sequence to obtain a concatenated video.
[0067] In summary, the video concatenation method based on text description provided in the embodiments of the present application can find multiple matching video segments in the video library through text description, can accurately match the user's needs, and generate a customized concatenated video that meets the user's needs; by adjusting the resolution, aspect ratio, and frame rate of each video segment to be consistent, it can ensure that the concatenated video is visually and smoothly played; by jointly fine-tuning the video frame interpolation model using the public dataset and the dataset in the video library, it can not only improve its performance in specific video segments, but also prevent the model from having the problem of catastrophic forgetting through joint fine-tuning, ensuring the video smoothness of the generated concatenated video.
[0068] Such as Figure 2As shown in the figure, it shows a flowchart of a video splicing method based on text description provided by an embodiment of the present application. The video splicing method based on text description can be applied to a computer device. The video splicing method based on text description may include:
[0069] Step 201: Obtain the text description of the spliced video, and encode the text description using a text encoder to obtain a text vector.
[0070] The text description is a text used to describe the semantic information of the spliced video.
[0071] The computer device can input the obtained text description into the text encoder, encode the text description using the text encoder, and output a text vector.
[0072] As Figure 3 shown, assume the text description is “This video shows the majestic beauty of a waterfall cascading down a cliff into a serene lake. The waterfall, with its powerful flow, is the central focus of the video.”, and the text encoder encodes the text description to obtain the text vector Emb I .
[0073] Step 202: Process each video segment in the video library using the MiniCPM model to obtain video tags.
[0074] MiniCPM is a pre-trained model for Chinese (Mini Chinese Pre-trained Model), which is used to extract the semantic information of video segments and generate video tags according to the semantic information.
[0075] As Figure 3 shown, assume there are n video segments, then the computer device inputs the prompt word and n video segments into MiniCPM, and MiniCPM extracts the video tags [Caption 1 , Caption 2 ……Caption n . Among them, the prompt word can be “Please describe this video”.
[0076] Step 203: Encode each video tag using the text encoder to obtain the tag vector of each video segment.
[0077] The computer device inputs each video tag into the text encoder, and the text encoder encodes each video tag to obtain multiple tag vectors.
[0078] As Figure 3 shown, the text encoder encodes the video tags [Caption 1 , Caption 2 ……Caption n to obtain the tags [Emb 1 , Emb 2 ……Embn n .
[0079] Step 204, calculate the cosine similarity between the text vector and each tag vector.
[0080] The computer device calculates the cosine similarity between the text vector and each tag vector.
[0081] Step 205, select multiple video segments that match the text description according to the cosine similarity.
[0082] The computer device can sort the cosine similarities in descending order and select the video segments corresponding to the top several cosine similarities, and determine these video segments as the video segments that match the text description. Alternatively, the computer device can compare each cosine similarity with a preset threshold, retain the video segments corresponding to the cosine similarities greater than the preset threshold, and determine these video segments as the video segments that match the text description.
[0083] As Figure 3 shown, the computer device selects multiple video segments [V 2 ……V n related to "waterfall" according to the text description, and the content of these video segments matches the waterfall landscape described in the text.
[0084] Step 206, select the target video segment with the highest resolution from the multiple video segments, and obtain the target resolution and target aspect ratio of the target video segment.
[0085] The computer device compares the resolutions of the multiple video segments, selects the video segment with the highest resolution from them, determines this video segment as the target video segment, determines the resolution of this target video segment as the target resolution, and determines the aspect ratio of this target video segment as the target aspect ratio.
[0086] Step 207, perform black edge filling on the other video segments so that the aspect ratios of the filled video segments are equal to the target aspect ratio.
[0087] For each video segment other than the target video segment, the computer device uses the black border filling technology to adjust the aspect ratio of the video segment to the target aspect ratio.
[0088] Step 208, use interpolation upsampling or interpolation downsampling to adjust the resolution of each video segment so that the resolution of each adjusted video segment is equal to the target resolution.
[0089] Specifically, using interpolation upsampling or interpolation downsampling to adjust the resolution of each video segment may include: calculating the target resolution ratio according to the target resolution; for each video segment, using black border filling to adjust the resolution of the video segment to an intermediate resolution, and the resolution ratio of the intermediate resolution is equal to the target resolution ratio; if the intermediate resolution is lower than the target resolution, use interpolation upsampling to proportionally adjust the intermediate resolution to the target resolution; if the intermediate resolution is higher than the target resolution, use interpolation downsampling to proportionally adjust the intermediate resolution to the target resolution.
[0090] For example, the resolution of video segment A is 480×720, and the resolution of video segment B is 200×1440. According to 480:720, the target resolution ratio is calculated to be 2:3, and 480×720>200×1440. Then it is determined that the resolution of video segment B needs to be adjusted. The adjustment process is to first use black border filling to adjust 200×1440 to 960×1440 (because 960:1440 = 2:3). Since 960×1440>480×720, then use interpolation downsampling to proportionally adjust 960×1440 to 480×720, so that the resolutions of video segment A and video segment B are both 480×720.
[0091] Similarly, assume that the resolution of video segment A is 480×720, and the resolution of video segment B is 256×256. According to 480:720, the target resolution ratio is calculated to be 2:3, and 480×720>256×256. Then it is determined that the resolution of video segment B needs to be adjusted. The adjustment process is to first use black border filling to adjust 256×256 to 256×384 (because 256:384 = 2:3). Since 256×384 is less than 480×720, then use interpolation upsampling to proportionally adjust 256×384 to 480×720, so that the resolutions of video segment A and video segment B are both 480×720.
[0092] Step 209, for each video segment, use the fine-tuned video frame interpolation model to insert interpolation video frames between the original video frames in the video segment so that the frame rates of multiple video segments are the same. The fine-tuned video frame interpolation model is obtained by jointly fine-tuning the original video frame interpolation model using the public dataset and the dataset in the video library.
[0093] The original video frame interpolation model is used to insert video frames in a video clip to make the frame rates of each video clip the same. Among them, the video frame interpolation model can be Sparse Global Matching for Video Frame Interpolation (SGM-VFI), and its loss function includes warp loss and reconstruction loss. First, the fine-tuning process of the video frame interpolation model will be described below.
[0094] Specifically, multiple video clips are randomly selected from the video library; for each video clip, every three consecutive video frames are used as a data group, the middle frame in the data group is used as the label image, and the first and last frames are used as the input images to form the first training data; multiple groups of second training data are extracted from the public dataset, and each group of second training data includes three consecutive video frames, and the middle frame among the three consecutive video frames is the label image, and the first and last frames are the input images; the original video frame interpolation model is fine-tuned using the first training data, the second training data, and a preset loss function. Among them, the quantities of the first training data and the second training data are in proportion. For example, the quantity of the second training data: the quantity of the first training data = 2:1.
[0095] Figure 4 The fine-tuning dataset mentioned above refers to the dataset of the first training data, and the public dataset refers to the dataset of the second training data. When making the first training data, the video frames v n 、v n+1 and v n+2 in a video clip are used as a data group, the middle video frame v n+1 is used as the label image v GT , and the video frames v n and v n+2 are used as the input images; the video frames v n+3 、v n+4 and v n+5 are used as a data group, the middle video frame v n+4 is used as the label image v GT , and the video frames v n+3 and v n+5 are used as the input images. For each piece of the first training data, the first video frame and the third video frame are input into the SGM-VFI, and the SGM-VFI predicts the middle video frame v o , and according to the two loss functions, the video frame v o and the label image v GTPerform loss calculation and adjust the model parameters of SGM-VFI according to the calculation results. Similarly, the model parameters of SGM-VFI can be adjusted using the second training data.
[0096] Specifically, inserting interpolated video frames between the original video frames in the video clip using the fine-tuned video frame interpolation model may include:
[0097] (1) Select the video clip with the highest frame rate from multiple video clips and obtain the target frame rate of the video clip.
[0098] (2) For each video clip, if the target frame rate is an integer multiple of the frame rate of the video clip, use the fine-tuned video frame interpolation model to insert interpolated video frames between the original video frames in the video clip so that the frame rate of the video clip is equal to the target frame rate.
[0099] For example, if the frame rate of the original video clip is 15 frames per second and the target frame rate is 30 frames per second, and 30 is twice 15, the frame rate of the video clip can be directly adjusted to the target frame rate.
[0100] (3) If the target frame rate is not an integer multiple of the frame rate of the video clip, calculate the integer multiple frame rate that is greater than and closest to the target frame rate according to the frame rate of the video clip; use the fine-tuned video frame interpolation model to insert interpolated video frames between the original video frames in the video clip so that the frame rate of the video clip is equal to the integer multiple frame rate; calculate the difference N between the integer multiple frame rate and the target frame rate, and delete the last N video frames played per second in the video clip so that the frame rate of the video clip is equal to the target frame rate. Here, N is a positive integer.
[0101] For example, if the frame rate of the original video clip is 8 frames per second and the target frame rate is 30 frames per second, and 30 is not an integer multiple of 8, the frame rate of the video clip can be first adjusted to 32 frames per second, and then the last 2 video frames played per second are deleted to adjust the frame rate of the video clip to 30 frames per second.
[0102] Step 210, splice multiple video clips to obtain a spliced video.
[0103] After obtaining each video clip with the same resolution, aspect ratio, and frame rate, the computer device can splice these video clips in sequence to obtain a spliced video.
[0104] In summary, the video splicing method based on text description provided by the embodiments of the present application can find multiple matching video segments in the video library through the text description, accurately match the user's needs, and generate a customized spliced video that meets the user's needs. By adjusting the resolution, aspect ratio, and frame rate of each video segment to be consistent, it can ensure that the spliced video is visually and smoothly played. By jointly fine-tuning the video frame interpolation model using the public dataset and the dataset in the video library, it can not only improve its performance under specific video segments but also prevent the model from suffering from catastrophic forgetting through joint fine-tuning, ensuring the video smoothness of the generated spliced video.
[0105] As Figure 5 shown, it shows a structural block diagram of a video splicing device based on text description provided by an embodiment of the present application. The video splicing device based on text description can be applied to a computer device. The video splicing device based on text description may include:
[0106] A search module 510, configured to search for multiple matching video segments in the video library according to the text description of the spliced video;
[0107] An adjustment module 520, configured to adjust multiple video segments to the same resolution and aspect ratio;
[0108] An interpolation module 530, configured to, for each video segment, insert interpolation video frames between the original video frames in the video segment by using the fine-tuned video frame interpolation model, so that the frame rates of multiple video segments are the same. The fine-tuned video frame interpolation model is obtained by jointly fine-tuning the original video frame interpolation model using the public dataset and the dataset in the video library;
[0109] A splicing module 540, configured to splice multiple video segments to obtain a spliced video.
[0110] In an optional embodiment, the adjustment module 520 is further configured to:
[0111] Select a target video segment with the highest resolution from multiple video segments, and obtain the target resolution and target aspect ratio of the target video segment;
[0112] Perform black edge filling on other video segments so that the aspect ratio of each filled video segment is equal to the target aspect ratio;
[0113] Adjust the resolution of each video segment by upsampling or downsampling using the interpolation method so that the resolution of each adjusted video segment is equal to the target resolution.
[0114] In an optional embodiment, the adjustment module 520 is further configured to:
[0115] Calculate the target resolution ratio according to the target resolution;
[0116] For each video segment, use black border filling to adjust the resolution of the video segment to the intermediate resolution, and the resolution ratio of the intermediate resolution is equal to the target resolution ratio;
[0117] If the intermediate resolution is lower than the target resolution, use interpolation upsampling to adjust the intermediate resolution to the target resolution proportionally;
[0118] If the intermediate resolution is higher than the target resolution, use interpolation downsampling to adjust the intermediate resolution to the target resolution proportionally.
[0119] In an alternative embodiment, the frame interpolation module 530 is further configured to:
[0120] Select the video segment with the highest frame rate from multiple video segments, and obtain the target frame rate of the video segment;
[0121] For each video segment, if the target frame rate is an integer multiple of the frame rate of the video segment, use the fine-tuned video frame interpolation model to insert interpolation video frames between the original video frames in the video segment, so that the frame rate of the video segment is equal to the target frame rate;
[0122] If the target frame rate is not an integer multiple of the frame rate of the video segment, calculate the integer multiple frame rate that is greater than and closest to the target frame rate according to the frame rate of the video segment; use the fine-tuned video frame interpolation model to insert interpolation video frames between the original video frames in the video segment, so that the frame rate of the video segment is equal to the integer multiple frame rate; calculate the difference N between the integer multiple frame rate and the target frame rate, and delete the last N video frames played per second in the video segment, so that the frame rate of the video segment is equal to the target frame rate.
[0123] In an alternative embodiment, the search module 510 is further configured to:
[0124] Obtain the text description of the spliced video, and use the text encoder to encode the text description to obtain a text vector;
[0125] Obtain the label vectors of each video segment in the video library;
[0126] Calculate the cosine similarity between the text vector and each label vector;
[0127] Select multiple video segments that match the text description according to the cosine similarity.
[0128] In an alternative embodiment, the search module 510 is further configured to:
[0129] Use the MiniCPM model to process each video segment in the video library to obtain video labels;
[0130] Encode each video tag using a text encoder to obtain the tag vectors of each video segment.
[0131] In an optional embodiment, the device further includes a training module for:
[0132] Randomly select multiple video segments from the video library;
[0133] For each video segment, take every three consecutive video frames as a data group, use the middle frame in the data group as the tag image, and the front and back frames as the input images to form the first training data;
[0134] Extract multiple groups of second training data from the public dataset. Each group of second training data includes three consecutive video frames, and the middle frame among the three consecutive video frames is the tag image, and the front and back frames are the input images;
[0135] Fine-tune the original video frame interpolation model using the first training data, the second training data, and a preset loss function.
[0136] In summary, the video splicing device based on text description provided by the embodiments of the present application can find multiple matching video segments in the video library through text description, can accurately match the user's needs, and generate a customized spliced video that meets the user's needs; by adjusting the resolution, aspect ratio, and frame rate of each video segment to be consistent, it can ensure that the spliced video is visually and playback-smoothly consistent; by jointly fine-tuning the video frame interpolation model using the public dataset and the dataset in the video library, it can not only improve its performance under specific video segments, but also prevent the model from having the problem of catastrophic forgetting through joint fine-tuning, ensuring the video smoothness of the generated spliced video.
[0137] An embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned video splicing method based on text description.
[0138] An embodiment of the present application provides a computer device, and the computer device includes any of the above-mentioned video splicing devices based on text description.
[0139] It should be noted that when the video splicing device based on text description provided in the above embodiments performs video splicing based on text description, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the video splicing device based on text description is divided into different functional modules to complete all or part of the functions described above. In addition, the video splicing device based on text description provided in the above embodiments and the embodiments of the video splicing method based on text description belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0140] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk or an optical disc, etc.
[0141] The above does not intend to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.
Claims
1. A video splicing method based on text description, characterized in that: The method comprises: Searching for multiple matching video clips in a video library according to the text description of the spliced video; resizing the plurality of video clips to the same resolution and aspect ratio; For each video segment, interpolating an interpolated video frame between original video frames in the video segment using a fine-tuned video frame interpolation model so that the frame rates of the multiple video segments are the same, wherein the fine-tuned video frame interpolation model is obtained by fine-tuning the original video frame interpolation model using a public dataset and a dataset in the video library; The multiple video clips are spliced together to obtain the spliced video.
2. The video splicing method based on text description according to claim 1, characterized in that: The step of adjusting the plurality of video clips to have the same resolution and aspect ratio comprises: Selecting a target video segment with the highest resolution from the multiple video segments, and obtaining a target resolution and a target aspect ratio of the target video segment; Filling other video clips with black borders so that the aspect ratio of each filled video clip is equal to the target aspect ratio; The resolution of each video segment is adjusted by using interpolation up-sampling or interpolation down-sampling, so that the adjusted resolution of each video segment is equal to the target resolution.
3. The video splicing method based on text description according to claim 2 is characterized in that: The step of adjusting the resolution of each video segment by upsampling or downsampling using an interpolation method includes: Calculating a target resolution ratio according to the target resolution; For each video segment, adjusting the resolution of the video segment to an intermediate resolution by using black border padding, wherein the resolution ratio of the intermediate resolution is equal to the target resolution ratio; If the intermediate resolution is lower than the target resolution, adjusting the intermediate resolution to the target resolution in proportion by using interpolation upsampling; If the intermediate resolution is higher than the target resolution, the intermediate resolution is proportionally adjusted to the target resolution by using interpolation downsampling.
4. The video splicing method based on text description according to claim 1, characterized in that: The step of inserting the interpolated video frame between the original video frames in the video clip by using the fine-tuned video frame interpolation model includes: Selecting a video segment with the highest frame rate from the multiple video segments, and obtaining a target frame rate of the video segment; For each video segment, if the target frame rate is an integer multiple of the frame rate of the video segment, interpolating interpolated video frames between original video frames in the video segment using the fine-tuned video frame interpolation model so that the frame rate of the video segment is equal to the target frame rate; If the target frame rate is not an integer multiple of the frame rate of the video clip, then calculate a frame rate that is greater than and closest to the integer multiple of the target frame rate based on the frame rate of the video clip; use the fine-tuned video frame interpolation model to insert interpolated video frames between the original video frames in the video clip so that the frame rate of the video clip is equal to the integer multiple of the frame rate; calculate the difference N between the integer multiple of the frame rate and the target frame rate, and delete the last N video frames played per second in the video clip so that the frame rate of the video clip is equal to the target frame rate.
5. The video splicing method based on text description according to claim 1, characterized in that: The step of searching for multiple matching video clips in a video library according to the text description of the spliced video includes: Obtaining a text description of the spliced video, and encoding the text description using a text encoder to obtain a text vector; Obtaining a label vector for each video clip in the video library; Calculating the cosine similarity between the text vector and each label vector; A plurality of video clips matching the text description are selected according to the cosine similarity.
6. The video splicing method based on text description according to claim 5, characterized in that: The step of obtaining a label vector of each video clip in the video library includes: Using the MiniCPM model to process each video clip in the video library to obtain a video tag; Each video label is encoded using a text encoder to obtain a label vector for each video clip.
7. The video splicing method based on text description according to any one of claims 1 to 6, characterized in that: The method further comprises: randomly selecting a plurality of video clips from the video library; For each video clip, every three consecutive video frames are used as a data group, the middle frame in the data group is used as a label image, and the two preceding and following frames are used as input images to form the first training data; Extracting multiple sets of second training data from a public data set, each set of second training data includes three consecutive video frames, and the middle frame of the three consecutive video frames is a label image, and the preceding and following frames are input images; The original video frame interpolation model is fine-tuned using the first training data, the second training data and a preset loss function.
8. A video splicing device based on text description, characterized in that: The device comprises: A search module, used for searching for multiple matching video clips in a video library according to the text description of the spliced video; An adjustment module, used for adjusting the multiple video clips to the same resolution and aspect ratio; a frame interpolation module, configured to, for each video segment, insert an interpolated video frame between original video frames in the video segment using a fine-tuned video frame interpolation model so as to make the frame rates of the multiple video segments the same, wherein the fine-tuned video frame interpolation model is obtained by fine-tuning the original video frame interpolation model using a public data set and a data set in the video library; The splicing module is used to splice the multiple video clips to obtain the spliced video.
9. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the video splicing method based on text description as described in any one of claims 1 to 7.
10. A computer device, characterized in that: The computer device includes: the video splicing device based on text description as described in claim 8.
Citation Information
Patent Citations
Video processing method and device, electronic equipment, storage medium and program product
CN113099132A
Video frame insertion model determination method and device and video frame insertion method and device
CN116156218A
Presentation method of signals with different frame rates and related product
CN116366782A
Automatic video mixing and cutting method based on text-video retrieval
CN116614672A
Model training method, data processing method, equipment and storage medium
CN117520842A