Video generation method, electronic device and storage medium

By extracting frames of the original video and merging similar video frames, the time-consuming and labor-intensive problem of traditional video production is solved, and multiple highly relevant e-commerce product promotion videos are efficiently generated.

CN119893207BActive Publication Date: 2025-08-22GUANGZHOU LAILA SMART TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510118607.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-08-22
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

Traditional video production methods are time-consuming and labor-intensive. It is impossible to produce multiple e-commerce product promotion videos in a short time, and cannot meet the needs of efficient and fast video generation.

Method used

By extracting the original video frames, filtering the target video frames, searching for similar video frames in the media library, merging and generating new videos, filtering and comparing video frames using technologies such as text annotation, image recognition and feature matching, and controlling the scene similarity of adjacent video frames to improve the smoothness of the video.

Benefits of technology

It realizes the rapid generation of multiple videos with high correlation with the target object, improves video production efficiency and reduces the time cost of video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119893207B_ABST
    Figure CN119893207B_ABST
Patent Text Reader

Abstract

The present invention provides a video generation method, electronic device and storage medium. The video generation method includes the following steps: S1. Obtaining a video frame set, wherein the video frame set is each video frame contained in the original video; S2. Filtering a target frame set from the video frame set based on keywords, pictures or videos, wherein the target frame set contains at least one target video frame, and the target video frame contains a target object determined based on the keywords, pictures or videos; S3. In the media library, searching for corresponding similar video frames according to each target video frame to obtain multiple material frame sets, wherein the material frame sets contain one target video frame and at least one similar video frame; S4. Selecting the target video frame and / or similar video frame from each material frame set respectively, and merging them to generate at least one new video. The present invention can quickly generate multiple videos based on a certain video, thereby improving the efficiency of video production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and more particularly to a video generation method, electronic equipment, and storage medium. Background Art

[0002] With the rapid development of the e-commerce industry, market competition is becoming increasingly fierce. Merchants need to use various means to attract consumers' attention and increase product exposure and sales. The rise of social media and short video platforms such as Douyin, Kuaishou, and Xiaohongshu has provided new channels for e-commerce product promotion. These platforms have a large user base and concentrated traffic. By creating video content suitable for these platforms, product exposure and promotion can be effectively improved. Short videos, as an intuitive and vivid presentation method, can better attract consumer interest and enhance purchasing intention. Therefore, e-commerce product promotion videos have become an inevitable choice for merchants.

[0003] However, e-commerce product promotion videos need to be produced efficiently and quickly to meet promotion and marketing needs. Traditional video production methods are time-consuming and labor-intensive, making it impossible to produce multiple videos in a short period of time. Summary of the Invention

[0004] In order to overcome the technical problems existing in the above-mentioned prior art, the present invention provides a video generation method, an electronic device and a storage medium, which can generate videos efficiently and quickly, thereby improving the efficiency of video production.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] In a first aspect, the present invention provides a video generation method, comprising the following steps:

[0007] S1 obtains a video frame set, the video frame set is each video frame contained in the original video;

[0008] S2. Based on keywords, pictures or videos, a target frame set is filtered out from the video frame set, the target frame set includes at least one target video frame, the target video frame includes a target object determined based on the keyword, picture or video;

[0009] S3 in the media library, according to each of the target video frame search corresponding similar video frame, to obtain a plurality of material frame set, the material frame set includes a target video frame and / or at least one similar video frame;

[0010] S4. Select target video frames and / or similar video frames from each set of material frames, and merge them to generate at least one new video.

[0011] In the technical solution of the present invention, the original video is first split into multiple video frames to obtain a video frame set, and then the video frames in the video frame set are screened according to keywords, pictures or videos, and target video frames containing target objects are extracted to form a target frame set. In this way, video clips related to the target object (i.e., keywords) can be quickly obtained.

[0012] Secondly, by searching for similar video frames in the media library, similar video frames containing the target object can be obtained. The scenes of the similar video frames can be the same or similar to the scenes of the original video, or they can be different. These similar video frames constitute a material frame set, and each target video frame corresponds to a material frame set, thereby obtaining multiple material frame sets, each of which contains at least one similar video frame.

[0013] Finally, target video frames and / or similar video frames are selected from each set of source frames and combined to generate at least one new video. Generally, the selected video frames are arranged and combined in the order in which the target video frames appear in the original video to generate multiple new videos.

[0014] Through the above technical solution, the present invention can quickly generate multiple videos based on a certain video, thereby improving the efficiency of video production, and the videos are highly relevant to the subject.

[0015] The video frame mentioned in the present invention refers to: the smallest unit of video, which is a static image; for example, when playing video information, a picture frozen at any moment is a video frame.

[0016] In this context, a video scene refers to a video typically consisting of shots of different scenes, with the content of each shot typically changing continuously. Each shot typically represents a scene, and within each scene, the presentation of the target object (e.g., a product) changes continuously, but the target object typically remains within the same or similar environment and background.

[0017] In S1, the video frame set may be single or multiple; frames are extracted from different original videos to obtain multiple video frame sets.

[0018] When the present invention performs screening and searching based on keywords, images, or videos, for methods using keywords, a method of directly performing keyword searches by matching text annotations with keywords can be used, or the keywords can be converted into images representing the keywords and the images used for comparison. For methods using videos, the first video frame of the video is first extracted and used for comparison. Preferably, images are used for screening and searching.

[0019] Optionally, in S2, the method of filtering out a target frame set from the video frame set includes:

[0020] When keywords are used, the target video frame can be searched from the video frame set using text annotation and keyword matching methods, or image recognition and keyword association methods; when pictures are used, one or more methods of feature extraction and matching, image segmentation and comparison, and template matching are used to compare the picture with each video frame in the video frame set, and output the target video frame.

[0021] Optionally, in S3, the media library is a video library stored locally, on the Internet, on a mobile network, or in other cloud storages.

[0022] Generally, the target video frames in the target frame set are arranged in the order in which they appear in the original video.

[0023] Optionally, in S3, the method for searching for corresponding similar video frames based on each target video frame includes: obtaining a video to be searched from the media library, performing frame extraction on the video to be searched to obtain the video frames to be searched; and comparing the target video frame with each of the video frames to be searched to obtain similar video frames; wherein the comparison method is one or more of feature extraction and matching, image segmentation and comparison, and template matching. By combining multiple comparison methods, the accuracy of the comparison can be improved.

[0024] In the present invention, the principles and steps of the comparison method used are as follows:

[0025] The text annotation and keyword matching process involves first annotating the video frames with text, including information such as character dialogue, scene descriptions, actions, and objects. Then, the annotated text is searched for keywords, and the video frames containing the keywords are identified as target video frames. Furthermore, to search for similar video frames, a text similarity calculation method (such as cosine similarity) can be used to calculate the similarity between the text annotation of the target video frame and the text annotations of other video frames, identifying video frames with high similarity as similar video frames.

[0026] Image recognition and keyword association: Image recognition technology is used to identify objects, scenes, people, and other elements in video frames, and these recognition results are associated with keywords. For example, if the keyword is "flower," video frames containing flowers can be selected as target video frames. In addition, by comparing the characteristics, color, shape, and other attributes of objects in different video frames, feature matching algorithms (such as SIFT and SURF) can be used to extract feature points in the video frames, and then match and compare them to find similar video frames.

[0027] Feature extraction and matching: Extract key features (such as color histogram, texture features, and shape features) from the target video frame. Use the same feature extraction method to extract features from each frame in the searched video frame. These features are then matched against the key features of the target video frame. The searched video frame with the highest feature similarity is selected as the similar video frame. Additionally, feature matching algorithms (such as FLANN) can be used to accelerate the search.

[0028] The image segmentation and comparison process involves segmenting the target video frame to obtain multiple regions, performing the same segmentation on the video frame to be searched, comparing the features (e.g., color, texture, etc.) of each region, and selecting similar video frames to be searched as similar video frames based on the similarity of the region features.

[0029] Template matching uses the target video frame as a template and performs template matching on the search video frame. The similarity between the template and the search video frame is calculated, and the search video frame with the highest similarity is selected as the similar video frame. Template matching can use a variety of algorithms (such as squared difference matching and normalized cross correlation).

[0030] The similar video frame is a video frame whose similarity to the target video frame is greater than a set threshold.

[0031] Optionally, the material frame set includes at least one similar video frame; or, if no similar video frame is found, the material frame set includes one target video frame; or, the material frame set includes one target video frame and at least one similar video frame.

[0032] Optionally, in S4, the selection is manual selection or automatic random selection. The automatic random selection is to randomly select similar video frames from each material frame set to combine and generate multiple new videos.

[0033] In order to avoid the video scenes switching too frequently in a short period of time and giving the viewer a noticeable sense of disconnection, the smoothness of the video is improved, that is, the smoothness of the transition between each video frame is enhanced. The present invention provides the following embodiments:

[0034] In one embodiment, in S4, the selection method includes: determining a first video frame, sorting similar video frames in a next material frame set based on scene similarity with the first video frame, and selecting one of the top m similar video frames with higher scene similarity as the second video frame; repeating this step to determine the nth video frame, sorting similar video frames in an n+1th material frame set based on scene similarity with the nth video frame, and selecting one of the top m similar video frames with higher scene similarity as the n+1th video frame; and until the last video frame is selected. n is an integer greater than or equal to 1.

[0035] The m can be preset. m is an integer greater than or equal to 1; optionally, m is 1, 2, 3, 4, 5, 6, 7 or 8.

[0036] By controlling the scene similarity of adjacent video frames and combining video frames with high scene similarity, the new video generated by the merger can reduce the sense of video fragmentation and improve the smoothness of the transition between video frames.

[0037] Furthermore, in S4, the selection method includes: determining a first video frame, comparing a similar video frame in the next material frame set with the first video frame, obtaining a similar video frame with the highest scene similarity to the first video frame, and selecting it as the second video frame; repeating this step to determine the nth video frame, comparing a similar video frame in the n+1th material frame set with the nth video frame, obtaining a similar video frame with the highest scene similarity to the nth video frame, and selecting it as the n+1th video frame; and until the last video frame is selected. Through the above technical solution, a new video with the highest smoothness can be obtained.

[0038] Optionally, when determining the first video frame to be selected, the first video frame may be selected randomly or based on scene similarity. A similar video frame with low scene similarity is selected as the first video frame, and the resulting new video has a significant difference in scene from the original video, thereby improving the distinctiveness of the new video.

[0039] The number of new videos can be set according to a preset number value. Optionally, in S4, multiple new videos are generated by merging according to the preset number value.

[0040] E-commerce product promotional videos must be produced efficiently and quickly to meet the demands of large-scale marketing. For one thing, during the initial stages of a product launch, merchants cannot predict which video will be most appealing to consumers, so they often produce multiple promotional / introduction videos for the same product. Furthermore, distributors or sellers of the same product may reprocess the manufacturer's promotional / introduction videos, extracting the desired content and creating different videos. In both cases, traditional video production methods are time-consuming and labor-intensive, making it impossible to produce multiple videos in a short period of time.

[0041] To address the above issues, this invention provides a method for batch video generation. This method can be applied to creating product videos, such as e-commerce product introduction videos. Starting with a product video, the method extracts target video frames and searches for similar video frames of the same product. Similar video frames from different scenes can be selected. The original video scene is then replaced with the new scene.

[0042] Alternatively, a product introduction video can be split into multiple product introduction videos by adopting the technical solution of the present invention. By adopting the technical solution of the present invention, it is possible to avoid directly copying other people's videos and greatly increase the originality of the video.

[0043] In a second aspect, the present invention provides an electronic device. The electronic device includes at least one memory and at least one processor; the at least one memory is coupled to the at least one processor, the at least one memory is configured to store a computer program, and the at least one processor is configured to invoke the computer program, wherein the computer program includes instructions that, when executed by the at least one processor, cause the electronic device to perform the video generation method described in the first aspect.

[0044] In a third aspect, the present invention provides a computer storage medium comprising computer instructions, which, when executed on an electronic device, causes the electronic device to execute the video generation method according to the first aspect.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] The present invention primarily extracts target video frames from the original video, finds similar video frames based on the target video frames, classifies them into multiple sets of source frames, and then merges these sets of source frames to generate a new video. This method can quickly generate multiple videos, improving video production efficiency, and ensures that the video content is highly relevant to the target object. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flowchart of the video generation method described in an embodiment of the present application.

[0048] Figure 2 This is a schematic diagram of obtaining a video frame set from a video in step S1 of an embodiment of the present application.

[0049] Figure 3 This is a schematic diagram of screening out a target frame set from a video frame set in step S2 of an embodiment of the present application.

[0050] Figure 4 This is a schematic diagram of obtaining a material frame set in step S3 of an embodiment of the present application.

[0051] Figure 5 This is a schematic diagram of step S4 of an embodiment of the present application, in which video frames are selected from a set of material frames and merged to generate a new video.

[0052] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] To make the objectives, technical solutions, and advantages of this application more clear, this application will be further described in detail below with reference to the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0054] In the description of the embodiments of this application, unless otherwise specified, "multiple" means two or more. The number of video frames included in the video is listed here for the convenience of illustration only. It may be more or less than the actual number and should not be regarded as a limitation of the technical solution.

[0055] Figure 1 FIG. 4 is a flow chart of a video generation method according to an embodiment of the present invention. The video generation method may include the following steps:

[0056] Step S1: Obtain a video frame set, where the video frame set is each video frame contained in the original video.

[0057] Before obtaining a set of video frames, first obtain the original video that needs to be processed and fissioned, as follows:

[0058] For example, the URL (Uniform Resource Locator) corresponding to the original video is input into the electronic device, and the electronic device can then download the original video from the corresponding resource server according to the URL for processing. Alternatively, the electronic device can respond to the request and retrieve a stored video from a memory as the original video for processing, etc. This embodiment does not limit the method for obtaining the original video.

[0059] The video frame set may be single or multiple; for multiple original videos, frames are extracted from different original videos respectively to obtain multiple video frame sets.

[0060] like Figure 2 As shown, a single original video is subjected to frame extraction and split into multiple video frames, which are used to form a video frame set to obtain a video frame set.

[0061] For example, if an original video to be processed contains 20 video frames, frame extraction is performed on the video to obtain 20 video frames, that is, 20 static images (pictures), forming a video frame set.

[0062] Step S2: Filter out a target frame set from the video frame set based on keywords, pictures or videos.

[0063] When filtering and searching by keywords, images, or videos, for keywords, you can directly search by matching text annotations with keywords, or convert keywords into images representing the keywords and use these images for comparison. For videos, the first frame of the video is first extracted and used for comparison.

[0064] When keywords are used, a text annotation and keyword matching method, or an image recognition and keyword association method may be used to search for a target video frame from the video frame set.

[0065] When a picture is used, one or more methods of feature extraction and matching, image segmentation and comparison, and template matching are used to compare the picture with each video frame in the video frame set, and the target video frame is output.

[0066] As one of the embodiments, a keyword search and screening method is used, such as Figure 3 As shown, the keyword is first converted into a picture 1 representing the keyword through an algorithm or artificial intelligence model recognition. For example, if the keyword is a water cup, the keyword is converted into a picture representing the water cup, that is, the picture 1 presents an image of a water cup.

[0067] By comparing Image 1 with each video frame in the video frame collection, the video frame containing the water cup is selected as the target video frame and output as the target video frame. For example, if the video frame collection has 10 video frames, and the second, third, fourth, and eighth video frames all contain the target object of the water cup, the second, third, fourth, and eighth video frames are selected as the target video frames and output as the target video frames. Therefore, the target frame set selected from the video frame collection includes the second, third, fourth, and eighth video frames.

[0068] As a second embodiment, a method of image search and screening is adopted. For example, Picture 2 is an image of a teapot, that is, the teapot is the target object. The video frame set is directly compared with Picture 2. Based on the comparison between Picture 2 and each video frame in the video frame set, the video frame containing the teapot is screened out as the target video frame, and the target video frame is output. For example, the video frame set has 20 video frames, of which the 1st to 5th video frames and the 12th to 15th video frames all contain the target object of the teapot. Then the 1st to 5th video frames and the 12th to 15th video frames are used as the target video frames and the target video frames are output. Therefore, the target frame set screened out from the video frame set includes the 1st to 5th video frames and the 12th to 15th video frames.

[0069] As a third embodiment, a video search and screening method is used. The first video frame of the video is first extracted and used for comparison. For example, the first video frame extracted from the video is picture 3. Picture 3 may contain one object or multiple objects. If picture 3 contains multiple objects, the multiple objects are used as target objects. For example, if picture 3 contains a water cup and a teapot, the water cup and teapot are used as target objects. Picture 3 is compared with each video frame in the video frame set, and the video frames containing the water cup and / or teapot are selected as target video frames, and the target video frames are output. For example, if the video frame set has 20 video frames, where frames 1-3 contain a water cup, frames 10-13 contain a teapot, and frames 16-20 contain both a water cup and a teapot, frames 1-3, 10-13, and 16-20 are used as target video frames and output. Therefore, the target frame set selected from the video frame set includes the 1st to 3rd video frames, the 10th to 13th video frames, and the 16th to 20th video frames.

[0070] Step S3, such as Figure 4 As shown, in the media library, a corresponding similar video frame is searched for each target video frame to obtain a plurality of material frame sets, each of which includes one target video frame and / or at least one similar video frame.

[0071] The media library is a video library stored locally, on the Internet, in a mobile network, or in other cloud storages. As one embodiment, the media library is a video library stored locally, where the local refers to the electronic device executing the video generation method or a server or storage in the same local area network as the electronic device.

[0072] The method for searching for corresponding similar video frames according to each target video frame includes: obtaining a video to be searched from a media library, performing frame extraction on the video to be searched, and obtaining a video frame to be searched; comparing the target video frame with each video frame to be searched, and obtaining similar video frames; the comparison method is one or more of feature extraction and matching, image segmentation and comparison, and template matching.

[0073] For example, the target frame set contains 20 target video frames, and the local media library has 3 videos. The target video frames in the target frame set are arranged in the order in which they appear in the original video. The search starts from the first target video frame in the target frame set, and the three videos in the media library are searched in turn with the first target video frame as the target. When searching the video, the video to be compared is decomposed into multiple video frames, and the decomposed video frames are compared with the first target video frame one by one. The video frames whose similarity with the target video frame is greater than a set threshold (for example, the interval value is 0-100%, the higher the value, the higher the similarity) are selected as similar video frames, and the obtained similar video frames are classified into the first material frame set. Furthermore, depending on the situation, the first target video frame can be placed in the first material frame set.

[0074] In this way, the video frames contained in the first material frame set will have the following conditions:

[0075] In the first case, the material frame set contains at least one similar video frame;

[0076] In the second case, no similar video frame is found in the three videos in the local media library, and the material frame set includes a target video frame, namely the first target video frame, which is used for subsequent combination with the video frame;

[0077] In a third case, the material frame set includes the first target video frame and at least one similar video frame.

[0078] Taking the first case as an example, if the first video in the local media library contains 20 video frames A1-A20, of which A1, A2, and A3 are similar video frames; the second video contains 16 video frames B1-16, of which B5, B6, B10, and B11 are similar video frames; the third video contains 25 video frames C1-25, and no similar video frames are found; therefore, the first material frame set contains similar video frames A1, A2, A3, B5, B6, B10, and B11.

[0079] Next, with the second target video frame as the target, the three videos in the media library are searched in sequence to obtain the second material frame set. The similar video frames in the second material frame set may be partially the same as the similar video frames in the first material frame set, or they may be completely different.

[0080] Next, with the third target video frame as the target, the three videos in the media library are searched in sequence to obtain the third material frame set.

[0081] This process is repeated until the last target video frame, that is, the 20th target video frame, is taken as the target, and the 20th material frame set is obtained.

[0082] After step S3, 20 material frame sets are obtained based on the target frame set (including 20 target video frames). Among the 20 material frame sets, some material frame sets contain multiple similar video frames, some material frame sets contain only one similar video frame, and some material frame sets contain no similar video frames but only the corresponding target video frame.

[0083] Step S4: Select target video frames and / or similar video frames from each material frame set respectively, and merge them to generate at least one new video.

[0084] The first implementation method: manually selecting target video frames and / or similar video frames from each material frame set, and merging the selected similar video frames to generate a new video.

[0085] For example, Figure 5 As shown, after step S3, 4 material frame sets are obtained, the first material frame set includes similar video frames such as A1, A2, A3, A4, the second material frame set includes similar video frames such as B1, B2, B3, B4, the third material frame set includes similar video frames such as C0, C1, C2, C3, C4, and the fourth material frame set includes similar video frames such as D0, D1, D2, D3, D4, among which C0 and D0 are target video frames.

[0086] Manually select the four video frames A1, B2, C3, and D4 and merge them to generate the first new video.

[0087] You can continue by selecting the four video frames A1, B1, C2, and D2 and merging them to generate the second new video.

[0088] You can continue by selecting the four video frames A3, B4, C0, and D0 and merging them to generate the third new video.

[0089] Repeat this process to get other new videos.

[0090] Since the materials for making videos have been classified and organized, video producers only need to select video frames from the material frame set to merge and generate videos, which improves the efficiency of video production and saves time.

[0091] The second implementation method: using an automatic random selection method can further improve efficiency and save time.

[0092] Automatic random selection randomly selects similar video frames from each source frame set and combines them to generate multiple new videos. The number of new videos can be set according to a preset value. For example, if the preset value is 10, 10 new videos will be output.

[0093] A third implementation method is as follows: manually or automatically and randomly select the first video frame from the first material frame set, determine the selected first video frame, sort similar video frames in the next material frame set according to the scene similarity with the first video frame, and select one of the similar video frames with higher scene similarity and ranked in the top m as the second video frame; repeat this step to determine the selected n-th video frame, sort similar video frames in the n+1-th material frame set according to the scene similarity with the n-th video frame, and select one of the similar video frames with higher scene similarity and ranked in the top m as the n+1-th video frame; until the last video frame is selected.

[0094] n is an integer greater than or equal to 1. m can be preset. m is an integer greater than or equal to 1, and m can be selected from 1, 2, 3, 4, 5, 6, 7, or 8.

[0095] For example, after step S3, 8 material frame sets are obtained, and the first video frame is manually selected or automatically randomly selected. After determining the first video frame, the similar video frames in the second material frame set are sorted according to the scene similarity with the first video frame, and one of the similar video frames with higher scene similarity and ranked in the top 4 is randomly selected as the second video frame. After determining the second video frame, the similar video frames in the third material frame set are sorted according to the scene similarity with the second video frame, and one of the similar video frames with higher scene similarity and ranked in the top 4 is randomly selected as the third video frame. This process is repeated until the eighth video frame is selected. Based on the obtained 8 video frames, a new video is generated by merging.

[0096] Next, while keeping the first selected video frame unchanged (the same first video frame), randomly select one of the top four similar video frames with high scene similarity from the second set of source frames as the second video frame. Similarly, randomly select one of the top four similar video frames with high scene similarity from the third set of source frames as the third video frame. Repeat this process until the eighth video frame is selected. The resulting eight video frames are then merged again to generate a second new video.

[0097] Repeat the above steps to generate a third and fifth new video, respectively, until the preset number of new videos is reached. Alternatively, if the generated video's content overlaps with previous videos due to a small number of similar frames in the source frame set, the generation of new videos will cease. By controlling the scene similarity between adjacent video frames and combining frames with high scene similarity, these eight new videos have a more fragmented feel and smoother transitions between frames than videos generated by random merging.

[0098] The fourth implementation method is as follows: from the first material frame set, determine to select the first video frame, compare the similar video frame in the next material frame set with the first video frame, obtain the similar video frame with the highest scene similarity to the first video frame, and select it as the second video frame; repeat this step to determine the selected n-th video frame, compare the similar video frame in the n+1-th material frame set with the n-th video frame, obtain the similar video frame with the highest scene similarity to the n-th video frame, and select it as the n+1-th video frame; until the last video frame is selected.

[0099] For example, after step S3, 8 material frame sets are obtained, and the first video frame is selected manually or automatically and randomly. After determining the first video frame, the similar video frame in the second material frame set is compared with the first video frame, and the similar video frame with the highest scene similarity to the first video frame is obtained based on the scene similarity with the first video frame, and is selected as the second video frame; after determining the second video frame, the similar video frame with the highest scene similarity to the second video frame is obtained from the third material frame set based on the scene similarity with the second video frame, and is selected as the third video frame; and this is repeated until the eighth video frame is selected. Based on the obtained 8 video frames, a new video is generated by merging. Through the above implementation, only a new video is obtained, but the smoothness of the transition between the video frames of this new video should be the best, and the sense of fragmentation should be the lowest.

[0100] A fifth implementation method: Determine and select the first video frame from the first set of material frames. When determining the first video frame, the method for selecting the first video frame is to select based on the degree of scene similarity. A similar video frame with low scene similarity is selected as the first video frame. After determining the first video frame, the method of the third or fourth implementation method is used to select the second video frame, the third video frame, and so on, until the last video frame is selected. Merge and generate a new video. The generated new video has a significant difference in scene from the original video, thereby improving the difference of the new video.

[0101] On the contrary, when determining the first video frame, a video frame with high scene similarity can be selected as the first video frame.

[0102] Based on the same inventive concept as the above-mentioned embodiment of the video generation method, an electronic device is also provided in the embodiment of the present application.

[0103] like Figure 6 As shown, a specific structural block diagram of an electronic device 100, the electronic device 100 includes: at least one memory 101, and at least one processor 102; the at least one memory 101 is coupled to the at least one processor 102, the at least one memory 101 is used to store a computer program, and the at least one processor 102 is used to call the computer program, and the computer program includes instructions, and when the instructions are executed by the at least one processor, the electronic device executes the above-mentioned video generation method.

[0104] Based on the same inventive concept as the above-mentioned embodiments of the video generation method and electronic device, a computer storage medium is also provided in the embodiments of the present application.

[0105] The following describes a computer storage medium containing the above-mentioned video generation method.

[0106] A computer storage medium includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device is enabled to perform the above-mentioned video generation method.

[0107] The electronic device 100 and computer storage medium provided in this embodiment are based on the same concept as the above-mentioned video generation method. The specific implementation process is detailed in the full text of the specification and will not be repeated here.

[0108] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A video generation method, characterized in that: The following steps are involved: S1 obtains a video frame set, the video frame set is each video frame contained in the original video; S2. Based on keywords, pictures or videos, a target frame set is filtered out from the video frame set, the target frame set includes at least one target video frame, the target video frame includes a target object determined based on the keyword, picture or video; S3 in the media library, according to each of the target video frame search corresponding similar video frame, to obtain a plurality of material frame set, the material frame set includes a target video frame and / or at least one similar video frame; S4. Select target video frames and / or similar video frames from each material frame set respectively, and merge them to generate at least one new video; the selection method includes: determining the nth video frame, sorting the similar video frames in the n+1th material frame set according to the scene similarity with the nth video frame, and selecting one of the similar video frames with higher scene similarity and ranked in the top m as the n+1th video frame; repeating the above steps until the last video frame is selected; n is an integer greater than or equal to 1.

2. The video generation method according to claim 1, wherein: In S2, the method for filtering out a target frame set from the video frame set includes: when using keywords, a text annotation and keyword matching method, or an image recognition and keyword association method can be used to search for target video frames from the video frame set; when using pictures, one or more methods of feature extraction and matching, image segmentation and comparison, and template matching are used to compare the pictures with each video frame in the video frame set, and output the target video frame.

3. The video generation method according to claim 1, wherein: In S3, the media library is a video library stored locally, on the Internet, on a mobile network, or in other cloud locations.

4. The video generation method according to claim 1, wherein: In S3, the method for searching for corresponding similar video frames according to each of the target video frames includes: obtaining a video to be searched from the media library, performing frame extraction on the video to be searched, and obtaining a video frame to be searched; and comparing the target video frame with each of the video frames to be searched, and obtaining similar video frames.

5. The video generation method according to claim 1, wherein: In S4, the selection is manual selection or automatic random selection.

6. The video generation method according to claim 1, wherein: In S4, the selection method includes: determining the nth video frame, comparing the similar video frame in the n+1th material frame set with the nth video frame, obtaining the similar video frame with the highest scene similarity to the nth video frame, and selecting it as the n+1th video frame; repeating the above steps until the last video frame is selected; n is an integer greater than or equal to 1.

7. The video generation method according to claim 1, characterized in that: In S4, when determining the first video frame to be selected, the first video frame is selected by random selection or selection based on the level of scene similarity.

8. An electronic device, characterized in that: comprising at least one memory and at least one processor; The at least one memory is coupled to the at least one processor, the at least one memory is used to store a computer program, the at least one processor is used to call the computer program, the computer program includes instructions, and when the instructions are executed by the at least one processor, the electronic device executes the video generation method as described in any one of claims 1 to 7.

9. A computer storage medium, characterized in that The method comprises computer instructions, which, when executed on an electronic device, enable the electronic device to execute the video generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video generation method and device, computer device and storage medium

    CN110072120A

  • Video processing method and device and electronic equipment

    CN112070047A

  • Method for quickly generating short video

    CN112911399A

  • Video production method and system based on machine learning algorithm

    CN114915841A

  • Advertisement material production method and device, storage medium and computer equipment

    CN117676048A