A method, apparatus, electronic device and medium for extracting video segments
By using similarity calculation and threshold judgment in video clip extraction, the start and end frames of video clips are directly determined, which solves the problem of large amount of calculation in the prior art and realizes efficient and accurate video clip extraction.
Patent Information
- Application Number
- CN202310312994.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-03-28
AI Technical Summary
In the prior art, video clip extraction efficiency is low, and full-frame calculation and scoring of video frame images are required, resulting in large amounts of calculation.
By obtaining all frame images of the total video, the similarity calculation is performed with the preset keyframe images in turn, the start and end frames of the video clip are determined using the similarity threshold, and the target video clip is directly intercepted, avoiding segmentation and multiple scoring of the video content.
The efficiency of video clip extraction is improved, the calculation amount is reduced, and the accuracy of extraction is improved through similarity judgment.
Smart Images

Figure CN116310994B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of video extraction, and in particular to a method, device, electronic device and medium for extracting video clips. Background Art
[0002] The extraction of video segments may be performed on any one or more relatively short video segments in the video. For example, to extract wonderful video segments in a video, one or more video segments whose contents are more wonderful than those of other video segments in the video may be extracted.
[0003] In the related art, video segment extraction requires that the video be completely acquired, divided into multiple video segments according to the content of the video, and each video segment needs to be scored, and the video segment extraction is performed based on the score of each video segment. However, this method calculates the score of the video segment based on all the frame images of the video segment, and the extraction efficiency is extremely low.
[0004] Therefore, how to improve the efficiency of video segment extraction is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention
[0005] In order to improve the efficiency of video segment extraction, the present application provides a video segment extraction method, device, electronic device and medium.
[0006] In a first aspect, the present application provides a video segment extraction method, which adopts the following technical solution:
[0007] A video clip extraction method, comprising:
[0008] Get all frame images of the total video;
[0009] Calculate the similarity of each frame image in all the frame images with the preset first key frame image in turn, and obtain the similarity value corresponding to each frame image in turn;
[0010] When the similarity values obtained in sequence are greater than the preset similarity threshold for the first time, obtaining the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time;
[0011] Get the duration of the target video segment;
[0012] The target video segment is determined according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment.
[0013] By adopting the above technical solution, by obtaining all the frame images of the total video, calculating the similarity between each frame image in all the frame images of the total video and a preset first key frame image in sequence, obtaining the similarity value corresponding to each frame image in sequence, when the similarity value obtained in sequence is greater than the preset similarity threshold for the first time, determining the target video segment from the total video based on the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the obtained target video segment. By calculating the similarity between all the frame images of the obtained total video and the preset first key frame image at once, and using the frame image whose similarity value is greater than the preset similarity threshold for the first time as the initial image of the target image, it is not necessary to divide the total video content into multiple video segments and score each video segment, reducing a large amount of computational complexity and improving the efficiency of video segment extraction.
[0014] In a preferred example of the present application, it can be further configured that: determining the target video segment according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment includes:
[0015] Obtaining an initial target video segment based on the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment;
[0016] Obtaining the first frame image of the initial target video segment, where the first frame image is the last frame image of the initial target video segment;
[0017] Calculating the similarity between the first frame image and a preset second key frame image to obtain a first similarity value;
[0018] If the first similarity value is greater than the preset first similarity threshold, determining the initial target video segment as the target video segment.
[0019] By adopting the above technical solution, by calculating the similarity between the last frame image of the initial target video segment obtained based on the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment and a preset second key frame image to obtain a first similarity value, when the first similarity value is not greater than the preset first similarity threshold, determining the initial target video segment as the target video segment, and by using the preset second key frame image and the last frame image of the initial target video for judgment, the accuracy of video segment extraction is improved.
[0020] In a preferred example of the present application, it can be further configured that: after calculating the similarity between the first frame image and a preset second key frame image to obtain a first similarity value, it further includes:
[0021] If the first similarity value is not greater than a preset similarity threshold, obtain the time corresponding to the frame image of the first similarity value;
[0022] Determine a first time interval according to the time corresponding to the frame image of the first similarity and a preset time;
[0023] Obtain all frame images of the first time interval;
[0024] Perform similarity calculation on all frame images of the first time interval and a preset second key frame image to obtain a plurality of second similarity values;
[0025] Obtain the time corresponding to the frame image of the maximum similarity value among all second similarity values;
[0026] Determine a target video segment according to the time corresponding to the frame image of the maximum similarity value and the duration of the target video segment.
[0027] By adopting the above technical solution, when the first similarity value is not greater than the preset similarity threshold, determine a first time interval according to the time corresponding to the frame image of the first similarity and a preset time, and perform similarity calculation on all frame images within the obtained first time interval and a preset second key frame image to obtain a plurality of similarity values, obtain the maximum similarity value among the plurality of similarity values, and determine a target video segment according to the time of the frame image corresponding to the maximum similarity value and the total duration of the target video. By determining the first time interval according to the time corresponding to the frame image of the first similarity and a preset time, and determining the last frame image of the target video segment according to the preset second key frame image and all frame images within the first interval, the accuracy of video segment extraction is improved.
[0028] In a preferred example of the present application, it can be further configured that: after determining the target video segment according to the time corresponding to the frame image where the similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment, it further includes:
[0029] Perform image cropping on a preset first key frame image to obtain a first key image;
[0030] Determine a plurality of frame images corresponding to the first key image according to the first key image and all frame images of the total video;
[0031] Judge whether each frame image corresponding to the first key image is a target image;
[0032] If so, obtain the time corresponding to each target image;
[0033] Determine each first target video segment according to the time corresponding to each target image and a preset time;
[0034] Sort all the first target video segments to obtain a second target video segment.
[0035] By adopting the above technical solution, a plurality of frame images corresponding to the first key image are determined from the first key image obtained by cropping the preset first key frame image and all the frame images of the total video, and it is determined whether each frame image corresponding to the first key image is a target image. If so, according to the time corresponding to each obtained target image and the preset time, each first target video segment is determined, and all the first target video segments are sorted to obtain a second target video segment. By determining the target image from the first key image obtained by cropping the preset first key frame image and all the frame images of the total video, and determining the second target video segment according to the target image, the efficiency of sorting the target video segment is improved.
[0036] In a preferred example of the present application, it can be further configured that: the determining whether each frame image corresponding to the first key image is a target image includes:
[0037] Obtain the time of each frame image corresponding to the first key image;
[0038] Determine a plurality of third video segments according to the time of each frame image corresponding to the first key frame and the preset first time;
[0039] Calculate each third video segment to obtain the score of each third video segment;
[0040] If there is a score greater than the preset score threshold among the scores of the plurality of third video segments, determine the frame images with the scores of the third video segments greater than the preset score threshold as target images.
[0041] By adopting the above technical solution, a plurality of third video segments are determined according to the time of each frame image corresponding to the first key frame obtained and the preset first time, and it is determined whether the score of each third video segment is greater than the preset score threshold. If so, determine the frame image as a target image. By determining the third video segment according to the preset time and determining whether the frame image is a target image according to the score of the third video segment, the accuracy of determining whether each frame image corresponding to the first key image is a target image is improved by judging all the video content within the preset time.
[0042] In a preferred example of the present application, it can be further configured that: the calculating each third video segment to obtain the score of each third video segment includes:
[0043] Input the third video segment into a pre-trained semantic model to obtain all user interest points in the third video segment;
[0044] Calculate the score of the third video clip based on all user interest points in the third video clip and their respective weights.
[0045] By adopting the above technical solution, by inputting the third video clip into a pre-trained semantic model and calculating according to all obtained user interest points and their respective weights, the score of the third video clip is obtained. By using the semantic model to determine all user interest points and determining the score of the third video clip according to all user interest points and their respective weights, the accuracy of calculating the score of each third video clip is improved.
[0046] In a preferred example of the present application, it can be further configured that: the training process of the semantic model includes:
[0047] Obtain a large number of sample videos, where the sample videos include videos and various user interest points marked in the videos;
[0048] Input the sample videos into the semantic model to be trained to obtain the semantic model annotation results;
[0049] Input the semantic model annotation results and each user interest point marked in the video into a preset loss function to obtain a loss value;
[0050] Iteratively train the semantic model to be trained according to the loss value and the sample videos until the loss value reaches a preset loss threshold, and determine the semantic model to be trained that reaches the preset loss threshold as the semantic model.
[0051] By adopting the above technical solution, the present application embodiment provides a training process of a semantic model. Compared with the related art, the present application embodiment uses the obtained sample videos and all marked interest points to iteratively train the semantic model until the accuracy of the semantic model meets the user requirements, improving the accuracy of predicting all user interest points of the third video clip.
[0052] In the second aspect, the present application provides a video clip extraction device, adopting the following technical solution:
[0053] A video clip extraction device includes
[0054] A first acquisition module: used to acquire all frame images of the total video;
[0055] A calculation module: used to calculate the similarity between each frame image in all frame images and a preset first key frame image in sequence, and obtain the similarity value corresponding to each frame image in sequence;
[0056] The second acquisition module: configured to, when the similarity values obtained in sequence are greater than a preset similarity threshold for the first time, acquire the time corresponding to the frame image for which the similarity value is greater than the preset similarity threshold for the first time;
[0057] The third acquisition module: configured to acquire the duration of the target video segment;
[0058] The first determination module: configured to determine the target video segment based on the time corresponding to the frame image for which the similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment.
[0059] By adopting the above technical solution, by acquiring all the frame images of the total video, calculating the similarity of each frame image in all the frame images of the total video with a preset first key frame image in sequence, obtaining the similarity value corresponding to each frame image in sequence, when the similarity values obtained in sequence are greater than the preset similarity threshold for the first time, determining the target video segment from the total video based on the time corresponding to the frame image for which the similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment acquired, by calculating the similarity of all the frame images of the acquired total video with the preset first key frame image at one time, and using the frame image for which the similarity value is greater than the preset similarity threshold for the first time as the initial image of the target image, it is not necessary to divide the total video content into multiple video segments and score each video segment, reducing a large amount of computational complexity and improving the efficiency of video segment extraction.
[0060] In a third aspect, the present application provides an electronic device, adopting the following technical solution:
[0061] At least one processor;
[0062] A memory;
[0063] At least one application program, where at least one application program is stored in the memory and is configured to be executed by at least one processor, and the at least one application program is configured to: execute the above video segment extraction method.
[0064] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution:
[0065] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed on a computer, the computer is made to execute the above video segment extraction method.
[0066] In summary, the present application includes at least one of the following beneficial technical effects:
[0067] By obtaining all the frame images of the total video, calculating the similarity between each frame image in all the frame images of the total video and a preset first key frame image in sequence, obtaining the similarity value corresponding to each frame image in sequence, when the similarity value obtained in sequence is greater than the preset similarity threshold for the first time, determining the target video segment from the total video based on the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the obtained target video segment. By calculating the similarity between all the frame images of the obtained total video and the preset first key frame image at one time, and using the frame image whose similarity value is greater than the preset similarity threshold for the first time as the initial image of the target image, it is not necessary to divide the total video content into multiple video segments and score each video segment, reducing a large amount of computational complexity and improving the efficiency of video segment extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 FIG. is a schematic flowchart of a video segment extraction method provided by an embodiment of the present application;
[0069] Figure 2 FIG. is a schematic structural diagram of a video segment extraction device provided by an embodiment of the present application;
[0070] Figure 3 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] The following further describes the present application in detail Figure 1 with reference to the attached Figure 3 drawings.
[0072] This specific embodiment is only an explanation of the present application and does not limit the present application. Those skilled in the art can make modifications without creative contributions to this embodiment after reading this specification, but as long as they are within the scope of the claims of the present application, they are protected by the Patent Law.
[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0074] In addition, the term "and / or" in this text is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.
[0075] The embodiments of the present application will be further described in detail below with reference to the accompanying drawings of the specification.
[0076] In the related art, for the extraction of video segments of a video, it is necessary to completely obtain the video, divide it into multiple video segments according to the content of the video, and score each video segment, and extract the video segments based on the scores of each video segment. However, this method calculates the scores of video segments based on all the frame images of the video segments, and the extraction efficiency is extremely low.
[0077] To solve the above technical problems, the present application provides a method, device, electronic device and medium for extracting video segments. By obtaining all the frame images of the total video, calculating the similarity between each frame image in all the frame images of the total video and a preset first key frame image in sequence, and obtaining the similarity value corresponding to each frame image in sequence. When the similarity value obtained in sequence is greater than the preset similarity threshold for the first time, determine the target video segment from the total video based on the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the obtained target video segment duration. By calculating the similarity between all the frame images of the obtained total video and the preset first key frame image at one time, and using the frame image whose similarity value is greater than the preset similarity threshold for the first time as the initial image of the target image, it is not necessary to divide the total video content into multiple video segments and score each video segment, reducing a large amount of calculation and improving the efficiency of video segment extraction.
[0078] The embodiments of the present application provide a method for extracting video segments, which is executed by an electronic device. The electronic device can be a server or a terminal device. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not limit this here.
[0079] Combined Figure 1 , Figure 1 is a schematic flowchart of a method for extracting video segments provided by an embodiment of the present application. As Figure 1As shown, the method includes step S101, step S102, step S103, step S104, and step S105, where:
[0080] Step S101: Obtain all frame images of the total video.
[0081] Among them, the total video is all videos. Video segment extraction is to extract some video segments required by users based on the total video. The video segments extracted from the total video can be one or several.
[0082] Among them, the method for obtaining all frame images of the total video can be frame extraction, and the total video is divided into frame images corresponding to each frame by using frame extraction.
[0083] Step S102: Calculate the similarity between each frame image in all frame images and a preset first key frame image in sequence, and obtain the similarity value corresponding to each frame image in sequence.
[0084] Among them, the preset first key frame image is pre-stored in the electronic device and represents the frame image at the beginning of the target video segment. After obtaining all frame images of the total video, calculate the similarity with the preset first key frame image in sequence according to the order corresponding to all frame images to obtain the similarity value. Among them, the embodiments of the present application do not limit the similarity calculation method, which can be any one of the Jaccard similarity coefficient, cosine similarity, Pearson correlation coefficient, and Manhattan distance.
[0085] Step S103: When the similarity value obtained in sequence is greater than the preset similarity threshold for the first time, obtain the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time.
[0086] Among them, the embodiments of the present application do not limit the preset similarity threshold, and the user can customize the setting according to experience. After obtaining the similarity value in sequence, judge the similarity value with the preset similarity threshold. When the similarity value is greater than the preset similarity threshold, use the image corresponding to the first time greater than the preset similarity threshold as the start frame image of the target video segment, and obtain the time of the frame image corresponding to the similarity value greater than the preset similarity threshold for the first time in the total video.
[0087] Step S104: Obtain the duration of the target video segment.
[0088] The duration of the target video segment is pre-stored in the electronic device. After obtaining the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time, the electronic device automatically obtains the duration of the target video segment.
[0089] Step S105: Determine the target video segment according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment.
[0090] After determining the time corresponding to the frame image where the similarity for the first time is greater than the preset similarity threshold, that is, the time of the start frame image of the target video in the total video, and the duration of the target video segment, the video segment can be directly intercepted in the total video starting from the time of the start frame image in the total video until the intercepted duration is equal to the duration of the target video segment, and the intercepted video segment is used as the target video segment.
[0091] In the embodiment of the present application, by obtaining all frame images of the total video, calculating the similarity of each frame image in all frame images of the total video with a preset first key frame image in sequence, obtaining the similarity value corresponding to each frame image in sequence, when the similarity value obtained in sequence is greater than the preset similarity threshold for the first time, determining the target video segment from the total video according to the time corresponding to the frame image where the similarity value is greater than the preset similarity threshold for the first time and the duration of the obtained target video segment. By calculating the similarity of all frame images of the obtained total video with the preset first key frame image once, and using the frame image where the similarity value is greater than the preset similarity threshold for the first time as the initial image of the target image, it is not necessary to divide the total video content into multiple video segments and score each video segment, reducing a large amount of calculation and improving the efficiency of video segment extraction.
[0092] A possible implementation manner of the embodiment of the present application for determining the target video segment according to the time corresponding to the frame image where the similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment includes:
[0093] Obtaining an initial target video segment according to the time corresponding to the frame image where the similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment;
[0094] Obtaining the first frame image of the initial target video segment, where the first frame image is the last frame image of the initial target video segment;
[0095] Calculating the similarity between the first frame image and a preset second key frame image to obtain a first similarity value;
[0096] If the first similarity value is greater than the preset first similarity threshold, determining the initial target video segment as the target video segment.
[0097] Among them, the first frame image of the target video segment is pre - saved in the electronic device. The first frame image is the last frame image of the initial target video segment. The frame image whose similarity is greater than the preset similarity threshold for the first time is the start frame image of the target video segment. When the similarity value between the last frame image of the target video segment and the preset second key - frame image is greater than the preset first similarity threshold, it is determined that the last frame image of the initial target video segment is the last frame image of the target video segment. When the similarity value between the last frame image of the initial target video and the preset second key - frame image is not greater than the preset first similarity threshold, it is determined that the last frame image of the initial target video is not the last frame image of the target video segment. In the embodiments of the present application, the preset first similarity threshold is not limited, and the user can customize the setting according to actual needs.
[0098] In the embodiments of the present application, by calculating the similarity between the last frame image of the initial target video segment obtained according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the target video segment duration and the preset second key - frame image, a first similarity value is obtained. When the first similarity value is not greater than the preset first similarity threshold, it is determined that the initial target video segment is the target video segment. By using the preset second key - frame image and the last frame image of the initial target video for judgment, the accuracy of video segment extraction is improved.
[0099] A possible implementation manner of the embodiments of the present application, after calculating the similarity between the first frame image and the preset second key - frame image to obtain the first similarity value, further includes:
[0100] If the first similarity value is not greater than the preset similarity threshold, obtain the time corresponding to the frame image of the first similarity value;
[0101] Determine the first time interval according to the time corresponding to the frame image of the first similarity and the preset time;
[0102] Obtain all frame images of the first time interval;
[0103] Calculate the similarity between all frame images of the first time interval and the preset second key - frame image to obtain a plurality of second similarity values;
[0104] Obtain the time corresponding to the frame image of the maximum similarity value among all second similarity values;
[0105] Determine the target video segment according to the time corresponding to the frame image of the maximum similarity value and the target video segment duration.
[0106] Among them, when the first similarity value is not greater than the preset similarity threshold, it is determined that the last frame image of the initial target video is not the last frame image of the target video segment. All frame images within the first time interval of the last frame image of the initial target video segment can be obtained and the similarity is calculated with the preset second key frame image to obtain the second similarity value. The embodiments of the present application do not limit the preset time, and the user can customize the setting according to the actual situation.
[0107] Among them, if the second similarity value is greater than the preset first similarity threshold, it is determined that the frame image with the second similarity value greater than the preset similarity threshold is the last frame image of the target video segment. However, due to the influence of the frame rate, there are multiple consecutive frame images within 1S, and the similarity of each frame image is greater than the preset first similarity threshold. Therefore, the image corresponding to the maximum similarity in the second similarity threshold can be obtained, and the frame image with the maximum similarity is used as the last frame image of the target video segment to determine the target video segment. Among them, the embodiments of the present application do not limit the preset second similarity threshold, and the user can customize the setting according to actual needs.
[0108] In the embodiments of the present application, when the first similarity value is not greater than the preset similarity threshold, the first time interval is determined according to the time corresponding to the frame image of the first similarity and the preset time, and the similarity is calculated with the preset second key frame image for all frame images within the obtained first time interval to obtain multiple similarity values. The maximum similarity value among the multiple similarity values is obtained, and the target video segment is determined according to the time of the frame image corresponding to the maximum similarity value and the total duration of the target video. By determining the first time interval according to the time corresponding to the frame image of the first similarity and the preset time, and determining the last frame image of the target video segment according to the preset second key frame image and all frame images within the first interval, the accuracy of video segment extraction is improved.
[0109] A possible implementation manner of the embodiments of the present application further includes, after determining the target video segment according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment:
[0110] Performing image cropping on the preset first key frame image to obtain the first key image;
[0111] Determining multiple frame images corresponding to the first key image according to the first key image and all frame images of the total video;
[0112] Judging whether each frame image corresponding to the first key image is a target image;
[0113] If so, obtaining the time corresponding to each target image;
[0114] Determine each first target video segment according to the time corresponding to each target image and the preset time;
[0115] Sort out all the first target video segments to obtain the second target video segment.
[0116] Wherein, when the target video segment is a combination of multiple video segments, image cropping can be performed according to the preset first key-frame image to obtain the first key image, where the first key image is a representative image in the preset first key-frame image, for example: a person image, a plant image, an animal image, or other representative images.
[0117] After determining the first key image, determine whether there is an image identical to the first key image among all the frame images of the total video. If so, determine whether all the frame images corresponding to the first key image are target images, where the target image is the image required for the target video segment. When the frame images corresponding to the first key image are target images, determine each first target video segment according to the time corresponding to each target image and the preset time, and sort out and combine all the first target video segments to determine the target video segment.
[0118] In the embodiment of the present application, multiple frame images corresponding to the first key image are determined by the first key image obtained by image cropping of the preset first key-frame image and all the frame images of the total video, and it is determined whether each frame image corresponding to the first key image is a target image. If so, according to the time corresponding to each obtained target image and the preset time, each first target video segment is determined, and all the first target video segments are sorted out to obtain the second target video segment. By determining the target image by the first key image obtained by image cropping of the preset first key-frame image and all the frame images of the total video, and determining the second target video segment according to the target image, the efficiency of sorting out the target video segment is improved.
[0119] A possible implementation manner of the embodiment of the present application for determining whether each frame image corresponding to the first key image is a target image includes:
[0120] Obtain the time of each frame image corresponding to the first key image;
[0121] Determine multiple third video segments according to the time of each frame image corresponding to the first key frame and the preset first time;
[0122] Calculate each third video segment to obtain the score of each third video segment;
[0123] If there is a score greater than the preset score threshold among the scores of multiple third video segments, determine the frame images with the scores of the third video segments greater than the preset score threshold as target images.
[0124] Among them, after determining the frame image corresponding to the first key image in the total video, the time of the frame image corresponding to the first key image is obtained. In the embodiments of the present application, the method for calculating the score based on the third video segment is not limited. It can be to calculate the score of each third video segment using big data, or each third video segment can be input into a pre-trained semantic model to label all the points of interest, and the score of each video segment is determined according to the different categories of the points of interest and the weight corresponding to each category. Among them, the embodiments of the present application do not limit the preset score threshold, and the user can customize the setting according to the actual situation.
[0125] In the embodiments of the present application, multiple third video segments are determined by comparing the time of each frame image corresponding to the first key frame obtained with a preset first time, and it is determined whether the score of each third video segment is greater than the preset score threshold. If it is greater, the frame image is determined as the target image. By determining the third video segment according to the preset time and determining whether the frame image is the target image according to the score of the third video segment, and by judging all the video content within the preset time, the accuracy of determining whether each frame image corresponding to the first key image is the target image is improved.
[0126] A possible implementation manner of the embodiments of the present application for calculating the score of each third video segment includes:
[0127] Input the third video segment into a pre-trained semantic model to obtain all the user points of interest in the third video segment;
[0128] Calculate according to all the user points of interest in the third video segment and their respective corresponding weights to obtain the score of the third video segment.
[0129] Among them, a pre-trained semantic model is pre-stored in the electronic device. After determining the third video segment, after inputting the third video segment into the pre-trained semantic model, all the user points of interest in the third video segment are obtained. The embodiments of the present application do not limit the structure of the semantic model, and the user can customize the selection according to the actual situation.
[0130] Among them, the weights corresponding to each point of interest are pre-stored in the electronic device. For example, the points of interest can be human expressions, human actions, human ornaments, etc. After determining all the user points of interest in the third video segment, determine the category of each user point of interest, and determine the score of the third video segment according to the category of each user point of interest and their respective corresponding weights.
[0131] In an embodiment of the present application, by inputting the third video segment into a pre-trained semantic model and calculating based on all the obtained user interest points and their respective corresponding weights, the score of the third video segment is obtained. By using the semantic model to determine all the user interest points and determining the score of the third video segment according to all the user interest points and their respective corresponding weights, the accuracy of calculating the score of each third video segment is improved.
[0132] A possible implementation manner of the embodiment of the present application, the training process of the semantic model includes:
[0133] Obtain a large number of sample videos, where the sample videos include videos and various user interest points annotated in the videos;
[0134] Input the sample videos into the semantic model to be trained, and obtain the annotation results of the semantic model;
[0135] Input the annotation results of the semantic model and the various user interest points annotated in the videos into a preset loss function to obtain a loss value;
[0136] Iteratively train the semantic model to be trained according to the loss value and the sample videos until the loss value reaches a preset loss threshold, and determine the semantic model to be trained that reaches the preset loss threshold as the semantic model.
[0137] Specifically, obtain sample videos through a web crawler. The sample videos include videos and all user interest points annotated in the videos, where the annotation method can be manual annotation.
[0138] In an embodiment of the present application, input the sample videos into the semantic model to be trained to obtain the user interest points of the predicted sample videos. According to all the user interest points annotated in the sample videos and the user interest points of the predicted sample videos, use a preset loss function to determine the loss value. The smaller the loss value, the smaller the difference between the predicted value and the standard value, and the more accurate the detection result. The embodiment of the present application does not limit the preset loss function, and the user can select according to actual needs
[0139] Specifically, the embodiment of the present application provides a training process of a semantic model. Compared with the related technology, the embodiment of the present application uses the obtained sample videos and all the annotated interest points to iteratively train the semantic model until the accuracy of the semantic model meets the user requirements, improving the accuracy of predicting all user interest points of the third video segment.
[0140] The above embodiment introduces a video segment extraction method from the perspective of the method flow. The following embodiment introduces a video segment extraction device from the perspective of virtual modules or virtual units. For details, see the following embodiment.
[0141] The embodiments of the present application provide a video clip extraction device 200, as Figure 2 shown, Figure 2 which is a schematic structural diagram of a video clip extraction device provided by the embodiments of the present application. The video clip extraction device 200 may specifically include:
[0142] A first acquisition module 201: configured to acquire all frame images of the total video;
[0143] A calculation module 202: configured to calculate the similarity between each frame image in all the frame images and a preset first key frame image in sequence, and obtain the similarity value corresponding to each frame image in sequence;
[0144] A second acquisition module 203: configured to, when the similarity values obtained in sequence are greater than a preset similarity threshold for the first time, acquire the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time;
[0145] A third acquisition module 204: configured to acquire the duration of the target video clip;
[0146] A first determination module 205: configured to determine the target video clip according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video clip.
[0147] For the embodiments of the present application, by acquiring all frame images of the total video, calculating the similarity between each frame image in all the frame images of the total video and a preset first key frame image in sequence, and obtaining the similarity value corresponding to each frame image in sequence, when the similarity values obtained in sequence are greater than a preset similarity threshold for the first time, determining the target video clip from the total video according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the acquired target video clip, by calculating the similarity between all the frame images of the acquired total video and a preset first key frame image at one time, and using the frame image whose similarity value is greater than the preset similarity threshold for the first time as the initial image of the target image, it is not necessary to divide the total video content into multiple video clips and score each video clip, reducing a large amount of calculation and improving the efficiency of video clip extraction.
[0148] In a possible implementation manner of the embodiments of the present application, when the first determination module 205 executes to determine the target video clip according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video clip, it is specifically configured to:
[0149] Obtain an initial target video clip according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video clip;
[0150] Obtain the first frame image of the initial target video segment, where the first frame image is the last frame image of the initial target video segment;
[0151] Calculate the similarity between the first frame image and the preset second key frame image to obtain the first similarity value;
[0152] If the first similarity value is greater than the preset first similarity threshold, determine the initial target video segment as the target video segment.
[0153] In a possible implementation manner of the embodiments of the present application, the video segment extraction device 200 further includes:
[0154] A second determination module: configured to, if the first similarity value is not greater than the preset similarity threshold, obtain the time corresponding to the frame image of the first similarity value;
[0155] Determine the first time interval according to the time corresponding to the frame image of the first similarity and the preset time;
[0156] Obtain all frame images of the first time interval;
[0157] Calculate the similarity between all frame images of the first time interval and the preset second key frame image to obtain multiple second similarity values;
[0158] Obtain the time corresponding to the frame image of the maximum similarity value among all the second similarity values;
[0159] Determine the target video segment according to the time corresponding to the frame image of the maximum similarity value and the target video segment duration.
[0160] In a possible implementation manner of the embodiments of the present application, the video segment extraction device 200 further includes:
[0161] A judgment module: configured to crop the preset first key frame image to obtain the first key image;
[0162] Determine multiple frame images corresponding to the first key image according to the first key image and all frame images of the total video;
[0163] Judge whether each frame image corresponding to the first key image is a target image;
[0164] If so, obtain the time corresponding to each target image;
[0165] Determine each first target video segment according to the time corresponding to each target image and the preset time;
[0166] Sort all the first target video segments to obtain the second target video segment.
[0167] In a possible implementation manner of the embodiment of the present application, when the judgment module executes the judgment on whether each frame image corresponding to the first key image is a target image, it specifically is used for:
[0168] Obtain the time of each frame image corresponding to the first key image;
[0169] Determine a plurality of third video segments according to the time of each frame image corresponding to the first key frame and a preset first time;
[0170] Perform calculations on each third video segment to obtain the score of each third video segment;
[0171] If there is a score greater than a preset score threshold among the scores of the plurality of third video segments, determine the frame images with the scores of the third video segments greater than the preset score threshold as target images.
[0172] In a possible implementation manner of the embodiment of the present application, when the judgment module executes the calculation on each third video segment to obtain the score of each third video segment, it specifically is used for:
[0173] Input the third video segment into a pre-trained semantic model to obtain all user interest points in the third video segment;
[0174] Perform calculations according to all user interest points in the third video segment and their respective corresponding weights to obtain the score of the third video segment.
[0175] In a possible implementation manner of the embodiment of the present application, the training process of the semantic model includes:
[0176] Obtain a large number of sample videos, where the sample videos include videos and various user interest points annotated in the videos;
[0177] Input the sample videos into the semantic model to be trained to obtain the semantic model annotation results;
[0178] Input the semantic model annotation results and various user interest points annotated in the videos into a preset loss function to obtain a loss value;
[0179] Perform iterative training on the semantic model to be trained according to the loss value and the sample videos until the loss value reaches a preset loss threshold, and determine the semantic model to be trained that reaches the preset loss threshold as the semantic model.
[0180] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of a video segment extraction device 200 described above can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated herein.
[0181] In the embodiment of the present application, an electronic device is provided, such asFigure 3 As shown Figure 3 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Figure 3 The electronic device 300 shown includes: a processor 301 and a memory 303. Among them, the processor 301 and the memory 303 are connected, such as connected through a bus 302. Optionally, the electronic device 300 may further include a transceiver 304. It should be noted that in practical applications, the transceiver 304 is not limited to one, and the structure of the electronic device 300 does not constitute a limitation to the embodiments of the present application.
[0182] The processor 301 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 301 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0183] The bus 302 may include a path for transmitting information between the above components. The bus 302 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 302 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0184] The memory 303 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0185] The memory 303 is used to store the application program code for executing the solution of this application, and is controlled by the processor 301 to execute. The processor 301 is used to execute the application program code stored in the memory 303 to implement the content shown in the foregoing method embodiments.
[0186] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. It can also be a server, etc. Figure 3 The shown electronic device is only an example and should not impose any restrictions on the functions and usage scope of the embodiments of this application.
[0187] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, the computer can execute the corresponding content in the foregoing method embodiment. Compared with the related art, in the embodiment of the present application, by obtaining all frame images of the total video, calculating the similarity between each frame image in all frame images of the total video and a preset first key frame image in sequence, obtaining the similarity value corresponding to each frame image in sequence, when the similarity value obtained in sequence is greater than the preset similarity threshold for the first time, determining the target video segment from the total video according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the obtained target video segment, by calculating the similarity between all frame images of the obtained total video and the preset first key frame image at one time, and using the frame image whose similarity value is greater than the preset similarity threshold for the first time as the initial image of the target image, it is not necessary to divide the total video content into multiple video segments and score each video segment, reducing a large amount of calculation and improving the efficiency of video segment extraction.
[0188] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed at the same time, but can be executed at different times, and their execution order does not necessarily need to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0189] The above is only a partial implementation manner of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and retouches can be made, and these improvements and retouches should also be regarded as the protection scope of the present application.
Claims
1. A method for extracting video segments, characterized in that, Including: Obtain all frame images of the total video; Calculate the similarity of each frame image in all frame images with a preset first key frame image in sequence, and obtain the similarity value corresponding to each frame image in sequence; When the similarity values obtained in sequence are greater than the preset similarity threshold for the first time, obtain the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time; Obtain the duration of the target video segment; Determine the target video segment according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment; The determining the target video segment according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment includes: Obtain the initial target video segment according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment; Obtain the first frame image of the initial target video segment, and the first frame image is the last frame image of the initial target video segment; Calculate the similarity of the first frame image with a preset second key frame image to obtain the first similarity value; If the first similarity value is greater than the preset first similarity threshold, determine the initial target video segment as the target video segment; After calculating the similarity of the first frame image with a preset second key frame image to obtain the first similarity value, it further includes: If the first similarity value is not greater than the preset similarity threshold, obtain the time corresponding to the frame image of the first similarity value; Determine the first time interval according to the time corresponding to the frame image of the first similarity and the preset time; Obtain all frame images of the first time interval; Calculate the similarity of all frame images of the first time interval with a preset second key frame image to obtain multiple second similarity values; Obtain the time corresponding to the frame image with the maximum similarity value among all second similarity values; Determine the target video segment according to the time corresponding to the frame image with the maximum similarity value and the duration of the target video segment.
2. The video clip extraction method according to claim 1, characterized in that, After determining the target video segment according to the time corresponding to the frame image whose similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment, it further includes: Perform image cropping on the preset first key frame image to obtain the first key image; Determine multiple frame images corresponding to the first key image according to the first key image and all frame images of the total video; Judge whether each frame image corresponding to the first key image is a target image; If so, obtain the time corresponding to each target image; Determine each first target video segment according to the time corresponding to each target image and the preset time; Organize all first target video segments to obtain the second target video segment.
3. The video clip extraction method according to claim 2, wherein The judging whether each frame image corresponding to the first key image is a target image includes: Obtain the time of each frame image corresponding to the first key image; Determine multiple third video segments according to the time of each frame image corresponding to the first key frame and the preset first time; Calculate each third video segment to obtain the score of each third video segment; If there is a score greater than a preset score threshold among the scores of multiple third video segments, the frame images with scores of the third video segments greater than the preset score threshold are determined as target images.
4. The video clip extraction method according to claim 3, wherein The calculating the score of each third video segment includes: Inputting the third video segment into a pre-trained semantic model to obtain all user interest points in the third video segment; Calculating according to all user interest points in the third video segment and their respective corresponding weights to obtain the score of the third video segment.
5. The video clip extraction method according to claim 4, characterized in that, The training process of the semantic model includes: Obtaining a large number of sample videos, where the sample videos include videos and various user interest points annotated in the videos; Inputting the sample videos into the semantic model to be trained to obtain the annotation results of the semantic model; Inputting the annotation results of the semantic model and various user interest points annotated in the videos into a preset loss function to obtain a loss value; Iteratively training the semantic model to be trained according to the loss value and the sample videos until the loss value reaches a preset loss threshold, and determining the semantic model to be trained that reaches the preset loss threshold as the semantic model.
6. A video clip extraction device, characterized in that, including: A first obtaining module: used to obtain all frame images of the total video; A calculating module: used to calculate the similarity between each frame image in all frame images and a preset first key frame image in turn, and obtain the similarity value corresponding to each frame image in turn; A second obtaining module: used to obtain the time corresponding to the frame image when the similarity value obtained in turn is greater than the preset similarity threshold for the first time; A third obtaining module: used to obtain the duration of the target video segment; A first determining module: used to determine the target video segment according to the time corresponding to the frame image when the similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment; When executing the determination of the target video segment according to the time corresponding to the frame image when the similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment, specifically used for: Obtaining an initial target video segment according to the time corresponding to the frame image when the similarity value is greater than the preset similarity threshold for the first time and the duration of the target video segment; Obtaining the first frame image of the initial target video segment, where the first frame image is the last frame image of the initial target video segment; Calculating the similarity between the first frame image and a preset second key frame image to obtain a first similarity value; If the first similarity value is greater than the preset first similarity threshold, determining the initial target video segment as the target video segment; A second determining module: used to, if the first similarity value is not greater than the preset similarity threshold, obtain the time corresponding to the frame image of the first similarity value; Determining a first time interval according to the time corresponding to the frame image of the first similarity and a preset time; Obtaining all frame images of the first time interval; Calculating the similarity between all frame images of the first time interval and a preset second key frame image to obtain a plurality of second similarity values; Obtaining the time corresponding to the frame image with the maximum similarity value among all second similarity values; Determining the target video segment according to the time corresponding to the frame image with the maximum similarity value and the duration of the target video segment.
7. An electronic device, characterized in that, including: At least one processor; A memory; At least one application program, wherein the at least one application program is stored in the memory and is configured to be executed by the at least one processor, and the at least one application program is configured to: execute the video segment extraction method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed in a computer, the computer is caused to execute the video segment extraction method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Movie editing method and device
CN108566567A
Target video clip extraction method and device
CN111787356A