Video transition point determination method and device, equipment and medium
By clustering and sorting video frames and determining the video transition points, the problem of being unable to split videos in the prior art is solved, and more efficient video processing is achieved.
Patent Information
- Application Number
- CN202510409003.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
Smart Images

Figure CN120264041A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of multimedia processing, and particularly to a method, apparatus, device, and medium for determining video transition points. Background Art
[0002] With the development of multimedia technology, many video processing tasks such as plot understanding and video description require splitting the video. Therefore, it is necessary to determine the video transition points to effectively split the video so that the video processing tasks can achieve better results.
[0003] In related technologies, the video is split by algorithms for scene transition points and shot transition points. However, the above algorithms cannot effectively split a single-shot video because there is no pixel mutation between the frame-level pictures in a single-shot video (the change in the pictures in the video is gradual and smooth). That is to say, the existing method is to identify video frames with pixel mutations such as scene switching or shot switching in the video as video transition points. However, this method cannot determine the video transition points for a single-shot video where there is no pixel mutation between the frame-level pictures, resulting in the inability to accurately split the single-shot video. Summary of the Invention
[0004] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, apparatus, device, and medium for determining video transition points.
[0005] An embodiment of the present disclosure provides a method for determining video transition points. The method includes: obtaining a video to be processed, and performing frame extraction on the video to be processed at a preset frame extraction frequency to obtain a set of video frames to be processed; performing clustering processing on the set of video frames to be processed based on a preset clustering algorithm to obtain the clustering category corresponding to each video frame to be processed; obtaining the time position of each video frame to be processed in the video to be processed, and sorting each video frame to be processed based on the time position to obtain a video frame sorting result; determining that two adjacent video frames to be processed correspond to different clustering categories based on the video frame sorting result and the clustering category corresponding to each video frame to be processed, and determining the target video frame with the earlier time position sorting among the two adjacent video frames to be processed as the video transition point.
[0006] Optionally, the video to be processed is split based on the video transition point.
[0007] Optionally, the clustering process of the set of video frames to be processed based on a preset clustering algorithm to obtain the clustering category corresponding to each video frame to be processed includes: processing each video frame to be processed in the set of video frames to be processed through a pre-trained image feature extraction model to obtain the image features of each video frame to be processed; calculating the similarity scores between all the video frames to be processed based on the image features of each video frame to be processed; taking at least two video frames to be processed with similarity scores greater than a preset similarity threshold as the same clustering category to obtain the clustering category corresponding to each video frame to be processed.
[0008] Optionally, the calculating the similarity scores between all the video frames to be processed based on the image features of each video frame to be processed includes: obtaining the target features corresponding to each target based on the image features of each video frame to be processed; calculating the similarity based on each target feature and the weight coefficient corresponding to each target to obtain the similarity scores between all the video frames to be processed.
[0009] Optionally, the method further includes: obtaining video frame samples, where each video frame sample has a corresponding feature label; processing the video frame samples through a pre-constructed image feature extraction model to be trained to obtain video frame features, and determining a loss value based on the similarity between the video frame features and the feature labels and a preset loss function to adjust the model parameters to obtain the image feature extraction model.
[0010] The method further includes: obtaining the video frame time points between adjacent target video frames, and determining the time difference between adjacent target video frames based on the video frame time points; when the time difference is greater than a preset time threshold, determining one of the adjacent target video frames as the video transition point.
[0011] Optionally, the method further includes: obtaining the frame extraction time based on the frame extraction frequency; determining the time threshold based on the frame extraction time.
[0012] An embodiment of the present disclosure also provides a video transition point determination device, which includes: an acquisition module, configured to acquire a video to be processed, and perform frame extraction on the video to be processed at a preset frame extraction frequency to obtain a set of video frames to be processed; a clustering module, configured to perform clustering processing on the set of video frames to be processed based on a preset clustering algorithm to obtain a clustering category corresponding to each video frame to be processed; a sorting module, configured to obtain the time position of each video frame to be processed in the video to be processed, and sort each video frame to be processed based on the time position to obtain a video frame sorting result; a processing module, configured to determine that two adjacent video frames to be processed correspond to different clustering categories based on the video frame sorting result and the clustering category corresponding to each video frame to be processed, and determine the target video frame with the earlier sorted time position among the two adjacent video frames to be processed as the video transition point.
[0013] An embodiment of the present disclosure also provides an electronic device, which includes: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the video transition point determination method provided by the embodiment of the present disclosure.
[0014] An embodiment of the present disclosure also provides a computer-readable storage medium, which stores a computer program, and the computer program is used to execute the video transition point determination method provided by the embodiment of the present disclosure.
[0015] An embodiment of the present disclosure also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the video transition point determination method provided by the embodiment of the present application.
[0016] The above technical solution provided by the embodiment of the present disclosure clusters all video frames corresponding to the video to obtain the clustering category of each video frame, and sorts each video frame according to the time position in the video, so that the sorting result of the clustering category of each video frame can be obtained. Therefore, according to the video frame sorting result, the target video frame with a changed clustering category can be determined as the video transition point, so that the video cut point can be accurately determined to split the video to be processed, improving the video processing effect and efficiency.
[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0018] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0019] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic flowchart of a method for determining video transition points provided by an embodiment of the present disclosure;
[0021] Figure 2 It is a schematic flowchart of a method for determining video transition points provided by an embodiment of the present disclosure;
[0022] Figure 3 It is a schematic structural diagram of a device for determining video transition points provided by an embodiment of the present disclosure;
[0023] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0024] In order to be able to more clearly understand the above objects, features, and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0025] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.
[0026] In existing multimedia application scenarios, such as during program recording, multiple shooting devices are used for recording. Each shooting device takes a long - time shot of a scene within a certain range through a fixed focal length to obtain a video, so that multiple single - take videos can be acquired. It can be understood that the frame - level picture change of a single - take video is gradual and smooth, that is, there is no pixel mutation. If the existing scene transition point and shot transition point algorithms are used to segment the video, it will lead to ineffective segmentation of the single - take video, resulting in less than ideal effects for subsequent tasks such as plot understanding and video description.
[0027] In view of the above problems, embodiments of the present disclosure propose a method for determining video transition points. By clustering all video frames corresponding to a video, the clustering categories of each video frame are obtained, and each video frame is sorted according to the time points in the video, so that the sorting result of the clustering categories of each video frame can be obtained. Therefore, according to the sorting result, the video frames with changed clustering categories can be determined as target video frames, so that the video cut points can be accurately determined to segment the video to be processed, improving the video processing effect and efficiency. The following is a detailed explanation:
[0028] Figure 1 FIG. is a flowchart of a method for determining video transition points provided by an embodiment of the present disclosure. This method can be applied to an electronic device, such as a computer, a mobile phone, a tablet computer, a television, a server, etc., which is not limited herein. As Figure 1 shown, the method mainly includes the following steps S102 to S108:
[0029] Step S102, obtain the video to be processed, and perform frame extraction on the video to be processed according to a preset frame extraction frequency to obtain a set of video frames to be processed.
[0030] Among them, the video to be processed refers to the video that needs to be segmented into video segments; in the embodiments of the present disclosure, the video to be processed is usually a video with a gradual and smooth change in the frame-level picture, that is, a video without pixel mutations.
[0031] In the embodiments of the present disclosure, there are many ways to obtain the video to be processed. As an example, obtain the video obtained by shooting a fixed range with a shooting device with a fixed focal length as the video to be processed; as another example, receive candidate videos uploaded by other platforms or terminals, obtain the pixel differences between the video frames in the candidate videos, and use the candidate videos with pixel differences less than a preset difference threshold as the video to be processed; the above are only examples of the ways to obtain the video to be processed, and the embodiments of the present disclosure do not make specific limitations.
[0032] Among them, the frame extraction frequency is set in advance, such as extracting one frame every two seconds, etc. The frame extraction frequency can be adjusted according to factors such as processing efficiency and video frame picture differences, such as magnifying the degree of video picture change by reducing the frame extraction frequency, etc., to further meet the flexibility of video transition point determination.
[0033] In the embodiments of the present disclosure, performing frame extraction on the video to be processed according to a preset frame extraction frequency to obtain a set of video frames to be processed can be understood as obtaining each video frame to be processed from the video to be processed according to the frame extraction frequency, so as to obtain a set of video frames to be processed; among them, the number of videos to be processed in the set of video frames to be processed obtained by different frame extraction frequencies is different, and it is specifically selected and set according to the actual application scenario.
[0034] Specifically, during the determination of the video transition point, the video to be processed is obtained, and the video to be processed is frame-extracted according to the frame extraction frequency, and a plurality of video frames to be processed are obtained as a set of video frames to be processed.
[0035] Step S104: Perform clustering processing on the set of video frames to be processed based on a preset clustering algorithm, and obtain the clustering category corresponding to each video frame to be processed.
[0036] Among them, the clustering category refers to at least one category obtained by performing clustering processing on the set of video frames to be processed. In the embodiments of the present disclosure, clustering algorithms such as SMLR (A Simple Image Clustering Script using CLIP and Hierarchical Clustering, a simple image clustering script using CLIP (Contrastive Language-Image Pre-training) and hierarchical clustering), density clustering, etc. are preset, and specific settings are selected according to actual applications.
[0037] In the embodiments of the present disclosure, there are many ways to perform clustering processing on the set of video frames to be processed based on a preset clustering algorithm and obtain the clustering category corresponding to each video frame to be processed. As an example, in the first step, one or more video frames to be processed are selected from the set of video frames to be processed as the initial center points. In the second step, for each video frame to be processed in the set of video frames to be processed, calculate its similarity score with each center point, and assign the video frame to be processed to the clustering category represented by the center point with the highest similarity score. In the third step, for each clustering category, recalculate its center point, and repeat the second step and the third step until the center point no longer changes or reaches the preset number of iterations, so as to obtain the clustering category corresponding to each video frame to be processed.
[0038] As another example, each video frame to be processed in the set of video frames to be processed is processed by a pre-trained image feature extraction model to obtain the image feature of each video frame to be processed. Calculate the similarity scores between all video frames to be processed based on the image features of each video frame to be processed, and use at least two video frames to be processed with similarity scores greater than the preset similarity threshold as the same clustering category, so as to obtain the clustering category corresponding to each video frame to be processed. The above are only two ways to perform clustering processing on the set of video frames to be processed based on a preset clustering algorithm and obtain the clustering category corresponding to each video frame to be processed. The embodiments of the present disclosure do not specifically limit the way to perform clustering processing on the set of video frames to be processed based on a preset clustering algorithm and obtain the clustering category corresponding to each video frame to be processed.
[0039] Specifically, a set of video frames to be processed corresponds to at least one clustering category. For example, all the video frames to be processed belong to the same clustering category, or there are multiple clustering categories through clustering processing, such as three clustering categories. Thus, a part of the video frames to be processed belongs to the first clustering category, a part of the video frames to be processed belongs to the second clustering category, and a part of the video frames to be processed belongs to the third clustering category, etc.
[0040] Step S106: Obtain the time position of each video frame to be processed in the video to be processed, and sort each video frame to be processed based on the time position to obtain a video frame sorting result.
[0041] Step S108: Based on the video frame sorting result and the clustering category corresponding to each video frame to be processed, determine that two adjacent video frames to be processed correspond to different clustering categories, and determine the target video frame with the earlier sorted time position among the two adjacent video frames to be processed as the video transition point.
[0042] In the embodiments of the present disclosure, the time position refers to the playback time point of the video frame to be processed in the video to be processed; it can be understood that playing the set of video frames to be processed in chronological order can realize the playback of the video to be processed. Therefore, each video frame to be processed can be sorted based on the time position of each video frame to be processed in the video to be processed to obtain a video frame sorting result, that is, sorting in the playback order of each video frame to be processed in the video to be processed. The video frame sorting result refers to the sorting result of the playback order of each video frame to be processed.
[0043] In the embodiments of the present disclosure, the target video frame refers to a video frame where there is a clustering category switch relative to the next video frame of this video frame; the video transition point can be used as a video cut position to split the video to be processed, and the video transition point refers to the time position corresponding to the target video frame in the video to be processed.
[0044] In the embodiments of the present disclosure, obtaining the target video frame based on the video frame sorting result and the clustering category corresponding to each video frame to be processed can be understood as sorting the clustering categories corresponding to each video frame to be processed according to the time position corresponding to each video frame to be processed, so as to determine the video frames where the clustering category changes, and then using the video frame where there is a clustering category switch relative to the next video frame of this video frame as the target video frame, and determining the target video frame as the video transition point.
[0045] Specifically, each video frame to be processed is sorted based on the time point in the video to be processed, and the sorting result of each video frame to be processed is obtained. That is to say, the sorting result of the clustering category corresponding to each video frame to be processed can be obtained, so that it can be determined that two adjacent video frames to be processed correspond to different clustering categories. The video frame to be processed with a time point sorted earlier among the two adjacent video frames to be processed is used as the target video frame. For example, if the sorting result of the clustering category by time point is AAAABBBB, the fourth video frame to be processed is used as the target video frame.
[0046] In some embodiments, the video to be processed is segmented based on the video transition point. Specifically, the time point corresponding to the target video frame in the video to be processed is used as the video segmentation point to segment the video to be processed, and multiple video segments are obtained for subsequent analysis and processing, further meeting the requirements for determining the video transition point.
[0047] For example, the set of videos to be processed includes five video frames to be processed. The five videos to be processed are sorted according to the time points of the video frames to be processed in the video to be processed. For example, the video frame to be processed X1 (category Y1), the video frame to be processed X2 (category Y1), the video frame to be processed X3 (category Y1), the video frame to be processed X4 (category Y2), and the video frame to be processed X5 (category Y2). Thus, the sorting result of the clustering category of each video frame can be obtained, and it can be seen which video frames to be processed have a change in the clustering category. Thus, it can be determined that the video frame to be processed X3 is the target video frame, and the time point corresponding to the video frame to be processed X3 is used as the video transition point to segment the video to be processed.
[0048] In summary, the video transition point determination solution of the embodiments of the present disclosure includes obtaining a video to be processed, and performing frame extraction processing on the video to be processed according to a preset frame extraction frequency to obtain a set of video frames to be processed; performing clustering processing on the set of video frames to be processed based on a preset clustering algorithm to obtain the clustering category corresponding to each video frame to be processed; obtaining the time point of each video frame to be processed in the video to be processed, and sorting each video frame to be processed based on the time point to obtain a video frame sorting result; determining that two adjacent video frames to be processed correspond to different clustering categories based on the video frame sorting result and the clustering category corresponding to each video frame to be processed, and determining the target video frame with an earlier sorted time point among the two adjacent video frames to be processed as the video transition point. Thus, by clustering all the video frames corresponding to the video, the clustering category of each video frame is obtained, and each video frame is sorted according to the time point in the video, so that the sorting result of the clustering category of each video frame can be obtained. Therefore, according to the sorting result, the target video frame with a change in the clustering category can be determined as the video transition point, so that the video segmentation point can be accurately determined to segment the video to be processed, improving the video processing effect and efficiency.
[0049] In some embodiments, a video obtained by a shooting device with a fixed focal length shooting a fixed range is used as a video to be processed; and / or, the pixel difference between each video frame in the candidate video is obtained, and a candidate video with a pixel difference less than a preset difference threshold is used as the video to be processed.
[0050] In the embodiments of the present disclosure, the video to be processed is usually a video with a gradual and smooth frame-level picture change, that is, a video without pixel mutation. Therefore, a one-shot video obtained by setting a shooting device with a fixed focal length to shoot a fixed range is used as the video to be processed; or for a candidate video, the pixel difference between each video frame in each candidate video is calculated, and a candidate video with a pixel difference less than a preset difference threshold is determined as the video to be processed; wherein, the difference threshold is selected and set according to actual application needs. A pixel difference less than the difference threshold indicates that the picture change between video frames is small. It can also be calculated by means of the pixel average difference between each video frame, etc., to further improve the calculation accuracy, thereby improving the accuracy of judgment and further improving the subsequent processing effect.
[0051] In the above solution, a video with a gradual and smooth frame-level picture change, that is, a video without pixel mutation, can be used as the video to be processed, so that the target video frames in such a video can be accurately selected, thereby realizing effective segmentation of the video and improving the determination effect and efficiency of the video transition point.
[0052] In some embodiments, clustering processing is performed on the set of video frames to be processed based on a preset clustering algorithm to obtain the clustering category corresponding to each video frame to be processed, including: processing each video frame to be processed in the set of video frames to be processed through a pre-trained image feature extraction model to obtain the image feature of each video frame to be processed; calculating the similarity score between all video frames to be processed based on the image features of each video frame to be processed; taking at least two video frames to be processed with a similarity score greater than a preset similarity threshold as the same clustering category to obtain the clustering category corresponding to each video frame to be processed.
[0053] Among them, the image feature extraction model can be pre-set, such as a pre-trained CLIP model based on contrastive text-image pairs or other multi-modal pre-trained models, to extract features from each video frame to be processed, so as to obtain the image features of each video frame to be processed.
[0054] In some embodiments, a video frame sample is obtained, where each video frame sample has a corresponding feature label. The video frame sample is processed through a pre-constructed image feature extraction model to be trained to obtain video frame features, and the loss value is determined based on the similarity between the video frame features and the feature labels and a pre-set loss function to adjust the model parameters, so as to obtain the image feature extraction model.
[0055] Specifically, the video frame samples can be input into a pre-constructed image feature extraction model to be trained, and video frame features can be obtained. Then, the video frame features and feature labels are subjected to operations such as dot product, and after obtaining the similarity between the video frame features and the feature labels, they are input into a pre-set loss function for calculation to obtain a loss value. It can be understood that when the loss value is less than a certain threshold, the training can be stopped. If the loss value is greater than or equal to the threshold, the model parameters of the image feature extraction model to be trained need to be adjusted, and then the model is continuously trained until the loss value is less than the threshold to obtain the image feature extraction model. Thus, the image feature extraction model to be trained can be accurately and quickly obtained, thereby improving the accuracy of subsequent image feature extraction.
[0056] In some embodiments, calculating the similarity scores between all the video frames to be processed based on the image features of each video frame to be processed includes: obtaining the target features corresponding to each target based on the image features of each video frame to be processed, and calculating the similarity based on each target feature and the weight coefficient corresponding to each target to obtain the similarity scores between all the video frames to be processed.
[0057] In the embodiments of the present disclosure, various targets in the video frame such as people and objects can be determined for the shooting scene, and different weight coefficients can be set according to the importance of the targets in the shooting scene. Thus, the target features corresponding to each target can be obtained from the image features of each video frame to be processed, and the similarity is calculated by weighted summation of each target feature and the weight coefficient corresponding to each target, so as to obtain the similarity scores between all the video frames to be processed, and clustering processing is performed according to the similarity scores and a preset similarity threshold.
[0058] Specifically, for a single-shot video scene, a relatively low weight coefficient can be set for targets such as the background of the video frame, and a relatively high weight coefficient can be set for targets such as people in the video frame. Thus, the video frame can be divided into similarity calculations between different targets, so as to be able to determine the changes between the target features, thereby improving the accuracy of determining the video transition points.
[0059] As an example, for the scenario of shooting a person in one continuous shot, the set of video frames to be processed includes five video frames to be processed. The five video frames to be processed are sorted according to their time points in the video to be processed. The sorting result is video frame to be processed X1, video frame to be processed X2, video frame to be processed X3, video frame to be processed X4, and video frame to be processed X5. Then, the image features corresponding to the five video frames to be processed are extracted, and the target features corresponding to each target in each video frame to be processed are obtained in sequence. For example, video frame to be processed X1 includes the number-of-people feature a1, the person-action feature a2, and the background feature a3, and video frame to be processed X2 includes the number-of-people feature a4, the person-action feature a5, and the background feature a6. The weight coefficients r1 for the number-of-people feature, r2 for the person-action feature, and r3 for the background feature are preset; where r2 is greater than or equal to r3 and greater than r1, and r1 + r2 + r3 = 1; the target features corresponding to each target in video frame to be processed X1 and video frame to be processed X2 are calculated for similarity and weighted summation is performed in combination with the corresponding weight coefficients to obtain the similarity score between video frame to be processed X1 and video frame to be processed X2. Then, the similarity score is compared with the preset similarity threshold. For example, if the similarity score is greater than the preset similarity threshold, it means that the changes in the number of people, person actions, etc. between video frame to be processed X1 and video frame to be processed X2 are not significant. Therefore, video frame to be processed X1 and video frame to be processed X2 are taken as the same clustering category Y1. Similarly, continue to calculate that the similarity score between video frame to be processed X2 and X3 is also greater than the preset similarity threshold, and video frame to be processed X3 is taken as the same clustering category Y1.
[0060] For another example, in actual shooting, there is a scenario where an additional person is added in video frame to be processed X4 and the action of one existing person changes. Video frame to be processed X3 includes the number-of-people feature a7, the person-action feature a8, and the background feature a9, and video frame to be processed X4 includes the number-of-people feature a10, the person-action feature a11, and the background feature a12. Based on the above preset weight coefficients r1 for the number-of-people feature, r2 for the person-action feature, and r3 for the background feature, the target features corresponding to each target in video frame to be processed X3 and video frame to be processed X4 are calculated for similarity and weighted summation is performed in combination with the corresponding weight coefficients to obtain the similarity score between video frame to be processed X3 and video frame to be processed X4, which is less than the preset similarity threshold, indicating that the number of people, person actions, etc. between video frame to be processed X3 and video frame to be processed X4 have changed. Therefore, video frame to be processed X3 and video frame to be processed X4 are not taken as the same clustering category, and thus video frame to be processed X4 is taken as the new clustering category Y2. Similarly, continue to calculate that the similarity score between video frame to be processed X4 and X5 is greater than the preset similarity threshold, and video frame to be processed X5 is taken as the same clustering category Y2.
[0061] Thus, for a one-shot video scene where the changes between video frames are not significant, clustering can be performed by calculating the similarity scores for each target such as characters and objects between video frames and comparing them with a similarity threshold based on pre-set weight coefficients. This can achieve clustering based on the number of characters or the degree of change in character actions in the video frames, improve the accuracy of clustering categories, and thus improve the accuracy of determining subsequent video transition points, ultimately enhancing the effects of subsequent video segmentation and other processing.
[0062] Specifically, according to the image features of each video frame to be processed, calculate the similarity scores between all video frames to be processed. For example, calculate the feature cosine value, Euclidean distance, etc. as the similarity scores between the video frames to be processed. Consider at least two video frames to be processed with similarity scores greater than the preset similarity threshold as the same clustering category, and obtain the clustering category corresponding to each video frame to be processed. Among them, the similarity threshold can be set according to actual application needs.
[0063] Specifically, consider each video frame to be processed as an independent clustering category. Select two video frames to be processed with the maximum similarity score from the pairwise similarity scores calculated between all video frames to be processed to form a clustering category, and merge the two closest clustering categories into a new clustering category, and update the distance matrix. Continue to calculate the pairwise similarity scores between the remaining video frames to be processed and the similarity scores between the video frames to be processed and the clustering categories, and then merge the video frames to be processed or clustering categories with the maximum similarity score together until the number of clustering categories or other set conditions are met, and obtain the clustering category corresponding to each video frame to be processed.
[0064] In the above solution, by obtaining the image features of each video frame to be processed and clustering all video frames to be processed based on the maximum similarity score between the image features of each video frame to be processed, the clustering category corresponding to each video frame to be processed can be accurately obtained, thereby improving the accuracy of determining the target video frame subsequently and further enhancing the effect of determining the video transition point.
[0065] In some embodiments, the method further includes: obtaining the video frame time points between adjacent target video frames, and determining the time difference between adjacent target video frames based on the video frame time points; when the time difference is greater than the preset time threshold, determine a target video frame from the adjacent target video frames as the video transition point.
[0066] In the embodiments of the present disclosure, usually two or more target video frames are determined in the set of video frames to be processed, for example, five target video frames are determined, the first target video frame and the second target video frame can be used as adjacent target video frames, the second target video frame and the third target video frame can be used as adjacent target video frames, the third target video frame and the fourth target video frame can be used as adjacent target video frames; the fourth target video frame and the fifth target video frame can be used as adjacent target video frames.
[0067] In the embodiments of the present disclosure, a time threshold is set in advance. In one embodiment, the frame extraction time can be obtained based on the frame extraction frequency, and the time threshold is determined based on the frame extraction time to further improve the accuracy of processing. For example, one frame is extracted every two seconds, and the frame extraction time is determined to be two seconds, then the time threshold is set to four seconds, that is, the time threshold is determined according to the set multiple and the frame extraction time, thereby further improving the processing flexibility.
[0068] In the disclosed embodiment, each target video frame has a corresponding video frame time point. By processing the video frame time point between adjacent target video frames, the time difference between adjacent target video frames can be obtained, and when the time difference is greater than a time threshold, the adjacent target video frames are determined to be video transition points. If the time difference is less than or equal to the time threshold, the target video frame with the time point sorted in front needs to be deleted, and only the target video frame with the time point sorted in the back is used as the video transition point.
[0069] Specifically, for example, the first target video frame and the second target video frame can be used as adjacent target video frames, so that the video frame time point corresponding to the second target video frame is subtracted from the video frame time point corresponding to the first target video frame to obtain the time difference between the first target video frame and the second target video frame. For example, the time difference is greater than the time threshold, which indicates that the first target video frame and the second target video frame have normal changes in the picture, and the first target video frame and the second target video frame are both used as video transition points; for example, the time difference is less than or equal to the time threshold, which indicates that the first target video frame and the second target video frame may be rapid and short-term changes in the picture, and no processing is performed, and the first target video frame may not be used as a video transition point.
[0070] In the above scheme, some target video frames with fast and short-term changes in the picture can be deleted, so as to further ensure the accuracy of the video cutting point, improve the accuracy of video cutting, and improve the effect and efficiency of video cutting.
[0071] In some embodiments, the method further includes: determining, based on the video frame sorting result and the clustering category corresponding to each video frame to be processed, that there are video frames to be processed belonging to other clustering categories between the video frames to be processed belonging to the same clustering category, and deleting the video frames to be processed belonging to other clustering categories; for example, if the sorting result of the clustering category by time point is AAAABAAA, the fifth video frame to be processed can be used as the target video frame, but there are video frames to be processed belonging to other clustering category B between the video frames to be processed belonging to the same clustering category A. Therefore, the fifth video frame to be processed can be deleted and not used as the target video frame for deletion processing.
[0072] In the above solution, the target video frame can be accurately determined, and some perturbed video frames can be deleted, further ensuring the accuracy of the video cut point, improving the accuracy of video segmentation, and enhancing the video segmentation effect and efficiency.
[0073] Figure 2 The flowchart of a method for determining video transition points provided by an embodiment of the present disclosure mainly includes the following steps S202 to step S408:
[0074] Step S202: Obtain the video to be processed, and perform frame extraction on the video to be processed according to a preset frame extraction frequency to obtain a set of video frames to be processed.
[0075] Step S204: Process each video frame to be processed in the set of video frames to be processed through a pre-trained image feature extraction model, obtain the image features of each video frame to be processed, and calculate the similarity scores between all video frames to be processed based on the image features of each video frame to be processed.
[0076] Step S206: Use at least two video frames to be processed with similarity scores greater than a preset similarity threshold as the same clustering category to obtain the clustering category corresponding to each video frame to be processed.
[0077] Step S208: Obtain the time point of each video frame to be processed in the video to be processed, and sort each video frame to be processed based on the time point to obtain a video frame sorting result. Determine that two adjacent video frames to be processed correspond to different clustering categories based on the video frame sorting result and the clustering category corresponding to each video frame to be processed, and use the video frame to be processed with the earlier time point sorting among the two adjacent video frames to be processed as the target video frame.
[0078] Step S210: Obtain the video frame time points between adjacent target video frames, and determine the time difference between adjacent target video frames based on the video frame time points. When the time difference is less than or greater than the time threshold, determine the video transition point based on the adjacent target video frames.
[0079] Step S212: Segment the video to be processed based on the video transition points.
[0080] Specifically, for a video to be processed with a long one-shot duration, such as those that need to be effectively segmented according to changes in the main character or the number of characters to achieve better results for the task. Therefore, the set of video frames to be processed of the video to be processed can be clustered through a clustering algorithm, and then the target video frames can be determined based on the clustering categories for dynamic segmentation. This can accurately select the target video frames such as changes in the number of characters and character actions in the video to determine the video transition points for use in downstream tasks, further meeting the video processing requirements.
[0081] Specifically, cluster the set of video frames to be processed of the video to be processed through SMLR. According to the clustering results, perform strategy processing at the time points where the clustering categories change, and finally output the target video frames to determine the video transition points. More specifically, extract frames from the one-shot video to be processed at a frequency of one frame every two seconds to obtain the set of video frames to be processed; use the SMLR clustering algorithm for the set of video frames to be processed, that is, first use the CLIP module to extract the image features of the video frames to be processed, and then obtain the clustering categories of the video frames to be processed through hierarchical clustering; sort according to the time (the time points at which they appear in the video to be processed) of the video frames to be processed, and record the clustering category of each video frame to be processed; extract the time points (video frames to be processed) where the clustering category changes, such as video frame 1 (category A), video frame 2 (category A), video frame 3 (category B), etc. Take video frame 2 as the clustering category change frame, that is, the target video frame. Finally, obtain a series of target video frames where the clustering category changes, and use the time points corresponding to the target video frames as the video transition points to segment the video to be processed.
[0082] It should be noted that since there may be disturbances in the video to be processed (such as rapid and short-term changes in the picture), for target video frames with a time difference (i.e., time interval) less than a certain time threshold, such as 4 seconds, they can be removed, and finally the target video frames of the one-shot video to be processed are obtained.
[0083] Thus, by extracting frames, the degree of change in the picture is amplified by reducing the sampling frequency to obtain a set of video frames to be processed with high efficiency and large differences, and by using image clustering to solve the selection of target video frames for one-shot videos, thereby achieving effective segmentation of the video and improving the video processing effect and efficiency.
[0084] Corresponding to the foregoing video transition point determination method, the embodiments of the present disclosure further provide a video transition point determination device. Figure 3The following is a schematic structural diagram of a video transition point determination device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and is applied to an electronic device. The device includes:
[0085] An acquisition module 302, configured to acquire a video to be processed, and perform frame extraction processing on the video to be processed according to a preset frame extraction frequency to obtain a set of video frames to be processed;
[0086] A clustering module 304, configured to perform clustering processing on the set of video frames to be processed based on a preset clustering algorithm to obtain a clustering category corresponding to each video frame to be processed;
[0087] A sorting module 306, configured to obtain the time position of each video frame to be processed in the video to be processed, and sort each video frame to be processed based on the time position to obtain a video frame sorting result;
[0088] A processing module 308, configured to determine that adjacent two video frames to be processed correspond to different clustering categories based on the video frame sorting result and the clustering category corresponding to each video frame to be processed, and determine the target video frame with the earlier sorted time position among the adjacent two video frames to be processed as the video transition point.
[0089] The above device provided by the embodiment of the present disclosure clusters all video frames corresponding to the video to obtain the clustering category of each video frame, and sorts each video frame according to the time position in the video, so that the sorting result of the clustering category of each video frame can be obtained. Therefore, according to the sorting result, the video frame with the changed clustering category can be determined as the target video frame, so that the video cut point can be accurately determined to cut the video to be processed, improving the video transition point determination effect and efficiency.
[0090] In some embodiments, the video to be processed is segmented based on the video transition point.
[0091] In some embodiments, the clustering module 304 includes: a processing unit, configured to process each video frame to be processed in the set of video frames to be processed through a pre-trained image feature extraction model to obtain the image feature of each video frame to be processed; a calculation unit, configured to calculate the similarity score between all video frames to be processed based on the image feature of each video frame to be processed; a clustering unit, configured to use at least two video frames to be processed with a similarity score greater than a preset similarity threshold as the same clustering category to obtain the clustering category corresponding to each video frame to be processed.
[0092] In some embodiments, the computing unit is specifically configured to: obtain target features corresponding to each target based on the image features of each video frame to be processed; perform similarity calculation based on each of the target features and the weight coefficient corresponding to each target to obtain similarity scores between all the video frames to be processed.
[0093] In some embodiments, the method device further includes: a training module, configured to obtain video frame samples, where each video frame sample has a corresponding feature label, process the video frame samples based on a pre-constructed image feature extraction model to be trained to obtain video frame features, and determine a loss value based on the similarity between the video frame features and the feature labels and a preset loss function to adjust model parameters, so as to obtain the image feature extraction model.
[0094] In some embodiments, the device further includes: an acquisition and determination module, configured to acquire video frame time points between adjacent target video frames, and determine a time difference between the adjacent target video frames based on the video frame time points; a determination and deletion module, configured to determine the video transition point from the adjacent target video frames when the time difference is less than or equal to a preset time threshold.
[0095] Optionally, the device further includes: a determination and acquisition module, configured to obtain a frame extraction time based on the frame extraction frequency, and determine the time threshold based on the frame extraction time.
[0096] The video transition point determination device provided by the embodiments of the present disclosure can execute the video transition point determination method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
[0097] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the device embodiments described above can refer to the corresponding process in the method embodiments, and will not be described in detail here.
[0098] The embodiments of the present disclosure provide an electronic device, which includes: a storage device storing a computer program thereon; a processing device configured to execute the computer program in the storage device to implement the steps of any method in the present disclosure.
[0099] Next, refer to Figure 4, which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0100] As Figure 4 shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 402 or the programs loaded from the storage device 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0101] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 400 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively.
[0102] Particularly, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.
[0103] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run by a processor, cause the processor to execute the image processing method provided by the embodiments of the present disclosure. The computer program products may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0104] In addition, embodiments of the present disclosure may also be computer-readable storage media, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the video transition point determination method provided by the embodiments of the present disclosure.
[0105] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0106] Embodiments of the present disclosure also provide a computer program product, including a computer program / instructions, which when executed by a processor implement the video transition point determination method in the embodiments of the present disclosure.
[0107] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0108] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure's technical solution based on the prompt message.
[0109] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may, for example, be in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0110] It can be understood that the above process of notifying and obtaining user authorization is merely illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0111] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0112] The above are only specific implementation manners of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for determining video transition points, characterized in that, The method includes: Obtain a video to be processed, and perform frame extraction on the video to be processed according to a preset frame extraction frequency to obtain a set of video frames to be processed; Perform clustering processing on the set of video frames to be processed based on a preset clustering algorithm, and obtain the clustering category corresponding to each video frame to be processed; Obtain the time position of each video frame to be processed in the video to be processed, and sort each video frame to be processed based on the time position to obtain a video frame sorting result; Based on the video frame sorting result and the clustering category corresponding to each video frame to be processed, determine that two adjacent video frames to be processed correspond to different clustering categories, and determine the target video frame with the earlier sorted time position among the two adjacent video frames to be processed as the video transition point.
2. The method according to claim 1, wherein The method further includes: Split the video to be processed based on the video transition point.
3. The method according to claim 1, wherein The performing clustering processing on the set of video frames to be processed based on a preset clustering algorithm, and obtaining the clustering category corresponding to each video frame to be processed includes: Process each video frame to be processed in the set of video frames to be processed through a pre-trained image feature extraction model to obtain the image feature of each video frame to be processed; Calculate the similarity scores between all video frames to be processed based on the image features of each video frame to be processed; Take at least two video frames to be processed with similarity scores greater than a preset similarity threshold as the same clustering category, and obtain the clustering category corresponding to each video frame to be processed.
4. The method according to claim 3, wherein The calculating the similarity scores between all video frames to be processed based on the image features of each video frame to be processed includes: Obtain target features corresponding to each target based on the image features of each video frame to be processed; Perform similarity calculation based on each target feature and the weight coefficient corresponding to each target to obtain the similarity scores between all video frames to be processed.
5. The method according to claim 3, wherein The method further includes: Obtain video frame samples, where each video frame sample has a corresponding feature label; Process the video frame samples through a pre-constructed image feature extraction model to be trained to obtain video frame features, and determine a loss value based on the similarity between the video frame features and the feature labels and a preset loss function to adjust the model parameters to obtain the image feature extraction model.
6. The method according to claim 1, wherein The method further includes: Obtain the time positions of video frames between adjacent target video frames, and determine the time difference between adjacent target video frames based on the time positions of the video frames; When the time difference is greater than a preset time threshold, determine one of the adjacent target video frames as the video transition point.
7. The method according to claim 6, characterized in that, The method further includes: Obtain the frame extraction time based on the frame extraction frequency; Determine the time threshold based on the frame extraction time.
8. A video transition point determination device, characterized in that, The device includes: An obtaining module, configured to obtain a video to be processed, and perform frame extraction on the video to be processed according to a preset frame extraction frequency to obtain a set of video frames to be processed; A clustering module, configured to perform clustering processing on the set of video frames to be processed based on a preset clustering algorithm, and obtain a clustering category corresponding to each video frame to be processed; A sorting module, configured to obtain the time position of each video frame to be processed in the video to be processed, and sort each video frame to be processed based on the time position to obtain a video frame sorting result; A processing module, configured to determine that two adjacent video frames to be processed correspond to different clustering categories based on the video frame sorting result and the clustering category corresponding to each video frame to be processed, and determine the target video frame with the earlier sorted time position among the two adjacent video frames to be processed as a video transition point.
9. An electronic device, characterized in that, The electronic device includes: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the video transition point determination method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is configured to execute the video transition point determination method according to any one of claims 1-7 above.