Method, device, equipment and storage medium for identifying video clips

The method automates the identification of video clip positions in TV dramas by determining frame pairs based on similarity and time differences, addressing inefficiencies in manual annotation and enhancing accuracy.

JP7680633B2Active Publication Date: 2025-05-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024523262
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-03-08
Filing Date
2022-11-29
Publication Date
2025-05-20
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

Manual annotation of opening and ending positions in TV dramas is time-consuming and inefficient, requiring significant human resources.

Method used

A method and apparatus for identifying video clips by determining video frame pairs based on similarity conditions and fusing them based on appearance time differences to automatically identify target video clips within a specified time range, utilizing machine learning techniques and computer devices.

Benefits of technology

Enables efficient and automated identification of video clip positions without human intervention, reducing time and resource consumption while improving accuracy in determining openings and endings of videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680633000002
    Figure 0007680633000002
  • Figure 0007680633000003
    Figure 0007680633000003
  • Figure 0007680633000004
    Figure 0007680633000004
Patent Text Reader

Abstract

This application discloses a video clip identification method, device, equipment and storage medium, which can be applied to video clip identification, artificial intelligence and in-vehicle applications in computer technology. According to the technical solution provided by the embodiment of the present application, a video frame feature of a first video and at least one video frame feature of a second video are obtained; a plurality of video frame pairs are determined according to the video frame feature of the first video and the at least one video frame feature of the second video, the video frame pair including a first video frame and a second video frame whose similarity meets a similarity condition, the first video frame belonging to the first video, and the second video frame belonging to at least one second video (201); a first video frame of the plurality of video frame pairs is fused according to an appearance time difference of the plurality of video frame pairs to obtain at least one candidate video clip of the first video, the appearance time difference being a numerical difference between the appearance times of two video frames in the video frame pair (202); a target time range is obtained; and at least one target video clip in the first video is determined according to the at least one candidate video clip and the target time range, the target video clip being within the target time range of the first video (203).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application seeks priority from a Chinese patent application filed with the China Patent Office on March 8, 2022, bearing application number 202210219364.3 and entitled "Video Clip Identification Method, Apparatus, Device, and Storage Medium," the entire contents of which are incorporated herein by reference.

[0002] The present application relates to the field of computer technology, and more particularly to a method, apparatus, device and storage medium for identifying video clips. [Background technology]

[0003] With the development of computer technology, video is experiencing a rapid growth, and more and more users are connected to the Internet to watch videos. Video also includes TV dramas, which usually have openings and endings, and in order to make it easier for users to watch TV dramas, a video platform can provide a function of determining the positions of the openings and endings in a TV drama and skipping the openings and endings.

[0004] In the related art, the opening and ending positions of a TV drama are all determined by manual annotation, that is, after watching a TV drama, the opening and ending positions of the TV drama are manually marked.

[0005] However, manual annotation requires consuming a large amount of time and human resources, which makes it less efficient to determine the opening and ending positions of a TV drama. Summary of the Invention [Problem to be solved by the invention]

[0006] The embodiments of the present application provide a method, device, apparatus, storage medium and computer program product for identifying a video clip, and the technical solutions thereof are as follows: [Means for solving the problem]

[0007] 1. A method for identifying a video clip, the method comprising: Obtaining video frame features of a first video and at least one video frame feature of a second video, and determining a plurality of video frame pairs based on the video frame features of the first video and the at least one video frame feature of the second video, wherein the video frame pairs include a first video frame and a second video frame whose similarity meets a similarity condition, and the first video frame belongs to the first video and the second video frame belongs to the at least one second video; Fusing a first video frame of the plurality of video frame pairs based on an appearance time difference of the plurality of video frame pairs to obtain at least one candidate video clip of the first video, the appearance time difference being a numerical difference between appearance times in a video of two video frames in the video frame pair; The method includes obtaining a target time range and determining at least one target video clip in the first video based on the at least one candidate video clip and the target time range, wherein the target video clip is within the target time range of the first video.

[0008] 1. An apparatus for identifying a video clip, the apparatus comprising: a video frame pair determination module for obtaining video frame features of a first video and at least one video frame feature of a second video, and determining a plurality of video frame pairs based on the video frame features of the first video and the at least one video frame feature of the second video, the video frame pairs including a first video frame and a second video frame whose similarity meets a similarity condition, the first video frame belonging to the first video, and the second video frame belonging to the at least one second video; a fusion module for fusing a first video frame of the plurality of video frame pairs based on an appearance time difference of the plurality of video frame pairs to obtain at least one candidate video clip of the first video, the appearance time difference being a numerical difference between appearance times in a video of two video frames in a video frame pair; and a target video clip determination module for obtaining a target time range and determining at least one target video clip in the first video based on the at least one candidate video clip and the target time range, the target video clip being within the target time range of the first video.

[0009] A computer device comprising one or more processors and one or more memories, at least one computer program stored in the one or more memories, the computer program being loaded and executed by the one or more processors to implement the method for identifying the video clip.

[0010] A computer readable storage medium is provided having stored thereon at least one computer program, the computer program being loaded and executed by the processor to implement the method for identifying video clips.

[0011] A computer program product is provided that includes a computer program which, when executed by a processor, implements the method for identifying video clips as described above. [Brief description of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic diagram of an implementation environment of a video clip identification method provided by an embodiment of the present application. [Diagram 2] FIG. 2 is a flow diagram of a method for identifying a video clip provided by an embodiment of the present application. [Diagram 3] FIG. 2 is a flow diagram of a method for identifying a video clip provided by an embodiment of the present application. [Figure 4] FIG. 2 is a flow diagram of a method for extracting video frame features provided by an embodiment of the present application. [Diagram 5] FIG. 2 is a schematic diagram of a first sub-clip and a second sub-clip provided by an embodiment of the present application. [Figure 6] 1A to 1C are schematic diagrams of first sub-clips in different overlay methods provided by embodiments of the present application; [Figure 7] FIG. 2 is a schematic diagram of fusing candidate video clips provided by an embodiment of the present application; [Figure 8] FIG. 2 is a flow diagram of a method for identifying a video clip provided by an embodiment of the present application. [Figure 9] FIG. 1 is a flow diagram of a clip mining system provided by an embodiment of the present application. [Figure 10] FIG. 2 is a flow diagram of a method for obtaining the opening and ending of a TV drama provided by an embodiment of the present application. [Figure 11] 2 is a schematic diagram of a storage method in a clip database provided by an embodiment of the present invention. FIG. [Figure 12] FIG. 2 is a flow diagram of a method for obtaining the opening and ending of a TV drama provided by an embodiment of the present application. [Figure 13] FIG. 2 is a flow diagram of a method for identifying infringing videos provided by an embodiment of the present application. [Figure 14]FIG. 2 is a flow diagram of a method for identifying a video clip provided by an embodiment of the present application. [Figure 15] 1 is a structural schematic diagram of a video clip identification device provided by an embodiment of the present application; [Figure 16] 1 is a schematic structural diagram of a terminal provided by an embodiment of the present application; [Figure 17] FIG. 2 is a schematic diagram of a server structure provided by an embodiment of the present application; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] The terms "first," "second," and the like, used in this application are used to distinguish between identical or similar items having essentially the same actions and functions, and it should be understood that there is no logical or chronological dependency between "first," "second," and "nth," and there are no limitations on the quantity or order of execution.

[0014] Artificial Intelligence (AI) is the theory, methods, techniques and application systems that use digital computers or digital computer control to machine-simulate, extend and expand human intelligence, sense the environment, gain knowledge and use that knowledge to achieve optimal results.

[0015] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines, including probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It is the specialized study of how computers can simulate or realize human learning behavior, acquire new knowledge or skills, reorganize existing knowledge submodels, and continually improve their own performance.

[0016] Hamming Distance: Used to measure the distance between binary features. It is achieved by taking the statistical value of the distance between different feature bits. For example, the Hamming distance between (1000) and (0011) is 3.

[0017] Furthermore, all of the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, data to be stored, data to be presented, etc.) and signals referred to in this application have been authorized by the user or fully authorized through various parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0018] FIG. 1 is a schematic diagram of an implementation environment of a video clip identification method provided by an embodiment of the present application. Referring to FIG. 1, the implementation environment may include a terminal 110 and a server 140 .

[0019] The terminal 110 is connected to the server 140 via a wireless or wired network. The terminal 110 may optionally be, but is not limited to, an in-vehicle terminal, a smartphone, a tablet PC, a laptop, a desktop computer, a smart speaker, a smart watch, a smart TV, etc. An application program supporting identification of video clips is implemented and operated in the terminal 110.

[0020] The server 140 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud memory, network services, cloud communications, middleware services, domain services, security services, Content Delivery Network (CDN), and big data, artificial intelligence platform, etc. The server 140 provides back-end services for application programs operated by the terminal 110.

[0021] In this embodiment, the number of terminals 110 and servers 140 is not limited.

[0022] After the implementation environment of the present embodiment has been introduced, the following will introduce the application scenario of the present embodiment by combining the above implementation environment. In the following explanation process, the terminal is the terminal 110 in the above implementation environment, and the server is the server 140 in the above implementation environment.

[0023] The video clip identification method provided in the present embodiment can be applied to identifying scenes that open and close a video, for example, to identifying scenes that open and close a TV drama, or to identifying scenes that open and close a documentary film, or to identifying scenes that open and close a short video collection, etc.

[0024] Take the video clip identification method provided in the embodiment of the present application as an example for application to the opening and ending identification scenes of a TV drama. An engineer selects a TV drama for which opening and ending need to be identified through a terminal, and the TV drama includes multiple videos, each video being an episode in the TV drama. After selecting the TV drama through the terminal, the server adopts the technical solutions provided by the embodiment of the present application to perform processing based on the multiple videos in the TV drama to obtain the opening and ending in the multiple videos. In the process of processing the multiple videos, the server determines multiple video frame pairs based on the video frame features of the first video and the video frame features of at least one second video, each video frame pair includes a first video frame and a second video frame whose similarity meets a similarity condition, the first video frame belongs to the first video, and the second video frame belongs to the at least one second video, that is, each video frame pair includes one video frame in the first video and one video frame in the second video, and both the first video and the at least one second video frame belong to the multiple videos. The server fuses a first video frame of the plurality of video frame pairs according to the appearance time difference of the plurality of video frame pairs to obtain at least one candidate video clip of the first video. The appearance time difference refers to a numerical difference between the appearance time of two video frames of the video frame pair in a video, that is, the numerical difference between the appearance time of the first video frame of the video frame pair in the first video and the appearance time of the second video frame of the video frame pair in the second video. The server determines at least one target video clip in the first video according to the at least one candidate video clip and a target time range, and when applied to a scene of identifying the opening and ending of a TV drama, the target time range is also the time range where the opening and ending are located, so the determined target video clip is also the opening or ending of the first video.

[0025] In addition, the above describes an example of applying the video clip identification method provided in the embodiments of the present application to identifying scenes of the opening and ending of a television drama. However, since the implementation process of the other application scenes mentioned above also belongs to the same inventive concept as the above description, the implementation process will not be described in detail here.

[0026] In addition, the video clip identification method provided in the present embodiment can be applied to identifying scenes that identify the opening and ending of the above-mentioned TV drama, the opening and ending of a documentary film, and the opening and ending of a short video collection, as well as to identifying scenes that identify the opening and ending of other types of videos, and the present embodiment is not limited thereto.

[0027] After the implementation environment and application scenarios of the embodiment of the present application are introduced, the video clip identification method provided by the embodiment of the present application will be described below. Referring to FIG. 2, the technical solution provided by the embodiment of the present application is executed by a computer device, and the computer device is executed as a terminal or a server, and the technical solution provided by the embodiment of the present application can also be executed by the terminal or the server together. In the following embodiment of the present application, an example in which the execution subject is a server is described. As can be easily understood, although the following embodiment is described as a server, each embodiment of the present application can also be executed by a terminal. That is, the technical solution provided by each embodiment of the present application is actually executed by a computer device.

[0028] Methods for identifying video clips include:

[0029] 201: A server obtains video frame features of a first video and at least one video frame feature of a second video, and determines a plurality of video frame pairs based on the video frame features of the first video and the at least one video frame feature of the second video, where the video frame pairs include a first video frame and a second video frame whose similarity meets a similarity condition, and the first video frame belongs to the first video and the second video frame belongs to the at least one second video.

[0030] The first video and at least one second video belong to the same video set, for example, the first video and the second video are different episodes of the same TV drama. The video frame feature is an embedding feature of the video frame, for example, a depth hash feature. The similarity between the first video frame and the second video frame is determined by the video frame feature of the first video frame and the video frame feature of the second video frame. Each video frame pair includes one first video frame and one second video frame, and if the similarity between the first video frame and the second video frame of the video frame pair meets a similarity condition, that is, the first video frame and the second video frame of the video frame pair are two video frames with a relatively high similarity.

[0031] 202: The server fuses a first video frame of the plurality of video frame pairs based on an appearance time difference of the plurality of video frame pairs to obtain at least one candidate video clip of the first video, where the appearance time difference refers to a numerical difference between appearance times in a video of two video frames in the video frame pair.

[0032] The first video frame of the video frame pair is a video frame having a relatively high similarity with the second video frame, and the candidate video clip is obtained by fusing the first video frames of the multiple video frame pairs, so that the candidate video clip is also a video clip having overlapping content with at least one of the first videos and the second video. The appearance time difference can reflect the appearance time deviation of the first video frame and the second video frame in the first video and the second video.

[0033] 203: The server obtains a target time range, and determines at least one target video clip in the first video based on the at least one candidate video clip and the target time range, where the target video clip is within the target time range of the first video.

[0034] The target time range refers to a time range in the video, and the target time range is set by the engineer according to the actual situation, so that the embodiment of the present application does not limit it.

[0035] According to the technical solution provided by the embodiment of the present application, a video frame pair including similar video frames is determined based on the similarity between the video frame features. A first video frame of the video frame pair is fused based on the appearance time difference to obtain at least one candidate video clip. Finally, a target video clip within a target time range is determined from the at least one candidate video clip. The process of determining the target clip does not require human intervention, and can be automatically performed by a computer device directly based on the first video and at least one second video, which is efficient.

[0036] The above steps 201-203 are a brief introduction to the video clip identification method provided by the embodiment of the present application, and the video clip identification method provided by the embodiment of the present application will be described in more detail below with reference to some examples. Referring to Fig. 3, the technical solution provided by the embodiment of the present application can be executed by a terminal or a server, or can be implemented by a terminal and a server together. In the embodiment of the present application, the implementation entity is a server, and the method includes the following:

[0037] 301: A server performs feature extraction on a first video and at least one second video to obtain video frame features of the first video and at least one video frame features of the second video.

[0038] In one possible embodiment, the server inputs a first video and at least one second video into a feature extraction model, and performs feature extraction on the first video and the at least one second video via the feature extraction model to obtain video frame features of the first video and video frame features of the at least one second video.

[0039] The process of the server performing feature extraction on the first video and at least one second video through the feature extraction model is also a process of performing feature extraction on a first video frame of the first video and a second video frame of the second video, and in this case, the feature extraction model is an image feature extraction model.

[0040] In this type of embodiment, feature extraction is performed on the first video and the at least one second video via a feature extraction model to obtain video frame features of the first video and video frame features of the at least one second video, and abstract representations of the first video and the at least one second video are realized, thereby improving computational efficiency in subsequent processes.

[0041] To explain the above embodiment, the following describes the above embodiment through three examples.

[0042] Example 1: A server inputs the first video and the at least one second video into a feature extraction model, and convolves and pools a plurality of first video frames and a plurality of second video frames through the feature extraction model to obtain video frame features of the plurality of first video frames and video frame features of the plurality of second video frames, where the plurality of first video frames are video frames of the first video, and the plurality of second video frames are at least one video frame of the second video.

[0043] The following describes a method for the server to perform feature extraction on a first video. The server inputs a plurality of first video frames of the first video into a feature extraction model, and performs convolution on the plurality of first video frames through a convolution layer of the feature extraction model to obtain a feature image of the plurality of first video frames. The server performs any one of max pooling or average pooling on the feature image of the plurality of first video frames through a pooling layer of the feature extraction model to obtain video frame features of the plurality of first video frames. In some embodiments, the server represents the first video frames in a matrix format and the video frame features in a vector format, and the process of performing convolution on the first video frames is realized in a form of adopting a convolution kernel to slide on the first video frames.

[0044] In some embodiments, the feature extraction model is a feature extractor based on Convolutional Neural Networks (CNN), for example, a neural network Resnet-101 (residual network 101) that adopts a large open source dataset to pre-train on imagenet (image network), and the structure of the neural network Resnet-101 is shown in Table 1. The output result of the pooling layer of the neural network Resnet-101 is the video frame feature, where 101 refers to the number of layers of the model, and the video frame feature is a 1×2048 vector.

[0045] [Table 1] In the table, Layer name is the name of each layer in the feature extraction model ResNet-101, Output size is the size of the output feature image, max pool refers to maximum pooling, stride refers to stride, blocks refers to layer, a layer may contain multiple convolution kernels, Conv refers to convolution layer, Pool refers to pooling layer, Class refers to classification layer, and full connection refers to full connection. In the process of extracting the above video frame features, the Class layer is not used.

[0046] Note that the above description is given using an example in which the feature extraction model is ResNet-101; however, in other possible embodiments, the feature extraction model may have other structures, and this is not limited to the embodiments of the present application.

[0047] In addition, the above feature extraction process is realized based on convolution, and the obtained video frame features are used to represent the image texture features of the video frame, and such video frame features are also referred to as lower-level features of the video frame. In other possible embodiments, the feature extraction model can further extract semantic features of the video frame, and the obtained video frame features can reflect the semantics of the video frame. The following describes how the server extracts semantic features of the video frame through the feature extraction model.

[0048] Example 2: the server inputs the first video and the at least one second video into a feature extraction model, and codes a plurality of first video frames and a plurality of second video frames according to an attention mechanism through the feature extraction model to obtain video frame features of the plurality of first videos and video frame features of the plurality of second video frames, where the plurality of first video frames are video frames of the first video, and the plurality of second video frames are video frames of at least one second video, and the video frame features obtained through the feature extraction model are also semantic features of corresponding video frames. In this kind of embodiment, the feature extraction model is a semantic feature encoder, for example, a Transformer encoder.

[0049] The following describes a method in which the server performs feature extraction on multiple first videos. The server inputs multiple first video frames of the first video into a feature extraction model, and performs code embedding on the multiple first video frames through the feature extraction model to obtain multiple embedding vectors, where each embedding vector corresponds to a first video frame, and the embedding vector is used to represent the position of the first video frame in the first video and the content of the first video frame. The server inputs the multiple embedding vectors into the feature extraction model, and linearly transforms the multiple embedding vectors through three linear transformation matrices of the feature extraction model to obtain a query vector, a key vector, and a value vector corresponding to each first video frame. The server obtains attention weights of the multiple first video frames based on the query vectors and the key vectors corresponding to the multiple first video frames through the feature extraction model. The server obtains attention coding vectors of each first video frame based on the attention weights of each first video frame and the value vectors of each first video frame through the feature extraction model, where the attention coding vectors are also video frame features of the first video frames.

[0050] For example, the server multiplies each embedding vector by three linear transformation matrices through the feature extraction model to obtain a query vector, a key vector and a value vector corresponding to each first video frame, respectively. For a first first video frame in the plurality of first video frames, the server determines a plurality of attention weights between the first first video frame of the plurality of first video frames according to the query vector of the first first video frame and the key vector of the plurality of first video frames through the feature extraction model. For a first first video frame in the plurality of first video frames, the server performs weighted addition of the attention weights for the first first video frame of the plurality of first video frames with the value vectors of the plurality of first video frames through the feature extraction model to obtain an attention coding vector of the first first video frame, which is also a video frame feature of the first first video frame.

[0051] In the above Examples 1 and 2, the sub-level features and semantic features of the video frame are extracted using the feature extraction model, respectively. In other possible embodiments, the server can further simultaneously obtain the sub-level features and semantic features of the video frame through the feature extraction model, which is described below as Example 3.

[0052] Example 3: A server inputs the first video and the at least one second video into a feature extraction model, and performs convolution and pooling on a plurality of first video frames and a plurality of second video frames through the feature extraction model to obtain underlayer features of the plurality of first video frames and underlayer features of the plurality of second video frames, where the plurality of first video frames are video frames of the first video, and the plurality of second video frames are at least one video frame of the second video. The server codes the plurality of first video frames and the plurality of second video frames based on an attention mechanism through the feature extraction model to obtain semantic features of the plurality of first video frames and semantic features of the plurality of second video frames. The server fuses the underlayer features and the semantic features of each first video frame to obtain video frame features of each first video frame. The server fuses the underlayer features and the semantic features of each second video frame to obtain video frame features of each second video frame.

[0053] For example, the feature extraction model includes a first sub-model and a second sub-model, the first sub-model is used to extract the sub-layer features of the video frames, and the second sub-model is used to extract the semantic features of the video frames. The server inputs the first video and the at least one second video into the feature extraction model, and then obtains the sub-layer features of the first video frames and the sub-layer features of the second video frames through the first sub-model, and obtains the semantic features of the first video frames and the semantic features of the second video frames through the second sub-model. When the server fuses the sub-layer features and the semantic features of each video frame, a weighted summation method can be adopted, and the weight of the weighted summation can be set by the engineer according to the actual situation, for example, 0.5, but is not limited thereto in the embodiment of the present application. The method of the server obtaining the sub-layer features and the semantic features of the video frames through the first sub-model and the second sub-model is the same as that of the above example 1 and example 2, respectively, and will not be described in detail here.

[0054] In addition, the above describes an example of extracting sub-level features and semantic features of a video frame using a feature extraction model. However, with the development of science and technology, the server can also adopt feature extraction models with other structures to extract video frame features, and this is not limited to the embodiments of the present application.

[0055] In some embodiments, the first video and at least one second video are videos belonging to the same video set, the first video is a video to be determined as a target video clip, and the at least one second video is all videos in the video set other than the first video, or the at least one second video is a video extracted from the video set, and the first video is masked when extracted. If the at least one second video is a video extracted from the video set, the server randomly extracts the second video of the target video quantity from the video set, and during the extraction process, the first video is masked, that is, the first video is not included in the extracted second video of the target video quantity, and the target video quantity is set by the engineer according to the actual situation, and is not limited thereto in the embodiment of the present application. The server respectively forms at least one video pair by the first video and the at least one second video, and each video pair includes the first video and one of the at least one second videos.

[0056] For example, if the video set contains 46 videos, for each first video i, the server randomly extracts 10 second videos r from the excess videos in the video set, and constructs 10 video pairs each consisting of the first video i and the 10 second videos r. In the subsequent processing steps, the video pair is taken as a unit, and 10 is the number of target videos.

[0057] In some embodiments, the server performs frame extraction on the first video and the at least one second video before performing feature extraction on the first video and the at least one second video to obtain a plurality of first video frames of the first video and a plurality of second video frames of each second video. By performing frame extraction on the videos, the amount of calculation in the subsequent feature extraction process can be reduced, and the efficiency of the feature extraction can be improved.

[0058] Taking the first video as an example, the server performs frame extraction from the first video at a target interval to obtain a plurality of first video frames of the first video, where the target interval refers to a target playback time length of the first video, for example, 1s, or the target interval refers to a target number of frame intervals, for example, 25 frames. If the target interval refers to a target playback time length of the first video, the server extracts one frame from the first video for each target playback time length as the first video frame. If the first video is 6s and the target playback time length is 1s, the server extracts six first video frames from the first video. If the target time interval refers to a target number of frame intervals, the server performs extraction from the first video for each target number of video frames to obtain a plurality of first video frames. If the first video includes 100 video frames and the target number is 10, the server extracts 10 first video frames from the first video. 4, for example, a server performs frame extraction from a first video 400 at a target interval to obtain a plurality of first video frames 401 of the first video. The server inputs the plurality of first video frames 401 of the first video into a feature extraction model 402, and outputs video frame features 403 of the plurality of first video frames 401 through the feature extraction model 402.

[0059] It should be noted that the above step 301 is an optional step, which can be executed in advance by the server, or can be executed when the server implements the technical solution provided by the embodiment of the present application, and is not limited thereto in the embodiment of the present application.

[0060] 302, the server determines a plurality of video frame pairs based on the video frame features of the first video and the video frame features of the at least one second video, the video frame pairs including a first video frame and a second video frame whose similarity meets a similarity condition, the first video frame belonging to the first video, and the second video frame belonging to the at least one second video.

[0061] In one possible embodiment, a server determines a similarity between video frame features of a plurality of first videos and video frame features of a plurality of second video frames, and the server determines the first video frames and the second video frames whose similarity meets a target condition as a video frame pair, where each video frame pair includes one first video frame and one second video frame.

[0062] The similarity between the video frame features is determined by Euclidean distance or cosine similarity, but is not limited thereto in the present embodiment.

[0063] In this type of embodiment, the server determines multiple video frame pairs based on the similarity between the first video frame and the second video frame, and the video frames in the video frame pairs are video frames with relatively high similarity in different videos, so that similar video clips can be quickly determined subsequently based on the video frame pairs, and finally the target video clip can be determined, so that the efficiency is relatively high.

[0064] When the similarity is Euclidean distance, the server determines the Euclidean distance between the video frame features of the first video frames and the video frame features of the second video frames. The server determines the first video frame and the second video frame whose Euclidean distance is equal to or less than the distance threshold as a video frame pair. The distance threshold is set by the engineer according to the actual situation, and is not limited in the embodiment of the present application. When the distance threshold is 0.5, if the Euclidean distance between the video frame features of any one of the first video frames and the video frame features of any one of the second video frames is equal to or less than 0.5, the server determines the first video frame and the second video frame as a video frame pair.

[0065] If the similarity is a cosine similarity, the server determines a cosine similarity between the video frame features of the first video frames and the video frame features of the second video frames. The server determines the first video frames and the second video frames whose cosine similarity is equal to or greater than a similarity threshold as a video frame pair. If the similarity threshold is 0.8, if the cosine similarity between the video frame features of any one of the first video frames and the video frame features of any one of the second video frames is equal to or greater than 0.8, the server determines the first video frame and the second video frame as a video frame pair.

[0066] In some embodiments, when the server constructs at least one video pair by a first video and at least one second video, the server determines a similarity between the video frame features of the first video and the video frame features of the second video in the video pair for each video pair to determine a plurality of video frame pairs in the video pair. For example, for a video pair (i, r), the server determines a similarity between the video frame features of the first video i and the video frame features of the second video r. The server determines the first and second video frames whose similarity meets the target condition as one video frame pair. That is, for each first video frame j in the first video i, the server determines the Euclidean distance between the video frame features of the first video frame j and each second video frame in the second video r. The server determines the Euclidean distance between the first video frame j and the video frame features of the second video frame in the second video r when the Euclidean distance is t 0 A second video frame that is less than 1 second is regarded as a similar frame of the first video frame j, and the first video frame j and the similar frame constitute one video frame pair. The server stores the obtained similar frames of the first video frame j in a first list, which is also referred to as a similar frame list (sim-id-list). In some embodiments, the server stores a frame identifier in the first list, and the frame identifier is used to indicate the video to which the frame belongs and the position of the frame in the video. For example, for a first video frame j=1, if the similar frame list sim-id-list is [1,2,3], it indicates that the video frames corresponding to 1st, 2nd, and 3rd seconds of the second video r are similar frames, and j=1 indicates the video frame corresponding to the 1st second in the first video.

[0067] In some embodiments, a plurality of video frame pairs are determined based on the video frame features of a first video and the video frame features of at least one second video, which includes: obtaining a video set, and determining a video of a target video clip to be determined in the video set as a first video; determining at least one video in the video set that is different from the first video as at least one second video; forming at least one video pair with the first video and the at least one second video, the video pair including the first video and one second video of the at least one second video; performing a similarity calculation of video frame features on the first video and the second video in the same video pair to obtain a similarity calculation result; and according to the similarity calculation result, determining a pair of video frame pairs of a first video frame and a second video frame whose similarity meets a similarity condition in the same video pair; the first video frame belongs to the first video, and the second video frame belongs to the at least one second video.

[0068] Optionally, after step 302, if the determined number of video frame pairs is zero, the server determines that the target video clip is not present in the first video.

[0069] 303, the server determines the occurrence time differences of a plurality of pairs of video frames.

[0070] In one possible embodiment, the server subtracts the appearance time of the first video frame in the plurality of video frame pairs in the first video from the appearance time of the second video frame in the plurality of video frame pairs in the second video to obtain the appearance time difference of the plurality of video frame pairs. In some embodiments, the server stores the appearance time difference of the plurality of video frame pairs in a second list, which is also called an appearance time difference list (diff-time-list), and can directly use the corresponding appearance time difference from the second list in subsequent processing steps. For example, for the first video frame with j=1, if the similar frame list sim-id-list is [1,2,3], the corresponding appearance time difference list diff-time-list is [0,1,2].

[0071] 304. The server divides the plurality of video frame pairs into a plurality of video frame groups based on appearance time differences of the plurality of video frame pairs, where video frame pairs in the same video frame group correspond to the same appearance time difference, and the appearance time difference refers to a numerical difference between the appearance times in the video of two video frames in the video frame pair.

[0072] In one possible embodiment, for any one video frame pair in the multiple video frame pairs, the server determines a first appearance time of a first video frame and a second appearance time of a second video frame in the video frame pair, the first appearance time refers to the time when the first video frame appears in the first video, and the second appearance time refers to the time when the second video frame appears in the second video, the server subtracts the second appearance time of the second video frame from the first appearance time of the first video frame in the video frame pair to obtain an appearance time difference of the video frame pair, the server classifies the video frame pairs with the same appearance time difference as one initial video frame group, and the appearance time difference of the video frame pair in the initial video frame group is the appearance time difference corresponding to the initial video frame group. The server merges the multiple initial video frame groups according to the appearance time differences corresponding to the multiple initial video frame groups to obtain the multiple video frame groups.

[0073] The initial video frame group includes a plurality of video frame pairs having the same appearance time difference, and different initial video frame groups correspond to different appearance time differences, and the appearance time difference corresponding to an initial video frame group refers to the appearance time difference of a video frame pair in the initial video frame group.

[0074] In one possible embodiment, before dividing the multiple video frame pairs into multiple video frame groups based on the appearance time difference of the multiple video frame pairs, the method further includes: for any one video frame pair of the multiple video frame pairs, subtracting a second appearance time of a second video frame of the video frame pair from a first appearance time of a first video frame of the video frame pair to obtain an appearance time difference of the video frame pair, wherein the first appearance time refers to the time when the first video frame appears in the first video and the second appearance time refers to the time when the second video frame appears in the second video.

[0075] In this type of embodiment, video frames of a video frame pair with the same occurrence time difference will likely constitute a complete video clip, and by combining the video frame pairs into a video frame group, it becomes easier to determine similar video clips in subsequent processes.

[0076] For example, the server obtains certain configuration information, and sorts the multiple initial video frame groups according to a target order in the configuration information; for any two adjacent candidate video frame groups in the multiple candidate video frame groups, if the matching time difference between the two adjacent candidate video frame groups meets a matching time difference condition, the two adjacent candidate video frame groups are merged into one video frame group, and the matching time difference refers to the numerical difference between the appearance time difference corresponding to the two adjacent candidate video frame groups.

[0077] The server sorts the initial video frame groups according to a target order in the predetermined configuration information to obtain a plurality of candidate video frame groups, and if a matching time difference between any two adjacent candidate video frame groups in the plurality of candidate video frame groups meets a matching time difference condition, the server merges the two adjacent candidate video frame groups into one video frame group, and the matching time difference refers to a numerical difference between the appearance time difference corresponding to the two adjacent candidate video frame groups.

[0078] In order to more clearly explain the technical process mentioned in the above example, the above example will be further described in two parts below.

[0079] First, the server sorts the initial set of video frames according to a target order to obtain candidate sets of video frames.

[0080] In one possible embodiment, the server sorts the initial video frames according to the order of corresponding appearance time difference from smallest to largest to obtain the candidate video frames, where the target order refers to the order of appearance time difference from largest to smallest. In some embodiments, for any one initial video frame, the server sorts according to the appearance time before or after the first video frame in the video frame pair.

[0081] In this type of embodiment, the server sorts the initial video frames according to order from largest to smallest, and in the obtained candidate video frames, if the appearance time differences corresponding to any two candidate video frames are relatively close to each other, the subsequent fusion process is simplified.

[0082] For example, if the initial video frame groups are [3,5], [11,12], [2,4], [4,6], [6,9], [7,10], [10,11], each bracket represents one video frame pair [i,r], the first number in the bracket is the identifier of the first video frame i, the second number is the identifier of the second video frame r, and the identifier is the appearance time of the video frame in the video. For the video frame pair [3,5], the appearance time difference is 5-3=2, and for the video frame pair [6,9], the appearance time difference is 9-6=3. The server sorts the initial video frame groups according to the order of the corresponding appearance time difference from smallest to largest, and obtains a plurality of candidate video frame groups [10,11], [11,12], [2,4], [3,5], [4,6], [6,9], [7,10].

[0083] In one possible embodiment, the server sorts the initial video frames according to the order of corresponding appearance time difference from smallest to largest to obtain the candidate video frames, where the target order refers to the order of appearance time difference from smallest to largest. In some embodiments, for any one initial video frame, the server sorts according to the appearance time before or after the first video frame in the video frame pair.

[0084] In this kind of embodiment, the server sorts the multiple initial video frame groups according to an order from smallest to largest, and in the obtained multiple candidate video frame groups, if the appearance time differences corresponding to any two candidate video frame groups are relatively close to each other, the subsequent fusion process is simplified.

[0085] In some embodiments, when a first list is adopted to store video frame pairs and a second list is adopted to store occurrence differences, the server generates a third list based on the first list and the second list, the third list is used to store the video frame pairs and the occurrence differences, and the third list can store multiple initial video frame groups, for example, the format of the third list is third list(match-dt-list):{d:{count, start-id, match-id-list}, ...}, where d is the occurrence time difference, d:{count, start-id, match-id-list} indicates the initial video frame group with the occurrence time difference d, count is the number of video frame pairs in the initial video frame group, start-id is the smallest identifier of the first video frame, and match-id-list is the video frame pair.

[0086] Second part, if the matching time difference between any two adjacent candidate video frame groups in the plurality of candidate video frame groups meets a matching time difference condition, the server merges the two adjacent candidate video frame groups into one video frame group.

[0087] In one possible embodiment, the two adjacent candidate video frame groups include a first candidate video frame group and a second candidate video frame group, and if the matching time difference between the appearance time difference corresponding to the first candidate video frame group and the appearance time difference corresponding to the second candidate video frame group is less than or equal to a matching difference threshold, the server adds the video frame pairs in the first candidate video frame group to the second candidate video frame group to obtain the video frame group.

[0088] The fusing of the plurality of candidate video frame groups into a plurality of video frame groups includes a plurality of iterative processes, and after fusing the first candidate video frame group and the second candidate video frame group into one video frame group, the server further determines the matching time difference between the newly merged video frame group and the next candidate video frame group, and if the matching time difference meets the matching time difference condition, the newly merged video frame group can be merged again with the next candidate video frame group, and the fusion process belongs to the same invention concept as the process of fusing the first candidate video frame group and the second candidate video frame group, so the implementation process is not described in detail. Of course, if the matching time difference does not meet the matching time difference condition, the server further determines the matching time difference between the next candidate video frame group and the next candidate video frame group, and performs further processing according to the matching time difference. The matching difference threshold is set by the technician according to the actual situation, and is not limited in the embodiment of the present application.

[0089] In this kind of embodiment, the candidate video frames are fused based on the occurrence time difference, so that the quantity of the candidate video frames can be reduced, the computation amount of the subsequent processing is reduced, and the computation efficiency is improved.

[0090] In one possible embodiment, the two adjacent candidate video frame groups include a first candidate video frame group and a second candidate video frame group, and fusing the two adjacent candidate video frame groups into one video frame group includes: adding a video frame pair in the first candidate video frame group to a second candidate video frame group when a matching time difference between the first candidate video frame group and the second video frame group is less than or equal to a matching difference threshold; and replacing the target video frame with a reference second video frame based on an appearance time difference corresponding to the second candidate video frame group to obtain a video frame group, wherein the target second video frame is a second video frame newly added to the second candidate video frame group, the reference second video frame is a second video frame of a second video having an appearance time difference between itself and the target first video frame that is a target numerical difference, the target numerical difference being the appearance time difference corresponding to the second candidate video frame group, and the target first video frame is a first video frame of the video frame pair to which the target second video frame belongs.

[0091] For example, the server determines a matching time difference between the first candidate video frame group and the second candidate video frame group, and if the matching time difference is equal to or less than a matching difference threshold, the server replaces the target second video frame with a reference second video frame according to an appearance time difference corresponding to the second candidate video frame group to obtain the video frame group, the target second video frame is a second video frame newly added in the second candidate video frame group, the reference second video frame is a second video frame of the second video whose appearance time difference with the target first video frame is an appearance time difference corresponding to the second candidate video frame group, and the target first video frame is a first video frame of a video frame pair to which the target second video frame belongs.

[0092] In this type of embodiment, after adding the video frame pair in the first candidate video frame group to the second candidate video frame group, the server further adjusts the video frames newly added to the second candidate video frame group according to the appearance time difference of the second candidate video frame group, so that the appearance time difference of the adjusted video frame pair is the same as that of the second candidate video frame group, and maintains the consistency between the appearance time difference of the video frame pair and the appearance difference of the video frame group.

[0093] For clearer description, the following takes as an example a first candidate video frame group whose corresponding appearance time difference is 3 and includes two video frame pairs of [6,9] and [7,10], a second candidate video frame group whose corresponding appearance time difference is 2 and includes three video frame pairs of [2,4], [3,5] and [4,6], and the matching difference threshold is 3. Because the matching time difference between the first candidate video frame group and the second candidate video frame group is 1, the server determines that the matching time difference is less than the matching difference threshold, and must merge the first candidate video frame group with the second candidate video frame group. The server adds two video frame pairs [6,9] and [7,10] in the first candidate video frame group to the second candidate video frame group, and the second candidate video frame group changes to [2,4], [3,5], [4,6], [6,9], [7,10], and the appearance time difference corresponding to the second candidate video frame group is 2, so the server adjusts the second video frames in the two video frame pairs [6,9] and [7,10] added to the second candidate video frame group according to the appearance time difference 2 to obtain two new video frame pairs [6,8] and [7,9]. After adjusting the second video frames newly added to the second candidate video frame group, the second candidate video frame group changes to [2,4], [3,5], [4,6], [6,8], [7,9], and the appearance time difference of each video frame pair is 2.

[0094] Note that the above describes an example in which the server adds a video frame in a first candidate video frame group to a second candidate video frame group, but in other possible embodiments, the server can also add a video frame pair in a second candidate video frame group to the first candidate video frame group.

[0095] In some embodiments, the server determines whether to add the video frame pairs in the first candidate video frame group to the second candidate video frame group or to add the video frame pairs in the second candidate video frame group to the first candidate video frame group based on the quantity of video frame pairs in the first candidate video frame group and the second candidate video frame group. For example, if the quantity of video frame pairs in the first candidate video frame group is greater than the quantity of video frame pairs in the second candidate video frame group, the server adds the video frame pairs in the second candidate video frame group to the first candidate video frame group. If the quantity of video frame pairs in the second candidate video frame group is greater than the quantity of video frame pairs in the first candidate video frame group, the server adds the video frame pairs in the first candidate video frame group to the second candidate video frame group. If the quantity of video frame pairs in the second candidate video frame group is equal to the quantity of video frame pairs in the first candidate video frame group, the server adds the video frame pairs in the first candidate video frame group to the second candidate video frame group. Alternatively, if the quantity of video frame pairs in the second complementary video frame group is equal to the quantity of video frame pairs in the first candidate video frame group, the server adds the video frame pairs in the second candidate video frame group to the first candidate video frame group.

[0096] In this case, the server determines a method for merging the candidate video frame groups according to the number of video frame pairs in the candidate video frame groups, and adds the candidate video frame group containing a smaller number of video frames to the video frame group containing a larger number of video frames, thereby reducing the amount of calculation and improving efficiency.

[0097] 305, for any one of the plurality of video frame groups, the server merges a first video frame of a video frame pair in the video frame group into one candidate video clip according to an appearance time of the first video frame of the video frame pair in the video frame group in the first video.

[0098] In one possible embodiment, the server compares the appearance times of the first video frames of any two adjacent video frame pairs in the video frame group to obtain an appearance time difference between the two adjacent video frame pairs. If the numerical difference between the appearance times of the first video frames of the two adjacent video frame pairs in the first video meets an appearance time condition, the server adds the two adjacent video frame pairs to a temporary frame list. If the numerical difference between the appearance times of the first video frames of the two adjacent video frame pairs in the first video does not meet an appearance time condition, the server merges the video frames in the temporary frame list into one reference video clip. The server determines the at least one candidate video clip based on multiple reference video clips.

[0099] The temporary frame list is used to store the video frame pairs whose numerical difference between the appearance times meets the appearance time condition. In some embodiments, the numerical difference between the appearance times meets the appearance time condition refers to the numerical difference between the appearance times being equal to or less than the appearance time difference threshold, where the appearance time difference threshold is set by the engineer according to the actual situation, for example, 8s, and is not limited thereto in the present embodiment.

[0100] In order to describe the above embodiment more clearly, the following description of the above embodiment is divided into four parts.

[0101] First, the server compares the occurrence times in the first video of a first video frame of any pair of two adjacent video frames in the set of video frames.

[0102] In some embodiments, the server may use the appearance time of the first video frame in the first video as the identifier of the first video frame, and the appearance time of the second video frame in the second video as the identifier of the second video frame, in which case the server only needs to compare the appearance times of the first video frames of any two adjacent video frame pairs in the first video by comparing the identifiers of the two first video frames. For example, if the video frame set includes video frame pairs [2,4], [3,5], [4,6], [6,8], and [7,9], the server may sequentially compare the appearance times of the first video frames of the video frame pairs in the first video. In the first comparison process, the server may compare the appearance times of the first video frame 2 of the first video frame pair [2,4] and the first video frame 3 of the second video frame pair [3,5] in the first video.

[0103] Second part: if a numerical difference between the appearance times in the first video of the first video frames of the two adjacent video frame pairs meets an appearance time condition, the server adds the two adjacent video frame pairs to a temporary frame list.

[0104] In one possible embodiment, if the numerical difference between the appearance times of the first video frames of the two adjacent video frames in the first video is less than or equal to the appearance time difference threshold, the server adds the two adjacent video frames to a temporary frame list. For example, further taking the video frame pairs [2,4], [3,5], [4,6], [6,8], and [7,9] of the video frame group as an example, for video frame pairs [2,4] and [3,5], if the appearance time difference threshold is 3, the appearance time difference of the first video frames in [2,4] and [3,5] in the first video is 3-2=1, so the server adds the two video frame pairs to a temporary frame list (Tmplist), Tmplist=[[2,4], [3,5]].

[0105] The server's decision to add the video frame pair to the temporary frame list includes multiple iterations, and in any one iteration, the server compares the appearance time difference in the first video of the first video frame of the current video frame pair and the first video frame of the previous video frame pair, where the current video frame pair refers to the video frame pair currently being processed, and the previous video frame pair refers to the video frame pair processed in the previous iteration. For example, after the server adds the video frame pair [2,4] and [3,5] to the temporary frame list, the server further determines the relationship between the appearance time difference in the first video of the first video frames of the video frame pair [3,5] and [4,6] and the appearance time difference threshold value, and the appearance time difference in the first video of the first video frames of [3,5] and [4,6] is 4-3=1, so the server adds the video frame pair [4,6] to the temporary frame list (Tmplist), and Tmplist=[[2,4], [3,5], [4,6]]. After multiple iterations, a temporary frame list Tmplist = [[2,4], [3,5], [4,6], [6,8], [7,9]] is obtained.

[0106] Third part: if the numerical difference between the appearance times in the first video of the first video frames of the two adjacent video frame pairs does not meet an appearance time condition, the server merges the video frame pair in the temporary frame list into a reference video clip.

[0107] The reference video clip includes a first sub-clip and a second sub-clip, where the first sub-clip is composed of a first video frame of the video frame pair and the second sub-clip is composed of a second video frame of the video frame pair.

[0108] In one possible embodiment, if the numerical difference between the appearance times in the first video of the first video frames of the two adjacent video frame pairs is greater than an appearance time difference threshold, the server merges the first video frame in the temporary frame list into a first sub-clip, and merges the second video frame in the temporary frame list into a second sub-clip, and the first sub-clip and the second sub-clip constitute the reference video clip. Since the first video frame and the second video frame of the video frame pair are video frames with a relatively high similarity, the first sub-clip and the second sub-clip are also clips with a relatively high similarity. For example, referring to FIG. 5, the format of a first sub-clip 501 and a second sub-clip 502 is shown, in which the first video frame at the beginning of the first sub-clip 501 and the first video frame at the beginning of the second sub-clip 502 constitute one video frame pair, and the first video frame at the end of the first sub-clip 501 and the first video frame at the end of the second sub-clip 502 constitute another video frame pair. In some embodiments, the first sub-clip and the second sub-clip within a reference video clip are also referred to as matching sections.

[0109] For example, if the two adjacent video frame pairs are [9,11] and [2,4], the numerical difference between the appearance times of the first video frames of the two video frame pairs in the first video is 9-2=7, so the server merges the first video frames in the temporary frame list into one reference video clip. For example, if the temporary frame list Tmplist=[[2,4], [3,5], [4,6], [6,8], [7,9]], the server merges the first video frames [2,], [3,], [4,], [6,], [7, ] is merged into the first subclip (2,7), and the second video frames [,4], [,5], [,6], [,8], and [,9] in the temporary frame list are merged into the second subclip (4,9), where the first subclip (2,7) and the second subclip (4,9) constitute the reference video clip (2,7,4,9), which has the format (src-startTime,src-endTime,ref-startTime,ref-endTime), where src-startTime refers to the beginning of the first subclip, i.e., the first video frame with the smallest serial number in the temporary frame list, and src-endTime refers to the end of the first subclip. , ref-startTime points to the beginning of the second subclip, i.e., the first video frame with the highest serial number in the temporary frame list, ref-startTime points to the beginning of the second subclip, i.e., the second video frame with the lowest serial number in the temporary frame list, ref-endTime points to the end of the second subclip, i.e., the second video frame with the highest serial number in the temporary frame list, and the serial number refers to an identifier of a video frame and indicates a position of the video frame in a video, where a smaller serial number indicates an earlier position of the video frame in the video and a smaller serial number indicates a later position of the video frame in the video. In some embodiments, the server stores the reference video clip in a matching section list match-duration-list.When determining the video frame pairs, all video frames in the first and second videos are traversed, which may result in a situation where a video frame is similar to multiple video frames, resulting in time overlap between the two reference video clips in the match-duration-list.

[0110] In some embodiments, traverse the video frames in the video frame group to determine a current video frame pair currently being traversed and a previous video frame pair previously traversed, the current video frame pair and the previous video frame pair being two adjacent video frame pairs in the video frame group, compare the appearance times in the first video of a first video frame of the current video frame pair and the previous video frame pair to obtain a numerical difference in the appearance times of the first video frames, and if the numerical difference in the appearance times of the first video frames meets the appearance time condition, add the current video frame pair and the previous video frame pair to the temporary frame list, and if the numerical difference in the appearance times of the first video frames meets the appearance time condition, add the current video frame pair and the previous video frame pair to the temporary frame list. If not, merge the video frame pair in the temporary frame list into the reference video clip, clear the temporary frame list after merging, determine the next video frame pair to be traversed, and return to the step of comparing the appearance times of the first video frames of the current video frame pair and the previous video frame pair in the first video, and continue to the last video frame pair to be traversed, and if there is a video frame pair in the temporary frame list, merge the video frame pair in the temporary frame list into the reference video clip, and determine at least one candidate video clip based on the multiple reference video clips. The numerical difference in appearance times of the first video frames refers to the numerical difference in appearance times of the first video frames of two adjacent video frame pairs in the video frame group. In some embodiments, the reference video clip can further carry information such as the appearance time difference corresponding to the first sub-clip, the duration of the first sub-clip, and the number of video frames included in the first sub-clip, which is convenient for server utilization.

[0111] In addition to the method provided in the third part above, the present embodiment also provides a method that adopts another method for merging video frame pairs in a temporary frame list with a reference video clip.

[0112] In one possible embodiment, if the currently processed video frame pair is the last video frame pair in the video frame group, the server adds the video frame pair to a temporary frame list, and merges the video frame pair in the temporary frame list with the reference video clip. For example, if the video frame group includes five video frame pairs, i.e., [2,4], [3,5], [4,6], [6,8], and [7,9], and the server processes the video frame pair [7,9], since the video frame pair [7,9] is the last video frame pair in the video frame group, the server adds the video frame pair [7,9] to a temporary frame list, and merges the video frame pair in the temporary frame list with the reference video clip, and the merging process shall refer to the description of the above embodiment, and will not be described in detail here.

[0113] Since video frames with a relatively small appearance time difference may constitute one relatively perfect video clip, by fusing video frames with a relatively small appearance time difference, a relatively perfect reference video can be obtained. Compared with determining a target video clip using miscellaneous video frames, the embodiment of the present application can subsequently easily determine one more perfect target video clip based on the relatively perfect reference video.

[0114] A fourth part, the server determines the at least one candidate video clip based on a plurality of reference video clips.

[0115] The plurality of reference video clips include a first overlay video clip and / or a second overlay video clip, where the first overlay video clip refers to a reference video clip belonging to a first reference video clip in the plurality of reference video clips, and the second overlay video clip refers to a reference video clip partially overlaid with a second reference video clip in the plurality of reference video clips.

[0116] When a first overlay video clip belongs to the first reference video clip, it means that the content of the first overlay video clip is completely contained in the first reference video clip, or that the first reference video clip completely contains the first overlay video clip.

[0117] In order to more clearly explain the content of the above fourth part, the following describes how the server determines the first overlay video clip from multiple reference video clips.

[0118] In one possible embodiment, the server determines a first overlay video clip from the plurality of reference video clips based on an occurrence time in a first video of a first sub-clip in the plurality of reference video clips.

[0119] The first sub-clip is a video clip constituted by a first video frame, and the appearance time includes the start time and end time of the first sub-clip in the first video.

[0120] For example, reference video clip A in the plurality of reference video clips 1 and Reference Video Clip B 1 The server receives the reference video clip A 1 The appearance time of the first sub-clip in the first video and the reference video clip B 1 and comparing the appearance time of the first sub-clip in the first video with that of the reference video clip B. 1 The appearance time of the first sub-clip in the first video of the reference video clip A 1 If the first sub-clip B has a first video, the first sub-clip B has a first video. 1 For example, referring to FIG. 6, the plurality of reference video clips include reference video clip A. 1 and Reference Video Clip B 1 The server references the video clip A. 1 First subclip of m 1The appearance time in the first video and the reference video clip B 1 First subclip of n 1 The appearance time of the first sub-clip n in the first video is compared with that of the first sub-clip n. 1 The start time of the first subclip is m 1 and the first sub-clip n 1 The end time of the first subclip is m 1 If it is before the reference video clip B, the server 1 is determined as the first overlay video clip, and the reference video clip A 1 This is the first reference video clip.

[0121] After describing how the server determines a first overlay video clip from a plurality of reference video clips, the following describes how the server determines a second overlay video clip from a plurality of reference video clips.

[0122] In one possible embodiment, the server determines a second overlay video clip from the plurality of reference video clips based on an occurrence time in a first video of a first sub-clip in the plurality of reference video clips.

[0123] For example, reference video clip A in the plurality of reference video clips 2 and Reference Video Clip B 2 The server receives the reference video clip A 2 The appearance time of the first sub-clip in the first video and the reference video clip B 2 The appearance time of the first sub-clip of the reference video clip B in the first video is compared with the appearance time of the first sub-clip of the reference video clip A in the first video, and if there is an intersection between the appearance time of the first sub-clip of the reference video clip B in the first video and the appearance time of the first sub-clip of the reference video clip A in the first video, the reference video clip A or the reference video clip B, which has a shorter duration, is determined as the second overlay video clip. 2and Reference Video Clip B 2 The server references the video clip A. 2 First subclip of m 2 The appearance time in the first video and the reference video clip B 2 First subclip of n 2 The appearance time of the first sub-clip n in the first video is compared with that of the first sub-clip n. 2 The start time of the first subclip is m 2 After the start time and before the end time of the first sub-clip n 2 The end time of the first subclip is m 2 or after the first subclip n 2 The start time of the first subclip is m 2 before the first sub-clip n 2 The end time of the first subclip is m 2 before the end time and after the start time of reference video clip B 2 The duration of reference video clip A 2 If the reference video clip B is smaller than 2 is determined to be the second overlay video clip, and the reference video clip A 2 This is the second reference video clip mentioned above.

[0124] After the introduction of the method in which the server determines the first overlay video clip and the second overlay video clip is completed, the steps provided by the above fourth part will be described below.

[0125] In one possible embodiment, if the first overlay video clip is included in the plurality of reference video clips, the server removes the first overlay video clip to obtain the at least one candidate video clip.

[0126] In this kind of embodiment, the server can remove the duplicated first overlay video clips from the multiple reference video clips to reduce the quantity of the obtained candidate video clips, so that the amount of calculation is reduced and the calculation efficiency is improved.

[0127] In one possible embodiment, if the second overlay video clip is included in the multiple reference video clips, the server removes the overlapping portion between the second overlay video clip and the second reference video clip to obtain the at least one candidate video clip.

[0128] In this kind of embodiment, the server can remove the overlap portion between the second overlay video clip and the second reference clip to reduce the length of the resulting candidate video clip, thereby reducing the amount of computation and improving computation efficiency.

[0129] Based on the above embodiment, optionally, the server can further perform the following steps:

[0130] In some embodiments, after deleting the overlapping portion between the second overlay video clip and the second reference clip, the server compares the duration of the third-class reference video clip with a target duration, and the third-class reference video clip refers to the second overlay video clip from which the overlapping portion has been deleted. If the duration of the third-class reference video clip is equal to or greater than the target duration, the server reserves the third-class reference video clip. If the duration of the third-class reference video clip is less than the target duration, the server deletes the third-class reference video clip.

[0131] The target duration can be set by the engineer according to the actual situation, and is not limited in the embodiment of the present application. If the server reserves the third type reference video clip, the third type reference video clip is used to replace the original second overlay video clip.

[0132] The above embodiment will be described below through two examples.

[0133] Example 1: Reference video clip A among the multiple reference video clips 2 and Reference Video Clip B 2 For the reference video clip A 2 First subclip of m2 and the reference video clip B 2 First subclip of n 2 has partial overlap, and the first sub-clip m 2 The start time of the first subclip is n 2 If it is earlier than n, the server 2 Set the start time of the first subclip m 2 Move to the end of the subclip 1 and obtain the sub-clip l 1 is the first sub-clip of the third type reference video clip. 1 If the duration of the sub-clip is equal to or less than the target duration, the server 1 and at the same time delete the subclip 1 Delete the third-class reference video clip to which the sub-clip belongs. 1 If the duration of the sub-clip is longer than the target duration, the server 1 At the same time, the sub-clip 1 This section reserves the right to use Class 3 reference video clips to which the above definition applies.

[0134] Example 2: Reference video clip A among the multiple reference video clips 2 and Reference Video Clip B 2 For the reference video clip A 2 First subclip of m 2 and the reference video clip B 2 First subclip of n 2 has partial overlap, and the first sub-clip n 2 The start time of the first subclip is m 2 If it is earlier than n, the server 2 Set the end time of the first subclip m 2 Go to the start time of the subclip 2 and obtain the sub-clip l 2 is the first sub-clip of the third type reference video clip. 2 If the duration of the sub-clip is equal to or less than the target duration, the server 2and at the same time delete the subclip 2 Delete the third-class reference video clip to which the sub-clip belongs. 2 If the duration of the sub-clip is longer than the target duration, the server 2 At the same time, the sub-clip 2 This section reserves the right to use Class 3 reference video clips to which the above definition applies.

[0135] If the duration of the third type reference video clip is shorter than the target duration, it can be recognized that the third type reference video clip contains a relatively small number of video frames and is likely to be an erroneously generated reference video clip. Therefore, by deleting the reference video clip, the accuracy of the target video clip generated based on the subsequent surplus reference video clips can be improved.

[0136] 306, the server determines, based on the at least one candidate video clip, at least one target candidate video clip, where the target candidate video clip has an appearance frequency in the at least one candidate video clip that meets a frequency condition.

[0137] In one possible embodiment, the server determines at least one reference candidate video clip according to the at least one candidate video clip, and the server determines the number of occurrences of each reference candidate video clip in the at least one reference candidate video clip, and the server determines the reference candidate video clip whose number of occurrences meets the number of occurrences condition as the target candidate video clip.

[0138] The occurrence number of the reference candidate video clip in the at least one reference candidate video clip refers to the quantity of the reference candidate video clip in the at least one reference candidate video clip. For example, when the at least one reference candidate video clip is 1, 2, 3, 1, 4, 5, when referring to the reference candidate video clip 1, the occurrence number is 2.

[0139] In order to explain the above embodiment, the following description will be divided into three parts.

[0140] First, the server determines at least one reference candidate video clip based on the at least one candidate video clip.

[0141] The at least one candidate video clip includes a third overlay video clip and / or a fourth overlay video clip, where the third overlay video clip refers to a candidate video clip belonging to a first candidate video clip in the at least one candidate video clip, and the fourth overlay video clip refers to a candidate video clip that partially overlays with a second candidate video clip in the at least one candidate video clip.

[0142] In order to more clearly explain the contents of the above first part, the following describes how the server determines the third overlay video clip from at least one candidate video clip.

[0143] In one possible embodiment, the server determines a third overlay video clip from the at least one candidate video clip based on an occurrence time in the first video of a first sub-clip in the at least one candidate video clip.

[0144] The candidate video clips include a first sub-clip and a second sub-clip, where the first sub-clip is composed of a first video frame in the video frame pair and the second sub-clip is composed of a second video frame in the video frame pair.

[0145] For example, if the at least one candidate video clip is two candidate video clips, then candidate video clip C 1 and candidate video clip D 1 The server then selects the candidate video clip C 1The appearance time of the first sub-clip in the first video of the candidate video clip D 1 and comparing the appearance time of the first sub-clip of the candidate video clip D 1 The appearance time of the first sub-clip of the candidate video clip C 1 If the candidate video clip D 1 is determined to be the third overlay video clip.

[0146] For example, the at least one candidate video clip is two candidate video clips, and candidate video clip C 1 and candidate video clip D 1 If the candidate video clip C 1 1st subclip of 1 the appearance time in the first video and the candidate video clip D 1 First subclip of p 1 The first sub-clip p 1 The start time of the first subclip 1 After the first sub-clip p 1 The end time of the first subclip 1 If the candidate video clip D 1 is determined as the third overlay video clip, and the candidate video clip C 1 This is the first candidate video clip mentioned above.

[0147] After describing how the server determines the third overlay video clip from the at least one candidate video clip, the following describes how the server determines the fourth overlay video clip from the at least one candidate video clip.

[0148] In one possible embodiment, the server determines a fourth overlay video clip from the at least one candidate video clip based on an occurrence time in a first video of a first sub-clip in the at least one candidate video clip.

[0149] For example, if the at least one candidate video clip is two candidate video clips, then candidate video clip C 2 and candidate video clip D 2 The server then selects the candidate video clip C 2 The appearance time of the first sub-clip in the first video of the candidate video clip D 2 and comparing the appearance time of the first sub-clip of the candidate video clip D 2 The appearance time of the first sub-clip of the candidate video clip C 2 If there is an intersection between the appearance time of the first sub-clip of the candidate video clip C 2 and candidate video clip D 2 The candidate video clip having the shorter duration is determined as the fourth overlay video clip.

[0150] For example, the at least one candidate video clip is two candidate video clips, and candidate video clip C 2 and candidate video clip D 2 If the candidate video clip C 2 1st subclip of 2 the appearance time in the first video and the candidate video clip D 2 First subclip of p 2 The first sub-clip p 2 The start time of the first subclip 2 After the start time and before the end time of the first sub-clip p 2 The end time of the first subclip 2 or after the first sub-clip p2 The start time of the first subclip 2 and the first sub-clip p 2 The end time of the first subclip 2 before the end time and after the start time of candidate video clip D 2 The duration of the candidate video clip C 2 If the candidate video clip D 2 is determined as the fourth overlay video clip, and the candidate video clip C 2 This is the second candidate video clip mentioned above.

[0151] After the introduction of the method in which the server determines the third and fourth overlay video clips is completed, the steps provided by the above first part will be described below.

[0152] In one possible embodiment, if the at least one candidate video clip includes the third overlay video clip, the server deletes the third overlay video clip to obtain the at least one reference candidate video clip. In some embodiments, before deleting the third overlay video clip, the server accumulates the number of occurrences of the third overlay video clip in the first candidate video clip. Since the third overlay video clip is completely contained in the first candidate video clip, accumulating the number of occurrences of the third overlay video clip in the first candidate video clip can increase the weight of the first candidate video clip in subsequent processing.

[0153] In this kind of embodiment, if the server removes the overlapping third overlay video clip from at least one candidate video clip, the amount of obtained reference candidate video clips is reduced, the amount of calculation is reduced, and the calculation efficiency is improved.

[0154] A specific example will be given below.

[0155] Candidate video clip D 1 1st subclip of 1The candidate video clip C 1 First subclip of p 1 , and the first subclip o 1 duration of > 0.5 * first subclip p 1 If so, the server 1 At the same time, delete the candidate video clip D 1 Also delete the candidate video clip D 1 The number of occurrences of the candidate video clip C 1 Accumulate to.

[0156] Based on the above embodiment, optionally, before accumulating the number of occurrences of the third overlay video clip in the first candidate video clip, the server can further determine the duration of the third overlay video clip and the duration of the first candidate video clip, and determine whether to accumulate the number of occurrences of the third overlay video clip in the first candidate video clip based on the duration of the third overlay video clip and the duration of the first candidate video clip.

[0157] For example, the server determines the duration of the third overlay video clip and the duration of the first candidate video clip, the server determines a first comparison value between the duration of the third overlay video clip and the duration of the first candidate video clip, and if the first comparison value is greater than or equal to a comparison value threshold, the server accumulates the appearance count of the third overlay video clip in the first candidate video clip, and if the first comparison value is less than the comparison value threshold, the server does not accumulate the appearance count of the third overlay video clip in the first candidate video clip, and the comparison value threshold is set by the engineer according to actual circumstances, for example, 0.5, and is not limited thereto in the embodiment of the present application.

[0158] In one possible embodiment, if the at least one candidate video clip includes the fourth overlay video clip, and the overlapping degree between the fourth overlay video clip and the second candidate video clip meets an overlapping degree condition, the server determines the number of occurrences of the fourth overlay video clip. The server determines the at least one reference candidate video clip based on the number of occurrences of each fourth overlay video clip whose overlapping degree meets the overlapping degree condition.

[0159] The degree of overlap refers to a comparison value between the duration of the overlapped video clip and the duration of the video clip to be compared. For example, for the fourth overlaid video clip and the second candidate video clip, if the second candidate video clip is the video clip to be compared, the degree of overlap between the fourth overlaid video clip and the second candidate video clip can be determined by dividing the duration of the overlapped video clip between the fourth overlaid video clip and the second candidate video clip by the duration of the second candidate video clip. The degree of overlap meets the degree of overlap condition means that the degree of overlap is equal to or greater than the degree of overlap threshold.

[0160] In the following, two embodiments will be described to explain how the server in the above embodiment determines the at least one reference candidate video clip based on the appearance frequency of the fourth overlay video clip.

[0161] In the first embodiment, if the occurrence number of the fourth overlay video clip is equal to or greater than the first occurrence number threshold, the server merges the fourth overlay video clip with the second candidate video clip to obtain the at least one reference candidate video clip. In some embodiments, the fourth overlay video clips, each of which has an overlap degree that meets the overlap degree condition, are merged with the corresponding second candidate video clip to obtain the at least one reference candidate video clip. In some embodiments, before the fourth overlay video clip is merged with the second candidate video clip, the server accumulates the occurrence number of the fourth overlay video clip with the second candidate video clip.

[0162] The first occurrence number threshold is set by a technician according to actual circumstances, for example, set to 3, and is not limited thereto in the embodiment of the present application. If the occurrence number is equal to or greater than the first occurrence number threshold, it indicates that the fourth overlay video clip cannot be ignored, and needs to be further processed to improve the accuracy of the acquired target video clip.

[0163] The following describes how the server in the above embodiment merges the fourth overlay video clip with the second candidate video clip.

[0164] In some embodiments, when the duration of the fourth overlay video clip is shorter than that of the second candidate video clip, the server deletes the overlapping portion between the fourth overlay video clip and the second candidate video clip, and adds the redundant portion onto the second candidate video clip to obtain one candidate video clip. For example, referring to FIG. 7, the duration of the fourth overlay video clip 701 is shorter than that of the second candidate video clip 702, and the duration of the fourth overlay video clip 704 is also shorter than that of the second candidate video clip 705. If the end time of the fourth overlay video clip 701 is later than that of the second candidate video clip 702, the server merges the fourth overlay video clip 701 and the second candidate video clip 702 to obtain one candidate video clip 703. If the start time of the fourth overlay video clip 704 is earlier than the start time of the second candidate video clip 705 , the server merges the fourth overlay video clip 704 and the second candidate video clip 705 to obtain one candidate video clip 706 .

[0165] By fusing the fourth overlay video clip with the second candidate video clip, the number of video clips can be reduced, so that the amount of calculations is reduced and the calculation efficiency is improved.

[0166] In embodiment 2, if the appearance count of the fourth overlay video clip is less than the first appearance count threshold, the server deletes the fourth overlay video clip to obtain at least one reference candidate video clip, and the server accumulates the appearance count of the fourth overlay video clip into the second candidate video clip.

[0167] If the number of occurrences is less than the first occurrence threshold, this indicates that the fourth overlay video clip may be ignored, and the server need only delete the fourth overlay video clip.

[0168] By deleting a part of the fourth overlay video clips, the number of video clips can be reduced, the amount of calculations is reduced, and the calculation efficiency is improved.

[0169] In one possible embodiment, when the at least one candidate video clip includes the fourth overlay video clip, and the overlapping degree between the fourth overlay video clip and the second candidate video clip does not meet the overlapping degree condition, the server deletes the fourth overlay video clip to obtain the at least one reference candidate video clip. In some embodiments, before deleting the fourth overlay video clip, the server accumulates the occurrence number of the fourth overlay video clip to the second candidate video clip.

[0170] In one possible embodiment, if the at least one candidate video clip includes the fourth overlay video clip, and the duration of the fourth overlay video clip is less than that of the second candidate video clip, the server deletes the fourth overlay video clip to obtain the at least one reference candidate video clip. In some embodiments, before deleting the fourth overlay video clip, the server accumulates the occurrence number of the fourth overlay video clip to the second candidate video clip.

[0171] In some embodiments, at least one reference candidate video clip is stored in a match-list by the server and utilized.

[0172] By deleting the fourth overlay video clip whose overlap rate does not meet the overlap rate condition or whose duration is less than that of the second candidate video clip, the number of video clips can be reduced, the amount of calculations in subsequent processes can be reduced, and calculation efficiency can be improved.

[0173] In a second part, the server determines the number of occurrences of the reference candidate video clip in the at least one reference candidate video clip.

[0174] Through the above-mentioned first processing step, the server determines at least one reference candidate video clip based on at least one candidate video clip, and in the determination process, the server merges and deletes related occurrence counts, and the server re-determines the occurrence counts of the at least one reference candidate video clip. In some embodiments, the server can store the occurrence counts of the at least one reference candidate video clip in an occurrence count list for use.

[0175] For example, when determining a target video clip in a first video, the server adopts three second videos for mining, and for the sake of simplicity, the first video is named i, and the three second videos are named vid1, vid2, and vid3. After adopting the above steps, the server determines two candidate video clips [(2,7,4,9), (10,11,11,12)] based on the first video i and the second video vid1, determines one candidate video clip [(2,7,4,9)] based on the first video i and the second video vid2, and determines one candidate video clip [(2,7,4,10)] based on the first video i and the second video vid3. The server takes statistics of the four candidate video clips and determines that the number of occurrences of the candidate video clip (2,7,4,9) is 2, the number of occurrences of the candidate video clip (2,7,4,10) is 1, and the number of occurrences of the candidate video clip (10,11,11,12) is 1. After fusing the four candidate video clips according to the above first method, two reference candidate video clips [(2,7,4,9), (10,11,11,12)] are obtained, and the number of occurrences of the reference candidate video clip (2,7,4,9) is 3, and the number of occurrences of the reference candidate video clip (10,11,11,12) is 1, which is stored in the count list (count-list) as count-list=[3,1].

[0176] Third, the server determines the reference candidate video clip whose occurrence count meets the occurrence count condition as the target reference candidate video clip.

[0177] In one possible embodiment, the server determines the reference candidate video clips whose occurrence count is equal to or greater than a second occurrence count threshold as target candidate video clips.

[0178] The second occurrence threshold is positively correlated with the quantity of the at least one reference candidate video clip, i.e., the more the quantity of the at least one reference candidate video clip, the higher the second occurrence threshold, and the less the quantity of the at least one reference candidate video clip, the lower the second occurrence threshold. In some embodiments, the second occurrence threshold is a product of a target comparison value and the quantity of the at least one reference candidate video clip, and the target comparison value is a positive number less than 1.

[0179] For example, if the two obtained reference candidate video clips are [(2,7,4,9), (10,11,11,12)], and the occurrence count of the reference candidate video clip (2,7,4,9) is 3, the occurrence count of the reference candidate video clip (10,11,11,12) is 1, and the second occurrence count threshold is 3, the server deletes the reference candidate video clip (10,11,11,12), and finally reserves the reference candidate video clip (2,7,4,9) and the occurrence count of 3. When this is stored in the matching list (match-list) and the count list (count-list), match-list=(2,7,4,9) and count-list=[3].

[0180] 307, for any one target candidate video clip, if the appearance time of the target candidate video clip in the first video is within a target time range, the server determines the target candidate video clip as a target video clip in the first video.

[0181] The target time range is set by the engineer according to the actual situation, for example, when the technical solution provided in the embodiment of the present application is applied to the scene of identifying the opening and ending of a video, the target time range is the time range in which the opening and ending of the video may exist, in which the target time range includes a first time range and a second time range, the first time range is the range in which the opening may exist, and the second time range is the range in which the ending may exist. For example, if the first 1 / 5 of the time of the video is set as the opening time, that is, the first time range, and the last 1 / 5 of the time is set as the ending time, that is, the second time range, then for a 10-minute video, the opening will probably only appear in the first 2 minutes, and the ending will only appear in the last 2 minutes. 1 / 5 is set by the engineer according to the actual situation, and can be adjusted accordingly for different types of videos, for example, 1 / 5 can be adopted for a children's animation of about 15 minutes, and 1 / 8 can be adopted for a 45-minute TV drama.

[0182] Note that the above steps 301 to 307 are described using an example in which the server determines a target video clip for a first video, but if the first video and at least one second video belong to the same video set, the server can employ a method similar to the above steps 301 to 307 to determine target video clips for other videos in the video set, where the other videos refer to videos other than the first video.

[0183] The technical solution provided by the embodiment of the present application will be described below in combination with FIG.

[0184] Referring to Fig. 8, in the present embodiment, the server performs matching based on the similarity between video frame features to obtain a plurality of video frame pairs. The server divides the plurality of video frame pairs into a plurality of initial video frame groups based on the appearance time difference. The server merges the plurality of initial video frame groups into a plurality of candidate video frame groups based on the appearance time difference. The server merges the plurality of candidate video frame groups into a plurality of video frame groups. The server outputs a target video clip of the first video based on the plurality of video frame groups.

[0185] In some embodiments, the above steps 301-307 can be implemented by a clip mining system, and when the technical solution provided by the embodiments of the present application is applied to the scene of identifying the opening and ending of a video, the clip mining system is an opening and ending mining system. Referring to FIG. 9, the video clip mining system is provided with the following functions, namely, a function of extracting video frame features of a plurality of videos, a function of forming a video pair for each video by the video and another video among the plurality of videos, a function of performing matching based on the plurality of video pairs to obtain a plurality of video frame pairs, a function of fusing the plurality of video frame pairs to obtain a plurality of video frame groups, a function of determining the position of a target video clip in the video based on the plurality of video frame groups, and a function of obtaining the target video clip based on the position of the target video clip in the video. When the technical solution provided by the embodiments of the present application is applied to the scene of identifying the opening and ending of a video, the target video clip is the opening or ending of the video.

[0186] Referring to Figure 10, when the technical solution provided by the embodiments of the present application is applied to identifying the opening and ending scenes of a TV drama, a TV drama is obtained, and the TV drama includes a plurality of videos. The plurality of videos are input into a clip mining system, and the openings and endings of the plurality of videos are output through the clip mining system. In some embodiments, the clip mining system can output the timestamps of the openings and endings of the plurality of videos.

[0187] 308, the server stores the target video clip of the first video in the clip database.

[0188] In one possible embodiment, the server performs feature extraction on the target video clip of the first video to obtain video frame features of the target video clip. The server stores the video frame features of the target video clip in the clip database. In some embodiments, the server associates the video frame features of the target video clip with the first video. For example, the server sets an identifier of the video frame features of the target video clip as an identifier of the first video. If the first video belongs to a video set, the server associates the identifier of the first video as an identifier of the video set, which facilitates subsequent query processes.

[0189] Performing feature extraction on the target video clip to obtain the video frame features of the target video clip belongs to the same inventive concept as step 301 above, and the implementation process can refer to the description of step 301 above, so it will not be described in detail here.

[0190] For example, if the target video clip is (2,7), the server obtains a target video clip corresponding to 2 to 7 seconds from the first video, and extracts a number of reference video frames from the target video clip. The server performs feature extraction on the number of reference video frames to obtain video frame features of the number of reference video frames. The server stores the video frame features of the number of reference video frames in a clip database. The server associates the video frame features of the number of reference video frames with an identifier Vid1 of the first video, and associates the identifier Vid1 of the first video with an identifier Cid1 of a video set to which the first video belongs. FIG. 11 shows a storage format of the clip database. Referring to FIG. 11, in a database 1100, em1 to emN are video frame features, vid1 to vidK are identifiers of different videos, and N and K are both positive integers.

[0191] After the server stores the target video clip of the first video in the clip database, the server can further use the clip database to search for the video clip, and the method is as follows.

[0192] In one possible embodiment, the server performs feature extraction on a plurality of target video frames of the target video to be identified to obtain video frame features of the plurality of target video frames, and the server determines at least one target video clip of the target video based on the video frame features of the plurality of target video frames, the video frame features of the first video frame, and the video frame features of the at least one second video.

[0193] The process of the server performing feature extraction on the target video frames of the target video to obtain the video frame features of the target video frames belongs to the same inventive concept as the above step 301, and the implementation process can refer to the description of the above step 301, and therefore will not be described in detail here. The process of the server determining at least one target video clip of the target video based on the video frame features of the target video frames, the video frame features of the first video frame, and the video frame features of the at least one second video belongs to the same inventive concept as the above steps 302-307, and the implementation process can refer to the description of the above steps 302-307, and therefore will not be described in detail here. In some embodiments, performing the search for video clips in the clip database is realized by a video search system. In some embodiments, the video frame features of the first video frame and the video frame features of the at least one second video are stored in a clip database.

[0194] By designing a time domain matching algorithm, a similar video section matching method based on image embedding features is realized, which supports matching of similar video sections with length changes (embodied in the matching logic, when matching frames are merged in the time domain under the same occurrence time difference, it is not required that the merged frames are consecutive one after the other) and position changes (embodied in the matching logic, when the occurrence time difference is 0, there is no change in the position, and when the occurrence time difference is greater than 0, there is a change in the position). The method is less time-consuming and has good performance.

[0195] The opening and ending mining scheme generated based on the method of matching video time domain can realize the identification and location of complex video openings and endings with complex changes in length and position, and can solve difficult situations that cannot be solved by existing technical schemes.

[0196] By combining the opening-ending search scheme based on time domain matching, a real-time (within 10 minutes) opening-ending mining scheme can be realized, which is excellent in application.

[0197] The above-mentioned video clip identification method can be applied to scenes of identifying the opening and ending of a video clip, and can also be applied to scenes of identifying infringing videos. In the following, these two application scenes will be introduced respectively.

[0198] When the video clip retrieval method is applied to a scene for retrieving the opening or ending of a video clip, the target video to be identified is input to the video retrieval system, and the video retrieval system performs feature extraction on the target video to obtain video frame features of the target video frames. The video retrieval system performs matching in a clip database based on the video frame features of the target video frames to obtain a target video clip of the target video, which is the opening or ending of the target video.

[0199] Taking the case of identifying the opening and ending of a newly updated video in a TV drama as an example, if the TV drama has already been updated with 10 episodes, the opening and ending of the 10th episode are obtained by the above steps 301 to 307, and the opening and ending of the 10th episode are stored in the clip database by the above step 308. When the 11th episode of the TV drama is updated, the 11th episode is input as the target video to the video retrieval system, and the video retrieval system performs feature extraction on the target video to obtain video frame features of the multiple target video frames. Through the video retrieval system, matching is performed in the clip database based on the video frame features of the multiple target video frames, and a target video clip in the target video is obtained, and the target video clip is the opening or ending of the target video. When the video frame features are associated with the identifier of a video and the identifier of a video set in the clip database, matching can be performed within a finite range based on the identifier of the video set to improve the efficiency of determining the target video clip, and the video set is the TV drama.

[0200] Further explanation will be given below in combination with FIG.

[0201] A TV drama to be identified for opening and ending is determined, and a number of videos in the TV drama are obtained. The number of videos are input to a clip mining system 1201, and the clip mining system 1201 outputs the openings and endings of the number of videos, and the openings and endings of the number of videos are stored in a clip database 1202. If the TV drama updates a target video, the target video is input to a video search system 1203, and the video search system 1203 adopts the target video to perform a search in the clip database 1202 to obtain the opening and ending of the target video. In the technical solution provided by the embodiment of the present application, when mining openings and endings for videos in the same video set, a method of searching the same time region of the video is adopted, that is, for the same video set, the same video clips are found through search and chronological positioning, and are used as the mined openings and endings. Cross-duplicate elimination refers to finding duplicated video clips from videos in a video set through cross-search. The purpose of de-duplication searching of videos is to search for video clips that are identical to the stored video in the first video.

[0202] Furthermore, one video may have multiple openings or endings that meet the above requirements, which is a normal situation; however, for a TV drama with the type of opening song + main story highlights + the same ad insertion + main story, the opening song and ad insertion can be matched in multiple videos, but the highlights are different for each episode and cannot be matched, resulting in the appearance of two openings.

[0203] When the video clip retrieval method is applied to the identification scene of the infringing video, the target video to be identified is input to the video retrieval system, and the video retrieval system performs feature extraction on the target video to obtain the video frame features of the target video frames. The target video is the video to be identified. The video retrieval system performs matching in a clip database based on the video frame features of the target video frames to obtain a target video clip of the target video, and the target video clip is the opening or ending of the target video. The target video clip is deleted from the target video, and infringement identification is performed based on the target video after the target video clip is deleted, and the purpose of the infringement identification is to determine whether the target video after the target video clip is deleted is the same as the content of the specified video. The infringement identification is realized by the infringement identification system, which performs duplicate elimination in the infringement protection video database for the query video, and if duplicates are found, it indicates infringement. However, only the main content needs to be protected, and the opening and ending sequences of ordinary movies and TV dramas are not within the scope of infringement duplication elimination. Therefore, by adopting the technical solution provided in the embodiments of the present application, it is possible to realize the identification of the opening and ending sequences of movies and TV dramas.

[0204] Further explanation will be given below in combination with FIG.

[0205] A television drama to be identified as a target of infringement is determined, a number of videos in the television drama are obtained, and the number of videos are stored in an infringement protection video database 1301. The number of videos are input into a clip mining system 1302, and the openings and endings of the number of videos are output by the clip mining system 1302, and the openings and endings of the number of videos are stored in a clip database 1303. If it is necessary to perform infringement identification on a target video, the target video is input into a video retrieval system 1304, and the video retrieval system 1304 employs the target video to perform a search in the clip database 1303 to obtain the opening and ending of the target video. The opening and ending of the target video are deleted, and an infringement result of the target video is output through an infringement identification system 1305, and the infringement result includes infringement and non-infringement.

[0206] In some embodiments, after performing a query in the clip database for a target video based on the above method, if multiple target video clips of the target video are obtained, the server determines the longest target video clip among the multiple target video clips as the final target video clip; when the technical solution provided by the embodiments of the present application is applied to identifying the opening and ending of a video, the target video clip is the opening and ending of the target video, and the process is as shown in FIG. 14.

[0207] In addition, the video search system and the clip mining system can simultaneously provide an external interface, i.e., a search database storage and a mining database storage, and simultaneously open specific functions that the user specifies as necessary to use. Alternatively, they can provide only one identification interface, and determine whether to perform search or mining depending on whether the backend already has the opening and ending of the TV drama corresponding to the video identifier in the database, and launch the specific functions that the backend should use, including search and mining.

[0208] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described in detail here.

[0209] Through the technical solution provided by the embodiment of the present application, a pair of video frames including similar video frames is determined based on the similarity between the video frame features. A first video frame of the pair of video frames is fused based on the appearance time difference to obtain at least one candidate video clip. Finally, a target video clip within a target time range is determined from the at least one candidate video clip. The process of determining the target clip does not require human intervention, and can be automatically performed by a computer device directly based on the first video and at least one second video, so it is efficient.

[0210] Through the above video section matching algorithm design, a method for matching similar video clips based on video frame features is realized, which can support matching similar video sections with length changes (embodied in the matching logic, when merging video frame pairs with the same occurrence time difference, the frames to be merged are not required to be consecutive) and position changes (embodied in the matching logic, when the occurrence time difference is 0, there is no change in position, and when the occurrence time difference is greater than 0, there is a change in position). The method is less time-consuming and has good performance.

[0211] FIG. 15 is a structural schematic diagram of a video clip identification device provided in an embodiment of the present application. Referring to FIG. 15, the device includes: a video frame pair determination module 1501, a fusion module 1502 and a target video clip determination module 1503.

[0212] The video frame pair determination module 1501 is configured to determine a plurality of video frame pairs based on video frame features of a first video and at least one video frame feature of a second video, the video frame pairs including a first video frame and a second video frame whose similarity meets a similarity condition, the first video frame belonging to the first video, and the second video frame belonging to the at least one second video.

[0213] The fusion module 1502 is configured to fuse a first video frame of the plurality of video frame pairs to obtain at least one candidate video clip of the first video based on an appearance time difference of the plurality of video frame pairs, the appearance time difference referring to a numerical difference between the appearance times in the video of two video frames in the video frame pair.

[0214] The target video clip determination module 1503 is configured to determine at least one target video clip in the first video based on the at least one candidate video clip and a target time range, the target video clip being within the target time range of the first video.

[0215] In one possible embodiment, the fusing module 1502 is configured to divide the plurality of video frame pairs into a plurality of video frame groups based on appearance time differences of the plurality of video frame pairs, where video frame pairs in the same video frame group correspond to the same appearance time difference, and for any one of the plurality of video frame groups, fuse a first video frame of a video frame pair in the video frame group into one of the candidate video clips according to an appearance time of the first video frame of the video frame pair in the video frame group in the first video.

[0216] In one possible embodiment, the fusion module 1502 is configured to, for any one of the multiple video frame pairs, subtract a second appearance time of a second video frame of the video frame pair from a first appearance time of a first video frame of the video frame pair to obtain an appearance time difference of the video frame pair, where the first appearance time refers to the appearance time of the first video frame in the first video, and the second appearance time refers to the appearance time of the second video frame in the second video, and divide the video frame pairs having the same appearance time difference into one initial video frame group, and the appearance time difference of the video frame pairs in the initial video frame group is the appearance time difference corresponding to the initial video frame group. According to the appearance time differences corresponding to the multiple initial video frame groups, the multiple initial video frame groups are fused to obtain the multiple video frame groups.

[0217] In one possible embodiment, the merging module 1502 is configured to sort the initial video frames according to a target order to obtain a plurality of candidate video frames. For any two adjacent candidate video frames in the plurality of candidate video frames, if the matching time difference between the two adjacent candidate video frames meets a matching time difference condition, the two adjacent candidate video frames are fused into one video frame. The matching time difference refers to the numerical difference between the appearance time difference corresponding to the two adjacent candidate video frames.

[0218] In one possible embodiment, the two adjacent candidate video frame groups include a first candidate video frame group and a second candidate video frame group, and the fusion module 1502 is configured to add a video frame pair in the first candidate video frame group to the second candidate video frame group to obtain the video frame group if a matching time difference between an appearance time difference corresponding to the first candidate video frame group and an appearance time difference corresponding to the second candidate video frame group is less than or equal to a matching difference threshold.

[0219] In one possible embodiment, the two adjacent candidate video frames include a first candidate video frame group and a second candidate video frame group, and the merging module 1502 is configured to add a video frame pair in the first candidate video frame group to the second candidate video frame group when a matching time difference between the first candidate video frame group and the second candidate video frame group is equal to or less than a matching difference threshold, and replace a target second video frame with a reference second video frame according to an appearance time difference corresponding to the second candidate video frame group to obtain the video frame group, where the target second video frame is a second video frame newly added in the second candidate video frame group, the reference second video frame is a second video frame in the second video whose appearance time difference with the target first video frame is an appearance time difference corresponding to the second candidate video frame group, and the target first video frame is a first video frame in the video frame pair to which the target second video frame belongs.

[0220] In one possible embodiment, the fusing module 1502 is configured to traverse video frame pairs in the video frame set to determine a current video frame pair currently being traversed and a previous video frame pair previously traversed, the current video frame pair and the previous video frame pair being two adjacent video frame pairs in the video frame set, and compare the appearance times in the first video of the first video frames of the current video frame pair and the previous video frame pair to obtain a numerical difference in the appearance times of the first video frames. If the numerical difference in the appearance time of the first video frames meets the appearance time condition, add the current video frame pair and the previous video frame pair to a temporary frame list; if the numerical difference in the appearance time of the first video frames does not meet the appearance time condition, merge the video frame pair in the temporary frame list into the reference video clip, and clear the temporary frame list after merging, determine the next video frame pair to be traversed, set the next video frame pair to be traversed as a new current video frame pair, and continue to execute until the last video frame pair to be traversed, and determine the at least one candidate video clip based on the multiple reference video clips.

[0221] In one possible embodiment, the plurality of reference video clips includes a first overlay video clip, which refers to a reference video clip belonging to a first reference video clip in the plurality of reference video clips, and the fusion module 1502 is configured to, when the plurality of reference video clips includes the first overlay video clip, delete the first overlay video clip to obtain the at least one candidate video clip.

[0222] In one possible embodiment, the plurality of reference video clips includes a second overlay video clip, which refers to a reference video clip that is partially overlaid with a second reference video clip in the plurality of reference video clips, and the fusion module 1502 is configured to, when the plurality of reference video clips includes the second overlay video clip, remove the overlapping portion between the second overlay video clip and the second reference clip to obtain the at least one candidate video clip.

[0223] In one possible embodiment, the merging module 1502 is further configured to compare the duration of the third type reference video clip with a target duration, where the third type reference video clip refers to the second overlay video clip from which the overlay portion has been deleted, and if the duration of the third type reference video clip is greater than or equal to the target duration, to retain the third type reference video clip, and if the duration of the third type reference video clip is less than the target duration, to delete the third type reference video clip.

[0224] In one possible embodiment, the target video clip determination module 1503 is configured to determine at least one target candidate video clip based on the at least one candidate video clip, where the number of occurrences of the target candidate video clip in the at least one candidate video clip meets a number condition.

[0225] For any one of the target candidate video clips, if the appearance time of the target candidate video clip in the first video is within the target time range, the target candidate video clip is determined as the target video clip in the first video.

[0226] In one possible embodiment, the target video clip determination module 1503 is configured to determine at least one reference candidate video clip based on the at least one candidate video clip, determine the number of occurrences of each reference candidate video clip in the at least one reference candidate video clip, and determine a reference candidate video clip whose number of occurrences meets the occurrence number condition as a target candidate video clip.

[0227] In one possible embodiment, the at least one candidate video clip includes a third overlay video clip, which refers to a candidate video clip belonging to a first candidate video clip in the at least one candidate video clip, and the target video clip determination module 1503 is configured to delete the third overlay video clip to obtain the at least one reference candidate video clip if the at least one candidate video clip includes the third overlay video clip.

[0228] In one possible embodiment, the at least one candidate video clip includes a fourth overlay video clip, which refers to a candidate video clip that is partially overlaid with a second candidate video clip in the at least one candidate video clip, and the target video clip determination module 1503 is configured to determine the number of occurrences of the fourth overlay video clip when the at least one candidate video clip includes the fourth overlay video clip and the degree of overlap between the fourth overlay video clip and the second candidate video clip meets an overlap degree condition, and determine the at least one reference candidate video clip based on the number of occurrences corresponding to each of the fourth overlay video clips whose degree of overlap meets the overlap degree condition.

[0229] In one possible embodiment, the at least one candidate video clip includes a fourth overlay video clip, which refers to a candidate video clip that is partially overlaid with a second candidate video clip in the at least one candidate video clip, and the target video clip determination module 1503 is configured to delete the fourth overlay video clip to obtain the at least one reference candidate video clip when the at least one candidate video clip includes the fourth overlay video clip and the overlap degree between the fourth overlay video clip and the second candidate video clip does not meet the overlap degree condition.

[0230] In one possible embodiment, the at least one candidate video clip includes a fourth overlay video clip, which refers to a candidate video clip that is partially overlaid with a second candidate video clip in the at least one candidate video clip, and the target video clip determination module 1503 is configured to delete the fourth overlay video clip and determine the at least one reference candidate video clip if the at least one candidate video clip includes the fourth overlay video clip and the duration of the fourth overlay video clip is less than that of the second candidate video clip.

[0231] In one possible embodiment, the target video clip determination module 1503 is configured to, for a fourth overlay video clip that meets an overlay condition of any one of the at least one candidate video clip, merge the fourth overlay video clip with a second candidate video clip to obtain the at least one reference candidate video clip if the number of occurrences of the fourth overlay video clip is greater than or equal to a first occurrence threshold.

[0232] In one possible embodiment, the target video clip determination module 1503 is configured to, for a fourth overlay video clip that meets an overlay condition of any one of the at least one candidate video clip, delete the fourth overlay video clip if the number of occurrences of the fourth overlay video clip is less than a first occurrence threshold to obtain the at least one reference candidate video clip.

[0233] In one possible embodiment, the device further comprises: a feature extraction module for performing feature extraction on a plurality of target video frames of a target video to be identified to obtain video frame features of the plurality of target video frames; The target video clip determination module 1503 is further configured to determine at least one target video clip of the target video based on the video frame features of the multiple target video frames, the video frame features of the first video frame, and the video frame features of the at least one second video.

[0234] It should be noted that, when the video clip identification device provided in the above embodiment identifies a video clip, the above-mentioned functional modules are only divided into examples, but in actual application, the above functions can be distributed to be achieved by different functional modules as required, that is, the internal structure of a computer device can be divided into different functional modules to achieve all or part of the above-mentioned functions.In addition, since the video clip identification device provided in the above embodiment belongs to the same inventive concept as the embodiment of the video clip identification method, the details of the specific implementation process are the same as those in the embodiment of the method, and will not be described in detail here.

[0235] According to the technical solution provided by the embodiment of the present application, a video frame pair including similar video frames is determined based on the similarity between the video frame features. A first video frame in the video frame pair is fused based on the appearance time difference to obtain at least one candidate video clip. Finally, a target video clip within a target time range is determined from the at least one candidate video clip. The process of determining the target clip does not require human intervention, and can be automatically performed by a computer device directly based on the first video and at least one second video, which is efficient.

[0236] In the present embodiment, a computer device for executing the above method is provided, and the computer device can be realized as a terminal or a server. The structure of the terminal is introduced below.

[0237] FIG. 16 is a structural schematic diagram of a terminal provided by an embodiment of the present application.

[0238] Typically, the terminal 1600 includes one or more processors 1601 and one or more memories 1602 .

[0239] The processor 1601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1601 may be implemented by adopting at least one of the following hardware formats: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1601 may include a main processor and a coprocessor, where the main processor is a processor for processing data in a wake-up state, also referred to as a CPU (Central Processing Unit), and the coprocessor is a low-power consumption processor for processing data in an idle state. In some embodiments, the processor 1601 may be integrated with a GPU (Graphics Processing Unit), where the GPU is responsible for rendering and creating content to be displayed on a display. In some embodiments, the processor 1601 may further include an AI (Artificial Intelligence) processor, where the AI ​​processor is used to process computational operations related to machine learning.

[0240] The memory 1602 may include one or more computer readable storage media, which may be non-transitory. The memory 1602 may further include high speed random access memory, and non-volatile memory, such as one or more magnetic disk memory devices, flash memory devices, etc. In some embodiments, the non-transitory computer readable storage media in the memory 1602 may be used to store at least one computer program, which may be executed by the processor 1601 to implement the method for identifying video clips provided by the method embodiments of the present application.

[0241] In some embodiments, the terminal 1600 further optionally includes a peripheral interface 1603 and at least one peripheral device. The processor 1601, the memory 1602, and the peripheral interface 1603 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral interface 1603 via a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1604, a display 1605, a camera component 1606, an audio frequency circuit 1607, and a power source 1608.

[0242] The peripheral interface 1603 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1601 and the memory 1602. In some embodiments, the processor 1601, the memory 1602, and the peripheral interface 1603 are integrated on the same chip or circuit board, and in some other embodiments, any one or two of the processor 1601, the memory 1602, and the peripheral interface 1603 can be realized on a single chip or circuit board, which is not limited in this embodiment.

[0243] The radio frequency circuit 1604 is used to receive and transmit RF (Radio Frequency) signals, also referred to as electromagnetic signals. The radio frequency circuit 1604 communicates with communication networks and other communication devices via electromagnetic signals. The radio frequency circuit 1604 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 1604 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, an encoding and decoding chip set, a user ID module card, and the like.

[0244] The display 1605 is used to display a UI (User Interface). The UI may include graphs, text, patterns, videos, and any combination thereof. If the display 1605 is a touch panel, the display 1605 is also capable of collecting touch signals on or above the surface of the display 1605. The touch signals may be input as control signals to the processor 1601 for processing. In this regard, the display 1605 may also be used to provide virtual buttons and / or a virtual keyboard, also referred to as soft buttons and / or a soft keyboard.

[0245] The camera component 1606 is used to collect images and videos, and optionally includes a front camera and a rear camera. Typically, the front camera is mounted on the front panel of the terminal, and the rear camera is mounted on the back of the terminal.

[0246] The audio frequency circuitry 1607 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment and convert the sound waves into electrical signals that can be input to the processor 1601 for processing or input to the radio frequency circuitry 1604 for audio communication.

[0247] The power source 1608 is used to power each component of the terminal 1600. The power source 1608 can be an AC power source, a DC power source, a disposable battery, or a rechargeable battery.

[0248] In some embodiments, terminal 1600 further includes one or more sensors 1609, including, but not limited to, an acceleration sensor 1610, a gyro sensor 1611, a pressure sensor 1612, an optical sensor 1613, and a proximity sensor 1614.

[0249] The acceleration sensor 1610 can detect the magnitude of acceleration on three coordinate axes in a coordinate system established by the terminal 1600 .

[0250] The gyro sensor 1611 can detect the angular velocity of the terminal 1600 in the body direction and rotation direction, and the gyro sensor 1611 can cooperate with the acceleration sensor 1610 to collect the user's 3D movement relative to the terminal 1600.

[0251] The pressure sensor 1612 can be installed on the side edge frame of the terminal 1600 and / or on the lower layer of the display 1605. When the pressure sensor 1612 is installed on the side edge frame of the terminal 1600, it detects a gripping signal of the user on the terminal 1600, and the processor 1601 can distinguish between left and right hands or perform a quick operation according to the gripping signal collected by the pressure sensor 1612. When the pressure sensor 1612 is installed on the lower layer of the display 1605, the processor 1601 realizes control over the operability control material on the UI interface according to the pressure operation on the display 1605 by the user.

[0252] The optical sensor 1613 is used to collect the ambient light intensity. In one embodiment, the processor 1601 can control the display brightness of the display 1605 according to the ambient light intensity collected by the optical sensor 1613.

[0253] The proximity sensor 1614 is used to collect the distance between the user and the front of the terminal 1600 .

[0254] As will be appreciated by those skilled in the art, the structure shown in FIG. 16 does not constitute a limitation on terminal 1600, which may include more or fewer components than shown, or may combine certain components, or may be configured employing different components.

[0255] The above computer device can further be realized as a server, and the structure of the server is introduced below.

[0256] 17 is a schematic diagram of the structure of a server provided by an embodiment of the present application. The server 1700 may include one or more processors (Central Processing Units, CPU) 1701 and one or more memories 1702, since the server 1700 may have relatively large differences due to differences in configuration or performance, and at least one computer program is stored in the one or more memories 1702, and the at least one computer program is loaded and executed by the one or more processors 1701 to realize the methods provided by implementing the above methods. Of course, the server 1700 may further include components such as a wired or wireless network interface, a keyboard, and an input / output interface, for convenient input and output, and the server 1700 may further include components for realizing other device functions, which will not be described in detail here.

[0257] In an exemplary embodiment, a computer-readable storage medium is further provided, which stores at least one computer program, and the computer program is loaded and executed by the processor to realize the method for identifying video clips in the above embodiment. For example, the computer-readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a soft disk, an optical data storage device, etc.

[0258] In an exemplary embodiment, a computer program product is further provided, comprising a computer program which, when executed by a processor, implements the method for identifying video clips.

[0259] In some embodiments, the computer programs referred to in the embodiments of the present application may be arranged to run on one computer device, or on multiple computer devices located at one location, or further, on multiple computer devices distributed across multiple locations and connected to each other via a communication network, and multiple computer devices distributed across multiple locations and connected to each other via a communication network may constitute a blockchain system.

[0260] As can be understood by those skilled in the art, the realization of all or part of the steps of the above embodiments can be achieved through hardware, or can be achieved by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0261] The above are merely selectable embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. 1. A method for identifying video clips implemented by a computing device, comprising: Obtaining video frame features of a first video and at least one video frame feature of a second video, and determining a plurality of video frame pairs based on the video frame features of the first video and the video frame features of the at least one second video, wherein the video frame pairs include a first video frame and a second video frame whose similarity meets a similarity condition, and the first video frame belongs to the first video and the second video frame belongs to the at least one second video; Fusing a first video frame of the plurality of video frame pairs based on an appearance time difference of the plurality of video frame pairs to obtain at least one candidate video clip of the first video, the appearance time difference being a numerical difference between appearance times in a video of two video frames in the video frame pair; A method for identifying video clips, comprising: obtaining a target time range; and determining at least one target video clip in the first video based on the at least one candidate video clip and the target time range, wherein the target video clip is within the target time range of the first video.

2. Fusing a first video frame of the plurality of video frame pairs based on an occurrence time difference of the plurality of video frame pairs to obtain at least one candidate video clip of the first video, Dividing the plurality of video frame pairs into a plurality of video frame groups based on appearance time differences of the plurality of video frame pairs, where video frame pairs in the same video frame group correspond to the same appearance time difference; 2. The method of claim 1, further comprising: for any one of the plurality of video frame groups, fusing a first video frame of a video frame pair among the video frame groups into one of the candidate video clips depending on an appearance time of the first video frame of the video frame pair among the video frame groups in the first video.

3. Prior to dividing the plurality of video frame pairs into a plurality of video frame groups based on occurrence time differences of the plurality of video frame pairs, the method further comprises: determining, for any one of the plurality of video frame pairs, a first appearance time of a first video frame and a second appearance time of a second video frame of the video frame pair, the first appearance time referring to a time when the first video frame appears in the first video and the second appearance time referring to a time when the second video frame appears in the second video; subtracting a second appearance time of a second video frame of the pair from a first appearance time of a first video frame of the pair to obtain an appearance time difference of the pair of video frames; The dividing the plurality of video frame pairs into a plurality of video frame groups based on the appearance time differences of the plurality of video frame pairs includes dividing the video frame pairs having the same appearance time difference into one initial video frame group, and setting the appearance time differences of the video frame pairs in the initial video frame group to the appearance time differences corresponding to the initial video frame group; The method of claim 2 , further comprising: fusing a plurality of initial video frames to obtain the plurality of video frames based on occurrence time differences corresponding to the plurality of initial video frames.

4. Fusing a plurality of initial video frames to obtain the plurality of video frames based on occurrence time differences corresponding to the plurality of initial video frames, obtaining predetermined configuration information, the configuration information including a target sequence; sorting the plurality of initial video frames according to the target order to obtain a plurality of candidate video frames; 4. The method of claim 3, comprising: for any two adjacent candidate video frame groups among the plurality of candidate video frame groups, if a matching time difference between the two adjacent candidate video frame groups meets a matching time difference condition, fusing the two adjacent candidate video frame groups into one video frame group, wherein the matching time difference refers to a numerical difference between appearance time differences corresponding to the two adjacent candidate video frame groups.

5. The two adjacent candidate video frame groups include a first candidate video frame group and a second candidate video frame group, and the fusing of the two adjacent candidate video frame groups into one video frame group includes: adding a video frame pair of the first candidate video frames to the second candidate video frames if a matching time difference between the first candidate video frames and the second candidate video frames is less than or equal to a matching difference threshold; 5. The method of claim 4, comprising: obtaining the video frame group by replacing a target second video frame with a reference second video frame based on an appearance time difference corresponding to the second candidate video frame group, wherein the target second video frame is a second video frame to be newly added to the second candidate video frame group, the reference second video frame is a second video frame having an appearance time difference between the second video frame and a target first video frame that is a target difference, the target difference is the appearance time difference corresponding to the second candidate video frame group, and the target first video frame is a first video frame in a video frame pair to which the target second video frame belongs.

6. The fusing of a first video frame of a video frame pair of the set of video frames into one of the candidate video clips in response to an appearance time in the first video of the first video frame of the video frame pair of the set of video frames, comprises: traversing video frame pairs of the set of video frames to determine a current video frame pair currently being traversed and a previous video frame pair previously traversed, the current video frame pair and the previous video frame pair being two adjacent video frame pairs of the set of video frames; comparing appearance times in the first video of first video frames of the current video frame pair and the previous video frame pair to obtain a numerical difference in appearance times of first video frames; adding the current video frame pair and the previous video frame pair to a temporary frame list if the numerical difference of the appearance times of the first video frames meets an appearance time condition; if the numerical difference of the appearance times of the first video frames does not meet an appearance time condition, merge the video frame pairs in the temporary frame list into a reference video clip, and clear the temporary frame list after merging; determining a next pair of video frames to be traversed, making the next pair of video frames a new current pair of video frames, and returning to the step of comparing the appearance times in the first video of the first video frames of the current pair of video frames and the previous pair of video frames, and continuing to perform the steps until a final pair of video frames is traversed; and determining the at least one candidate video clip based on a plurality of reference video clips.

7. The plurality of reference video clips includes a first overlay video clip, and the first overlay video clip refers to a reference video clip belonging to a first reference video clip among the plurality of reference video clips. Determining the at least one candidate video clip based on the plurality of reference video clips includes: The method of claim 6 , further comprising: if the first overlay video clip is included in the plurality of reference video clips, removing the first overlay video clip to obtain the at least one candidate video clip.

8. The plurality of reference video clips includes a second overlay video clip, and the second overlay video clip refers to a reference video clip that partially overlaps a second reference video clip among the plurality of reference video clips. Determining the at least one candidate video clip based on the plurality of reference video clips includes: The method of claim 6, further comprising: if the second overlay video clip is included in the plurality of reference video clips, removing an overlap portion between the second overlay video clip and the second reference video clip to obtain the at least one candidate video clip.

9. When the second overlay video clip is included in the plurality of reference video clips, after deleting the overlay portion between the second overlay video clip and the second reference video clip, the method further comprises: Comparing a duration of a third-class reference video clip with a target duration, the third-class reference video clip being the second overlay video clip from which an overlay portion has been deleted; If the duration of the third-class reference video clip is equal to or longer than the target duration, the third-class reference video clip is reserved; The method of claim 8 , further comprising: deleting the third-class reference video clip if the duration of the third-class reference video clip is less than the target duration.

10. determining at least one target video clip in the first video based on the at least one candidate video clip and the target time range includes: determining at least one target candidate video clip based on the at least one candidate video clip, the target candidate video clip having an appearance count in the at least one candidate video clip that meets a frequency condition; 2. The method of claim 1, further comprising: for any one of the target candidate video clips, determining the target candidate video clip as a target video clip in the first video if the appearance time of the target candidate video clip in the first video is within the target time range.

11. determining at least one target candidate video clip based on the at least one candidate video clip, determining at least one reference candidate video clip based on the at least one candidate video clip; determining a number of occurrences of each of the reference candidate video clips in the at least one reference candidate video clip; The method of claim 10 , further comprising: determining, as a target candidate video clip, a reference candidate video clip whose occurrence count meets the occurrence count condition.

12. The at least one candidate video clip includes a third overlay video clip, and the third overlay video clip refers to a candidate video clip that belongs to a first candidate video clip among the at least one candidate video clip, and the determining at least one reference candidate video clip based on the at least one candidate video clip includes: The method of claim 11 , further comprising: if the at least one candidate video clip includes the third overlay video clip, removing the third overlay video clip to obtain the at least one reference candidate video clip.

13. The at least one candidate video clip includes a fourth overlay video clip, the fourth overlay video clip being a candidate video clip partially overlaid with a second candidate video clip among the at least one candidate video clip, and the determining at least one reference candidate video clip based on the at least one candidate video clip includes: The method of claim 11, further comprising: determining a number of occurrences of the fourth overlay video clip when the at least one candidate video clip includes the fourth overlay video clip and the degree of overlay between the fourth overlay video clip and the second candidate video clip meets an overlay condition; and determining the at least one reference candidate video clip based on a number of occurrences corresponding to each of the fourth overlay video clips whose degree of overlay meets the overlay condition.

14. The at least one candidate video clip includes a fourth overlay video clip, the fourth overlay video clip being a candidate video clip partially overlaid with a second candidate video clip among the at least one candidate video clip, and determining at least one reference candidate video clip based on the at least one candidate video clip includes: The method of claim 11, further comprising: if the at least one candidate video clip includes the fourth overlay video clip and the overlap degree between the fourth overlay video clip and the second candidate video clip does not meet an overlap degree condition, deleting the fourth overlay video clip to obtain the at least one reference candidate video clip.

15. The at least one candidate video clip includes a fourth overlay video clip, the fourth overlay video clip being a candidate video clip partially overlaid with a second candidate video clip among the at least one candidate video clip, and the determining at least one reference candidate video clip based on the at least one candidate video clip includes: The method of claim 11, further comprising: if the at least one candidate video clip includes the fourth overlay video clip and the duration of the fourth overlay video clip is less than that of the second candidate video clip, deleting the fourth overlay video clip to obtain the at least one reference candidate video clip.

16. determining the at least one reference candidate video clip based on a number of occurrences corresponding to each of the fourth overlay video clips whose overlay degrees meet an overlay degree condition; The method of claim 13, further comprising: for a fourth overlay video clip that meets any one of the at least one candidate video clips and whose number of occurrences is greater than or equal to a first occurrence threshold, fusing the fourth overlay video clip with a second candidate video clip to obtain the at least one reference candidate video clip.

17. determining the at least one reference candidate video clip based on a number of occurrences corresponding to each of the fourth overlay video clips whose overlay degrees meet an overlay degree condition; The method of claim 13, further comprising: for a fourth overlay video clip that meets any one of the overlay degree conditions among the at least one candidate video clip, deleting the fourth overlay video clip if the number of occurrences of the fourth overlay video clip is less than a first occurrence threshold to obtain the at least one reference candidate video clip.

18. Obtaining a target video to be identified, and performing feature extraction of a plurality of target video frames in the target video to be identified to obtain video frame features of the plurality of target video frames; 2. The method of claim 1, further comprising: determining at least one target video clip of the target video based on video frame characteristics of the plurality of target video frames, video frame characteristics of the first video frame, and video frame characteristics of the at least one second video frame.

19. a video frame pair determination module for obtaining video frame features of a first video and at least one video frame feature of a second video, and determining a plurality of video frame pairs based on the video frame features of the first video and the at least one video frame feature of the second video, the video frame pairs including a first video frame and a second video frame whose similarity meets a similarity condition, the first video frame belonging to the first video, and the second video frame belonging to the at least one second video; a fusion module for fusing a first video frame of the plurality of video frame pairs based on an appearance time difference of the plurality of video frame pairs to obtain at least one candidate video clip of the first video, the appearance time difference being a numerical difference between appearance times in a video of two video frames in the video frame pair; A video clip identification device comprising: a target video clip determination module for obtaining a target time range and determining at least one target video clip in the first video based on the at least one candidate video clip and the target time range, wherein the target video clip is within the target time range of the first video.

20. A computing device comprising one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and wherein the computer program is loaded and executed by the one or more processors to implement the method for identifying video clips according to any one of claims 1 to 18.

21. A computer program which, when executed by a processor, implements the method for identifying video clips according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Similarity / dissimilarity decision device, similarity / dissimilarity decision method and similarity / dissimilarity decision program

    JP2008065400A

  • Broadcast program recording / reproducing device and broadcast program recording / reproducing method

    JP2008193585A

  • Image creation device, image creation method, and program

    JP2011151605A

  • Handling of video segments in a video stream

    US20200154165A1