Video duplicate checking method and device, electronic equipment, storage medium and product

Through lens segmentation and graph index optimization, the problem of low-stubble checking efficiency caused by redundant frames in the video library is solved, and efficient and accurate video stubble checking results are achieved.

CN120407853APending Publication Date: 2025-08-01BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510585961.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the existing video plagiarism checking technology, there are a large number of the same video frames in the video library, resulting in high repetition of feature extraction and low efficiency, and low search efficiency.

Method used

The lens slicing algorithm is used to divide the video into multiple segments, extract keyframes to build a graph index, and improve the similarity retrieval efficiency of video frames through feature extraction model and sparse matrix compression.

Benefits of technology

Through lens segmentation and graph index optimization, redundant frame processing is reduced, the efficiency and accuracy of video plagiarism checking is improved, and the computing and storage requirements are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407853A_ABST
    Figure CN120407853A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video duplicate checking method and device, electronic equipment, a storage medium and a product. The method comprises the following steps: acquiring at least one to-be-processed video frame in a video to be subjected to duplicate checking; determining at least one similar video frame associated with the to-be-processed video frame according to the image features of the to-be-processed video frame and a pre-constructed graph index; wherein the graph index comprises a plurality of nodes and edges connected with the plurality of nodes, the nodes correspond to frame features of the key frame, the edges are used for representing that the frame features of the two nodes meet preset similar attributes, and the key frame is a video frame extracted based on shot segmentation of a historical video; and according to the at least one target video to which the at least one similar video frame belongs, determining a duplicate checking result corresponding to the video to be subjected to duplicate checking. According to the technical scheme, the image index is determined based on the shot segmentation algorithm, so that the data volume of the image index is greatly reduced, and the processing efficiency can be improved when the target video is determined according to the image index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of data processing, and in particular, to a video duplicate checking method, apparatus, electronic device, storage medium, and product. Background Art

[0002] With the development of networking, more and more users can share corresponding videos on various platforms. When sharing videos, duplicate checking of video content has become an important task for relevant platforms.

[0003] Currently, the main method for a platform to check for duplicate videos is as follows: First, build a video library, which can store many uploaded videos or video features with copyright. Second, when checking for duplicate videos, it is possible to determine whether the video to be checked is a duplicate video based on the video to be checked and the historical videos stored in the video library.

[0004] When the inventor implemented the present technical solution based on the above method, the following problems were found:

[0005] The video frames in the video library are extracted from historical videos, and the video features corresponding to the extracted video frames are determined, and then the video library is constructed based on the video features. There are a large number of identical video frames among the extracted video frames, resulting in problems of high repetition degree and low efficiency in feature extraction during analysis and processing.

[0006] Furthermore, when checking for duplicate videos, there is a problem of low search efficiency when processing based on a large number of video frames in the video library. Summary of the Invention

[0007] Embodiments of the present invention provide a video duplicate checking method, apparatus, electronic device, storage medium, and product to achieve the effect of improving the efficiency of video duplicate checking.

[0008] In a first aspect, an embodiment of the present invention provides a video duplicate checking method, the method including:

[0009] Obtain at least one video frame to be processed in the video to be checked;

[0010] For the at least one video frame to be processed, determine at least one similar video frame associated with the video frame to be processed according to the image feature of the video frame to be processed and a pre-constructed graph index; wherein, the graph index includes a plurality of nodes and edges connecting the plurality of nodes, the nodes correspond to the frame features of key frames, the edges are used to represent that the frame features of two nodes satisfy a preset similarity attribute, and the key frames are video frames extracted after shot segmentation of historical videos;

[0011] Determine the duplicate check result corresponding to the video to be checked for duplicates according to at least one target video to which the at least one similar video frame belongs.

[0012] Further, the method further includes:

[0013] For multiple historical videos obtained, determine the transition frames in the historical videos based on a shot segmentation algorithm, and divide the historical videos into multiple video segments based on the transition frames;

[0014] Perform frame extraction on the video frames in the multiple video segments to obtain multiple key frames corresponding to the historical videos, and determine the graph index based on the multiple key frames of the multiple historical video frames.

[0015] Further, the shot segmentation algorithm corresponds to a frame type classification model, and determining the transition frames in the historical videos based on the shot segmentation algorithm includes:

[0016] For at least one video frame in the historical video frames, input the current video frame and the video frames of a preset number of frames before the current video frame into a pre-trained frame type classification model, and output the frame type identifier of the current video frame, and when the frame type identifier is a preset identifier, determine that the current video frame is a transition frame;

[0017] Correspondingly, dividing the historical video into multiple video segments based on the transition frames includes:

[0018] According to the playing time sequence of the transition video frames, use the video frames between two adjacent transition frames as a video segment.

[0019] Further, determining the graph index based on the multiple key frames of the multiple historical video frames includes:

[0020] Extract features from the multiple key frames based on a pre-trained feature extraction model to obtain the frame features corresponding to the multiple key frames; wherein, the frame features are represented by feature vectors of a preset dimension;

[0021] Based on the frame features of the multiple key frames, determine the similarity attributes between any two key frames, and when the similarity attributes meet the preset similarity attributes, establish connection information between the two key frames to obtain the graph index.

[0022] Further, after obtaining the graph index, the method further includes:

[0023] Compress the graph index based on a sparse matrix to obtain a compressed graph index; and / or,

[0024] By performing node clustering processing on the graph index, an updated graph index is obtained.

[0025] Further, determining at least one similar video frame associated with the video frame to be processed based on the image features of the video frame to be processed and a pre-constructed graph index includes:

[0026] Extract the image features of the video frame to be processed;

[0027] Based on the image features and the frame features corresponding to each node in the graph index, determine target frame features whose similarity to the image features is higher than a similarity threshold;

[0028] Use the video frames corresponding to the nodes of the target frame features as the similar video frames.

[0029] Further, the graph index is determined based on node clustering. Determining at least one similar video frame associated with the video frame to be processed based on the image features of the video frame to be processed and a pre-constructed graph index includes:

[0030] Based on the image features of the video frame to be processed and the clustering features corresponding to at least one clustering center in the graph index, determine at least one cluster associated with the image features, where the cluster includes multiple nodes and the clustering center after clustering of the multiple nodes;

[0031] Based on the similarity between the frame features of each node in the at least one cluster and the image features, determine the target frame features, and use the video frames corresponding to the nodes of the target frame features as the similar video frames.

[0032] Further, determining the duplicate check result corresponding to the video to be checked based on at least one target video to which the at least one similar video frame belongs includes:

[0033] Determine at least one target video to which the at least one similar video frame belongs;

[0034] For the at least one target video, determine the video similarity of the target video according to the total number of frames of the at least one video frame to be processed and the number of associated frames associated with the target video in the at least one video frame to be processed;

[0035] If there is a target video with a video similarity greater than a preset threshold, determine that the duplicate check result of the video to be checked is a duplicate video.

[0036] Further, when the duplicate check result is the result of whether the video to be checked is a duplicate video, the method further includes:

[0037] Display the target interface, and display a prompt message indicating that the video to be checked for duplication has not been successfully uploaded in the target interface, so as to determine the reason for the unsuccessful upload based on the prompt message.

[0038] In a second aspect, an embodiment of the present invention further provides a video duplication checking device, which includes:

[0039] A module for extracting video frames to be processed, configured to obtain at least one video frame to be processed in the video to be checked for duplication;

[0040] A module for determining similar video frames, configured to, for the at least one video frame to be processed, determine at least one similar video frame associated with the video frame to be processed according to the image features of the video frame to be processed and a pre-constructed graph index; wherein, the graph index includes a plurality of nodes and edges connecting the plurality of nodes, the nodes correspond to the frame features of key frames, the edges are used to represent that the frame features of two nodes satisfy a preset similarity attribute, and the key frames are video frames extracted after shot segmentation of historical videos;

[0041] A module for determining the duplication checking result, configured to determine the duplication checking result corresponding to the video to be checked for duplication according to at least one target video to which the at least one similar video frame belongs.

[0042] In a third aspect, an embodiment of the present invention provides an electronic device, which includes:

[0043] One or more processors;

[0044] A memory for storing one or more programs;

[0045] When the one or more programs are executed by the one or more processors, the one or more processors implement the video duplication checking method provided in any embodiment of the present invention.

[0046] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the video duplication checking method provided in any embodiment of the present invention.

[0047] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the video duplication checking method according to any one of the embodiments of the present invention.

[0048] The technical solution provided by the embodiments of the present invention can, when receiving a video to be checked for duplication, obtain at least one video frame to be processed in the video to be checked for duplication, and then determine at least one similar video frame associated with the video frame to be processed based on the image features of each video frame to be processed and the frame features corresponding to each node in the pre-created graph index. Further, based on the occurrence frequencies of the historical videos corresponding to all the similar video frames, the duplication check result corresponding to the video to be checked for duplication is determined, which solves the problems in the prior art that the constructed video library is based on a large number of identical video frames, resulting in a high degree of feature duplication and low efficiency when determining the duplication check result. Further, when determining the duplication check result based on a large number of video frames, there is a problem of low search efficiency. The present invention realizes that the index library is based on a limited number of video frames and the video frames are non-repetitive, thereby achieving the effect of improving the convenience of determining the target duplication check result of the video to be checked for duplication. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the introduced drawings are only the drawings of a part of the embodiments to be described by the present invention, rather than all the drawings. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a flowchart of a video duplication check method provided by an embodiment of the present invention;

[0051] Figure 2 It is a flowchart of a video duplication check method provided by an embodiment of the present invention;

[0052] Figure 3 It is a flowchart of a video duplication check method provided by an embodiment of the present invention;

[0053] Figure 4 It is a schematic structural diagram of a video duplication check device provided by an embodiment of the present invention;

[0054] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all the structures.

[0056] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only the parts related to the present invention rather than all the structures are shown in the drawings.

[0057] Before introducing the technical solutions provided in the embodiments of the present invention, an exemplary description of the application scenario can be given first.

[0058] In existing software with a video upload function, a target user can upload the video made by him / her to the software to achieve the effect of sharing the video. The uploaded video by the user can be used as the video to be checked for duplication. When uploading the video to be checked for duplication to the software, the video to be checked for duplication can be detected to determine whether the video to be uploaded is a duplicate uploaded video. Or, based on the video with copyright, a graph index is determined, and then after receiving the video to be checked for duplication, the target duplication check result of the video to be checked for duplication can be determined based on the video to be checked for duplication and the graph index.

[0059] The application scenario of the embodiments of the present invention can be any scenario that requires video duplication checking. For example, the functional module corresponding to the embodiments of the present invention is integrated in the existing application software and is displayed in the form of a triggerable control in the application software. The description information of the triggerable control can be video duplication checking description. Correspondingly, the display icon corresponding to the trigger control can be a video duplication checking button. When a trigger operation on the video duplication checking button is detected, the target duplication check result of the video to be checked for duplication can be determined based on the solution provided in the embodiments of the present invention.

[0060] Figure 1 The figure is a flowchart of a video duplication checking method provided for an embodiment of the present invention. This embodiment is applicable to the scenario of determining whether the video to be checked for duplication is a duplicate upload or duplicate processing. The video duplication checking method provided in this embodiment can be executed by the client, or by the server, or by the cooperation of the client and the server. The video duplication checking device integrated in the client and / or the server can be implemented in the form of software and / or hardware, and is integrated in an electronic device, which can be a mobile terminal or a PC terminal, etc.

[0061] As Figure 1 shown, the method specifically includes the following steps:

[0062] S110. Obtain at least one video frame to be processed in the video to be checked for duplication.

[0063] Among them, the video that needs to be processed currently can be used as the video to be checked for duplication. The video to be checked for duplication is composed of multiple video frames. Each video frame can be used as a video frame to be processed, or the video frames in the video to be checked for duplication can be frame-extracted according to actual needs to obtain the video frames to be processed that need to be processed.

[0064] Generally, for a certain video, there are similarities in the picture content of adjacent frames. If all video frames are processed, there will be a problem that the same content needs to be processed repeatedly, resulting in low processing efficiency. Based on this, the video frames in the video to be checked for duplication can be frame-extracted to reduce the data volume of the video frames to be processed, thereby improving the processing efficiency of the video to be checked for duplication.

[0065] Optionally, at least one video frame to be processed is extracted from the video to be checked for duplication according to a preset frame extraction condition; wherein, the preset frame extraction condition includes frame extraction according to a preset number of frames at intervals and / or frame extraction according to a preset duration.

[0066] Among them, the preset number of frames at intervals can be understood as extracting a video frame to be processed every preset number of video frames. The preset duration can be, based on the video to be checked for duplication, extracting a video frame to be processed every preset duration. For example, the preset number of frames at intervals is five frames, that is, starting from the first video frame of the video to be checked for duplication, extracting a video frame to be processed every five video frames until all the video frames in the video to be checked for duplication are traversed. The preset interval duration can be 200 ms, and it can be starting from the starting playback moment of the video to be checked for duplication, extracting a video frame to be processed every 200 ms.

[0067] It should be noted that the extraction conditions for at least one video frame to be processed can be set according to actual needs, and in this embodiment, there is no limitation on how to extract the video frames to be processed.

[0068] S120. For at least one video frame to be processed, at least one similar video frame associated with the video frame to be processed is determined according to the image feature of the video frame to be processed and the pre-constructed graph index.

[0069] Among them, the image feature is the feature obtained after extracting the features of the video frame to be processed based on an existing feature extraction model. The graph index is determined based on historical videos. The historical video graph can be a video processed by an application software or platform. The constructed graph index includes multiple nodes and edges connecting the multiple nodes. The nodes correspond to the frame features of key frames, and the edges are used to represent that the frame features of two nodes satisfy a preset similarity attribute. The key frames are video frames extracted after segmenting the historical video into shots.

[0070] It can be understood that for each video frame to be processed, the similarity between the frame features corresponding to each node in the graph index and the image features of the video frame to be processed can be determined. At least one target node can be determined based on the similarity. The frame features corresponding to the at least one target node are used as the similar video frames of the video frame to be processed. For each video frame to be processed, the number of corresponding similar video frames can include one or more, and the specific number is related to the preset similarity threshold parameter.

[0071] In this embodiment, determining at least one similar video frame associated with the video frame to be processed according to the image features of the video frame to be processed and the pre-constructed graph index includes: extracting the image features of the video frame to be processed; determining, based on the image features and the frame features corresponding to each node in the graph index, target frame features whose similarity to the image features is higher than the similarity threshold; and using the video frames corresponding to the nodes of the target frame features as the similar video frames.

[0072] Among them, an existing feature extraction model can be used to extract features from the video frame to be processed, and the obtained features are used as the image features. The similarity threshold is preset and is used as the basis for determining whether the key frame of the node in the graph index is similar to the video frame to be processed. The similarity between the image features and the frame features of each node in the graph index can be calculated. The frame features corresponding to the nodes whose similarity is higher than the preset similarity threshold are used as the target frame features. Correspondingly, the key frames corresponding to the nodes of the target frame features are used as the similar video frames.

[0073] In this embodiment, if the graph index is determined based on node clustering, determining at least one similar video frame associated with the video frame to be processed according to the image features of the video frame to be processed and the pre-constructed graph index includes: determining at least one cluster associated with the image features based on the image features of the video frame to be processed and the cluster features corresponding to at least one cluster center in the graph index, where a cluster includes multiple nodes and the cluster center after clustering of the multiple nodes; determining target frame features based on the similarity between the frame features of each node in the at least one cluster and the image features, and using the video frames corresponding to the nodes of the target frame features as the similar video frames.

[0074] Among them, the graph index is determined by clustering the frame features of at least one node. Then, there can be multiple clustering centers for the graph index. After obtaining the image features of the video frame to be processed, at least one target clustering center can be determined based on the distance information between the image features and the clustering centers. For each target clustering center, the number of associated nodes can include one or more. The similarity between the frame features of the video frame to be processed and the nodes associated with the target clustering center can be determined. Based on this similarity and a preset similarity threshold, the target frame features can be determined, and the video frame corresponding to the nodes of the target frame features is used as the target similar video frame.

[0075] That is to say, the similarity between the image features and the frame features can be determined based on the image features of at least one video frame to be processed and the frame features of all nodes in the graph index. The frame features with a similarity higher than the similarity threshold are used as the target frame features. Correspondingly, the nodes corresponding to the target frame features are used as the target nodes. The key frames corresponding to the target nodes are used as the similar video frames.

[0076] The advantage of the similar video frames determined based on the above method is that the graph index is determined based on a limited number of video frames. Therefore, when determining the similar video frames, the efficiency of determining the similar video frames can be improved.

[0077] S130. Determine the duplicate check result corresponding to the video to be checked according to at least one target video to which at least one similar video frame belongs.

[0078] Among them, the video corresponding to the similar video frame is used as the target video. The number of similar video frames includes at least one, so the number of target videos can also include one or more.

[0079] Specifically, after obtaining at least one similar video frame corresponding to each video frame to be processed, the duplicate check result corresponding to the video to be checked can be determined according to at least one target video to which at least one similar video frame belongs.

[0080] In this embodiment, determining the duplicate check result corresponding to the video to be checked based on at least one target video corresponding to at least one target similar video frame includes: determining at least one target video to which at least one similar video frame belongs; for at least one target video, determining the video similarity of the target video according to the total number of at least one video frame to be processed and the number of associated frames associated with the target video in at least one video frame to be processed; if there is a target video with a video similarity greater than the preset threshold, determining that the duplicate check result of the video to be checked is a duplicate video.

[0081] Among them, the video similarity is related to the degree of overlap of the target video with respect to the video to be checked for duplication. The total number of frames is the number of frames of the video frames to be processed extracted. The associated number of frames is the number of frames of the video frames to be processed that belong to a certain target video. The preset threshold is a constraint condition set in advance for whether the video to be checked for duplication is a duplicate video.

[0082] Specifically, at least one target video corresponding to each similar video frame can be determined. For each target video, the video similarity of the target video can be determined based on the total number of video frames to be processed and the occurrence frequency of the target video to which at least one video frame to be processed belongs, that is, the associated number of frames. The video similarity can be determined by the ratio of the associated number of frames to the total number of frames. If the video similarity is greater than the preset threshold, it indicates that the target video is similar to the video to be checked for duplication, that is, the duplication check result of the video to be checked for duplication is a duplicate video.

[0083] In this embodiment, when the duplication check result is that the video to be checked for duplication is a duplicate video, the method further includes: displaying a target interface, and displaying a prompt message indicating that the video to be checked for duplication has not been successfully uploaded in the target interface, so as to determine the reason for the unsuccessful upload based on the prompt message.

[0084] Among them, the target interface is an interface for displaying feedback information on whether the video to be checked for duplication has been successfully uploaded. If the duplication check result of the video to be checked for duplication is a non-duplicate video, feedback information indicating that the video to be checked for duplication has been successfully uploaded can be displayed in the target interface. If the duplication check result of the video to be checked for duplication is a duplicate video, feedback information indicating that the video to be checked for duplication has not been successfully uploaded can be displayed in the target interface.

[0085] At the same time, in order to facilitate the user to clearly understand the reason for the unsuccessful upload of the video to be checked for duplication, a prompt message can be displayed in the target interface. The description information in the prompt message is mainly the reason for the unsuccessful upload of the video to be checked for duplication.

[0086] The technical solution provided by the embodiment of the present invention can, when receiving a video to be checked for duplication, obtain at least one video frame to be processed in the video to be checked for duplication, and then determine at least one similar video frame associated with the video frame to be processed based on the image features of each video frame to be processed and the frame features corresponding to each node in the pre-created graph index. Further, based on the occurrence frequency of the historical videos corresponding to all similar video frames, the duplication check result corresponding to the video to be checked for duplication is determined, which solves the problems in the prior art that the constructed video library is based on a large number of identical video frames, resulting in a high degree of feature duplication and low efficiency when determining the duplication check result. Further, when determining the duplication check result based on a large number of video frames, there is a problem of low search efficiency, and realizes that the index library is based on a limited number of video frames and the video frames are non-repetitive, thereby improving the convenience of determining the target duplication check result of the video to be checked for duplication.

[0087] Figure 2 The figure shows a schematic flowchart of a video duplicate checking method provided by an embodiment of the present invention. Based on the foregoing embodiments, a figure index may be first determined, and then the duplicate checking result of the video to be checked may be determined according to the figure index. The specific manner of determining the figure index may refer to the detailed description of this embodiment. Technical terms that are the same or corresponding to those in the foregoing embodiments will not be elaborated herein.

[0088] As shown in FIG. 2, the method includes:

[0089] S210. For a plurality of acquired historical videos, determine transition frames in the historical videos based on a shot segmentation algorithm, so as to divide the historical videos into a plurality of video segments based on the transition frames.

[0090] It should be noted that the processing manner for each historical video is the same. Here, taking the processing of one of the historical videos as an example for illustration.

[0091] Among them, videos with a duplicate checking result of non-duplicate videos and that have been uploaded to a software or platform are used as historical videos. For a historical video, it may be composed of multiple shots, and the number of video frames collected by each shot may be the same. Here, a shot segmentation algorithm may be used to process the historical video to determine video segments, and further determine the video frames constituting the figure index.

[0092] The shot segmentation algorithm may be an algorithm for identifying whether a video frame in a historical video is a shot transition frame. A transition frame is a video frame in a historical video where a shot transition occurs. The historical video may be divided into a plurality of video segments based on the transition frames.

[0093] Specifically, for a plurality of acquired historical videos, frame segmentation processing may be respectively performed on the historical videos based on the shot segmentation algorithm to obtain a plurality of transition frames corresponding to each historical video. Based on the plurality of transition frames of each historical video, a plurality of video segments corresponding to each historical video may be determined.

[0094] In this embodiment, determining the transition frames in the historical video based on the shot segmentation algorithm includes: for at least one video frame in the historical video frames, inputting the current video frame and the video frames of a preset number of frames before the current video frame into a pre-trained frame type classification model, and outputting a frame type identifier of the current video frame, so as to determine the current video frame as a transition frame when the frame type identifier is a preset identifier; correspondingly, dividing the historical video into a plurality of video segments based on the transition frames includes: using the video frames between two adjacent transition frames as a video segment according to the playing time sequence of the transition frames in the historical video.

[0095] Among them, the lens segmentation algorithm corresponds to a frame type classification model. The frame type classification model refers to a model used to determine the type information of each video frame in a historical video. The frame type classification model is a binary classification model. The input of the frame type classification model can be at least one video frame, and the output can be a model of the class identifier indicating whether a certain video frame is a transition frame or a non-transition frame. The preset identifier is used to represent the identifier indicating that the frame type of the video frame is a transition frame. The preset number of frames can be a preset number of frames.

[0096] Specifically, the current video frame and the video frames of the preset number of frames before the current video frame are jointly input into the frame type classification model, and the frame type classification model can output the frame type identifier indicating whether the current video frame is a transition frame. If the frame type identifier is the preset identifier, it means that the current video frame is a transition frame. On the contrary, if the frame type identifier is not the preset identifier, it means that the current video frame is not a transition frame.

[0097] Correspondingly, after determining the transition frames in the historical video, according to the playback timing of the transition frames in the historical video, the video frames between two adjacent transition frames and including the transition frames can be used as a video segment.

[0098] In this embodiment, the advantage of determining the transition frames is that: generally, there is a certain degree of repetition or similarity in the content of the video frames within the same lens. The video segments can be determined based on the lens segmentation algorithm. Furthermore, when determining the key frames based on the divided video segments, there are fewer repeated video frames, that is, the quality of the key frames is higher. Further, when determining the graph index based on the key frames, the quality of the graph index can be improved. Furthermore, when determining the similar video frames based on the graph index, the processing efficiency can be improved.

[0099] S220. Perform frame extraction on the video frames in multiple video segments to obtain multiple key frames corresponding to the historical video, and determine the graph index based on the multiple key frames of the multiple historical video frames.

[0100] Among them, after obtaining at least one video segment corresponding to each historical video, frame extraction can be performed on the at least one video segment. The video frames corresponding to the frame extraction are used as key frames. The advantage of determining the key frames in this way is that the effectiveness of determining the key frames can be improved.

[0101] After obtaining the multiple key frames of all historical videos, feature extraction can be performed on the multiple key frames, and a graph index can be generated based on the extracted features.

[0102] Specifically, frame extraction is performed on the video segments according to a preset frame extraction method, and the extracted video frames are used as key frames. After obtaining the key frames of all historical videos, they can be processed to determine the graph index.

[0103] It should be noted that when initially constructing the graph index, multiple historical videos can be obtained to determine the graph index based on the above method. After the graph index is constructed, when historical videos are received again, the graph index can be updated based on the historical videos, and the video to be duplicate-checked can be processed based on the updated graph index.

[0104] In this embodiment, determining the graph index based on multiple key frames of multiple historical video frames includes: extracting features of the multiple key frames based on a pre-trained feature extraction model to obtain frame features corresponding to the multiple key frames; wherein, the frame features are represented by feature vectors of a preset dimension; based on the frame features of the multiple key frames, determining the similarity attribute between any two key frames, and when the similarity attribute meets the preset similarity attribute, establishing connection information between the two key frames to obtain the graph index.

[0105] Among them, the feature extraction model is a model pre-trained for extracting feature information. This feature extraction model can adopt any existing model architecture. The feature extraction model can be trained by means of metric learning. The features corresponding to the key frames are used as the frame features. The frame features can be represented by feature vectors. A similarity determination method can be adopted to calculate the similarity attribute between any two frame features. If the similarity attribute between two frame features is higher than the preset similarity attribute, it indicates that a connection can be established between the two key frames. Based on the above method, the graph index can be determined.

[0106] Specifically, extract the features of the key frames based on the pre-trained feature extraction model to obtain the frame features corresponding to each key frame. For any two key frames, the similarity attribute between the two frame features can be calculated. If the similarity attribute is greater than the preset similarity attribute, the connection information between the two key frames can be established. The connection information is the edge. The graph index can be obtained based on the above method. The nodes of the graph index are the frame features of the key frames.

[0107] Based on the above technical solution, after obtaining the graph index, the method further includes: compressing the graph index based on a sparse matrix to obtain a compressed graph index; and / or, obtaining an updated graph index by performing clustering processing on the nodes in the graph index.

[0108] It can be understood that: in order to reduce the data volume of the graph index, the graph index can be compressed by using a sparse matrix to obtain the graph index. At the same time, clustering processing can be performed on the frame features of the nodes in the graph index to obtain at least one clustering center. The updated graph index is obtained based on the graph index obtained from at least one clustering center.

[0109] In this embodiment, when determining similar video frames based on the graph index after clustering processing, the similarity between the frame features of the video frame to be processed and the central features of each cluster center in the graph index can be calculated. Based on the similarity, at least one cluster associated with the video frame to be processed can be determined. Based on the frame features of the nodes in at least one cluster and the image features of the video frame to be processed, similar video frames can be determined. In this way, the efficiency of determining similar video frames can be improved.

[0110] In this embodiment, the transition frames in the historical video can be determined through the shot segmentation algorithm, and then the historical video can be divided based on the playback order of the transition frames to obtain multiple video segments. By performing frame extraction on each video segment, key frames for determining the graph index can be obtained. Since the graph index is determined based on the key frames, the efficiency of determining the graph index can be improved. Further, when determining at least one similar video frame corresponding to the video frame to be processed based on the graph index, the effect of improving its efficiency can be achieved.

[0111] As an alternative embodiment of the above embodiment, the duplicate checking result of the video to be checked for duplicates can be determined from two dimensions. For example, the duplicate checking result can be determined from the methods of offline database building and online duplicate checking. Offline database building is mainly to build an index database, and online duplicate checking is mainly to implement retrieval and duplicate calculation.

[0112] Next, in combination with Figure 3 to further understand the technical solution provided by the embodiments of the present invention.

[0113] Previously, when constructing the graph index, a frame was intercepted at a fixed time interval or a frame was intercepted at a preset number of frames interval. At this time, the connection between video frames was ignored. At the same time, in some videos, there are situations where the picture changes very little within a few seconds or even more than ten seconds, that is, under one shot, if a frame is intercepted every second, a lot of redundant and repeated pictures will be intercepted, resulting in a large amount of computational waste. Therefore, in order to filter out redundant and repeated pictures, the shot segmentation algorithm can be used.

[0114] At this time, the shot segmentation algorithm can be regarded as a classification problem, that is, all video frames in the historical video are classified, and the identification of frame classification includes the identification of transition frames and non-transition frames. When predicting whether a certain video frame is a transition frame, the current video frame and a preset number of frames before the current video frame are used. Optionally, the preset number of frames can be 7 video frames before the current video frame. Fusion is performed on the RGB channels of the image. If there are less than seven video frames before the current video frame, zero-padding is performed, and the processed video frame is input into the neural network (MobileNetV4) of the (frame type classification model) for frame type classification. This method has been verified, and both the accuracy rate and the recall rate reach 95%.

[0115] The advantage of using the shot segmentation algorithm in the embodiments of the present invention is that: a fixed number of frames can be intercepted within each shot, redundant frames are removed, and the number of frames intercepted is reduced by 70% compared to intercepting frames at a fixed time interval (one frame per second), which reduces the computational amount and storage amount for subsequent steps by 70%. At the same time, each shot has intercepted frames, which also ensures the richness of information, has little impact on subsequent recall, and can improve the efficiency of determining the recall result.

[0116] After determining the key frames using the shot segmentation algorithm, training samples can be determined to determine the feature extraction model based on the training samples. In this embodiment, when constructing the feature extraction model, the feature extraction model can be trained in a metric learning manner.

[0117] A positive and negative sample data set is constructed using the above-mentioned shot segmentation algorithm. The positive samples are multiple video frames within the same shot, and the negative samples are multiple video frames from different shots. In addition, data augmentation methods, including color changes, geometric transformations, compression noise, etc., can be used to process the positive samples or negative samples to obtain updated positive and negative samples. After training the feature extraction model based on the positive and negative samples, 128-dimensional features of the key frames can be extracted.

[0118] In this embodiment, after obtaining the above features, the way to construct the graph index can be:

[0119] Using the shot segmentation algorithm and the feature extraction algorithm. Split the historical video frames in the video library to obtain multiple video segments. Perform frame extraction on the video segments to obtain multiple key frames. Extract the features of the key frames based on the feature extraction model to obtain 128-dimensional vectors. Perform normalization processing on them to ensure that the features are compared on the same scale.

[0120] Graph structure initialization. Use the feature vectors as the nodes in the graph index, and the feature of the node is the 128-dimensional vector; then define the edges between nodes according to a specific similarity metric (such as cosine similarity or Euclidean distance). A threshold can be set, and only node pairs with a similarity higher than this threshold are connected by edges.

[0121] Graph index construction. Use the adjacency to represent the neighbor nodes of each node for fast access and traversal, and then use a sparse matrix or other compression techniques to reduce memory usage. Or, use dimensionality reduction techniques to further optimize the representation of the 128-dimensional feature vector to improve computational efficiency, and then cluster the nodes to aggregate similar nodes together to form a higher-level index structure.

[0122] Index structure optimization will build a hierarchical index structure to support fast multi-level queries and achieve dynamic updates of the index to cope with data changes. Query acceleration uses approximate nearest neighbor search algorithms (such as ANN or LSH) to accelerate similarity queries, and then utilizes parallel computing technology to improve query speed. Through these steps, the 128-dimensional vector features of the video library can be effectively used to construct and optimize the graph index to support efficient similarity queries and analysis.

[0123] After the graph index is constructed, the graph index can be used to retrieve similar video frames.

[0124] Specifically, when performing duplicate checking on the video to be checked, the method of intercepting frames at fixed time intervals can be adopted, that is, one frame per second, and a feature extraction model is used to extract features to obtain image features. Assume that a total of K video frames to be processed are intercepted. Use the K video frames to retrieve in the graph index constructed above. Each video frame to be processed will recall the N most similar video frames. Assume that the N similar video frames belong to M videos, and a K x N matrix can be obtained. Set the preset video similarity threshold to 0.9. After verification, the accuracy of this preset video similarity threshold is above 95%. Vote on the M videos respectively to count how many of the K frames of the query video are similar to the recalled video frames up to the threshold. Assume there are J similar frames, then the repeatability of this query video and the recalled video is J / K. This method is simple and efficient. By using the recalled video feature vectors to calculate the repeatability again, only the similarity values returned during retrieval need to be reused, and the CPU can be used to quickly calculate the repeatability in milliseconds. At the same time, thanks to the highly robust and accurate feature extraction model, various partially repeated or fully repeated videos can be recalled.

[0125] The technical solution provided by the embodiments of the present invention can first intercept frames for all historical videos in the video library. At this time, the shot segmentation algorithm is mainly used to remove redundant frames while ensuring the extraction of key frames, and filter some simple frames, so that the total number of frames is greatly reduced, reducing the subsequent feature extraction calculation amount, and at the same time ensuring the integrity of the content information, with little impact on the recall loss; further, the feature extraction model of the present invention uses multiple data augmentation techniques and loss functions to train an efficient and robust feature extraction model, greatly improving the recall of various repeated videos. Next, use the graph index as the retrieval engine to implement an efficient retrieval system. When performing duplicate checking on the video to be checked, the same frame intercepting and feature extraction methods as those for offline library building can be adopted to determine the image features of the video frames to be processed. Based on the image features and the graph index for retrieval, so that each video frame to be processed recalls multiple similar video frames. Using the similarity threshold output during retrieval, finally select the threshold to vote and calculate the repeatability of the query video and the recalled video, so that the CPU can return the duplicate check result in milliseconds, which is simple and efficient.

[0126] The following is an embodiment of the video duplicate detection device provided by the embodiments of the present invention. This device and the video duplicate detection methods of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiment of the video duplicate detection device, reference may be made to the embodiments of the above video duplicate detection methods.

[0127] Figure 4 FIG. 5 is a schematic structural diagram of a video duplicate detection device provided by an embodiment of the present invention. The device specifically includes: a to-be-processed video frame extraction module 310, a similar video frame determination module 320, and a duplicate detection result determination module 330.

[0128] Among them, the to-be-processed video frame extraction module 310 is configured to obtain at least one to-be-processed video frame in the to-be-detected video; the similar video frame determination module 320 is configured to, for the at least one to-be-processed video frame, determine at least one similar video frame associated with the to-be-processed video frame according to the image features of the to-be-processed video frame and a pre-constructed graph index. Among them, the graph index includes a plurality of nodes and edges connecting the plurality of nodes. The nodes correspond to the frame features of key frames, and the edges are used to represent that the frame features of two nodes satisfy a preset similarity attribute. The key frames are video frames extracted after shot segmentation of historical videos; the duplicate detection result determination module 330 is configured to determine the duplicate detection result corresponding to the to-be-detected video according to at least one target video to which the at least one similar video frame belongs.

[0129] Based on the above technical solution, the device includes:

[0130] A video segment determination module, configured to, for a plurality of obtained historical videos, determine transition frames in the historical videos based on a shot segmentation algorithm, so as to divide the historical videos into a plurality of video segments based on the transition frames;

[0131] A graph index determination module, configured to perform frame extraction on the video frames in the plurality of video segments to obtain a plurality of key frames corresponding to the historical videos, and determine the graph index based on the plurality of key frames of the plurality of historical video frames.

[0132] Based on the above technical solutions, the shot segmentation algorithm corresponds to a frame type classification model. The video segment determination module includes:

[0133] The transition frame determination unit is configured to, for at least one video frame in the historical video frames, input the current video frame and the video frames with a preset number of frames before the current video frame into a pre-trained frame type classification model, and output the frame type identifier of the current video frame, so as to determine the current video frame as a transition frame when the frame type identifier is a preset identifier; correspondingly, the dividing the historical video into multiple video segments based on the transition frames includes: a video segment determination unit configured to, according to the playing time sequence of the transition video frames, use the video frames between two adjacent transition frames as one video segment.

[0134] Based on the above technical solutions, the graph index determination module includes:

[0135] A frame feature determination unit is configured to perform feature extraction on the multiple key frames based on a pre-trained feature extraction model to obtain the frame features corresponding to the multiple key frames; wherein, the frame features are represented by feature vectors of a preset dimension; a graph index determination unit is configured to, based on the frame features of the multiple key frames, determine the similarity attributes between any two key frames, and establish connection information between the two key frames when the similarity attributes meet the preset similarity attributes, so as to obtain the graph index.

[0136] Based on the above technical solutions, after determining the graph index, the graph index module includes:

[0137] A first compression unit is configured to compress the graph index based on a sparse matrix to obtain a compressed graph index; and / or, a graph index update unit is configured to obtain an updated graph index by performing clustering processing on the nodes in the graph index.

[0138] Based on the above technical solutions, the similar video frame determination module includes:

[0139] An image feature extraction unit is configured to extract the image features of the video frame to be processed;

[0140] A target frame feature determination unit is configured to determine target frame features with a similarity higher than a similarity threshold to the image features based on the image features and the frame features corresponding to each node in the graph index;

[0141] A similar video frame determination unit is configured to use the key frames corresponding to the nodes of the target frame features as the similar video frames.

[0142] Based on the above technical solutions, the similar video frame determination module includes:

[0143] A cluster determination unit, configured to determine at least one cluster associated with the image feature based on the image feature of the video frame to be processed and the cluster feature corresponding to at least one cluster center in the graph index, where the cluster includes multiple nodes and the cluster center after clustering of the multiple nodes;

[0144] A similar video frame determination unit, configured to determine the target frame feature based on the similarity between the frame feature of each node in the at least one cluster and the image feature, and use the video frame corresponding to the node with the target frame feature as the similar video frame.

[0145] Based on the above technical solutions, the duplicate check result determination module includes:

[0146] A target video determination unit, configured to determine at least one target video to which the at least one similar video frame belongs;

[0147] A video similarity determination unit, configured to, for the at least one target video, determine the video similarity of the target video according to the total number of frames of the at least one video frame to be processed and the number of associated frames associated with the target video in the at least one video frame to be processed;

[0148] A duplicate result determination unit, configured to, if there is a target video with a video similarity greater than a preset threshold, determine that the duplicate check result of the video to be checked for duplicates is a duplicate video.

[0149] Based on the above technical solutions, the apparatus further includes: a target interface display module, configured to display a target interface, and display a prompt message indicating that the video to be checked for duplicates has not been successfully uploaded in the target interface, so as to determine the reason for the unsuccessful upload based on the prompt message.

[0150] The technical solution provided by the embodiment of the present invention, when receiving a video to be checked for duplicates, can obtain at least one video frame to be processed in the video to be checked for duplicates, and then determine at least one similar video frame associated with the video frame to be processed based on the image feature of each video frame to be processed and the frame feature corresponding to each node in the pre-created graph index. Further, based on the occurrence frequency of the historical videos corresponding to all similar video frames, determine the duplicate check result corresponding to the video to be checked for duplicates, which solves the problem in the prior art that the constructed video library is based on a large number of identical video frames, resulting in a high degree of feature duplication and low efficiency when determining the duplicate check result. Further, when determining the duplicate check result based on a large number of video frames, there is a problem of low search efficiency, and realizes that the index library is based on a limited number of video frames and the video frames are not repeated, so as to improve the convenience of determining the target duplicate check result of the video to be checked for duplicates.

[0151] The video duplicate checking device provided by the embodiments of the present invention can execute the video duplicate checking method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the video duplicate checking method.

[0152] It should be noted that in the embodiments of the above video duplicate checking device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.

[0153] Figure 5 It is a schematic structural diagram of a server provided by an embodiment of the present invention. Figure 5 The block diagram of an exemplary electronic device 12 suitable for implementing the embodiments of the present invention is shown. Figure 5 The displayed electronic device 12 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0154] As Figure 5 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0155] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0156] The electronic device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0157] The system memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 5 not shown, commonly referred to as a "hard disk drive"). Although Figure 5Not shown, a disk drive for reading and writing a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical medium) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data medium interfaces. The system memory 28 can include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0158] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in the system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 42 generally perform the functions and / or methods in the embodiments described in the present invention.

[0159] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Also, the electronic device 12 can communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0160] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, for example, implementing the steps of the video duplicate checking method provided in the first embodiment of the present invention. The method includes:

[0161] Obtain at least one to-be-processed video frame in the video to be checked for duplicates;

[0162] For the at least one video frame to be processed, at least one similar video frame associated with the video frame to be processed is determined according to the image features of the video frame to be processed and a pre-constructed graph index; wherein, the graph index includes a plurality of nodes and edges connecting the plurality of nodes, the nodes correspond to the frame features of key frames, the edges are used to represent that the frame features of two nodes satisfy a preset similarity attribute, and the key frames are video frames extracted after segmenting historical videos into shots;

[0163] According to at least one target video to which the at least one similar video frame belongs, a duplicate-check result corresponding to the video to be duplicate-checked is determined.

[0164] Certainly, those skilled in the art can understand that the processor can also implement the technical solutions of the video duplicate-checking method provided in any embodiment of the present invention.

[0165] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the video duplicate-checking method provided in the foregoing embodiments of the present invention are implemented. The method includes:

[0166] Obtain at least one video frame to be processed in the video to be duplicate-checked;

[0167] For the at least one video frame to be processed, at least one similar video frame associated with the video frame to be processed is determined according to the image features of the video frame to be processed and a pre-constructed graph index; wherein, the graph index includes a plurality of nodes and edges connecting the plurality of nodes, the nodes correspond to the frame features of key frames, the edges are used to represent that the frame features of two nodes satisfy a preset similarity attribute, and the key frames are video frames extracted after segmenting historical videos into shots;

[0168] According to at least one target video to which the at least one similar video frame belongs, a duplicate-check result corresponding to the video to be duplicate-checked is determined.

[0169] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. The computer-readable storage media may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage media may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0170] The computer-readable signal media may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal media may also be any computer-readable media other than the computer-readable storage media, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0171] The program code contained on the computer-readable media may be transmitted using any appropriate media, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.

[0172] The computer program code for performing the operations of the embodiments of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0173] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A video duplicate checking method, characterized in that, Including: Obtaining at least one video frame to be processed in the video to be checked for duplication; For the at least one video frame to be processed, determining at least one similar video frame associated with the video frame to be processed according to the image feature of the video frame to be processed and a pre-constructed graph index; wherein, the graph index includes a plurality of nodes and edges connecting the plurality of nodes, the nodes correspond to the frame features of key frames, and the edges are used to represent that the frame features of two nodes satisfy a preset similarity attribute, and the key frames are video frames extracted after shot segmentation of a historical video; Determining a duplication check result corresponding to the video to be checked for duplication according to at least one target video to which the at least one similar video frame belongs.

2. The method according to claim 1, wherein The method further includes: For a plurality of obtained historical videos, determining transition frames in the historical videos based on a shot segmentation algorithm, so as to divide the historical videos into a plurality of video segments based on the transition frames; Performing frame extraction on the video frames in the plurality of video segments to obtain a plurality of key frames corresponding to the historical videos, and determining the graph index based on the plurality of key frames of the plurality of historical video frames.

3. The method according to claim 2, wherein The shot segmentation algorithm corresponds to a frame type classification model, and determining the transition frames in the historical videos based on the shot segmentation algorithm includes: For at least one video frame in the historical video frames, inputting the current video frame and video frames with a preset number of frames before the current video frame into a pre-trained frame type classification model, and outputting a frame type identifier of the current video frame, so as to determine that the current video frame is a transition frame when the frame type identifier is a preset identifier; Correspondingly, dividing the historical video into a plurality of video segments based on the transition frames includes: Regarding the video frames between two adjacent transition frames as a video segment according to the playing time sequence of the transition frames in the historical video.

4. The method according to claim 2, wherein Determining the graph index based on the plurality of key frames of the plurality of historical video frames includes: Performing feature extraction on the plurality of key frames based on a pre-trained feature extraction model to obtain frame features corresponding to the plurality of key frames; wherein, the frame features are represented by feature vectors with a preset dimension; Determining the similarity attribute between any two key frames based on the frame features of the plurality of key frames, and establishing connection information between the two key frames when the similarity attribute satisfies the preset similarity attribute, so as to obtain the graph index.

5. The method according to claim 4, wherein After obtaining the graph index, the method further includes: Compressing the graph index based on a sparse matrix to obtain a compressed graph index; and / or, Performing clustering processing on the nodes in the graph index to obtain an updated graph index.

6. The method according to claim 1, characterized in that, Determining at least one similar video frame associated with the video frame to be processed according to the image feature of the video frame to be processed and a pre-constructed graph index includes: Extracting the image feature of the video frame to be processed; Determining target frame features with a similarity higher than a similarity threshold to the image feature based on the image feature and the frame features corresponding to each node in the graph index; Regarding the key frames corresponding to the nodes of the target frame features as the similar video frames.

7. The method according to claim 1, characterized in that The figure index is determined based on node clustering. Determining at least one similar video frame associated with the video frame to be processed according to the image features of the video frame to be processed and the pre-constructed figure index includes: Based on the image features of the video frame to be processed and the clustering features corresponding to at least one clustering center in the figure index, determining at least one cluster associated with the image features, where the cluster includes a plurality of nodes and the clustering center after clustering of the plurality of nodes; Based on the similarity between the frame features of each node in the at least one cluster and the image features, determining the target frame features, and using the video frame corresponding to the node with the target frame features as the similar video frame.

8. The method according to claim 1, wherein Determining the duplicate check result corresponding to the video to be checked according to at least one target video to which the at least one similar video frame belongs includes: Determining at least one target video to which the at least one similar video frame belongs; For the at least one target video, determining the video similarity of the target video according to the total number of frames of the at least one video frame to be processed and the number of associated frames associated with the target video in the at least one video frame to be processed; If there is a target video with a video similarity greater than a preset threshold, determining that the duplicate check result of the video to be checked is a duplicate video.

9. The method according to claim 1, wherein When the duplicate check result is that the video to be checked is a duplicate video, the method further includes: Displaying a target interface and displaying a prompt message for the video to be checked that has not been successfully uploaded in the target interface to determine the reason for the unsuccessful upload based on the prompt message.

10. A video duplicate checking device, characterized in that, Including: A video frame to be processed extraction module, configured to obtain at least one video frame to be processed in the video to be checked; A similar video frame determination module, configured to, for the at least one video frame to be processed, determine at least one similar video frame associated with the video frame to be processed according to the image features of the video frame to be processed and the pre-constructed figure index; where the figure index includes a plurality of nodes and edges connecting the plurality of nodes, the nodes correspond to the frame features of key frames, the edges are used to represent that the frame features of two nodes satisfy a preset similarity attribute, and the key frames are video frames extracted after shot segmentation of a historical video; A duplicate check result determination module, configured to determine the duplicate check result corresponding to the video to be checked according to at least one target video to which the at least one similar video frame belongs.

11. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory, configured to store one or more programs; When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the video duplicate check method according to any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the video duplicate check method according to any one of claims 1-9.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the video duplicate check method according to any one of claims 1-9.

Citation Information

Cited By

  • Video frame determination method and apparatus, and electronic device

    CN120726548A