A highly robust method for analyzing and identifying infringement videos

By calculating the histogram feature distance between video frames and iteratively determine the number of clusters, combined with hash mapping and robust statistical representation, the problem of difficult to identify infringing videos generated by face swap or light editing in the prior art is solved, and highly robust infringing video recognition is achieved.

CN118397309BActive Publication Date: 2025-05-06ORIGINAL GUARD (WUXI) TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410497912.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2024-04-24
Publication Date
2025-05-06
Estimated Expiration
2044-04-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify infringing videos generated by face swaps or light editing, which makes it difficult for infringement to be accurately detected and identified.

Method used

By calculating the histogram feature distance between adjacent video frames, iteratively determine the number of clusters of video frame sequences, and using hash mapping to map the convolutional feature map of keyframes into vector sequences, robust statistical characterization is constructed to filter the anomaly data distribution, and finally a robust similarity measure is performed to identify infringing videos.

Benefits of technology

It realizes high robustness identification of infringing videos, can effectively identify infringing videos generated through content timing adjustment, and improves the accuracy and efficiency of infringing video detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118397309B_ABST
    Figure CN118397309B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of infringing video analysis, and discloses a highly robust infringing video analysis and identification method, the method comprising: segmenting the infringing video to be analyzed to obtain a video frame sequence, extracting the histogram features of the video frames in the video frame sequence, determining the number of clusters of the video frame sequence according to the histogram features for clustering processing, and determining the key frames of each cluster of the video frame sequence; converting the key frames of each cluster of the video frame sequence into key frame vectors; performing robust similarity measurement on the key frame vector sequence of the infringing video to be analyzed and the video key frame vector sequence in the video library, and obtaining the infringement analysis and identification result. The present invention constructs a robust statistic to characterize and filter the abnormal data distribution existing in the infringing video, uses the robust statistic to correct the key frame vector sequence and robustly measure the similarity, obtains a robust similarity measurement result that does not consider the temporal order of the key frame vector, and realizes highly robust infringing video analysis and identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of infringing video analysis, and in particular to a highly robust infringing video analysis and identification method. Background Art

[0002] With the rapid development of the Internet and digital media technology, a large number of videos have been uploaded and shared. However, there are many videos that infringe copyright, which brings economic losses and challenges to intellectual property protection to copyright holders. Some infringing videos evade infringing video detection by modifying the original video through face-changing or light editing. Such videos generated by light editing methods such as face-changing seriously affect the legitimate rights and interests of the original video. To address this problem, the present invention proposes a highly robust infringing video analysis and identification method, which realizes video content infringement detection through deep video understanding. Summary of the invention

[0003] In view of this, the present invention provides a highly robust infringing video analysis and identification method, the purpose of which is to: 1) calculate the histogram feature distance between adjacent video frames in combination with the numerical difference of histogram features in adjacent video frames and the complexity of histogram distribution, and iterate the number of clusters according to the histogram feature distance between adjacent video frames to obtain the number of clusters of the video frame sequence, and then perform fast clustering processing and key frame extraction processing on the video frame sequence, and use hash mapping to map the convolution feature map corresponding to the key frame to construct a key frame vector sequence of the video to be analyzed for infringement, and realize the key frame vector extraction of the video to be analyzed for infringement; 2) construct a robust statistic to characterize and filter the abnormal data distribution in the video to be analyzed for infringement, use the robust statistic to correct the key frame vector sequence, and perform robust similarity measurement on the key frame vector sequence of the video to be analyzed for infringement and the key frame vector sequence of the video in the video library, and obtain a robust similarity measurement result without considering the time sequence of the key frame vector, so as to effectively identify the infringing video with the content time sequence adjusted for the original video, and realize highly robust infringing video analysis and identification.

[0004] To achieve the above object, the present invention provides a highly robust infringing video analysis and identification method, comprising the following steps:

[0005] S1: Segment the video to be analyzed for infringement to obtain a video frame sequence, extract the histogram features of the video frames in the video frame sequence, and determine the number of clusters of the video frame sequence according to the histogram features;

[0006] S2: clustering the video frame sequence based on the number of clusters, and determining the key frame of each cluster of the video frame sequence;

[0007] S3: Construct a video key frame feature extraction model to convert the key frames of each cluster of video frame sequences into key frame vectors.

[0008] A sequence of key frame vectors constituting the video to be analyzed for infringement;

[0009] S4: Perform robust similarity measurement on the key frame vector sequence of the video to be analyzed for infringement and the key frame vector sequence of the video in the video library. If the robust similarity measurement result exceeds the specified threshold, it means that the content of the two is highly similar, and the video to be analyzed for infringement may contain video infringement.

[0010] As a further improvement method of the present invention:

[0011] Optionally, in step S1, segmenting the video to be analyzed for infringement to obtain a video frame sequence, and extracting histogram features of the video frames in the video frame sequence, includes:

[0012] The video to be analyzed for infringement is segmented to obtain a video frame sequence, where the video frame sequence is represented as follows:

[0013] I=(I 1 ,I 2 ,...,I n ,...,I N )

[0014] in:

[0015] I represents the video frame sequence of the video to be analyzed for infringement, I n represents the nth video frame of the video to be analyzed for infringement, and N represents the total number of video frames of the video to be analyzed for infringement;

[0016] Extract the histogram features of the video frames in the video frame sequence I, where video frame I n The histogram feature extraction process is as follows:

[0017] Calculate the video frame I n Any pixel I in row x and column y n Color value at (x,y):

[0018]

[0019]

[0020]

[0021] in:

[0022] h n (x, y) represents the video frame I n Any pixel I in row x and column y n The color value of (x, y);

[0023] In turn, pixel In (x, y) is the color value in the RGB color channel;

[0024] Count any color value i in video frame I n The number of times it appears in count n (i), forming a video frame I n Histogram features of , where i∈[0, 360];

[0025] The number of clusters of the video frame sequence is determined according to the histogram features of adjacent video frames.

[0026] Optionally, in step S1, determining the number of clusters of the video frame sequence according to the histogram features of adjacent video frames includes:

[0027] The number of clusters of the video frame sequence is determined according to the histogram features of adjacent video frames, wherein the process of determining the number of clusters of the video frame sequence I is as follows:

[0028] S11: Calculate the histogram feature distance between adjacent video frames, where video frame I n With video frame I n+1 The histogram feature distance between n :

[0029]

[0030]

[0031] in:

[0032] exp(·) represents an exponential function with a natural constant as the base;

[0033] Sum represents the total number of pixels in the video frame;

[0034] Represents video frame I n The probability measurement result of color value i in ;

[0035] S12: Calculate the average histogram feature distance dis ave :

[0036]

[0037] S13: Initialize the control parameter α and the number of clusters m, where the initial value of m is 0;

[0038] S14: Iterate the number of clusters m:

[0039]

[0040] S15: Select the number of clusters obtained by the final iteration as the number of clusters M of the video frame sequence I.

[0041] Optionally, in step S2, clustering the video frame sequence based on the number of clusters to determine the key frames of each cluster of the video frame sequence includes:

[0042] The video frame sequence is clustered based on the clustering number M to obtain several clustered video frame sequences, wherein the clustering process of the video frame sequence is as follows:

[0043] S21: Randomly select histogram features of M video frames as initial clustering centers;

[0044] S22: for the histogram features of the video frames that are not cluster centers, calculate the distance between the histogram features and the histogram features of the cluster centers, and add the histogram features of the video frames to the clusters where the nearest cluster centers are located;

[0045] S23: Calculate the histogram feature mean of each cluster, use the histogram feature mean as the cluster center of the cluster, and return to step S22;

[0046] S24: repeating steps S22-S23 until the M cluster centers no longer change, each cluster contains a number of video frames, and M cluster video frame sequences are constructed;

[0047] Determine the key frame of each cluster video frame sequence, where the video frame in the cluster video frame sequence that is closest to the cluster center histogram feature is the key frame in the cluster video frame sequence, then the key frame of the mth cluster video frame sequence is I m .

[0048] Optionally, in step S3, constructing a video key frame feature extraction model to convert key frames of each cluster of video frame sequences into key frame vectors includes:

[0049] A video key frame feature extraction model is constructed to convert the key frames of each cluster of video frame sequences into key frame vectors, and to form a key frame vector sequence of the video to be analyzed for infringement, wherein the video key frame feature extraction model includes an input layer, a convolution pooling layer, and a hash mapping layer, wherein the input layer is used to receive key frames, the convolution pooling layer is used to perform convolution pooling operations on the key frames to obtain convolution feature maps of the key frames, and the hash mapping layer is used to perform hash mapping processing on the convolution feature maps to generate key frame vectors, and the key frame I based on the video key frame feature extraction model is converted into a key frame vector. m The feature extraction process is:

[0050] S31: Input layer receives key frame I m , for key frame I m Grayscale processing is performed;

[0051] S32: Convolutional pooling layer for grayscale key frame I m Perform convolution pooling processing to obtain the convolution feature map of the key frame, where the convolution pooling processing formula is:

[0052]

[0053] in:

[0054] F 1 (I m ) represents key frame I m The convolution feature map of

[0055] W represents the convolution weight matrix, s represents the pooling step size, and T represents transpose;

[0056] Pooling(·) represents the average pooling operation;

[0057] g m Indicates key frame I m Grayscale processing result of

[0058] S33: Hash mapping layer for convolutional feature map F 1 (I m ) performs hash mapping processing to generate key frame vectors:

[0059] F m =(A·F 1 (I m )mod 2 b )rsh(br)

[0060] in:

[0061] F m Indicates key frame I m The corresponding key frame vector;

[0062] rsh(·) means right shift;

[0063] A represents an odd number, where 2 b-r <A<2 b ;

[0064] b represents the number of bits, set b to 32;

[0065] r represents the right shift step length;

[0066] mod represents the modulo operator.

[0067] Optionally, the key frame vector sequence constituting the video to be analyzed for infringement in step S3 includes:

[0068] The key frame vector sequence of the video to be analyzed for infringement is:

[0069] F=(F 1 , F 2 , ..., F m , ..., F M )

[0070] in:

[0071] F represents the key frame vector sequence of the video to be analyzed for infringement;

[0072] F m Indicates key frame I m The corresponding keyframe vector.

[0073] Optionally, in step S4, performing robust similarity measurement on the key frame vector sequence of the video to be analyzed for infringement and the key frame vector sequence of the video in the video library includes:

[0074] The key frame vector sequence F of the video to be analyzed for infringement is measured robustly with the key frame vector sequence G of the video in the video library, and the infringement of the video to be analyzed for infringement is identified based on the robust similarity measurement result, where G = (G 1 , G 2 , ..., G k ,...,G K ), G represents the video key frame vector sequence corresponding to any video in the video library, K represents the total number of key frames obtained by dividing any video in the video library, G k represents the key frame vector corresponding to the kth key frame in the video key frame vector sequence G, and the robust similarity measurement process is:

[0075] S41: Construct robust statistics:

[0076]

[0077]

[0078]

[0079] E = max{M, K}

[0080]

[0081] in:

[0082] Q M (F) represents the robust statistics of the key frame vector sequence F; Q K (G) represents the robust statistics of the video key frame vector sequence G; Q E (F, G) represents the robust statistics between the key frame vector sequence F and the video key frame vector sequence G;

[0083] β M , β K , β E are the adjustment coefficients of three robust statistics respectively;

[0084] It means that after sorting the differences between any two key frame vectors in the key frame vector sequence F in ascending order, the Key frame vector differences;

[0085] It means that after sorting the differences between any two key frame vectors in the video key frame vector sequence G in ascending order, the first Key frame vector differences;

[0086] It means that after sorting the differences between any two key frame vectors between the key frame vector sequence F and the video key frame vector sequence G in ascending order, the first Key frame vector differences;

[0087] S42: Based on the robust statistics, the robust similarity between different key frame vectors is calculated, and a robust similarity matrix is ​​constructed:

[0088]

[0089]

[0090]

[0091] Y=min{M,K}

[0092] in:

[0093] R F represents the robust similarity matrix of the key frame vector sequence F, R F (1, Y) represents the key frame vector F 1 With the key frame vector F Y Robust similarity between ;

[0094] R G represents the robust similarity matrix of the video key frame vector sequence G, G F (1, Y) represents the key frame vector G 1 With the key frame vector G Y Robust similarity between ;

[0095] R F,G represents the robust similarity matrix between the key frame vector sequence F and the video key frame vector sequence G, R F,G (1, Y) represents the key frame vector F 1 With the key frame vector GY Robust similarity between ;

[0096] In the embodiment of the present invention, the calculation formula of the robust similarity between different key frame vectors is as follows:

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106] in:

[0107] μ F , μ G They represent the key frame vector means of the key frame vector sequence F and the video key frame vector sequence G respectively;

[0108] R F (m, m') represents F m , F m′ The robust similarity between G (k, k') represents G k , G k′ Robust similarity between ;

[0109] R F,G (m, k) represents F m , G k Robust similarity between ;

[0110] S43: Calculate the robust similarity measurement result between the key frame vector sequence F of the video to be analyzed for infringement and the video key frame vector sequence G in the video library:

[0111]

[0112] in:

[0113] Sim(F, G) represents the robust similarity measurement result between the key frame vector sequence F of the video to be analyzed for infringement and the key frame vector sequence G of the video in the video library.

[0114] Optionally, in step S4, performing infringement identification of the video to be analyzed for infringement according to the robust similarity measurement result includes:

[0115] If the robust similarity measurement result Sim(F, G) exceeds the specified threshold, it means that the content of the video to be analyzed for infringement is highly similar to that of any video in the video library, and the video to be analyzed for infringement may contain video infringement.

[0116] In order to solve the above problem, the present invention provides an electronic device, the electronic device comprising:

[0117] A memory storing at least one instruction;

[0118] Communication interface, enabling electronic equipment to communicate; and

[0119] The processor executes the instructions stored in the memory to implement the highly robust infringement video analysis and identification method described above.

[0120] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is executed by a processor in an electronic device to implement the highly robust infringement video analysis and identification method described above.

[0121] Compared with the prior art, the present invention proposes a highly robust infringing video analysis and identification method, which has the following advantages:

[0122] First, this scheme proposes a method for calculating the number of clusters of a video frame sequence, extracting the histogram features of the video frames in the video frame sequence I, where video frame I n The histogram feature extraction process is as follows: Calculate the video frame I n Any pixel I in row x and column y n Color value at (x, y):

[0123]

[0124]

[0125]

[0126] Where: h n (x, y) represents the video frame I n Any pixel I in row x and column y n The color value of (x, y); In turn, pixel I n (x, y) is the color value in the RGB color channel; counts any color value i in video frame In The number of times it appears in count n (i), forming a video frame I n The histogram features of i∈[0,360] are used to determine the number of clusters of the video frame sequence according to the histogram features of adjacent video frames. The process for determining the number of clusters of the video frame sequence I is as follows: the histogram feature distance of adjacent video frames is calculated, where video frame I n With video frame I n+1 The histogram feature distance between n :

[0127]

[0128]

[0129] Where: exp(·) represents an exponential function with a natural constant as the base; Sum represents the total number of pixels in the video frame; Represents video frame I n The probability measurement result of the color value i in the equation is: the average histogram feature distance dis is calculated ave :

[0130] Initialize the control parameter α and the number of clusters m, where the initial value of m is 0; iterate the number of clusters m:

[0131]

[0132] The number of clusters obtained by the final iteration is selected as the number of clusters M of the video frame sequence I. This scheme combines the numerical difference of the histogram features in adjacent video frames and the complexity of the histogram distribution to calculate the histogram feature distance between adjacent video frames, and iterates the number of clusters according to the histogram feature distance between adjacent video frames to obtain the number of clusters of the video frame sequence, and then performs fast clustering processing and key frame extraction processing on the video frame sequence, and uses a hash mapping method to map the convolution feature map corresponding to the key frame, constructs a key frame vector sequence of the video to be analyzed for infringement, and realizes the key frame vector extraction of the video to be analyzed for infringement.

[0133] At the same time, this scheme proposes a robust similarity measurement method, which performs robust similarity measurement on the key frame vector sequence F of the video to be analyzed for infringement and the key frame vector sequence G of the video in the video library, and performs infringement identification of the video to be analyzed for infringement based on the robust similarity measurement result, where G = (G 1 , G 2 , ..., G k , ..., G K), G represents the video key frame vector sequence corresponding to any video in the video library, K represents the total number of key frames obtained by dividing any video in the video library, G k represents the key frame vector corresponding to the kth key frame in the video key frame vector sequence G. The robust similarity measurement process is as follows: construct a robust statistic:

[0134]

[0135]

[0136]

[0137] E = max{M, K}

[0138]

[0139] Where: Q M (F) represents the robust statistics of the key frame vector sequence F; Q K (G) represents the robust statistics of the video key frame vector sequence G; Q E (F, G) represents the robust statistics between the key frame vector sequence F and the video key frame vector sequence G; β M , β K , β E are the adjustment coefficients of three robust statistics respectively; It means that after sorting the differences between any two key frame vectors in the key frame vector sequence F in ascending order, the Key frame vector differences; It means that after sorting the differences between any two key frame vectors in the video key frame vector sequence G in ascending order, the first Key frame vector differences; It means that after sorting the differences between any two key frame vectors between the key frame vector sequence F and the video key frame vector sequence G in ascending order, the first The key frame vector differences are calculated based on the robust statistics to obtain the robust similarity between different key frame vectors, and the robust similarity matrix is ​​constructed:

[0140]

[0141]

[0142]

[0143] Y=min{M,K}

[0144] Where: R F represents the robust similarity matrix of the key frame vector sequence F, R F(1, Y) represents the key frame vector F 1 With the key frame vector F Y The robust similarity between G represents the robust similarity matrix of the video key frame vector sequence G, G F (1, Y) represents the key frame vector G 1 With the key frame vector G Y The robust similarity between F,G represents the robust similarity matrix between the key frame vector sequence F and the video key frame vector sequence G, R F,G (1, Y) represents the key frame vector F 1 With the key frame vector G Y The robust similarity between the key frame vector sequence F of the video to be analyzed for infringement and the video key frame vector sequence G in the video library is calculated as follows:

[0145]

[0146] in:

[0147] Sim(F, G) represents the robust similarity measurement result between the key frame vector sequence F of the video to be analyzed for infringement and the key frame vector sequence G of the video in the video library. If the robust similarity measurement result Sim(F, G) exceeds the specified threshold, it means that the content similarity between the video to be analyzed for infringement and any video in the video library is high, and the video to be analyzed for infringement may contain video infringement. This scheme constructs robust statistics to characterize and filter the abnormal data distribution in the video to be analyzed for infringement, such as abnormal key frame vectors caused by face-changing or light editing, and uses robust statistics to correct the key frame vector sequence. The key frame vector sequence of the video to be analyzed for infringement is robustly measured with the key frame vector sequence of the video in the video library to obtain a robust similarity measurement result that does not consider the temporal order of the key frame vectors, effectively identifying the infringing video that has adjusted the content timing of the original video, and achieving highly robust analysis and recognition of infringing videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0148] Figure 1 A flowchart of a highly robust infringing video analysis and identification method provided by an embodiment of the present invention;

[0149] Figure 2 A schematic diagram of the structure of an electronic device for implementing a highly robust infringing video analysis and identification method provided by an embodiment of the present invention.

[0150] In the figure: 1 electronic device, 10 processor, 11 memory, 12 program, 13 communication interface.

[0151] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0152] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0153] The embodiment of the present application provides a highly robust infringing video analysis and identification method. The execution subject of the highly robust infringing video analysis and identification method includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided by the embodiment of the present application. In other words, the highly robust infringing video analysis and identification method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.

[0154] Embodiment 1:

[0155] S1: Segment the video to be analyzed for infringement to obtain a video frame sequence, extract the histogram features of the video frames in the video frame sequence, and determine the number of clusters of the video frame sequence according to the histogram features.

[0156] In the step S1, the video to be analyzed for infringement is segmented to obtain a video frame sequence, and histogram features of the video frames in the video frame sequence are extracted, including:

[0157] The video to be analyzed for infringement is segmented to obtain a video frame sequence, where the video frame sequence is represented as follows:

[0158] I=(I 1 , I 2 , ..., I n , ..., I N )

[0159] in:

[0160] I represents the video frame sequence of the video to be analyzed for infringement, I n represents the nth video frame of the video to be analyzed for infringement, and N represents the total number of video frames of the video to be analyzed for infringement;

[0161] Extract the histogram features of the video frames in the video frame sequence I, where video frame I n The histogram feature extraction process is as follows:

[0162] Calculate the video frame I n Any pixel I in row x and column y n Color value at (x, y):

[0163]

[0164]

[0165]

[0166] in:

[0167] h n (x, y) represents the video frame I n Any pixel I in row x and column y n The color value of (x, y);

[0168] In turn, pixel I n (x, y) is the color value in the RGB color channel;

[0169] Count any color value i in video frame I n The number of times it appears in count n (i), forming a video frame I n Histogram features of , where i∈[0, 360];

[0170] The number of clusters of the video frame sequence is determined according to the histogram features of adjacent video frames.

[0171] In the step S1, the number of clusters of the video frame sequence is determined according to the histogram features of adjacent video frames, including:

[0172] The number of clusters of the video frame sequence is determined according to the histogram features of adjacent video frames, wherein the process of determining the number of clusters of the video frame sequence I is as follows:

[0173] S11: Calculate the histogram feature distance between adjacent video frames, where video frame I n With video frame I n+1 The histogram feature distance between n :

[0174]

[0175]

[0176] in:

[0177] exp(·) represents an exponential function with a natural constant as the base;

[0178] Sum represents the total number of pixels in the video frame;

[0179] Represents video frame I n The probability measurement result of color value i in ;

[0180] S12: Calculate the average histogram feature distance dis ave :

[0181]

[0182] S13: Initialize the control parameter α and the number of clusters m, where the initial value of m is 0;

[0183] S14: Iterate the number of clusters m:

[0184]

[0185] S15: Select the number of clusters obtained by the final iteration as the number of clusters M of the video frame sequence I.

[0186] S2: clustering the video frame sequence based on the number of clusters, and determining the key frames of each cluster of the video frame sequence.

[0187] In step S2, clustering the video frame sequence based on the number of clusters to determine the key frames of each cluster of the video frame sequence includes:

[0188] The video frame sequence is clustered based on the clustering number M to obtain several clustered video frame sequences, wherein the clustering process of the video frame sequence is as follows:

[0189] S21: Randomly select histogram features of M video frames as initial clustering centers;

[0190] S22: for the histogram features of the video frames that are not cluster centers, calculate the distance between the histogram features and the histogram features of the cluster centers, and add the histogram features of the video frames to the clusters where the nearest cluster centers are located;

[0191] S23: Calculate the histogram feature mean of each cluster, use the histogram feature mean as the cluster center of the cluster, and return to step S22;

[0192] S24: repeating steps S22-S23 until the M cluster centers no longer change, each cluster contains a number of video frames, and M cluster video frame sequences are constructed;

[0193] Determine the key frame of each cluster video frame sequence, where the video frame in the cluster video frame sequence that is closest to the cluster center histogram feature is the key frame in the cluster video frame sequence, then the key frame of the mth cluster video frame sequence is I m .

[0194] S3: Construct a video key frame feature extraction model to convert the key frames of each cluster of video frame sequences into key frame vectors to form a key frame vector sequence of the video to be analyzed for infringement.

[0195] In the step S3, a video key frame feature extraction model is constructed to convert the key frames of each cluster of video frame sequences into key frame vectors, including:

[0196] A video key frame feature extraction model is constructed to convert the key frames of each cluster of video frame sequences into key frame vectors, and to form a key frame vector sequence of the video to be analyzed for infringement, wherein the video key frame feature extraction model includes an input layer, a convolution pooling layer, and a hash mapping layer, wherein the input layer is used to receive key frames, the convolution pooling layer is used to perform convolution pooling operations on the key frames to obtain convolution feature maps of the key frames, and the hash mapping layer is used to perform hash mapping processing on the convolution feature maps to generate key frame vectors, and the key frame I based on the video key frame feature extraction model is converted into a key frame vector. m The feature extraction process is:

[0197] S31: Input layer receives key frame I m , for key frame I m Grayscale processing is performed;

[0198] S32: Convolutional pooling layer for grayscale key frame I m Perform convolution pooling processing to obtain the convolution feature map of the key frame, where the convolution pooling processing formula is:

[0199]

[0200] in:

[0201] F 1 (I m ) represents key frame I m The convolution feature map of

[0202] W represents the convolution weight matrix, s represents the pooling step size, and T represents transpose;

[0203] Pooling(·) represents the average pooling operation;

[0204] g m Indicates key frame I m Grayscale processing result of

[0205] S33: Hash mapping layer for convolutional feature map F 1 (I m ) performs hash mapping processing to generate a key frame vector:

[0206] F m =(A·F 1 (I m )mod 2 b )rsh(br)

[0207] in:

[0208] Fm Indicates key frame I m The corresponding key frame vector;

[0209] rsh(·) means right shift;

[0210] A represents an odd number, where 2 b-r <A<2 b ;

[0211] b represents the number of bits, set b to 32;

[0212] r represents the right shift step length;

[0213] mod represents the modulo operator.

[0214] The key frame vector sequence of the video to be analyzed for infringement in step S3 includes:

[0215] The key frame vector sequence of the video to be analyzed for infringement is:

[0216] F=(F 1 , F 2 , ..., F m , ..., F M )

[0217] in:

[0218] F represents the key frame vector sequence of the video to be analyzed for infringement;

[0219] F m Indicates key frame I m The corresponding keyframe vector.

[0220] S4: Perform robust similarity measurement on the key frame vector sequence of the video to be analyzed for infringement and the key frame vector sequence of the video in the video library. If the robust similarity measurement result exceeds the specified threshold, it means that the content of the two is highly similar, and the video to be analyzed for infringement may contain video infringement.

[0221] In the step S4, the key frame vector sequence of the video to be analyzed for infringement is subjected to robust similarity measurement with the key frame vector sequence of the video in the video library, including:

[0222] The key frame vector sequence F of the video to be analyzed for infringement is measured robustly with the key frame vector sequence G of the video in the video library, and the infringement of the video to be analyzed for infringement is identified based on the robust similarity measurement result, where G = (G 1 , G 2 , ..., G k , ..., G K), G represents the video key frame vector sequence corresponding to any video in the video library, K represents the total number of key frames obtained by dividing any video in the video library, G k represents the key frame vector corresponding to the kth key frame in the video key frame vector sequence G, and the robust similarity measurement process is:

[0223] S41: Construct robust statistics:

[0224]

[0225]

[0226]

[0227] E = max{M, K}

[0228]

[0229] in:

[0230] Q M (F) represents the robust statistics of the key frame vector sequence F; Q K (G) represents the robust statistics of the video key frame vector sequence G; Q E (F, G) represents the robust statistics between the key frame vector sequence F and the video key frame vector sequence G;

[0231] β M , β K , β E are the adjustment coefficients of three robust statistics respectively;

[0232] It means that after sorting the differences between any two key frame vectors in the key frame vector sequence F in ascending order, the Key frame vector differences;

[0233] It means that after sorting the differences between any two key frame vectors in the video key frame vector sequence G in ascending order, the first Key frame vector differences;

[0234] It means that after sorting the differences between any two key frame vectors between the key frame vector sequence F and the video key frame vector sequence G in ascending order, the first Key frame vector differences;

[0235] S42: Based on the robust statistics, the robust similarity between different key frame vectors is calculated, and a robust similarity matrix is ​​constructed:

[0236]

[0237]

[0238]

[0239] Y=min{M,K}

[0240] in:

[0241] R F represents the robust similarity matrix of the key frame vector sequence F, R F (1, Y) represents the key frame vector F 1 With the key frame vector F Y Robust similarity between ;

[0242] R G represents the robust similarity matrix of the video key frame vector sequence G, G F (1, Y) represents the key frame vector G 1 With the key frame vector G Y Robust similarity between ;

[0243] R F,G represents the robust similarity matrix between the key frame vector sequence F and the video key frame vector sequence G, R F,G (1, Y) represents the key frame vector F 1 With the key frame vector G Y Robust similarity between ;

[0244] S43: Calculate the robust similarity measurement result between the key frame vector sequence F of the video to be analyzed for infringement and the video key frame vector sequence G in the video library:

[0245]

[0246] in:

[0247] Sim(F, G) represents the robust similarity measurement result between the key frame vector sequence F of the video to be analyzed for infringement and the key frame vector sequence G of the video in the video library.

[0248] In the step S4, the infringement identification of the video to be analyzed for infringement is performed according to the robust similarity measurement result, including:

[0249] If the robust similarity measurement result Sim(F, G) exceeds the specified threshold, it means that the content of the video to be analyzed for infringement is highly similar to that of any video in the video library, and the video to be analyzed for infringement may contain video infringement.

[0250] Embodiment 2:

[0251] like Figure 2, is a schematic diagram of the structure of an electronic device for implementing a highly robust infringing video analysis and identification method provided by an embodiment of the present invention.

[0252] The electronic device 1 may include a processor 10 , a memory 11 , a communication interface 13 and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 10 , such as a program 12 .

[0253] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. The memory 11 may be an internal storage unit of the electronic device 1 in some embodiments, such as a mobile hard disk of the electronic device 1. The memory 11 may also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 1. Further, the memory 11 may also include both an internal storage unit of the electronic device 1 and an external storage device. The memory 11 may be used not only to store application software and various types of data installed in the electronic device 1, such as the code of the program 12, etc., but also to temporarily store data that has been output or is to be output.

[0254] In some embodiments, the processor 10 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes or executes programs or modules stored in the memory 11 (such as the program 12 for realizing highly robust infringing video analysis and identification), and calls data stored in the memory 11 to execute various functions of the electronic device 1 and process data.

[0255] The communication interface 13 may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices, and to achieve connection and communication between internal components of the electronic device.

[0256] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.

[0257] Figure 2 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 2 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0258] For example, although not shown, the electronic device 1 may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, etc. The electronic device 1 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.

[0259] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), or a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 1 and to display a visual user interface.

[0260] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.

[0261] The program 12 stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve:

[0262] Segmenting the video to be analyzed for infringement to obtain a video frame sequence, extracting histogram features of the video frames in the video frame sequence, and determining the number of clusters of the video frame sequence according to the histogram features;

[0263] The video frame sequence is clustered based on the number of clusters, and a key frame of each cluster of the video frame sequence is determined;

[0264] Construct a video key frame feature extraction model to convert the key frames of each cluster of video frame sequences into key frame vectors to form a key frame vector sequence of the video to be analyzed for infringement;

[0265] The key frame vector sequence of the video to be analyzed for infringement is robustly measured with the key frame vector sequence of the video in the video library. If the robust similarity measurement result exceeds the specified threshold, it means that the content of the two is highly similar, and the video to be analyzed for infringement may contain video infringement.

[0266] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to Figure 1 to Figure 2 The description of the relevant steps in the corresponding embodiments will not be repeated here.

[0267] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. And the terms "including", "comprising" or any other variants thereof in this article are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0268] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0269] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A highly robust infringing video analysis and identification method, characterized in that: The method comprises: S1: Segment the video to be analyzed for infringement to obtain a video frame sequence, extract the histogram features of the video frames in the video frame sequence, and determine the number of clusters of the video frame sequence according to the histogram features; The representation of a video frame sequence is: I=(I1,I2,...,I n ,...,I N ) in: I represents the video frame sequence of the video to be analyzed for infringement, I n represents the nth video frame of the video to be analyzed for infringement, and N represents the total number of video frames of the video to be analyzed for infringement; Extract the histogram features of the video frames in the video frame sequence I, where video frame I n The histogram feature extraction process is: Calculate the video frame I n Any pixel I in row x and column y n Color value at (x,y): in: h n (x,y) represents the video frame I n Any pixel I in row x and column y n The color value of (x,y); In turn, pixel I n (x,y) is the color value in the RGB color channel; Count any color value i in video frame I n The number of times it appears in count n (i), forming a video frame I n Histogram features of , where i∈[0,360]; Determining the number of clusters of the video frame sequence according to the histogram features of adjacent video frames; The process of determining the number of clusters of the video frame sequence I is as follows: S11: Calculate the histogram feature distance between adjacent video frames, where video frame I n With video frame I n+1 The histogram feature distance between n : in: exp(·) represents an exponential function with a natural constant as the base; Sum represents the total number of pixels in the video frame; Represents video frame I n The probability measurement result of color value i in ; S12: Calculate the average histogram feature distance dis ave : S13: Initialize the control parameter α and the number of clusters m, where the initial value of m is 0; S14: Iterate the number of clusters m: S15: selecting the number of clusters obtained by the final iteration as the number of clusters M of the video frame sequence I; S2: clustering the video frame sequence based on the number of clusters, and determining the key frame of each cluster of the video frame sequence; The video frame sequence is clustered based on the clustering number M to obtain several clustered video frame sequences, wherein the clustering process of the video frame sequence is as follows: S21: Randomly select histogram features of M video frames as initial clustering centers; S22: for the histogram features of the video frames that are not cluster centers, calculate the distance between the histogram features and the histogram features of the cluster centers, and add the histogram features of the video frames to the clusters where the nearest cluster centers are located; S23: Calculate the histogram feature mean of each cluster, use the histogram feature mean as the cluster center of the cluster, and return to step S22; S24: repeating steps S22-S23 until the M cluster centers no longer change, each cluster contains a number of video frames, and M cluster video frame sequences are constructed; Determine the key frame of each cluster video frame sequence, where the video frame in the cluster video frame sequence that is closest to the cluster center histogram feature is the key frame in the cluster video frame sequence, then the key frame of the mth cluster video frame sequence is I m ; S3: Construct a video key frame feature extraction model to convert the key frames of each cluster of video frame sequences into key frame vectors to form a key frame vector sequence of the video to be analyzed for infringement; S4: Performing robust similarity measurement on the key frame vector sequence of the video to be analyzed for infringement and the key frame vector sequence of the video in the video library. If the robust similarity measurement result exceeds a specified threshold, it indicates that the similarity between the two contents is high, and the video to be analyzed for infringement may contain video infringement behavior; The key frame vector sequence F of the video to be analyzed for infringement is measured robustly with the key frame vector sequence G of the video in the video library, and the infringement of the video to be analyzed for infringement is identified based on the robust similarity measurement result, where G = (G1, G2, ..., G k ,...,G K ), G represents the video key frame vector sequence corresponding to any video in the video library, K represents the total number of key frames obtained by dividing any video in the video library, G k represents the key frame vector corresponding to the kth key frame in the video key frame vector sequence G, and the robust similarity measurement process is: S41: Construct robust statistics: E=max{M,K} in: Q M (F) represents the robust statistics of the key frame vector sequence F; Q K (G) represents the robust statistics of the video key frame vector sequence G; Q E (F, G) represents the robust statistics between the key frame vector sequence F and the video key frame vector sequence G; β M ,β K ,β E are the adjustment coefficients of three robust statistics respectively; It means that after sorting the differences between any two key frame vectors in the key frame vector sequence F in ascending order, the Key frame vector differences; It means that after sorting the differences between any two key frame vectors in the video key frame vector sequence G in ascending order, the first Key frame vector differences; It means that after sorting the differences between any two key frame vectors between the key frame vector sequence F and the video key frame vector sequence G in ascending order, the first Key frame vector differences; S42: Based on the robust statistics, the robust similarity between different key frame vectors is calculated, and a robust similarity matrix is ​​constructed: Y=min{M,K} in: R F represents the robust similarity matrix of the key frame vector sequence F, R F (1,Y) represents the key frame vector F1 and the key frame vector F Y Robust similarity between ; R G represents the robust similarity matrix of the video key frame vector sequence G, G F (1,Y) represents the key frame vector G1 and the key frame vector G Y Robust similarity between ; R F,G represents the robust similarity matrix between the key frame vector sequence F and the video key frame vector sequence G, R F,G (1,Y) represents the key frame vector F1 and the key frame vector G Y Robust similarity between ; S43: Calculate the robust similarity measurement result between the key frame vector sequence F of the video to be analyzed for infringement and the video key frame vector sequence G in the video library: in: Sim(F,G) represents the robust similarity measurement result between the key frame vector sequence F of the video to be analyzed for infringement and the key frame vector sequence G of the video in the video library.

2. A highly robust infringing video analysis and identification method as claimed in claim 1, characterized in that: In the step S3, a video key frame feature extraction model is constructed to convert the key frames of each cluster of video frame sequences into key frame vectors, including: A video key frame feature extraction model is constructed to convert the key frames of each cluster of video frame sequences into key frame vectors, and to form a key frame vector sequence of the video to be analyzed for infringement, wherein the video key frame feature extraction model includes an input layer, a convolution pooling layer, and a hash mapping layer, wherein the input layer is used to receive key frames, the convolution pooling layer is used to perform convolution pooling operations on the key frames to obtain convolution feature maps of the key frames, and the hash mapping layer is used to perform hash mapping processing on the convolution feature maps to generate key frame vectors, and the key frame I based on the video key frame feature extraction model is converted into a key frame vector. m The feature extraction process is: S31: Input layer receives key frame I m , for key frame I m Grayscale processing is performed; S32: Convolutional pooling layer for grayscale key frame I m Perform convolution pooling processing to obtain the convolution feature map of the key frame, where the convolution pooling processing formula is: in: F1(I m ) represents key frame I m The convolutional feature map of W represents the convolution weight matrix, s represents the pooling step size, and T represents transpose; Pooling(·) represents the average pooling operation; g m Indicates key frame I m Grayscale processing result of S33: Hash mapping layer for convolutional feature map F1 (I m ) performs hash mapping processing to generate a key frame vector: F m =(A·F1(I m )mod 2 b )rsh(b―r) in: F m Indicates key frame I m The corresponding keyframe vector; rsh(·) means right shift; A represents an odd number, where 2 b―r <A<2 b ; b represents the number of bits, set b to 32; r represents the right shift step length; mod represents the modulo operator.

3. A highly robust infringing video analysis and identification method as claimed in claim 1, characterized in that: The key frame vector sequence of the video to be analyzed for infringement in step S3 includes: The key frame vector sequence of the video to be analyzed for infringement is: F=(F1,F2,...,F m ,...,F M ) in: F represents the key frame vector sequence of the video to be analyzed for infringement; F m Indicates key frame I m The corresponding keyframe vector.

4. A highly robust infringing video analysis and identification method as claimed in claim 1, characterized in that: In the step S4, the infringement identification of the video to be analyzed for infringement is performed according to the robust similarity measurement result, including: If the robust similarity measurement result Sim(F,G) exceeds the specified threshold, it means that the content of the video to be analyzed for infringement is highly similar to that of any video in the video library, and the video to be analyzed for infringement may contain video infringement.

Citation Information

Patent Citations

  • Data visualization color matching extraction method based on image clustering

    CN111768469A

  • Video poster automatic generation method

    CN112004164A

  • Video to-be-retrieved positioning method applied to video copyright protection

    CN112395457A

  • Video processing method and device, electronic equipment and readable storage medium

    CN113313065A