Video duplicate clip detection method and device, computing equipment, storage medium and program product

By combining feature extraction and similarity matrix analysis with Hough voting and maximum subarray algorithms, the problem of detecting repeated content in live streams and videos is solved, achieving high accuracy and high recall in the identification and localization of repeated segments. It is suitable for detecting carousel playback in live streams and videos.

CN120877192APending Publication Date: 2025-10-31SHANGHAI HODE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511115095.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

The problem of repetitive playback of the same content in live streaming and video playback leads to resource waste and a decline in user experience. Existing detection methods are not accurate or comprehensive enough.

Method used

By extracting features from the video to be detected, obtaining the first and second feature matrices, constructing a similarity matrix, filtering out elements with similarity greater than a preset threshold, using the Hough voting algorithm to select candidate segments, and using the maximum subarray sum algorithm to determine duplicate segments, the detection is performed by combining video and audio features.

Benefits of technology

It achieves accurate identification and location of repetitive segments in videos, improving detection accuracy and recall. It can promptly detect periodic repetitions and is suitable for various application scenarios, including detection of carousel playback between live streaming rooms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877192A_ABST
    Figure CN120877192A_ABST
Patent Text Reader

Abstract

The invention discloses a video repeated clip detection method and device, computing equipment, a storage medium and a program product, and the method comprises the steps: carrying out the feature extraction of a to-be-detected video, and obtaining a first feature matrix and a second feature matrix of the to-be-detected video; obtaining a similarity matrix based on the first feature matrix and the second feature matrix; in the similarity matrix, elements with the similarity larger than a preset threshold value are screened out, and an element set is constructed; candidate segments are determined according to the element set, and each candidate segment corresponds to a linear array in a similarity matrix; for each candidate segment, obtaining a sub-array with the maximum sum in the linear array corresponding to the candidate segment; and determining whether the candidate fragment is a duplicate fragment based on the sub-array with the maximum sum.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video playback technology, and in particular to a method, apparatus, computing device, storage medium, and program product for detecting repeated video segments. Background Technology

[0002] Live streaming and video playback have become important channels for disseminating and providing information. Users can not only learn about products and watch content of interest through these methods, but also learn knowledge and skills and participate in interactive exchanges. Therefore, ensuring high-quality live streaming and video playback content is a key aspect of meeting user needs.

[0003] However, in practical applications, live streaming often involves looping content, where a single live stream plays the same content in a loop, or different live streams play the same content across different streams. Similarly, in video playback, there are instances of long videos repeating the same content or different videos containing the same information.

[0004] For platforms, playing repetitive videos wastes resources such as bandwidth and computing power. This is especially true in live streaming scenarios, where replaying low-quality, repetitive content leads to resource waste and reduces platform efficiency. Furthermore, it results in a lack of interaction between streamers and users, and there are even instances of streamers artificially inflating their viewing time to fraudulently obtain platform subsidies.

[0005] For users, repetitive content lacks novelty and interactivity, and the repetition of content in multiple live streams or videos severely impacts the viewing experience.

[0006] The information disclosed in this background section is included only to enhance the understanding of the context of this disclosure, and therefore may contain information that does not constitute relevant technology currently known to those skilled in the art. Summary of the Invention

[0007] This application provides a method, apparatus, computing device, storage medium, and program product for detecting repeated video segments, in order to solve the problem of detecting repeated playback of the same content in videos or live streams.

[0008] The technical solution adopted in this application is as follows.

[0009] Firstly, this application provides a method for detecting duplicate video segments, including: Feature extraction is performed on the video to be detected to obtain the first feature matrix and the second feature matrix of the video to be detected. Based on the first feature matrix and the second feature matrix, a similarity matrix is ​​obtained; In the similarity matrix, elements with similarity greater than a preset threshold are selected to construct an element set; Based on the set of elements, candidate segments are determined, and each candidate segment corresponds to a linear array in the similarity matrix; For each candidate segment: obtain the subarray with the largest sum in its corresponding linear array; Based on the subarray with the largest sum, it is determined whether the candidate segment is a duplicate segment, so as to determine whether the video to be detected includes duplicate segments.

[0010] Thus, this application can effectively identify whether there is repeated content in a video, and select candidate segments based on the set of elements to improve detection accuracy. Furthermore, by selecting the subarray with the largest sum, the accuracy of detecting repeated segments can be further improved, and the start and end times of repeated segments can be precisely located.

[0011] In conjunction with the first aspect, in one possible implementation, determining the candidate fragment based on the set of elements includes: Obtain the time difference of each element in the set of elements; Group elements with the same time difference into the same category; Select the k categories with the most elements, and use the elements in the k categories as the candidate fragments, where k is a positive integer.

[0012] Thus, this application improves the accuracy of candidate segments by mapping elements to the time difference parameter space.

[0013] In conjunction with the first aspect, in one possible implementation, determining whether the candidate segment is a duplicate segment based on the subarray with the largest sum includes: If the average similarity of elements in the subarray is greater than a preset average, then the video to be detected is determined to contain duplicate segments.

[0014] Thus, this application improves detection readiness by setting a preset average value.

[0015] In conjunction with the first aspect, in one possible implementation, the method further includes: Obtain the start and end points of the subarray; Based on the start point and the end point, obtain the start time and end time of the repeating segment.

[0016] Thus, this application can not only detect whether the video to be detected contains repeated segments, but also accurately locate the start and end times of the repeated segments.

[0017] In conjunction with the first aspect, in one possible implementation, the video to be detected includes a live video from a live streaming room; The first feature matrix is ​​the first audio vector matrix of the live video during the time period [-t, 0], and the second feature matrix is ​​the second audio vector matrix of the live video during the time period [-T, -t]; or The first feature matrix is ​​the first image vector matrix of the live video in the [-t,0] time period, and the second feature matrix is ​​the second image vector matrix of the live video in the [-T,-t] time period or the third image vector matrix of the live video in the [-t,0] time period; The method further includes: When it is determined, based on the first audio vector matrix and the second audio vector matrix, that the video to be detected contains repeating segments, and based on the first image vector matrix and the second image vector matrix, it is determined that the live stream contains repeating segments; or When it is determined that the video to be detected contains repeated segments based on the first audio vector matrix and the second audio vector matrix, and also determined that the video to be detected contains repeated segments based on the first image vector matrix and the third image vector matrix, it is determined that the live broadcast room has a loop. Wherein, T and t are both preset time values, and the value of t is less than the value of T. The start time of the [-t,0] time period is the time obtained by shifting t forward from the current time, and the end time is the current time. The start time of the [-T,-t] time period is the time obtained by shifting T forward from the current time, and the end time is the time obtained by shifting t forward from the current time.

[0018] Thus, this application can achieve more comprehensive detection by combining video and audio features, and can be applied to various application scenarios such as in-segment rotation, in-segment rotation, and in-segment rotation between different live broadcast sessions.

[0019] In conjunction with the first aspect, in one possible implementation, the method further includes: At preset time intervals, feature extraction is performed on the video to be detected to obtain the first feature matrix and the second feature matrix.

[0020] Thus, this application enables periodic detection of the video to be tested, and can promptly detect repeated playback.

[0021] In conjunction with the first aspect, in one possible implementation, the video to be detected includes a first live video from a first live room during a preset time period and a second live video from a second live room during the preset time period, wherein the first feature matrix is ​​the image vector matrix of the first live video and the second feature matrix is ​​the image vector matrix of the second live video. The method further includes: When it is determined, based on the first live video and the second live video, that there are repeated segments in the video to be detected, it is determined that the first live room and the second live room are in rotation.

[0022] Thus, this application enables the detection of live stream rotation between live streaming rooms.

[0023] In conjunction with the first aspect, in one possible implementation, the method further includes: Obtain key identifiers from multiple live streaming rooms to be detected; Group live streams with the same key identifier into the same group; Each live room in the same group is paired with the other live rooms in the same group to form all possible live room pairs, and the two live rooms in each live room pair are respectively designated as the first live room and the second live room.

[0024] Thus, this application enables comprehensive detection of live stream rotation between live streaming rooms.

[0025] In conjunction with the first aspect, in one possible implementation, the video to be detected includes a first video and a second video, the first feature matrix is ​​the image vector matrix and / or audio vector matrix of the first video, and the second feature matrix is ​​the image vector matrix and / or audio vector matrix of the second video; the method further includes: When it is determined, based on the first video and the second video, that there are repeated segments in the video to be detected, it is determined that the first video and the second video contain repeated playback content.

[0026] Thus, this application can detect the rotation of live streams between different rooms, improving detection accuracy.

[0027] Secondly, this application also provides a video repeating segment detection device, comprising: The feature extraction module is used to extract features from the video to be detected, and obtain the first feature matrix and the second feature matrix of the video to be detected respectively. A similarity acquisition module is used to acquire a similarity matrix based on the first feature matrix and the second feature matrix; The element set construction module is used to filter out elements with a similarity greater than a preset threshold from the similarity matrix and construct an element set. A candidate determination module is used to determine candidate segments based on the set of elements, wherein each candidate segment corresponds to a linear array in the similarity matrix; The subarray acquisition module is used to acquire, for each candidate segment, the subarray with the maximum sum in its corresponding linear array; The duplicate segment determination module is used to determine whether the candidate segment is a duplicate segment based on the subarray with the maximum sum, so as to determine whether the video to be detected includes duplicate segments.

[0028] For more detailed implementation information on the video duplicate segment detection device, please refer to the description of any of the implementation methods in the first aspect above.

[0029] Thirdly, this application provides a computing device, characterized in that it includes a memory and a processor, wherein the memory is used to store computer programs or instructions; when the computer programs or instructions are executed by the processor, they implement the method in the first aspect or any possible implementation of the first aspect.

[0030] Fourthly, this application also provides a computer-readable storage medium storing a computer program or instructions that, when executed by a processor, implement the method in the first aspect or any possible implementation of the first aspect.

[0031] Fifthly, this application also provides a computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, it implements the method in the first aspect or any possible implementation of the first aspect.

[0032] The beneficial effects of aspects two through five above can be referenced to aspect one or any possible implementation thereof, and will not be elaborated upon here. Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations.

[0033] Other advantages, objectives and features of this application will be partly apparent from the description below, and partly understood by those skilled in the art through study and practice of this application. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram illustrating an exemplary operating environment provided in this application embodiment; Figure 2 This is a flowchart of the video duplicate segment detection method provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the process of determining candidate fragments based on a set of elements in an embodiment of this application; Figure 4 This is a schematic flowchart of video duplicate segment detection according to an embodiment of this application; Figure 5 This is a schematic diagram of the similarity matrix in an embodiment of this application; Figure 6 This is a flowchart of a live room carousel detection method according to an embodiment of this application; Figure 7 This is a schematic diagram of the process for detecting a single live room carousel according to an embodiment of this application; Figure 8 This is a flowchart of a live room carousel detection method according to another embodiment of this application; Figure 9 This is a block diagram of a video repeating segment detection device according to an embodiment of this application; Figure 10 This is a schematic block diagram of a computing device illustrated in an exemplary embodiment of this application. Detailed Implementation

[0036] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0037] The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items. In this application, "at least one" means one or more, and "more than one" means two or more. The terms "first," "second," and other ordinal terms used in this application may be used to describe various constituent elements, but these constituent elements are not limited by these terms. The purpose of using these terms is solely to distinguish one constituent element from others and should not be construed as indicating or implying relative importance. For example, without departing from the scope of this application, a first constituent element may be named a second constituent element, and similarly, a second constituent element may be named a first constituent element.

[0038] The execution order of the steps in this application is not unique and can be adjusted according to actual needs. Those skilled in the art will understand that the steps can be executed in the order described in the embodiments of this application, or in parallel, interleaved, or other suitable orders without departing from the spirit and substance of this application. This application does not limit the order of the method steps, and any adjustment of the order of the steps should be considered as falling within the protection scope of this application.

[0039] Before introducing the embodiments of this application, the technical terms and background technology involved in this application will be introduced first.

[0040] Hough Voting (HV) can map points in image space to parameter space. In the embodiments of the application, Hough Voting can be used to map elements in the similarity matrix to the temporal difference parameter space.

[0041] MSS (Maximum Subarray Sum). In a sequence of arrays A, where the elements are in the real number field, MSS refers to finding a contiguous subarray B in A such that the sum of the elements in B is maximized.

[0042] In scenarios such as live streaming and video playback, situations involving carousels, repetitive video content, or videos containing identical content not only waste resources but also severely impact the user's viewing experience. Among related technologies, carousel detection methods suffer from inaccurate and incomplete detection.

[0043] The video repeating segment detection method of this application can effectively identify whether there is repeated playback content in a video. In live streaming scenarios, this method can accurately detect not only the carousel in a single live streaming room, but also the carousel between different live streaming rooms. Furthermore, by combining video and audio features, this application achieves more comprehensive detection and is applicable to various application scenarios such as carousel within segments, carousel between segments, and carousel between different live streaming sessions. This application employs the HV algorithm, which can more robustly select candidate segments, improving detection accuracy. Moreover, this method is unaffected by the position of the carousel segments, helping to improve recall. This application also uses the MSS algorithm to accurately locate the start and end times of repeating segments, improving review efficiency.

[0044] The following will describe the video duplicate segment detection method of this application embodiment in conjunction with the implementation method.

[0045] First, one or more exemplary operating environments for embodiments of this application are described. Figure 1 A schematic diagram of an implementation environment is provided, which includes: client 101, server 102 and control platform 103.

[0046] Client 101 and server 102 can communicate via a network connection. The network can be wired or wireless. The network includes various network devices such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or similar devices. The network can include physical links, such as coaxial cable links, twisted-pair cable links, fiber optic links, or combinations thereof, or wireless links, such as cellular links, satellite links, and Wi-Fi links.

[0047] Client 102 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, various messaging devices, or other electronic devices. These computer devices can run various types and versions of software applications and operating systems. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants, etc. Wearable devices may include head-mounted displays (such as smart glasses). Client 102 may include input / output interfaces. Input interfaces may include touchpads, touchscreens, mice, keyboards, or other sensing elements. Input interfaces can be configured to receive user commands that enable the client to perform various operations, such as uploading or browsing data. Output interfaces are used to output information to the user, such as displaying information.

[0048] Client 102 can be divided into a capture end and a user end. The capture end is the party providing video or hosting a live stream. The user end is the party watching the video or live stream. The capture end may have modules such as a camera, microphone, and processor, which can capture audio and video signals, encode and compress them, and then push them to server 101. The capture end can also acquire video content through other means, such as through video editing software. The user end may run client software with video playback or live streaming functions, such as an application (App). This application can be a standalone application or a subroutine within an application.

[0049] Server 101 should be interpreted broadly, referring to an entity capable of responding to external requests and providing data, resources, or services. Server 101 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0050] Server 101 may consist of one or more computing devices. These computing devices may include virtualized computing instances. Virtualized computing instances may include virtual machines, such as emulations of computer systems, operating systems, etc. The computing devices may load virtual machines based on virtual images and / or other data defining specific software (e.g., operating systems, dedicated applications, servers) used for emulation. As the demand for different types of processing services changes, different virtual machines may be loaded and / or terminated on one or more computing devices. A hypervisor may be implemented to manage the use of different virtual machines on the same computing device. Server 101 may run one or more services or software applications that enable the execution of the methods according to this application. Server 101 may also provide other services or software applications, which may include non-virtualized environments and virtualized environments.

[0051] Server 101 may include one or more components that implement the functions performed by server 101. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. In some embodiments, server 101 may provide services such as storage, reading, writing, querying, and deleting, such as providing data processing services to clients. In other embodiments, a client user may sequentially use one or more client applications to interact with server 101 to utilize the services provided by these components.

[0052] In embodiments of this application, a control platform 103 is also included, which can monitor and manage video or live streaming playback. The control platform 103 can monitor, manage, and statistically analyze data from the server 101 and client 102 by running specific software logic. The control platform 103 can be a management server or a control server, and includes modules such as a processor and memory. In some embodiments, the control platform 103 and the server 101 may be the same device. The control platform 103 in this application embodiment can be used to execute the video repeating segment detection method of subsequent embodiments.

[0053] The technical solutions of this application are described below through several embodiments. It should be understood that these embodiments can be implemented in many different forms and should not be construed as being limited to the embodiments set forth herein.

[0054] See Figure 2 The video duplicate segment detection method of this application includes: S11: Extract features from the video to be detected, and obtain the first feature matrix and the second feature matrix of the video to be detected.

[0055] In this embodiment, the video to be detected may include various video data transmitted over a network or stored locally. In a live streaming scenario, the video data in the live streaming room is acquired, encoded, and transmitted in real time. In scenarios such as on-demand or long video playback, the video data is pre-recorded content stored on a server or local storage, which users can select and play as needed.

[0056] After acquiring the video data, it can be decoded to obtain video frames and audio data. Frame extraction is performed on the video frames to obtain multiple image frames, and audio segments corresponding to the time periods of these video frames are extracted simultaneously. Therefore, in subsequent embodiments of this application, the image feature vectors and audio feature vectors are time-aligned.

[0057] In the embodiments of this application, frame extraction can be performed within a time interval of [-T, 0], where the value of T can be set according to factors such as available resources and traffic volume. For example, when T is set to 24 hours, the amount of data processed is large, the processing time is long, and the CPU utilization rate is high. Therefore, T can be set to a value that is more suitable for actual needs, such as 30 minutes or 1 hour.

[0058] In some embodiments, frame extraction can be performed according to predetermined rules, such as interval sampling or keyframe extraction. For example, a frame can be extracted every 5 or 10 seconds. The specific frame extraction interval can be adjusted according to the amount of data that the system can process.

[0059] In some embodiments, the extracted video frames may also be preprocessed, such as scaling, normalization, noise reduction, and cropping, to improve the effect of subsequent feature extraction.

[0060] The extracted multiple video frames are input into an image feature extraction model to convert them into an image vector matrix. In embodiments of this application, the image feature extraction model can be a convolutional neural network (CNN), a visual Transformer, a temporal convolutional network, or other machine learning or deep learning-based image feature extraction model.

[0061] Audio segments corresponding to multiple video frames are input into an audio feature extraction model to convert the audio segments into an audio vector matrix. In embodiments of this application, the audio feature extraction model can be a machine learning or deep learning-based audio feature extraction model such as Mel frequency cepstral coefficients (MFCC), Mel spectrogram, or short-time Fourier transform (STFT).

[0062] In this embodiment, the video to be detected may originate from a single video dataset or from two video datasets. Specifically, the video to be detected may include a first video and a second video. The first video and the second video may be video segments from different or the same time period within the same video dataset, or video segments from different or the same time period within different video datasets. The first feature matrix may be the image vector matrix and / or audio vector matrix of the first video, and the second feature matrix may be the image vector matrix and / or audio vector matrix of the second video. Thus, when it is determined that there are repeated segments in the video to be detected based on the first video and the second video, it is determined that there is repeated playback content in the first video and the second video.

[0063] If the videos to be detected come from the same video data, i.e., to detect whether there are duplicate segments within the video data, then the first feature matrix and the second feature matrix can be different image vector matrices within the same time period of the video to be detected. In some embodiments, the first feature matrix and the second feature matrix are two different image vector matrices of the video to be detected within the time period [-t, 0]. Thus, duplicate segment detection within a video segment can be achieved.

[0064] Furthermore, the first feature matrix and the second feature matrix can also be image vector matrices for different time periods of the video to be detected. In some embodiments, the first feature matrix is ​​the image vector matrix of the video to be detected during the [-t, 0] time period, and the second feature matrix is ​​the image vector matrix of the video to be detected during the [-T, -t] time period. Additionally, the first feature matrix and the second feature matrix can also be audio vector matrices for different time periods of the video to be detected. In some embodiments, the first feature matrix is ​​the audio vector matrix of the video to be detected during the [-t, 0] time period, and the second feature matrix is ​​the audio vector matrix of the video to be detected during the [-T, -t] time period. Thus, duplicate segment detection between video segments can be achieved.

[0065] In the embodiments of this application, the start time of the [-t,0] time period is the time obtained by shifting t forward from the current time, and the end time is the current time; the start time of the [-T,-t] time period is the time obtained by shifting T forward from the current time, and the end time is the time obtained by shifting t forward from the current time.

[0066] If the video to be detected comes from two video datasets (such as a first video and a second video), i.e., to detect whether there are duplicate segments between the two video datasets, the first feature matrix and the second feature matrix can be image vector matrices from different or the same time periods in the two video datasets, respectively, or audio vector matrices from the same or different time periods in the two video datasets, respectively. For example, in some embodiments, the first feature matrix is ​​an image vector matrix or audio vector matrix of the first video during the [-T, 0] time period, and the second feature matrix is ​​an image vector matrix or audio vector matrix of the second video during the [-T, 0] time period. Thus, duplicate segment detection between video segments of two videos can be achieved.

[0067] S12: Obtain the similarity matrix based on the first feature matrix and the second feature matrix.

[0068] In this embodiment, features can be extracted from the video to be detected at preset time intervals to obtain a first feature matrix and a second feature matrix. The preset time interval can be set according to actual conditions, such as 1 hour, 5 hours, 10 hours, etc. This allows for periodic detection of the video to be detected, enabling timely detection of repeated playback.

[0069] In one embodiment of this application, the similarity matrix is ​​a cosine similarity matrix of the first feature matrix and the second feature matrix. For example, if the first feature matrix A(n, d) (n d-dimensional vectors) and the second feature matrix B(m, d) (m d-dimensional vectors), the similarity between each row in A and each row in B is obtained to form a similarity matrix, where the element (i, j) represents the similarity between the i-th vector in A and the j-th vector in B.

[0070] It should be understood that the similarity matrix can also be the Euclidean distance matrix, the Manhattan distance matrix, the Jaccard similarity coefficient matrix, the Pearson correlation coefficient matrix, etc. The specific similarity or distance measurement method can be selected according to the actual needs.

[0071] S13: In the similarity matrix, filter out elements with similarity greater than a preset threshold and construct an element set.

[0072] In one embodiment of this application, a set of all elements with a similarity greater than a preset threshold h is selected based on the similarity matrix. For example, if the similarity matrix S has dimensions A*B, and the elements in the matrix... If an element is greater than a preset threshold h (i.e., the element in the i-th row and j-th column is greater than the preset threshold h), then that element is selected and added to the element set. In the timeline, the element... The time difference between the segment at time i in the first video and the segment at time j in the second video is ji.

[0073] If the first and second videos play at the same speed, that is, aligned on the timeline, then... If the value is greater than a preset threshold h, then the segments of the two videos at time point i and time point j are the same or similar. In some embodiments of this application, the value of the preset threshold h can be set according to actual conditions; for example, it can be set to a value greater than 0.9.

[0074] S14: Based on the set of elements, determine the candidate segments, each of which corresponds to a linear array in the similarity matrix.

[0075] In one embodiment of this application, candidate segments are determined from the set of elements using Hough Voting (HV). Each candidate segment corresponds to a linear array in the similarity matrix.

[0076] See Figure 3 In one embodiment of this application, determining candidate fragments based on the set of elements includes: S31: Get the time difference of each element in the set of elements.

[0077] S32: Group elements with the same time difference into the same category.

[0078] S33: Select the k categories with the most elements, and use the elements in the k categories as candidate fragments.

[0079] Set of elements The elements in the array are those with a similarity greater than a preset threshold h. These elements are then mapped to time differences. The parameter space is obtained by acquiring the time difference of each element and then setting the time difference... Elements with similar characteristics are grouped into the same category. For example, S 12 S 23 S 34 S 45 If the similarity of these four elements is greater than h and the time difference is 1, then they are grouped into one category, which contains 4 elements. 36 S 47 S 58 If the similarity of these three elements is greater than h and their time difference is 3, they are grouped into one category with 3 elements. The k categories with the most elements are selected as candidate segments. For example, if the category with a time difference of 1 has 4 elements and the category with a time difference of 3 has 3 elements, and k is 1, then the category with a time difference of 1 is selected as the candidate segment. This candidate segment includes 4 elements S. 12 S 23 S 34 S 45 That is, the candidate segment is a linear array in the similarity matrix S.

[0080] In this embodiment, the category with the largest number of occurrences is selected as the candidate fragment by setting a value of k. K is a positive integer; for example, the value of k can be selected in the range of 5-10.

[0081] S15: For each candidate segment: obtain the subarray with the largest sum in its corresponding linear array.

[0082] In the embodiments of this application, elements with the same time difference are clustered into the same category, and these elements correspond to points on a straight line. As mentioned earlier, the K categories with the most elements in the same category are selected. That is, for each category, all elements in the similarity matrix that lie on the straight line are found, resulting in a linear array, i.e., candidate segments. Since not every element on this straight line belongs to a repeated segment, by finding the subarray with the largest sum, repeated segments can be obtained more accurately, and the start and end points of the repeated segments can be obtained.

[0083] In the embodiments of this application, the MSS (Maximum Sub-array Sum) algorithm can be used to find the subarray with the largest sum, that is, to find a continuous subarray in a linear array such that the sum of the elements in the subarray is maximized. For example, the linear array corresponding to the candidate segment is... The values ​​in this linear array represent similarity scores. The subarray with the largest sum among these linear values ​​is... .

[0084] S16: Based on the subarray with the largest sum, determine whether the candidate segment is a duplicate segment, so as to determine whether the video to be detected includes duplicate segments.

[0085] As mentioned above, for each subarray, the average similarity of its elements is calculated. In the example shown earlier, the subarray with the largest sum is... If the similarity of the elements in the subarray is 0.9, then the average similarity is 0.9.

[0086] In one embodiment of this application, a preset mean is set to compare the average similarity of elements in the subarray with the preset mean. If the average similarity is greater than the preset mean, the candidate segment is determined to be a duplicate segment, that is, there is a duplicate segment in the video to be detected; if the average similarity is less than the preset mean, the candidate segment is determined to be a non-duplicate segment, that is, there is no duplicate segment in the video to be detected.

[0087] The preset mean can be set according to the actual situation; for example, it can be set to a value greater than 0.9.

[0088] For multiple subarrays, the average similarity of elements in each subarray is compared with a preset average. If the average similarity of elements in any subarray is greater than the preset average, it is determined that there are duplicate segments in the video to be detected. If the average similarity of elements in all subarrays is less than the preset average, it is determined that there are no duplicate segments in the video to be detected.

[0089] See Figure 4 This is a flowchart illustrating the video duplicate segment detection process according to an embodiment of this application. In this embodiment, by combining two feature matrices of the video to be detected with the Hough voting and MSS algorithms, the detection accuracy can be improved, the start and end times of duplicate segments can be accurately located, and the review efficiency can be improved.

[0090] In the embodiments of this application, the starting and ending points of the subarray with the largest sum correspond to the start and end times of the repeated segments, respectively. The method of this application not only determines whether repeated segments exist in the video to be detected, but also locates the start and end times of the repeated segments. When checking for repeated segments, the start and end times of the repeated segments can be directly located from the video to be detected, which improves the accuracy of judgment and the efficiency of verification compared to related technologies.

[0091] As described above, the first feature matrix and the second feature matrix in step S11 can be an image vector matrix within a video segment, an image vector matrix between video segments, or an audio vector matrix between video segments.

[0092] In one embodiment of this application, to improve the accuracy of duplicate segment detection, detection within and between video segments can be combined. That is, in each detection, duplicate segment detection within video segments, duplicate segment detection between video segments, and audio detection between video segments are used to jointly determine whether the video to be detected contains duplicate segments. In some embodiments, it can be set so that when both audio detection and any video detection (i.e., duplicate segment detection within video segments or duplicate segment detection between video segments) determine that duplicate segments are included, the video to be detected is considered to contain duplicate segments and is being played repeatedly.

[0093] For example, in some embodiments of this application, when determining whether a video to be detected contains duplicate segments, steps S11 to S16 can be executed first based on the image vector matrix within the video segment to obtain a first duplicate segment detection result. Then, steps S11 to S16 are executed again based on the image vector matrix within the video segment to obtain a second duplicate segment detection result. Next, steps S11 to S16 are executed again based on the audio vector matrix to obtain a third duplicate segment detection result. Finally, the results of the three detections are combined to determine whether a video to be detected contains duplicate segments. For example, if it is determined that the video to be detected contains duplicate segments based on the audio vector matrix, and also determined that the video to be detected contains duplicate segments based on the image vector matrix within the video segment, then it is ultimately determined that the video to be detected contains duplicate segments.

[0094] It should be understood that, in the embodiments of this application, image vector matrices and audio vector matrices are combined to improve the accuracy of duplicate segment detection. In the above examples, whether it is duplicate segment detection within video segments, duplicate segment detection between video segments, or audio detection between video segments, and further, audio detection within video segments can also be combined, the detection method and execution order are not limited. Parallel or sequential execution can be selected according to the actual situation of hardware resources. If sequential execution is adopted, the execution order of each detection can also be flexibly arranged.

[0095] See Figure 5 Let be a similarity matrix of the video to be detected, where white dots represent all elements with a similarity greater than h. Hough voting is performed based on the differences in the horizontal and vertical coordinates of the elements (i.e., time differences), selecting the segments with the most time difference votes as candidate segments, represented by the green lines in the linear array. Then, the MSS algorithm is executed on each candidate segment to locate the start and end points (blue boxes) and the similarity. Finally, the mean similarity is used to determine if there are duplicate segments.

[0096] The video duplicate segment detection method of this application can utilize the HV algorithm to select candidate segments, which has better robustness, improves the accuracy of the algorithm, and is not affected by the position of the duplicate segment, thus improving the recall rate (for example, if the actual number of videos that are played repeatedly is X1, the number of videos that are detected as playing repeatedly is X2, and the number of correctly detected videos in X2 is X3, then the recall rate is X3 / X1); by finding the subarray with the highest similarity among the candidate segments, the accurate location of the start and end timestamps of the duplicate is achieved.

[0097] The video duplicate segment detection method of this application is applicable to various application scenarios, such as: duplicate segment (carousel) detection in live streaming scenarios, detection of the same content looping in long videos, and identification of the same content between different videos. The following will specifically describe the implementation method when the video to be detected is a live stream video, in conjunction with a live streaming scenario. The video duplicate segment detection method of this application can detect carousel segments within a single live stream, as well as carousel content between different live streams. For a single live stream, an online detection method is used, and detection is performed periodically at regular intervals. Each detection consists of three logics: detection within video segments, detection between video segments, and audio detection. When both audio detection and any video detection determine carousel playback, the live stream is considered to be carousel playback. It should be understood that audio detection can be audio detection between video segments or audio detection within video segments.

[0098] See Figure 6 The live stream carousel detection method of one embodiment of this application includes the following steps: S61: At preset time intervals, acquire the live video of the live room to be detected.

[0099] S62: Extract features from the video frames within the video segment of the live stream to be tested, obtain feature matrix A1 and feature matrix B1 respectively, and determine whether there are duplicate segments within the video segment of the live stream to be tested based on feature matrix A1 and feature matrix B1.

[0100] In the embodiments of this application, the video frames of the live stream to be detected can be video frames within a preset time period of the live stream to be detected. For the live stream to be detected, a carousel detection can be performed periodically at preset time intervals, that is, the video data of the live stream to be detected is periodically acquired at regular intervals and decoded to obtain video frames. For example, video data is acquired every 10 minutes, 1 hour, or 2 hours for carousel detection. The preset time period can be [-T, 0], and as mentioned above, the value of T can be set according to factors such as available resources and traffic volume.

[0101] As mentioned earlier, after obtaining the video data within the [-T,0] time period of the live broadcast room, it can be decoded, frame extracted, preprocessed, etc. Then, the image vector matrix and audio vector matrix of the video data can be obtained through the image feature extraction model and the audio feature extraction model.

[0102] In this embodiment of the application, feature matrix A1 is an image vector matrix of the live room to be detected during the time period [-t, 0], and feature matrix B1 is another image vector matrix of the live room to be detected during the time period [-t, 0], where t is a preset time value.

[0103] In this embodiment, by splitting the [-T, 0] time period into the [-T, -t] time period and the [-t, 0] time period, where t is set to distinguish between segments, the computational load can be reduced and the detection efficiency improved. This is because the image vector matrix and audio vector matrix within the [-T, 0] time period have a matrix dimension of T*T. After splitting into two time periods, the matrix dimensions are (Tt)*t and t*t, respectively. Thus, the matrix dimension can be reduced, saving computational resources and reducing computational complexity. It should be understood that, if the saving of computational resources and computational complexity are not considered, the two time periods can also be omitted, and each detection step can be performed directly within the [-T, 0] time period.

[0104] After obtaining feature matrix A1 and feature matrix B1, the detection results within the video segment are obtained according to the above steps S12-S16, that is, to determine whether there are duplicate segments within the video segment of the video to be detected.

[0105] S63: Extract features from video frames between video segments in the live stream to be tested, obtain feature matrix A2 and feature matrix B2 respectively, and determine whether there are duplicate segments between video segments in the live stream to be tested based on feature matrix A2 and feature matrix B2.

[0106] In one embodiment of this application, feature matrix A2 is the image vector matrix of the live stream to be detected during the [-t, 0] time period, and feature matrix B2 is the image vector matrix of the live stream to be detected during the [-T, -t] time period. That is, video duplicate segment detection is performed on the image vector matrices during the [-T, -t] time period and the image vector matrices during the [-t, 0] time period.

[0107] It should be understood that the process of obtaining feature matrix A2 and feature matrix B2 can refer to the description of obtaining feature matrix A1 and feature matrix B1 in the aforementioned step S52, and will not be repeated here.

[0108] After obtaining feature matrix A2 and feature matrix B2, follow steps S12-S16 above to obtain the detection results between video segments, that is, determine whether there are duplicate segments between video segments of the video to be detected.

[0109] S64: Extract features from the audio data between video segments in the live stream to be tested, obtain feature matrix A3 and feature matrix B3 respectively, and determine whether there are duplicate segments between video segments in the live stream to be tested based on feature matrix A3 and feature matrix B3.

[0110] S65: When it is determined from the audio data that the video segments of the live room under test include repeated segments, and from the video frames that the video segments of the live room under test include repeated segments, it is determined that the live room under test has a carousel.

[0111] See Figure 7 In this embodiment of the application, when performing live stream carousel detection, video detection within segments, video detection between segments, and audio detection between segments are performed separately. That is, by combining the detection of video frames and audio data, the accuracy of detection can be improved. Furthermore, Hough voting and MSS algorithms are used in the detection process to further improve the accuracy of detection and enable the location of the start and end times of repeated segments.

[0112] In the embodiments of this application, when it is determined that the live room to be detected is in carousel mode, further processing operations can be performed on the live room to be detected, such as sending reminder information, reporting live room information, closing the live room or restricting the playback permissions of the live room.

[0113] For live streams rotating between rooms, since pairwise testing is required and the data volume is large, an offline testing method can be used, with test results provided on T+1 day.

[0114] See Figure 8 Another embodiment of the live stream carousel detection method of this application includes the following steps: S81: Obtain the key identifiers of the live streaming room to be tested.

[0115] In this embodiment, there are two or more live streaming rooms to be detected. Key identifiers in the live stream images of different rooms can be detected using OCR (Optical Character Recognition) technology. These key identifiers can be associated with the live streaming scenario; different scenarios will have different key identifiers. Key identifiers can be user IDs, text information in the live stream image, or keywords, etc. Keywords can be defined according to requirements and the focus of the detection. For example, for game live streaming, keywords can be defined as game scene names, character names in the game, etc.

[0116] S82: Group live streams with the same key identifier into the same group.

[0117] S83: For each live room in the same group, it is paired with the other live rooms in the group to form all possible live room pairs, and the two live rooms in each live room pair are respectively designated as the first live room and the second live room.

[0118] In this embodiment, after grouping live streaming rooms according to the same key identifier, the probability of live streaming rooms in the same group being in rotation is relatively high because they have the same key identifier. By performing duplicate segment detection between video segments in pairwise pairing of live streaming rooms in the same group, the rotation of live streaming rooms in the same group can be detected, ensuring the accuracy and comprehensiveness of the detection.

[0119] S84: Extract features from the live videos of the first and second live rooms within a preset time period to obtain feature matrix A4 and feature matrix B4, and determine whether there is a carousel in the first and second live rooms based on feature matrix A4 and feature matrix B4.

[0120] It should be understood that in this embodiment of the application, the video to be detected comes from two video data sets: a first video from a first live stream within a preset time period and a second video from a second live stream within the preset time period. Feature matrix A4 is an image vector matrix obtained from the first video, and feature matrix B4 is an image vector matrix obtained from the second video.

[0121] The extraction method for feature matrix A4 and feature matrix B4 can be referred to the description in the foregoing embodiment. After obtaining feature matrix A4 and feature matrix B4, determine whether there are duplicate segments in the video to be detected according to the above steps S12-S16, that is, determine whether there are duplicate segments between the video segments of the first live broadcast room and the second live broadcast room.

[0122] Because the detection of carousel playback between live streaming rooms involves a large amount of data, one embodiment of this application achieves carousel detection by detecting video frames in the video data of the live streaming rooms. For example, if the preset time period is 12 hours, then the video data can be all video data within 12 hours.

[0123] In the embodiments of this application, when it is determined that the video to be detected contains repeated segments based on feature matrix A4 and feature matrix B4, it is determined that the first live broadcast room and the second live broadcast room are rotating. Within the same group, the detection method is run once for every two live broadcast rooms. Therefore, detection can be achieved across all live broadcast rooms within the same group, ensuring the comprehensiveness and accuracy of the detection.

[0124] In the embodiments of this application, when it is determined that there is a carousel between the live streaming rooms, further processing operations can be performed on the two live streaming rooms that are carouseling, such as sending reminder information, reporting live streaming room information, closing the live streaming room or restricting the playback permissions of the live streaming room.

[0125] It should be understood that the feature matrices A1-A4 and B1-B4 in the embodiments of this application can correspond to the first feature matrix and the second feature matrix in the foregoing embodiments, respectively.

[0126] Based on the same technical concept, see [link / reference] Figure 9 This application also provides a video repeating segment detection device, including: The feature extraction module 901 is used to extract features from the video to be detected, and obtain the first feature matrix and the second feature matrix of the video to be detected respectively. The similarity acquisition module 902 is used to acquire a similarity matrix based on the first feature matrix and the second feature matrix; The element set construction module 903 is used to filter out elements with similarity greater than a preset threshold from the similarity matrix and construct an element set. The candidate determination module 904 is used to determine candidate segments based on the set of elements, wherein each candidate segment corresponds to a linear array in the similarity matrix; The subarray acquisition module 905 is used to acquire, for each candidate segment, the subarray with the maximum sum in its corresponding linear array; The duplicate segment determination module 906 is used to determine whether a candidate segment is a duplicate segment based on the mean similarity of elements in a subarray with the largest sum, so as to determine whether the video to be detected includes duplicate segments.

[0127] The modules described above correspond to the method steps; further details can be found in the method embodiments. The above description involves various modules. It should be noted that the division of these modules in the description is for clarity. However, in actual implementation, the boundaries between modules may be blurred. For example, any or all functional modules in this application may share various hardware and / or software elements. As another example, any and / or all functional modules in this application may be wholly or partially implemented by a shared processor executing software instructions. Furthermore, various software sub-modules executed by one or more processors may be shared among various software modules. Accordingly, unless explicitly required, the scope of this application is not limited by mandatory boundaries between various hardware and / or software elements.

[0128] Based on the same technical concept, embodiments of this application also provide a computing device, see reference. Figure 10 It includes a memory 1001 and a processor 1002. The memory 1001 is used to store computer instructions. When the processor 1002 executes the computer instructions, it implements the method steps in any method embodiment.

[0129] The specific entity of a computing device can be a server, which can be a rack server, blade server, tower server, or cabinet server (including a standalone server or a server cluster composed of multiple servers), etc.

[0130] The memory 1001 includes at least one type of computer-readable storage medium, including flash memory, hard disk, multimedia card, random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of an electronic device, such as the hard disk or memory of the electronic device. In other embodiments, the computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, secure digital card (SD card), flash memory card, etc., equipped on the electronic device. Of course, the computer-readable storage medium may include both internal storage units and external storage devices of the electronic device. In this embodiment, the computer-readable storage medium is typically used to store the operating system and various application software installed on the electronic device, such as the program code of the video repeating segment detection method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various types of data that have been output or will be output.

[0131] In some embodiments, processor 1002 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chip. Processor 1002 is typically used to control the overall operation of the processing device, such as performing control and processing related to data interaction or communication with other entities. In this embodiment, processor 1002 is used to run program code stored in memory 1001 or process data.

[0132] Based on the same technical concept, this application also provides a computer-readable storage medium, which includes a computer program or instructions stored in the storage medium. When the computer program or instructions are executed by a processing device, they implement the method steps in any method embodiment. Further details can be found in the method embodiments, which will not be repeated here. In this embodiment, the computer-readable storage medium can be non-volatile or volatile. Computer-readable storage media include flash memory, hard disk, multimedia card, random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of an electronic device, such as the hard disk or memory of the electronic device. In other embodiments, the computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, secure digital card (SD card), flash memory card, etc., equipped on the electronic device. Of course, the computer-readable storage medium can also include both internal storage units and external storage devices of the electronic device. In this embodiment, the computer-readable storage medium is typically used to store the operating system and various application software installed on the electronic device, such as the program code of the video repeating segment detection method in the embodiment. In addition, computer-readable storage media can also be used to temporarily store various types of data that have been output or will be output.

[0133] Based on the same technical concept, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video repeating segment detection method provided in the above-described method embodiments.

[0134] It should be noted that the order of description of the embodiments in this application is not intended to limit the priority of the embodiments.

[0135] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application and in its specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.

[0136] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many forms under the guidance of this application without departing from the spirit and scope of protection of the claims. All equivalent transformations made under the inventive concept of this application using the content of this application's specification and drawings, or direct / indirect applications in other related technical fields, are included within the patent protection scope of this application.

Claims

1. A method for detecting duplicate segments in a video, characterized in that, include: Feature extraction is performed on the video to be detected to obtain the first feature matrix and the second feature matrix of the video to be detected. Based on the first feature matrix and the second feature matrix, a similarity matrix is ​​obtained; In the similarity matrix, elements with similarity greater than a preset threshold are selected to construct an element set; Based on the set of elements, candidate segments are determined, and each candidate segment corresponds to a linear array in the similarity matrix; For each candidate segment: obtain the subarray with the largest sum in its corresponding linear array; Based on the subarray with the largest sum, it is determined whether the candidate segment is a duplicate segment, so as to determine whether the video to be detected includes duplicate segments.

2. The method according to claim 1, characterized in that, The step of determining candidate fragments based on the set of elements includes: Obtain the time difference of each element in the set of elements; Group elements with the same time difference into the same category; Select the k categories with the most elements, and use the elements in the k categories as the candidate fragments, where k is a positive integer.

3. The method according to claim 1, characterized in that, Based on the subarray with the largest sum, determining whether the candidate segment is a duplicate segment includes: If the average similarity of elements in the subarray is greater than a preset average, then the candidate segment is determined to be a duplicate segment.

4. The method according to claim 3, characterized in that, The method further includes: Obtain the start and end points of the subarray; Based on the start point and the end point, obtain the start time and end time of the repeating segment.

5. The method according to any one of claims 1-4, characterized in that, The video to be detected includes a live stream video from a live streaming room; The first feature matrix is ​​the first audio vector matrix of the live video in the [-t,0] time period, and the second feature matrix is ​​the second audio vector matrix of the live video in the [-T,-t] time period. or The first feature matrix is ​​the first image vector matrix of the live video in the [-t,0] time period, and the second feature matrix is ​​the second image vector matrix of the live video in the [-T,-t] time period or the third image vector matrix of the live video in the [-t,0] time period; The method further includes: When it is determined, based on the first audio vector matrix and the second audio vector matrix, that the video to be detected contains repeating segments, and based on the first image vector matrix and the second image vector matrix, it is determined that the live stream contains repeating segments; or When it is determined that the video to be detected contains repeated segments based on the first audio vector matrix and the second audio vector matrix, and also determined that the video to be detected contains repeated segments based on the first image vector matrix and the third image vector matrix, it is determined that the live broadcast room has a loop. Wherein, T and t are both preset time values, and the value of t is less than the value of T. The start time of the [-t,0] time period is the time obtained by shifting t forward from the current time, and the end time is the current time. The start time of the [-T,-t] time period is the time obtained by shifting T forward from the current time, and the end time is the time obtained by shifting t forward from the current time.

6. The method according to any one of claims 1-4, characterized in that, The method further includes: At preset time intervals, feature extraction is performed on the video to be detected to obtain the first feature matrix and the second feature matrix.

7. The method according to any one of claims 1-4, characterized in that, The videos to be detected include a first live video from a first live room within a preset time period and a second live video from a second live room within the preset time period. The first feature matrix is ​​the image vector matrix of the first live video, and the second feature matrix is ​​the image vector matrix of the second live video. The method further includes: When it is determined, based on the first live video and the second live video, that there are repeated segments in the video to be detected, it is determined that the first live room and the second live room are in rotation.

8. The method according to claim 7, characterized in that, The method further includes: Obtain key identifiers from multiple live streaming rooms to be detected; Group live streams with the same key identifier into the same group; Each live room in the same group is paired with the other live rooms in the same group to form all possible live room pairs, and the two live rooms in each live room pair are respectively designated as the first live room and the second live room.

9. The method according to any one of claims 1-4, characterized in that, The video to be detected includes a first video and a second video, the first feature matrix is ​​the image vector matrix and / or audio vector matrix of the first video, and the second feature matrix is ​​the image vector matrix and / or audio vector matrix of the second video; The method further includes: When it is determined, based on the first video and the second video, that there are repeated segments in the video to be detected, it is determined that the first video and the second video contain repeated playback content.

10. A video duplicate segment detection device, characterized in that, include: The feature extraction module is used to extract features from the video to be detected, and obtain the first feature matrix and the second feature matrix of the video to be detected respectively. A similarity acquisition module is used to acquire a similarity matrix based on the first feature matrix and the second feature matrix; The element set construction module is used to filter out elements with a similarity greater than a preset threshold from the similarity matrix and construct an element set. A candidate determination module is used to determine candidate segments based on the set of elements, wherein each candidate segment corresponds to a linear array in the similarity matrix; The subarray acquisition module is used to acquire, for each candidate segment, the subarray with the maximum sum in its corresponding linear array; The duplicate segment determination module is used to determine whether the candidate segment is a duplicate segment based on the subarray with the maximum sum, so as to determine whether the video to be detected includes duplicate segments.

11. A computing device, characterized in that, It includes a memory and a processor, the memory being used to store computer programs or instructions; when the computer programs or instructions are executed by the processor, the method of any one of claims 1-9 is implemented.

12. A computer-readable storage medium, characterized in that, The storage medium stores a computer program or instructions, which, when executed by a processor, implement the method of any one of claims 1-9.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-9.