Video identification method and apparatus, computer equipment, and computer program

The method enhances video identification by combining local and global segment matching to accurately identify and skip similar segments like openings and endings, improving video editing and comparison efficiency.

JP7753624B2Active Publication Date: 2025-10-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024522040
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-06-20
Filing Date
2023-04-18
Publication Date
2025-10-15
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

Existing video identification methods struggle to accurately identify similar segments such as openings and endings in videos across different platforms, leading to inefficiencies in video editing and comparison processes.

Method used

A method and apparatus that utilize video frame matching to identify locally and globally similar segments within a video series and across a video platform, determining overall similar segments by combining local and global segment positions, with optional correction updates using correction keywords.

Benefits of technology

Improves the accuracy and efficiency of identifying and skipping similar segments like openings and endings, enhancing video playback and comparison processes by reducing redundant data and improving processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007753624000002
    Figure 0007753624000002
  • Figure 0007753624000003
    Figure 0007753624000003
  • Figure 0007753624000004
    Figure 0007753624000004
Patent Text Reader

Abstract

A video identification method executed by a computer device, comprising: a step (202) of acquiring a target video and a video set reference video in a video series video set, where the video series video set includes videos belonging to the same series; a step (204) of identifying a video set local similar segment in the target video to the video set reference video based on a first matching result obtained by video frame matching between the target video and the video set reference video; a step (206) of acquiring a platform reference video from a video platform to which the target video belongs; a step (208) of identifying a platform global similar segment in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video; and a step (210) of determining an overall similar segment in the target video to the video set reference video and the platform reference video based on the positions in the target video of the video set local similar segment and the platform global similar segment, respectively.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to a Chinese patent application filed with the China Patent Office on June 20, 2022, bearing application number 202210695301.5 and entitled "Video Identification Method and Apparatus, Computer Equipment, and Storage Medium," the entire contents of which are incorporated herein by reference.

[0002] The present application relates to the field of computer technology, and in particular to a video identification method and apparatus, computer equipment, storage medium and computer program product. [Background technology]

[0003] With the development of computer technology, various network video platforms have emerged, and in addition to film video resources on the network, people can also independently create various videos on the network video platforms, including various types of videos such as lecture series, knowledge sharing, literary classes, current affairs episodes, entertainment videos, etc., to meet the new viewing needs of viewers. Similar video segments such as openings and endings are usually created in videos on various network video platforms, and these video segments are not the content of the video itself, so when performing video comparison or video editing processes, these video segments need to be identified and removed, but currently the identification accuracy of similar video segments such as openings and endings in videos is low. Summary of the Invention

[0004] According to various embodiments provided herein, a video identification method and apparatus, a computer device, a computer-readable storage medium, and a computer program product are provided.

[0005] In a first aspect, the present invention provides a computer device implemented method for video identification, said method comprising: Obtaining a target video and a video set reference video in a video series video set, where the video series video set includes videos belonging to the same series; identifying a video set locally similar segment in the target video to the video set reference video based on a first matching result obtained by video frame matching between the target video and the video set reference video; obtaining a platform reference video from a video platform to which the target video belongs; identifying platform-global similar segments in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video; determining overall similar segments in the target video to the video set reference video and the platform reference video based on the positions in the target video of the video set local similar segments and the platform global similar segments, respectively.

[0006] In a second aspect, the present invention further provides a video identification apparatus, said apparatus comprising: a video set video acquisition module configured to acquire a target video and a video set reference video in a video series video set, where the video series video set includes videos belonging to the same series; a locally similar segment identification module configured to identify a video set locally similar segment in the target video to the video set reference video based on a first matching result obtained by video frame matching between the target video and the video set reference video; a platform video acquisition module configured to acquire a platform reference video from a video platform to which the target video belongs; a global similar segment identification module configured to identify a platform global similar segment in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video; and an overall similar segment determination module configured to determine overall similar segments in the target video to the video set reference video and the platform reference video based on the positions in the target video of the video set local similar segments and the platform global similar segments, respectively.

[0007] In a third aspect, the present invention further provides a computing device comprising: a memory having computer-readable instructions stored thereon; and a processor that, when executing the computer-readable instructions, implements the above-described video identification method.

[0008] In a fourth aspect, the present invention further provides a computer-readable storage medium having stored thereon computer-readable instructions that, when executed by a processor, implement the above-described video identification method.

[0009] In a fifth aspect, the present invention further provides a computer program product, the computer program product comprising computer readable instructions which, when executed by a processor, implement the above video identification method.

[0010] The details of one or more embodiments of the invention are set forth in the drawings and description which follow. Other features, objects, and advantages of the invention will become apparent from the description, drawings, and claims. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a diagram illustrating an application environment of a video identification method according to an embodiment. [Figure 2] 1 is a flowchart of a video identification method in one embodiment. [Figure 3] 10 is a flowchart of a process for identifying platform global similar segments in one embodiment. [Figure 4] 1 is a flowchart for creating a user compilation video in one embodiment. [Figure 5] 1 is a flowchart of a video comparison in one embodiment. [Figure 6] FIG. 1 is an interface diagram introducing the opening of the platform screen in one embodiment. [Figure 7] FIG. 1 is a schematic diagram of an interface for playing video main content in one embodiment. [Figure 8] FIG. 10 is a schematic diagram of an interface introducing the ending of a platform screen in one embodiment. [Figure 9] FIG. 10 is a schematic interface diagram of the video platform introduction screen for the first period in one embodiment. [Figure 10] FIG. 10 is a schematic interface diagram of the video platform introduction screen for the second period in one embodiment. [Figure 11] 1 is an overall flowchart of a method for identifying openings and endings in one embodiment. [Figure 12] FIG. 1 is a schematic block diagram of an opening and ending mining method according to one embodiment. [Figure 13] FIG. 10 is a process schematic diagram of opening correction in one embodiment. [Figure 14] FIG. 10 is a process schematic diagram of opening correction in one embodiment. [Figure 15] FIG. 10 is a schematic diagram of matching segment information in one embodiment. [Figure 16] FIG. 1 is a schematic diagram illustrating time periods in one embodiment. [Figure 17] FIG. 10 is a schematic diagram illustrating updating of end times when overlapping portions exist in time periods according to an embodiment. [Figure 18]FIG. 10 is a schematic diagram illustrating updating of the start point time when an overlapping portion exists in a time period in one embodiment. [Figure 19] FIG. 10 is a schematic diagram illustrating an example in which overlapping portions exist between time periods. [Figure 20] FIG. 10 is a schematic diagram of updating the count of recommended openings and endings in one embodiment. [Figure 21] 1 is a block diagram showing the configuration of a video identification device according to an embodiment; [Figure 22] FIG. 2 is a diagram illustrating the internal configuration of a computer device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] The present invention will be described in more detail below with reference to the accompanying drawings and examples to make the objectives, technical solutions and advantages of the present invention clearer. The specific examples described in this specification are merely for the purpose of interpreting the present invention, and are not intended to limit the present invention.

[0013] The video identification method provided by the embodiment of the present invention can be applied to an application environment such as that shown in FIG. 1 , where a terminal 102 communicates with a server 104 via a network. A data storage system can store data required for processing by the server 104. The data storage system can be integrated into the server 104, located on a cloud, or located on another server. The server 104 performs video frame matching between a target video in a video series video set and a video set reference video, identifies video set local similar segments in the target video to the video set reference video based on the obtained first matching result, performs video frame matching between the platform reference video of the video platform to which the target video belongs and the target video, identifies platform global similar segments in the target video to the platform reference video based on the obtained second matching result, and determines overall similar segments in the target video based on the positions of the video set local similar segments and the platform global similar segments in the target video. When the terminal 102 plays the target video, the server 104 sends segment information of the overall similar segments in the target video to the video set reference video and the platform reference video to the terminal 102. The terminal 102 can skip playing the overall similar segments in the target video based on the received segment information, and if the overall similar segments are the opening or ending, can skip playing the opening or ending, thereby improving the video playback efficiency of the terminal 102. Furthermore, the video identification method provided in the present invention may be performed independently by the terminal 102 or the server 104, or may be performed jointly by the terminal 102 and the server 104 to realize the video identification process. Here, the terminal 102 includes, but is not limited to, various desktop computers, laptops, smartphones, tablet PCs, Internet of Things devices, and portable wearable devices.The Internet of Things devices include smart voice interaction devices, smart home appliances such as smart TVs and smart air conditioners, smart in-vehicle devices, airplanes, etc. The portable wearable devices include smart watches, smart bands, head-mounted devices, etc. The server 104 can be implemented as an independent server, a server cluster consisting of multiple servers, or a cloud server.

[0014] In one embodiment, a video identification method is provided, as shown in Figure 2, which may be performed by an electronic device such as a terminal or a server alone, or may be performed by a terminal and a server together, and in the embodiment of the present invention, the method is described by taking an example in which the method is applied to the server in Figure 1. The method includes the following steps:

[0015] In step 202, a target video and a video set reference video in a video series video set are obtained, where the video series video set includes videos belonging to the same series.

[0016] Here, a video series video set is a set of multiple videos belonging to the same series, which can be divided according to different series dimensions based on actual needs. For example, a drama series can be considered to belong to the same series, in which case a set of television videos included in the drama can be considered a video series video set for the drama. Also, videos produced by the same creator can be considered to belong to the same series, in which case a set of videos produced by the creator can be considered a video series video set, even if the length of each video is different. Furthermore, the same series can also include videos about the same theme, videos produced in the same location, etc. A video series video set includes multiple videos, and the multiple videos can have similar segments. For example, for videos produced by the same creator, each video can begin with an opening introducing the creator and an ending summarizing the video. A video opening is generally used to indicate the beginning of a video, and a video ending is used to indicate the end of a video. Openings and endings can take various forms, including, but not limited to, audio-video material, text, logos, etc.

[0017] The target video is a video in a video series set that is to be identified, i.e., to identify video segments from the target video that are similar to other videos. For example, the opening and ending may be identified, and the opening and ending are video segments similar to other videos. The reference video is used as a reference for target video identification, i.e., to identify similar video segments in the target video based on the reference video. The video set reference video is a reference video sampled and extracted from the video series set. The video set reference video and the target video belong to the same video series set. Similar video segments may exist between each video belonging to the same video series set, allowing skipping during playback or accurate editing of the main video. The number of videos in the video set reference video can be set according to actual needs. For example, the number of video set reference videos can be set to a fixed number or based on the length of the target video and the number of videos included in the video series set. For example, the longer the target video is, the more video set reference videos are set, and the more videos included in the video series set, the more video set reference videos are set. Furthermore, the number of video set reference videos can be set to a fixed ratio of the number of videos included in the video series video set, for example, 50%, so that if the number of videos included in the video series video set is 20, the number of video set reference videos can be 10, that is, 10 videos other than the target video are extracted from the video series video set as video set reference videos.

[0018] Specifically, when a video identification event is triggered, it indicates that video identification processing is required. The server obtains a target video and a video set reference video in the video series video set. Specifically, the server determines a video series video set that is the target of the video identification event, queries the video series video set, determines a target video from the video series video set, and extracts a video set reference video from the video series video set, thereby obtaining the target video and video set reference video that belong to the same video series video set. Furthermore, after determining the target video, the server determines a video series video set divided by the target video, thereby obtaining the target video and the video set reference video from the video series video set.

[0019] In step 204, based on a first matching result obtained by video frame matching between the target video and the video set reference video, a video set locally similar segment in the target video to the video set reference video is identified.

[0020] Here, a video frame refers to each image frame in a video consisting of multiple video frames. That is, a video includes multiple video frames, and each video frame is an image. Video frame matching involves performing image matching on video frames belonging to different videos to determine matching video frames in the different videos. For example, it involves determining whether there are video frames with matching similarities or video frames with matching image content. For example, image matching can be performed on a first video frame extracted from a first video with a second video frame extracted from a second video, thereby determining a video frame from the first video that matches a video frame in the second video. This may be a video frame containing the same image content, such as a video frame containing both opening and ending content. The first matching result is an image matching result obtained by performing video frame matching between a target video and a video set reference video. Specifically, the first matching result may include matching video frames obtained by identifying the target video and the video set reference video. In the video frame matching process between the target video and the video set reference video, similarity matching is performed on the video frames in the target video and the video frames in the video set reference video, and a first matching result including matching video frames between the target video and the video set reference video is obtained based on the video frames corresponding to a similarity that meets the similarity threshold.

[0021] A similar segment refers to a video segment with a similar screen between different videos, and a video set locally similar segment refers to a video segment in a target video that is similar to a segment in a video set reference video. If a video set locally similar segment in a target video is similar to a segment in the reference video, the video set locally similar segment may be a video content overlapping the target video and the video set reference video, for example, a video content overlapping the target video and the video set reference video, specifically, an opening, an ending, an advertisement, a platform introduction information, etc.

[0022] Specifically, the server identifies the target video and the video set reference video, and identifies video segments in the target video that are similar to the video set reference video. The server performs video frame matching between the target video and the video set reference video, specifically by extracting video frames from each of the target video and the video set reference video, and performing image matching on the extracted video frames. For example, image similarity matching is performed to obtain a first matching result. The server identifies video set local similar segments in the target video to the video set reference video based on the first matching result. Specifically, the server determines the video set local similar segments based on time attributes of matching video frames between the target video and the video set reference video, such as the timestamp positions of matching frames in the target video frames. Obtaining the video set local similar segments refers to identifying and obtaining the target video based on the video set reference video of the video set of the video series to which the target video belongs. The similar segments are identified and obtained based on the local video, as opposed to being identified and obtained based on each video in the entire video platform.

[0023] For example, if in the obtained first matching result, the video frame at 1 second of the target video matches the video frame at 3 second of the video set reference video, the video frame at 2 second of the target video matches the video frame at 4 second of the video set reference video, the video frame at 3 second of the target video matches the video frame at 5 second of the video set reference video, and the video frame at 4 second of the target video matches the video frame at 6 second of the video set reference video, the server can determine that the video segments at 1 second to 4 seconds of the target video are video set locally similar segments to the video set reference video, thereby identifying and obtaining video set locally similar segments.

[0024] In step 206, a platform reference video from the video platform to which the target video belongs is obtained.

[0025] Here, a video platform refers to a platform that can provide video resources, and users can perform operations such as playing, watching, downloading, and favorite videos on the video platform. In a specific implementation, video creators distribute (distribute) their created videos to the video platform so that video viewers can watch them. The platform reference video is from the video platform to which the target video belongs, i.e., it belongs to the same video platform as the target video. Specifically, a video extracted from the video platform to which the target video belongs can be used as a reference video to identify the target video.

[0026] Specifically, the server obtains the platform reference video, and in implementation, the server determines the video platform to which the target video belongs and obtains the platform reference video belonging to the video platform. In specific applications, the platform reference video is the original platform video directly obtained from the video platform, i.e., the platform video without further processing, and the platform reference video may be a video that has been edited on the original platform video, for example, a video segment cut out from the original platform video.

[0027] In step 208, a platform global similar segment in the target video to the platform reference video is identified based on the second matching result obtained by video frame matching between the target video and the platform reference video.

[0028] Here, the second matching result is an image matching result obtained by performing video frame matching between the target video and the platform reference video. The second matching result may specifically include matching video frames obtained by identifying between the target video and the platform reference video, such as whether there are video frames with matching similarity or video frames with matching image content. The video frame matching process between the target video and the platform reference video can adopt the same processing method as the video frame matching between the target video and the video set reference video. The platform global similar segment refers to a video segment in the target video that is similar to a segment in the platform reference video.

[0029] Specifically, the server identifies the target video and the platform reference video, and identifies video segments in the target video that are similar to the platform reference video. The server performs video frame matching on the target video and the platform reference video, specifically, extracting video frames from each of the target video and the platform reference video, and performing image matching on the extracted video frames to obtain a second matching result. The server identifies platform-global similar segments in the target video that are similar to the platform reference video based on the second matching result. Obtaining platform-global similar segments refers to identifying and obtaining the target video using the platform reference video in the video platform to which the target video belongs, and is a similar segment obtained by performing global video identification based on each video in the entire video platform.

[0030] In step 210, based on the positions of the video set local similar segments and the platform global similar segments in the target video, respectively, a global similar segment in the target video is determined for the video set reference video and the platform reference video.

[0031] Here, the positions of the video set local similar segments and platform global similar segments in the target video refer to the timestamp positions of the video set local similar segments and platform global similar segments in the target video. For example, if the video set local similar segments are video segments from 2 seconds to 6 seconds, the positions of the video set local similar segments in the target video may be the timestamp positions from 2 seconds to 6 seconds, and if the platform global similar segments are video segments from 3 seconds to 8 seconds, the positions of the platform global similar segments in the target video may be the timestamp positions from 3 seconds to 8 seconds. The integrated similar segments are video identification results obtained by integrating the video set local similar segments and platform global similar segments.

[0032] Specifically, the server determines the positions of the video set local similar segments and the platform global similar segments in the target video, and determines the overall similar segments in the target video for the video set reference video and the platform reference video based on the positions. For example, if the video set local similar segments are located between 2 and 6 seconds and the platform global similar segments are located between 3 and 8 seconds, the server can merge the two positions and determine the video segment corresponding to the position between 2 and 8 seconds as the overall similar segment in the target video. Furthermore, the user can adjust the overall similar segment to obtain a more accurate overall similar segment.

[0033] In a specific application, after determining the overall similar segments in the target video to the video set reference video and the platform reference video, the overall similar segments may be overlapping video segments in the target video, such as video content such as openings, endings, advertisements, or platform information, and when playing the target video, the overall similar segments can be skipped and played, thereby improving playback efficiency.Furthermore, in the application scenario of video comparison, if each video in the video series video set has overlapping openings, endings, or advertisement content that is not required for comparison, the overall similar segments can be cut out from the target video, so that video comparison can be performed with other video segments in the target video, reducing the amount of data required for video comparison processing and improving the processing efficiency of video comparison.

[0034] The video identification method includes: performing video frame matching between a target video and a video set reference video in a video series video set; identifying video set local similar segments in the target video for the video set reference video based on the obtained first matching result; performing video frame matching between the target video and a platform reference video of a video platform to which the target video belongs; identifying platform global similar segments in the target video for the platform reference video based on the obtained second matching result; and determining overall similar segments in the target video based on the positions of the video set local similar segments and the platform global similar segments in the target video. The video set local similar segments are obtained by identification based on the video set reference video belonging to the same video series video set as the target video, and the platform global similar segments are obtained by identification based on the platform reference video belonging to the same video platform as the target video. The overall similar segments obtained based on the positions of the video set local similar segments and the platform global similar segments in the target video take into account video similarity characteristics in the video series video set and video similarity characteristics in the video platform, thereby improving the accuracy of identifying similar video segments in the video.

[0035] In one embodiment, the video identification method performs a correction update on the video set locally similar segments based on the correction segments in the target video that include the correction keywords, to obtain updated video set locally similar segments.

[0036] Here, the correction keyword is a keyword used to perform a correction process on the video identification of the target video to improve the accuracy of video identification. Specifically, the correction keyword can be various types of keywords, such as platform introduction information keywords, advertising keywords, video introduction keywords, etc. For example, if the display content of the video segment from 2 seconds to 4 seconds of a certain video A is the video introduction keyword "Episode N" or "Pure Fiction," the video segment can be considered to belong to a similar segment rather than the main video content of the target video. Furthermore, for example, if the display content of the video segment from 1 second to 2.5 seconds of a certain video B is platform introduction information for "XXX Video Platform," the video segment can be determined to belong to a similar segment overlapping with each video on the video platform rather than the main video content of the target video. The correction segment is a video segment in the target video that requires a correction process for video identification, specifically, a video segment in the target video that includes the correction keyword. For example, if the video segment from 1 second to 2.5 seconds in the above video B contains the correction keyword "XXX video platform," the video segment from 1 second to 2.5 seconds in video B can be determined as the correction segment.

[0037] Specifically, the server determines correction segments in the target video that include the correction keywords. When applied, the server can perform text recognition on video frames in the target video to identify correction segments in the video frames of the target video that include the correction keywords. The correction keywords can be preset according to actual needs and can include various types of keywords, such as keywords in platform introduction information, advertising keywords, or video introduction keywords. The server performs correction updates on the video collection locally similar segments based on the correction segments in the target video. Specifically, the server uses the distribution of the correction segments in the target video, such as the positions of the correction segments in the target video, to perform correction updates on the positions of the video collection locally similar segments in the target video, updating the positions of the video collection locally similar segments in the target video, and obtaining updated video collection locally similar segments. If the correction segments include the correction keywords, the correction segments should also be considered to belong to similar segments that are overlappingly used in each video, and the correction segments should also be included in the results of video recognition. For example, if the video set locally similar segment of a video C is the video segment from 2 seconds to 5 seconds, but the video C includes the correction segment of the correction keyword from 0 seconds to 2 seconds, the server can determine that the updated video set locally similar segment is the video segment from 0 seconds to 5 seconds, thereby performing a correction update on the video set locally similar segment based on the correction segment, and improving the accuracy of video identification.

[0038] Furthermore, the step of determining overall similar segments in the target video to the video set reference video and the platform reference video based on the positions of the video set local similar segments and the platform global similar segments, respectively, in the target video includes the step of determining overall similar segments in the target video to the video set reference video and the platform reference video based on the positions of the updated video set local similar segments and the platform global similar segments, respectively, in the target video.

[0039] Specifically, the server determines a comprehensive similar segment based on the updated video set local similar segment and the platform global similar segment. When applied, the server can determine the positions of the updated video set local similar segment and the platform global similar segment in the target video, and can determine a comprehensive similar segment in the target video with respect to the video set reference video and the platform reference video based on the positions.

[0040] In this embodiment, a correction update is performed on the local similar segments of the video collection using the correction segments containing the correction keywords in the target video, and comprehensive similar segments are determined based on the updated local similar segments of the video collection and the platform global similar segments. The correction keywords can be used to perform a correction update on the local similar segments of the video collection, so that video segments with overlapping correction keywords can be identified, and the accuracy of identifying similar video segments in videos can be improved.

[0041] In one embodiment, the step of performing a correction update on a video set locally similar segment based on a correction segment in the target video that includes the correction keyword and obtaining an updated video set locally similar segment includes the steps of determining a correction segment in the target video that includes the correction keyword, updating the timestamp position of the video set locally similar segment in the target video based on the timestamp position of the correction segment in the target video and obtaining an updated timestamp position, and determining an updated video set locally similar segment in the target video based on the updated timestamp position.

[0042] Here, the timestamp position refers to the timestamp position in the video to which the video segment belongs, for example, if a video duration is 2 minutes, the timestamp is from 00:00 to 02:00, and if a video segment in the video is from 23 seconds to 59 seconds, the timestamp position of the video segment in the video is from 00:23 to 00:59. Different video segments in a video have different timestamp positions, and corresponding video segments can be determined from the video according to the timestamp positions.

[0043] Specifically, the server determines a correction segment in the target video that includes the correction keyword. For example, the server can perform text recognition on video frames in the target video to determine a correction segment in the target video that includes the correction keyword. The server determines a timestamp position of the correction segment in the target video and a timestamp position of a video set locally similar segment in the target video. The server updates the timestamp position of the video set locally similar segment in the target video to obtain an updated timestamp position, and determines an updated video set locally similar segment in the target video based on the updated timestamp position.

[0044] For example, if the server determines that the correction segment in the target video containing the correction keyword is the video segment from 30 seconds to 31 seconds, the server can determine that the timestamp position of the correction segment is 00:30 to 00:31, and if the timestamp position of the video set locally similar segment in the target video is 00:26 to 00:30, the server obtains that the updated timestamp position is 00:26 to 00:31, that is, the updated video set locally similar segment in the target video is the video segment from 26 seconds to 31 seconds.

[0045] In this embodiment, the timestamp position of the correction segment in the target video is updated to the timestamp position of the video collection locally similar segment in the target video, and the updated video collection locally similar segment in the target video is determined based on the updated timestamp position, thereby enabling accurate correction updating of the video collection locally similar segment based on the timestamp position, ensuring the accuracy of the video collection locally similar segment, and improving the accuracy of identifying similar video segments in the video.

[0046] In one embodiment, the step of determining a correction segment in the target video that includes the correction keyword includes the steps of performing text identification on video frames in the target video to obtain a text identification result, matching the text identification result with the correction keyword to obtain a matching result, and determining a correction segment from the target video that includes the correction keyword based on a video frame associated with the matching result that matches.

[0047] Here, the correction keywords can be preset according to actual needs. For example, a keyword library can be established, and various types of correction keywords can be stored in the keyword library. The text identification results of the target video are matched with each type of correction keyword in the keyword library to determine whether the target video contains a correction segment containing the correction keyword.

[0048] Specifically, the server acquires video frames from the target video, for example, extracting multiple video frames at equal intervals. The server performs text recognition on each acquired video frame to obtain a text recognition result. The server acquires a preset correction keyword and matches the text recognition result of the target video with the correction keyword to obtain a matching result. The server screens the matching results, determines each video frame associated with the matching result, and determines a correction segment containing the correction keyword from each target video. For example, for the first 10 seconds of the target video, one video frame is extracted every 0.5 seconds to obtain 20 video frames. The server performs text recognition on each video frame and matches the text recognition result of each video frame with the correction keyword. If the video frame associated with the matching result is the 18th to 20th video frames, the server can determine that the correction segment in the target video is the video segment between the 18th and 20th video frames, specifically, the video segment from the 9th to 10th seconds in the target video.

[0049] In this embodiment, text recognition is performed on video frames in the target video, and based on the matching results obtained by matching the text recognition results with the correction keywords, correction segments containing the correction keywords are determined from the target video, and the correction segments in the target video can be accurately identified through a text search method. Furthermore, based on the correction segments, correction updates are performed on the video collection local similar segments to improve the accuracy of video recognition.

[0050] In one embodiment, the platform reference video includes platform public video segments obtained from the public video library of the video platform to which the target video belongs, and platform related videos obtained from the video platform. As shown in Figure 3, the process of identifying platform global similar segments, that is, the step of identifying platform global similar segments in the target video to the platform reference video based on the second matching result obtained by video frame matching between the target video and the platform reference video, includes the following steps:

[0051] In step 302, video frame matching is performed on the target video and the platform public video segment to obtain a public video matching result.

[0052] Here, the public video library is associated with a video platform and is used to store each platform-public video segment on the video platform. A platform-public video segment is a video segment publicly used for each video on the video platform. For example, in the case of video platform A, for a video uploaded to the video platform A, video platform A adds a video segment introducing video platform A to the uploaded video to indicate the origin of the video. Each video on the video platform publicly uses a video segment introducing video platform A, and this video segment is a platform-public video segment. There can be one or more platform-public video segments, and the length and content of the platform-public video segment can be set by the video platform according to actual needs. Each video on the video platform includes a platform-public video segment, which does not belong to the main content of the video but belongs to a similar video segment. It can be identified and deleted when editing the main content of the video or performing video comparison processing.

[0053] Platform-related videos are videos obtained from the video platform to which the target video belongs, specifically, videos obtained by sampling from the video platform. The acquisition method of platform-related videos can be set according to actual needs, for example, random sampling can be adopted to extract from the video platform. In addition, screening conditions such as release time, theme content, keywords, etc. can be set to screen each video on the video platform to obtain platform-related videos. Public video matching results are matching results obtained by performing video frame matching between the target video and the platform's public video segments.

[0054] Specifically, the platform reference video acquired by the server includes a platform public video segment acquired from the public video library of the video platform to which the target video belongs. For example, the server can determine the video platform to which the target video belongs, query the public video library of the video platform, and acquire the platform public video segment from the public video library. The server performs video frame matching between the target video and the platform public video segment to obtain a public video matching result.

[0055] In step 304, if no similar segments can be identified based on the public video matching results, video frame matching is performed on the target video and platform related videos to obtain related video matching results.

[0056] Here, the related video matching result is a matching result obtained by performing video frame matching between the target video and the platform-related video, and may include matching video frames identified and obtained from the target video and the platform-related video.

[0057] Specifically, the server identifies similar segments in the target video based on the public video matching result. If similar segments cannot be identified, it indicates that the target video does not have video segments that are publicly shared with the platform public video segments. In this case, the server performs video frame matching on the target video and the platform related videos to obtain related video matching results.

[0058] In step 306, platform-global similar segments to the platform-related videos in the target video are identified based on the related video matching results.

[0059] Specifically, the server identifies platform-global similar segments in the target video to the platform-related videos based on the related video matching results. For example, the server determines matching video frames in the target video based on the related video matching results, and identifies platform-global similar segments in the target video to the platform-related videos based on the timestamp positions of each video frame.

[0060] In this embodiment, the platform reference videos include platform public video segments obtained from the public video library of the video platform to which the target video belongs and platform-related videos obtained from the video platform, and the server first performs a recognition process on the target video using the platform public video segments. If similar segments cannot be identified, the server then performs a recognition process on the target video using the platform-related videos to obtain platform-global similar segments in the target video to the platform-related videos. By first performing the recognition process using the platform public video segments, the relevance of similar segment recognition can be improved, the data volume of similar segment recognition process can be reduced, and the processing efficiency of similar segment recognition can be improved. However, if similar segments cannot be identified using the platform public video segments, the platform-related videos can be used to perform the recognition process to ensure the accuracy of similar segment recognition.

[0061] In one embodiment, after a platform global similar segment to a platform-related video in the target video is identified based on the related video matching result, the video identification method further includes a step of updating the identification statistical parameters of the platform global similar segment to obtain updated identification statistical parameters; and if the updated identification statistical parameters satisfy the platform public judgment condition, updating the platform global similar segment as a platform public video segment in the public video library.

[0062] Here, the identification statistical parameters are parameters obtained by statistically calculating the platform global similar segment identification process. The parameter type of the identification statistical parameters can be set according to actual needs. For example, the identification statistical parameters may include the frequency of successful identification of platform global similar segments. For each identified platform global similar segment, the identification statistical parameters can be obtained by statistically calculating the platform global similar segment identification process. The platform public use judgment condition is a judgment condition for determining whether a platform global similar segment is to be used as a platform public video segment. For example, the identification statistical parameters exceed a predetermined parameter threshold, specifically, the frequency exceeds a frequency threshold.

[0063] Specifically, after identifying a platform global similar segment for a platform-related video in a target video, the server can query the identification statistical parameters of the platform global similar segment, where the identification statistical parameters reflect the statistical results of successful identification of the platform global similar segment. The server updates the identification statistical parameters of the platform global similar segment. For example, if the identification statistical parameters of the platform global similar segment include the frequency of successful identification, specifically 5 times, the server can add 1 to the frequency and update the frequency of the identification statistical parameters to 6 times. The server queries a predetermined platform public judgment condition, compares the updated identification statistical parameters with the platform public judgment condition, and if the updated identification statistical parameters satisfy the platform public judgment condition, the server uses the platform global similar segment as a platform public video segment and updates the platform global similar segment in the public video library, thereby realizing dynamic updates to the public video library. In subsequent video identification processes, the server can first perform video identification processes using the platform global similar segment as the platform public video segment.

[0064] In this embodiment, after successfully identifying the platform global similar segment, the server updates the identification statistical parameters of the platform global similar segment. If the updated identification statistical parameters meet the platform public judgment conditions, the server updates the platform global similar segment as a platform public video segment in the public video library, thereby realizing dynamic updates to the public video library and ensuring the effectiveness of the platform public video segments in the public video library, and improving the accuracy and processing efficiency of the video similar segment identification process.

[0065] In one embodiment, obtaining a platform reference video from the video platform to which the target video belongs includes obtaining a platform public video segment from a public video library of the video platform to which the target video belongs.

[0066] Here, the public video library is associated with a video platform and is used to store each platform-public video segment on the video platform, and the platform-public video segment is a video segment publicly used for each video on the video platform. Specifically, the platform reference video acquired by the server includes the platform-public video segment acquired from the public video library of the video platform to which the target video belongs. For example, the server determines the video platform to which the target video belongs, queries the public video library of the video platform, and acquires the platform-public video segment from the public video library. In a specific application, the server can acquire all platform-public video segments in the public video library and can also screen from the public video library. For example, the server can screen based on the release time, video theme, etc., to obtain platform-public video segments that meet the screening conditions.

[0067] Further, the step of identifying platform global similar segments in the target video to the platform reference video based on second matching results obtained by video frame matching between the target video and the platform reference video includes the step of identifying platform global similar segments in the target video to the platform public video segments based on second matching results obtained by video frame matching between the target video and the platform public video segments.

[0068] Specifically, the server may perform video frame matching between the target video and the platform public video segments to obtain a second matching result, which may include matching video frames identified in the target video and the platform public video segments. The server may identify platform global similar segments in the target video for the platform public video segments based on the second matching result. For example, the server may determine platform global similar segments in the target video based on the positions in the target video of each of the matching video frames identified.

[0069] In this embodiment, the platform reference video includes platform public video segments obtained from the public video library of the video platform to which the target video belongs, and the server performs identification processing using the platform public video segments, thereby improving the relevance of similar segment identification, reducing the amount of data in the similar segment identification processing, and improving the processing efficiency of similar segment identification.

[0070] In one embodiment, the step of obtaining a platform reference video from the video platform to which the target video belongs includes the steps of determining the video platform to which the target video belongs and the correction keywords contained in the video frames of the target video, querying platform-related videos in the video platform that have a related relationship with the correction keywords, and screening to obtain the platform reference video from the platform-related videos according to the reference video screening conditions.

[0071] Here, the platform-related video is a video having a correction keyword relationship obtained from the video platform to which the target video belongs. The relationship between each video on the video platform and the correction keyword can be established in advance. For example, when a video is uploaded to the video platform, text recognition is performed on the video frames of the video, and the correction keywords contained in the video are determined based on the text recognition results, and the relationship between the video and the correction keyword is established. The reference video screening conditions are predetermined screening conditions for obtaining a platform reference video by screening from the platform-related videos. For example, various screening conditions can be used, such as release time, video theme, etc.

[0072] Specifically, the server determines the video platform to which the target video belongs. Specifically, the server queries the video attribute information of the target video and determines the video platform to which the target video belongs based on the video attribute information in the video attribute information. The server determines correction keywords included in video frames of the target video. Specifically, the server can perform text recognition on the video frames of the target video and determine the correction keywords included in the video frames of the target video based on the text recognition results. The server queries platform-related videos that have a relationship with the correction keywords from the video platform. For example, the server can query and obtain platform-related videos that have a relationship with the correction keywords based on the relationship between each video and keyword on the video platform. The server queries predetermined reference video screening conditions, such as screening conditions of release time, and screens platform-related videos based on the reference video screening conditions to obtain platform reference videos that meet the reference video screening conditions from the platform-related videos. For example, if the release time of the target video is June 1, 2022, the reference video screening condition is that the release time is within one month of the target video release time, and the server will screen platform reference videos from platform-related videos whose release time is between May 1, 2022 and June 1, 2022.

[0073] In this embodiment, the platform reference videos include platform-related videos having related relationships with the correction keywords obtained from the video platform, and are screened according to the reference video screening conditions. Thus, by performing global video identification processing using various videos in the video platform and controlling the number of platform reference videos, the amount of data used to perform similar segment identification processing using the platform reference videos as a whole can be reduced, and the processing efficiency of similar segment identification can be improved while ensuring the accuracy of similar segment identification.

[0074] In one embodiment, the video identification method further includes the steps of performing video frame text identification in a platform video belonging to a video platform to obtain video keywords; performing matching in a keyword library based on the video keywords to determine target keywords that match with the video keywords; and establishing an association relationship between the platform video and the target keywords.

[0075] Here, the platform video refers to each video belonging to the video platform, and the video keyword is a keyword obtained by performing text recognition on the platform video. A keyword library stores various keywords, and the target keyword is a keyword that matches with the video keyword in the keyword library. Specifically, the server performs text recognition on the platform video belonging to the video platform, for example, performs text recognition on the video frame in the platform video, and obtains the video keyword contained in the video frame of the platform video. The server queries the keyword library, which stores various correction keywords. The keyword library can be preset according to actual needs and can be dynamically updated and maintained. The server matches the video keyword in the keyword library to determine the target keyword that matches with the video keyword, and establishes an association relationship between the platform video and the target keyword, thereby querying the corresponding platform video based on the keyword and the association relationship.

[0076] Furthermore, querying the video platform for platform-related videos having an association relationship with the remediation keyword includes querying the video platform for platform-related videos associated with the remediation keyword based on the association relationship.

[0077] Specifically, for each platform video in the video platform, the server determines its related relationship, and queries and obtains platform related videos related to the correction keyword based on the related relationship and the correction keyword.

[0078] In this embodiment, for each platform video on the video platform, an association relationship between the platform video and a keyword is established, and platform-related videos related to the corrected keyword on the video platform are determined based on the association relationship, thereby improving the accuracy and processing efficiency of querying platform-related videos and improving the accuracy and processing efficiency of identifying similar segments.

[0079] In one embodiment, the step of determining overall similar segments in the target video to the video set reference video and the platform reference video based on the positions of the video set local similar segments and the platform global similar segments in the target video, respectively, includes the steps of determining a first timestamp position of the video set local similar segment in the target video and a second timestamp position of the platform global similar segment in the target video, merging the first timestamp position and the second timestamp position to obtain an overall timestamp position, and determining overall similar segments in the target video to the video set reference video and the platform reference video based on the overall timestamp position.

[0080] Here, the first timestamp position refers to the timestamp position of the video collection local similar segment in the target video, and the second timestamp position refers to the timestamp position of the platform global similar segment in the target video. The overall timestamp position is the timestamp position obtained by merging the first timestamp position and the second timestamp position. Based on the overall timestamp position, an overall similar segment can be determined from the target video.

[0081] Specifically, the server determines a first timestamp position of the video aggregation local similar segment in the target video and a second timestamp position of the platform global similar segment in the target video. Specifically, the server determines each timestamp position of the target video for each segment time of the video aggregation local similar segment and the platform global similar segment. The server merges the first timestamp position and the second timestamp position to obtain a comprehensive timestamp position. In a specific implementation, the server can directly merge the first timestamp position and the second timestamp position to obtain a comprehensive timestamp position. For example, if the first timestamp position is from 00:05 to 0:15 and the second timestamp position is from 00:02 to 00:06, the server can directly merge the first timestamp position and the second timestamp position to obtain a comprehensive timestamp position from 00:02 to 00:15. Furthermore, the server can further perform partial merging according to actual needs to obtain a comprehensive timestamp position. For example, if the first timestamp position is from 00:05 to 00:15 and the second timestamp position is from 00:04 to 00:14, the server obtains a range of overall timestamp positions from 00:05 to 00:14 based on the intersection position of the first timestamp position and the second timestamp position. The server determines overall similar segments from the target video to the video set reference video and the platform reference video based on the obtained overall timestamp positions. For example, if the overall timestamp position is from 00:02 to 00:15, the server can determine the video segment from 2 seconds to 15 seconds from the target video as an overall similar segment to the video set reference video and the platform reference video.

[0082] In this embodiment, the first timestamp position of the video collection local similar segment in the target video is merged with the second timestamp position of the platform global similar segment in the target video, and the comprehensive similar segment in the target video to the video collection reference video and the platform reference video is determined based on the comprehensive timestamp position, thereby realizing the comprehensive processing of the video collection local similar segment and the platform global similar segment based on the timestamp position, so that the comprehensive similar segment integrates the video similar characteristics in the video series video collection and the video similar characteristics in the video platform, and improves the accuracy of identifying similar video segments in the video.

[0083] In one embodiment, the step of identifying a video set locally similar segment in the target video to the video set reference video based on a first matching result obtained by video frame matching between the target video and the video set reference video includes: performing image matching of video frames between the target video and the video set reference video to obtain a video frame pair, the video frame pair including a video frame to be identified belonging to the target video and further including a video set reference video frame image-matched with the video frame to be identified in the video set reference video; determining a time offset of the video frame pair based on the temporal attributes of the video frame to be identified and the video set reference video frame in the video frame pair; and screening video frame pairs with matching time offsets; and determining a video set locally similar segment in the target video to the video set reference video based on the temporal attributes of the video frame to be identified in the video frame pair obtained by screening.

[0084] Here, the video frame pair is an image pair composed of successfully matched video frames determined by image matching of video frames between the target video and the reference video. When the reference video is a video set reference video, the video frame pair includes a video frame to be identified belonging to the target video and a video set reference video frame in the video set reference video that is image-matched with the video frame to be identified, that is, the video frame to be identified and the video set reference video frame in the video frame pair are obtained by successful image matching, and the video frame to be identified in the video frame pair is from the target video, and the video set reference video frame is from the video set reference video.

[0085] The time attribute is used to indicate the time information of the corresponding video frame and can represent the position of the video frame in the video. Specifically, the time attribute may be the timestamp of the corresponding video frame in the video, the frame number of the video frame, etc. For example, if the time attribute of a video frame is 2.0 seconds, it can indicate that the video frame is the 2.0-second video frame of the associated video. Furthermore, if the time attribute of a video frame is 500, it can indicate that the video frame is the 500th frame of the associated video. The time attribute can tag the position of the video frame in the associated video, thereby determining the appearance time of the video frame in the associated video. A video is obtained by combining multiple video frames according to time information, and each video frame of the video is assigned a time attribute including time information. The time offset is used to represent the time interval between the appearance time of the video frame to be identified in the target video and the appearance time of the reference video frame in the reference video of a video frame pair. The time offset is obtained by the respective time attributes of the video frame to be identified and the reference video frame. For example, in a video frame pair, the time attribute of the video frame to be identified may be 2 seconds, i.e., the video frame to be identified is the video frame at 2 seconds in the target video frame, while the time attribute of the video set reference video frame may be 3 seconds, i.e., the video set reference video frame is the video frame at 3 seconds in the video set reference video, i.e., the video frame at 2 seconds in the target video and the video frame at 3 seconds in the video set reference video frame match, thereby obtaining that the time offset of the video frame pair is 1 second based on the difference between the time attribute of the video frame to be identified and the time attribute of the video set reference video frame.

[0086] Specifically, the server performs image matching of video frames between the target video and the video set reference video. Specifically, image matching can be performed between video frames in the target video and video set reference video frames. For example, matching can be performed based on image similarity, thereby determining a video frame pair based on the matching result. A video frame pair is an image pair consisting of video frames that have been successfully image matched. In a video frame pair determined by image matching based on similarity, the image similarity between the video frame to be identified and the video set reference video frame in the video frame pair is high. That is, the video frame to be identified in the target video and the video set reference video are similar, so they may belong to the same video content, for example, they may be video frames belonging to an opening or an ending. For the obtained video frame pair, the server determines the temporal attributes of the video frame to be identified and the video set reference video frame in the video frame pair. Specifically, the corresponding temporal attributes can be determined by querying the frame information of the video frame to be identified and the video set reference video frame. The server determines a time offset of the video frame pair based on the obtained temporal attributes of the video frame to be identified and the video set reference video frame. For example, if the time attribute is a quantified value, the server can obtain the time offset of the video frame pair based on the numerical difference between the time attribute of the video frame to be identified and the time attribute of the video set reference video frame. The server screens each video frame pair based on the time offset and screens video frame pairs with matching time offsets. Specifically, the server can screen video frame pairs with the same time offset value or a difference within a certain range.The server determines, based on the video frame pairs obtained by screening, temporal attributes of the video frames to be identified in the video frame pairs obtained by screening, and obtains a video set locally similar segment in the target video to the video set reference video based on the temporal attributes of the video frames to be identified. For example, after determining the temporal attributes of the video frames to be identified in the video frame pairs obtained by screening, the server can determine the start time and end time based on the numerical magnitude of the temporal attribute of each video frame to be identified, thereby determining a video set locally similar segment in the target video based on the start time and end time.

[0087] In a specific application, the server can group videos according to the magnitude of the time offset values ​​to obtain sets of video frame pairs corresponding to different time offsets, where each set of video frame pairs includes video frame pairs with matching time offsets. For example, if the obtained video frame pairs have three time offsets, 1 s, 4 s, and 5 s, the server can use the video frame pair with a time offset of 1 s as a first video frame pair set and determine local similar video set segments in the target video based on the temporal attributes of the identified video frames in the video frame pairs in the first video frame pair set. The server can further use the video frame pairs with time offsets of 4 s and 5 s as a second video frame pair set and determine local similar video set segments in the target video based on the temporal attributes of the identified video frames in the video frame pairs in the second video frame pair set. The server can determine each local similar video set segment according to the temporal attributes of the identified video frames in the video frame pairs in each video frame pair set, and determine and merge local similar video set segments based on each video frame pair set. For example, the server can delete overlapping video set locally similar segments or update partially overlapping video set locally similar segments, thereby obtaining video set locally similar segments in the target video for each video set reference video.

[0088] In this embodiment, image matching of video frames is performed for a target video and a video set reference video in a video series video set, to obtain a video frame pair including a target video frame belonging to the target video and a video set reference video frame image-matched with the target video frame, a time offset of the video frame pair is determined based on the temporal attributes of the target video frame and the video set reference video frame in the video frame pair, and video frame pairs with matching time offsets are screened, and a video set local similar segment is determined from the target video to the video set reference video based on the temporal attributes of the target video frame in the video frame pair obtained by screening, for the target video and the video set reference video in the video series video set, a time offset of the video frame pair is determined based on the temporal attributes of the image-matched target video frame and the temporal attributes of the video set reference video frame, and a video set local similar segment is determined from the target video to the video set reference video according to the temporal attributes of the target video frame in the video frame pair with matching time offsets obtained by screening, so that similar video segments with different times can be flexibly determined based on the image-matched video frame pairs, thereby improving the accuracy of identifying similar video segments in various videos.

[0089] In one embodiment, the step of screening video frame pairs with matching time offsets and determining a video set locally similar segment in the target video to the video set reference video based on the temporal attributes of the identified video frame in the video frame pairs obtained by screening includes the steps of: performing numerical matching for the time offset of each video frame pair; screening video frame pairs with numerically matching time offsets based on the numerical matching results; determining a start time and an end time based on the temporal attributes of the identified video frame in the video frame pairs obtained by screening; and determining a video set locally similar segment from the target video to the video set reference video based on the start time and the end time.

[0090] Here, the time offset represents the time interval between the appearance time of the video frame to be identified in the target video and the appearance time of the video set reference video frame in the video set reference video in the video frame pair. The specific format of the time offset is a quantified numerical value, for example, a numerical value in seconds, indicating the time difference in seconds between the appearance times of the video frame to be identified and the video set reference video frame in the video frame pair in their respective videos. Numeric matching refers to matching the magnitude of the numerical value of the time offset of each video frame pair to obtain a numerical matching result. The numerical matching result may include the numerical difference between the time offsets of each video frame pair, i.e., the numerical difference between the time offsets. The start time refers to the video start time of the video segment, and the end time refers to the video end time of the video segment. Based on the start time and the end time, the start time is set as the video start time and the end time is set as the video end time, thereby determining the span time of the video, and thereby determining the corresponding video segment.

[0091] Specifically, the server performs numerical matching on the time offsets of each video frame pair, specifically, performs numerical matching on the time offsets of every two video frame pairs to obtain a numerical matching result. The server determines video frame pairs whose time offsets numerically match based on the obtained numerical matching result. For example, the numerical matching result includes a numerical difference between the time offsets of each video frame pair, and the server can determine the time offsets whose difference between the time offsets of each video frame pair is smaller than a predetermined threshold as numerically matching time offsets. Thus, after obtaining the screened video frame pairs whose time offsets numerically match based on the screened video frame pairs obtained by the numerically matching time offsets, the server determines the time attributes of the video frames to be identified in the screened video frame pairs. Specifically, the server can query the frame information of each video frame to be identified, thereby obtaining the time attributes of the video frames to be identified. The server determines the start time and end time based on the time attributes of the video frames to be identified.

[0092] In a specific application, after obtaining the temporal attributes of the video frame to be identified in the video frame pair obtained by screening, the server may determine the temporal attribute with the smallest numerical value among them and determine the start time based on the smallest temporal attribute, and may determine the temporal attribute with the largest numerical value among them and determine the end time based on the largest temporal attribute. For example, in this application, if the sequence of the temporal attributes of the video frame to be identified in the video frame pair obtained by screening is {1, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15}, the server may determine 1 s as the start time and 15 s as the end time. The server may determine a video set locally similar segment in the target video to the video set reference video based on the start time and end time. For example, the server may determine the video segment from the start time to the end time in the target video as the video set locally similar segment. For example, if the server determines 1 s as the start time and 15 s as the end time, the server may determine the video segment from 1 second to 15 seconds in the target video as the video set locally similar segment to the video set reference video.

[0093] In this embodiment, numerical matching is performed on the time offsets of video frame pairs, and based on the numerical matching results, video frame pairs with numerically matching time offsets are screened. The start time and end time are determined based on the time attributes of the video frame to be identified in the video frame pairs obtained by screening, and video collection local similar segments in the target video are determined based on the start time and end time. Thus, video collection local similar segments are determined from the target video based on the video frame to be identified in the video frame pairs obtained by screening, so that similar video segments can be flexibly determined based on the video frame to be identified at the frame level, and can be applied to videos containing similar video segments of different times, thereby improving the accuracy of identifying similar video segments in videos.

[0094] In one embodiment, the step of performing numerical matching on the time offsets of each video frame pair and screening video frame pairs whose time offsets numerically match based on the numerical matching result includes the steps of numerically comparing the time offsets of each video frame pair to obtain a numerical comparison result, screening video frame pairs whose numerical difference in time offset from each video frame pair is less than a numerical difference threshold based on the numerical comparison result, and performing offset updating on video frame pairs whose numerical difference in time offset is less than the numerical difference threshold to obtain video frame pairs whose time offsets numerically match.

[0095] Here, the numerical comparison refers to comparing the magnitude of the numerical values ​​of the time offsets of each video frame pair to obtain a numerical comparison result, which may include a numerical difference between the time offsets of each video frame pair. For example, if the time offset of video frame pair 1 is 1 s and the time offset of video frame pair 2 is 2 s, the numerical difference between the time offsets of video frame pair 1 and video frame pair 2 is 1 s. That is, the numerical comparison result of the time offsets of video frame pair 1 and video frame pair 2 is 1 s. The numerical difference threshold can be flexibly set according to actual needs and used to match the time offsets of each video frame pair. Specifically, video frame pairs whose numerical difference in time offsets is smaller than the numerical difference threshold can be used as the screened and acquired video frame pairs. The offset update refers to updating the time offsets of video frame pairs whose numerical difference in time offsets is smaller than the numerical difference threshold to match the time offsets of the video frame pairs. For example, the time offsets of video frame pairs can be uniformly updated to the same time offset.

[0096] Specifically, the server performs a numerical comparison on the time offsets of each video frame pair to obtain a numerical comparison result. The numerical comparison result includes a numerical difference between the time offsets of each video frame pair, which can be obtained by the server by subtracting two of the time offsets of each video frame pair. The server determines a predetermined numerical difference threshold and, based on the numerical comparison result, screens out video frame pairs whose numerical difference in time offsets is less than the numerical difference threshold. Specifically, the server compares the numerical difference of the numerical comparison result with the numerical difference threshold to determine video frame pairs associated with time offsets whose numerical difference is less than the numerical difference threshold, and screens out the video frame pairs. The server performs offset updates on video frame pairs whose numerical difference in time offsets is less than the numerical difference threshold. Specifically, the server can unify and update the time offsets of video frame pairs to the same value. For example, the server updates the time offsets of video frame pairs whose numerical difference in time offsets is less than the numerical difference threshold to the smallest value, thereby obtaining video frame pairs whose time offsets match numerically. For example, if the numerical difference threshold is 2s and the video frame pairs whose numerical difference in time offset obtained by screening is less than the numerical difference threshold include two types of time offsets, 1s and 2s, the server updates the time offset of the video frame pair whose time offset is 2s to 1s, thereby obtaining each video frame pair whose time offset is 1s, i.e., obtaining video frame pairs whose time offsets match numerically.

[0097] In this embodiment, a numerical comparison is performed based on the time offset of each video frame pair to obtain a numerical comparison result, and video frame pairs whose numerical difference in time offset is less than a numerical difference threshold are screened from the video frame pairs, and offset updates are performed on the video frame pairs obtained by the screening to obtain video frame pairs whose time offsets are numerically matched, thereby obtaining video frame pairs for determining video set locally similar segments by the screening, and the video frame pairs obtained by the screening can accurately identify video set locally similar segments from the target video to the video set reference video.

[0098] In one embodiment, the step of determining the start time and the end time based on the temporal attributes of the video frames to be identified in the video frame pairs obtained by screening includes the steps of obtaining a video frame pair list consisting of the video frame pairs obtained by screening; sorting each video frame pair in the video frame pair list in ascending order according to the numerical value of its time offset, and for video frame pairs with the same time offset, sorting them in ascending order according to the numerical value of the timestamp of the video frame to be identified included therein, where the timestamp is determined by the temporal attribute of the video frame to be identified included therein; determining, in the video frame pair list, a temporal attribute distance between the temporal attributes of the video frames to be identified in adjacent video frame pairs; determining adjacent video frame pairs whose temporal attribute distance does not exceed a distance threshold as video frame pairs belonging to the same video segment; and determining the start time and the end time based on the timestamp of the video frames to be identified in the video frame pairs belonging to the same video segment.

[0099] Here, the video frame pair list is constructed by sorting the video frame pairs obtained by screening. In the video frame pair list, each video frame pair obtained by screening is sorted in ascending order of the time offset value, and video frame pairs with the same time offset are sorted in ascending order of the timestamp value of the included video frame to be identified. The timestamp is determined by the temporal attribute of the included video frame to be identified, and the timestamp is the appearance time of the included video frame to be identified in the target video. In the video frame pair list, the video frame pairs are sorted in ascending order of the time offset value, and if the time offsets are the same, they are sorted in ascending order of the timestamp value of the included video frame to be identified. That is, in the video frame pair list, the smaller the time offset, the earlier it is ranked, and for video frame pairs with the same time offset, the smaller the timestamp of the included video frame to be identified, the earlier it is ranked. The temporal attribute distance is determined for adjacent video frame pairs in the video frame pair list based on the temporal attribute of the included video frame to be identified, and represents the time interval between adjacent video frame pairs. The distance threshold is set in advance according to actual needs and is used to determine whether they belong to the same video segment. Specifically, adjacent video frame pairs whose time attribute distance does not exceed the distance threshold can be determined as video frame pairs belonging to the same video segment, thereby adaptively performing video segment aggregation processing for each video frame pair, and thereby determining the start time and end time.

[0100] Specifically, the server obtains a video frame pair list obtained by sorting according to the video frame pairs obtained by screening. In a specific application, after the server obtains the video frame pairs by screening, it sorts the video frame pairs obtained by screening in ascending order of the numerical values ​​of their time offsets. For video frame pairs with the same time offset, the server determines timestamps based on the temporal attributes of the video frames to be identified contained in the video frame pairs, and sorts the timestamps in ascending order of the numerical values ​​of the video frames to be identified, thereby obtaining a video frame pair list. In the video frame pair list, the server compares the temporal attributes of the video frames to be identified in adjacent video frame pairs, specifically calculating the difference between their respective temporal attributes to obtain the temporal attribute distance. The server determines a predetermined distance threshold, compares the time attribute distance with the distance threshold, and determines from the video frame pair list adjacent video frame pairs whose time attribute distance does not exceed the distance threshold based on the comparison result, and determines adjacent video frame pairs whose time attribute distance does not exceed the distance threshold as video frame pairs belonging to the same video segment. That is, if the time attribute distance of the video frames to be identified in adjacent video frame pairs is small, the adjacent video frame pairs are considered to belong to the same video segment, and are thus aggregated into video segments based on the video frames to be identified in the video frame pairs. The server determines timestamps of the video frames to be identified in the video frame pairs belonging to the same video segment, and determines start times and end times based on the timestamps of each video frame to be identified. For example, the server determines the start time based on the timestamp with the smallest numerical value and the timestamp with the largest numerical value as the end time, and the determined start times and end times are the start times and end times of the video segments to which the video frame pairs belonging to the same video segment both belong.

[0101] In this embodiment, based on the video frame pair list composed of the video frame pairs obtained by screening, the video frame pairs belonging to the same video segment are determined based on the temporal attribute distance between the temporal attributes of the video frames to be identified in adjacent video frame pairs, and the start time and end time are determined based on the timestamps of the video frames to be identified in the video frame pairs belonging to the same video segment, thereby realizing inference and mining of the video frames to be identified into video segments and accurately identifying segments from the target video.

[0102] In one embodiment, the step of determining the start time and the end time based on the timestamps of the video frames to be identified in the video frame pair belonging to the same video segment includes the steps of determining a start video frame pair and an end video frame pair from the video frame pair belonging to the same video segment based on the timestamps of the video frames to be identified in the video frame pair belonging to the same video segment, obtaining the start time based on the timestamps of the video frames to be identified in the start video frame pair, and obtaining the end time based on the timestamps of the video frames to be identified in the end video frame pair.

[0103] Here, the timestamp of the video frame to be identified is determined by the time attribute of the video frame to be identified, and the timestamp of the video frame to be identified represents the time when the video frame to be identified appears in the target video. The start video frame pair and the end video frame pair are determined by the magnitude of the timestamp of the video frame to be identified included in each video frame pair belonging to the same video segment. The timestamp of the video frame to be identified included in the start video frame pair may be the timestamp with the smallest numerical value among the timestamps of the video frames to be identified included in each video frame pair belonging to the same video segment, while the timestamp of the video frame to be identified included in the end video frame pair may be the timestamp with the largest numerical value, thereby determining the video frame to be identified included in the start video frame pair as the start video frame pair belonging to the same video segment, and the video frame to be identified included in the end video frame pair as the end video frame pair belonging to the same video segment.

[0104] Specifically, the server determines the timestamps of the video frames to be identified in a video frame pair belonging to the same video segment, and based on the magnitude of each timestamp value, the server determines a start video frame pair and an end video frame pair from the video frame pairs belonging to the same video segment. Specifically, the server determines the video frame pair to which the video frame to be identified with the smallest timestamp belongs as the start video frame pair, and determines the video frame pair to which the video frame to be identified with the largest timestamp belongs as the end video frame pair. The server obtains a start time based on the timestamp of the video frame to be identified in the start video frame pair. For example, the server can determine the time corresponding to the timestamp as the start time. The server obtains an end time based on the timestamp of the video frame to be identified in the end video frame pair. For example, the server can determine the time corresponding to the timestamp as the end time.

[0105] In this embodiment, the server determines a start video frame pair and an end video frame pair based on the timestamps of the video frames to be identified in a video frame pair belonging to the same video segment, and determines the start time and end time based on the video frames to be identified contained in the start video frame pair and the end video frame pair, respectively, thereby realizing inference and mining into video segments using the video frames to be identified that belong to the same video segment, and improving the accuracy of identifying similar video segments from the target video.

[0106] In one embodiment, the video identification method further includes a step of determining a segment overlap relationship between each video set locally similar segment based on the respective start time and end time of each video set locally similar segment, and a step of performing segment updating for each video set locally similar segment based on the segment overlap relationship to obtain updated video set locally similar segments in the target video for the video set reference video.

[0107] Here, if there are multiple segments in the video set locally similar segments identified in the target video for the video set reference video, each video set locally similar segment is updated based on the segment overlap relationship between each video set locally similar segment to obtain an updated video set locally similar segment. The segment overlap relationship refers to the overlap relationship between the video set locally similar segments. For example, if the time range of video set locally similar segment A is (2,5), i.e., from second 2 to second 5 of the target video, and the time range of video set locally similar segment B is (3,4), video set locally similar segment A completely covers video set locally similar segment B. In this case, video set locally similar segment B can be deleted and video set locally similar segment A can be retained. If the time range of the video collection locally similar segment C is (2,6) and the time range of the video collection locally similar segment D is (5,8), the video collection locally similar segment C and the video collection locally similar segment D partially overlap. In this case, the video collection locally similar segment C and the video collection locally similar segment D can be extended and updated to obtain the updated video collection locally similar segment CD(2,8). If the time range of the video collection locally similar segment E is (4,8) and the time range of the video collection locally similar segment F is (1,5), the video collection locally similar segment E and the video collection locally similar segment F partially overlap. In this case, the video collection locally similar segment E and the video collection locally similar segment F can be extended and updated to obtain the updated video collection locally similar segment EF(1,8). Furthermore, if there is no overlap between multiple video collection locally similar segments, for example, (2,5) and (7,10), in this case, it can be determined that all video collection locally similar segments without overlap are the video identification result without performing a merging process on each video collection locally similar segment. According to different segment overlapping relationships, different update methods can be set, thereby ensuring the accuracy of updating to the video set local similar segments.

[0108] Specifically, when a plurality of video set locally similar segments are obtained, the server can determine a segment overlap relationship between each video set locally similar segment based on the start time and end time of each video set locally similar segment, such as a segment overlap relationship of inclusion, partial overlap, or no overlap. The server performs segment update for each video set locally similar segment based on the segment overlap relationship between each video set locally similar segment, specifically, merging, deleting, suspending, etc. each video set locally similar segment to obtain an updated video set locally similar segment for the video set reference video in the target video.

[0109] In this embodiment, when multiple segments of a video set locally similar segments are identified and obtained, segment updating is performed based on the segment overlap relationship between each video set locally similar segment, thereby obtaining more accurate video set locally similar segments and improving the accuracy of identifying video set locally similar segments from the target video.

[0110] In one embodiment, the video set reference videos are at least two, and the step of screening video frame pairs with matching time offsets and determining video set local similar segments in the target video to the video set reference video based on the temporal attributes of the video frames to be identified in the video frame pairs obtained by screening includes the steps of screening video frame pairs with matching time offsets and determining intermediate similar segments in the target video to the video set reference video based on the temporal attributes of the video frames to be identified in the video frame pairs obtained by screening, and performing segment updates for each intermediate similar segment in the target video that has an overlapping relationship with each video set reference video, to obtain video set local similar segments in the target video to each video set reference video.

[0111] Here, there are at least two video set reference videos, i.e., at least two video set reference videos are used to perform video frame matching processing on the target video. The intermediate similar segments refer to similar segments identified and obtained for a single video set reference video in the target video. The overlapping relationship refers to the overlapping relationship that exists between the intermediate similar segments identified and obtained based on different video set reference videos, and is specifically determined based on the time endpoints (including start time and end time) of each identified intermediate similar segment.

[0112] Specifically, the server obtains two or more video set reference videos, performs video identification processing on the target video and each of the two or more video set reference videos, and obtains intermediate similar segments in the target video for each video set reference video. The server performs segment update on each intermediate similar segment in the target video that has an overlapping relationship with each video set reference video, thereby obtaining video set local similar segments in the target video for each video set reference video.

[0113] In this embodiment, video identification is performed on the target video using multiple video set reference videos, and segment updates are performed on each intermediate similar segment based on the overlapping relationship existing in each identified intermediate similar segment, and video set local similar segments are obtained in the target video for each of the video set reference videos, thereby further improving the accuracy of the video set local similar segments obtained by identification with reference to multiple video set reference videos, and improving the accuracy of similar segments from the target video.

[0114] In one embodiment, for intermediate similar segments in the target video with respect to each video set reference video, segment updates are performed for each intermediate similar segment that has an overlapping relationship, and the step of obtaining video set local similar segments in the target video with respect to each video set reference video includes the steps of performing segment position comparison for intermediate similar segments in the target video with respect to each video set reference video to obtain segment comparison results, determining each intermediate similar segment that has an overlapping relationship as a result of the segment comparison, and performing segment updates for each intermediate similar segment that has an overlapping relationship based on the overlapping time length and statistics of each intermediate similar segment that has an overlapping relationship, to obtain video set local similar segments in the target video with respect to each video set reference video.

[0115] Here, segment position comparison refers to comparing the positions in the target video of intermediate similar segments identified based on each video set reference video to obtain a segment comparison result. The segment comparison result may include whether an overlapping relationship exists between the intermediate similar segments. If an overlapping relationship exists, segment updating is performed on each intermediate similar segment with an overlapping relationship to obtain a video set local similar segment for each video set reference video in the target video. The overlapping duration refers to the duration of the overlapping segment between the intermediate similar segments with an overlapping relationship. For example, if the time range of intermediate similar segment A determined by the first video set reference video is (2,8) and the time range of intermediate similar segment B determined by the second video set reference video is (5,10), an overlapping relationship exists between intermediate similar segment A and intermediate similar segment B, the overlapping segment is (5,8), and the overlapping duration is 4 seconds between the 5th and 8th seconds. The statistics may include the number of times the same intermediate similar segment is identified in the target video among the intermediate similar segments for each video set reference video identification. The larger the value of the statistic, the more times the corresponding intermediate similar segment has been identified, and thus the higher the probability that the intermediate similar segment belongs to the video set local similar segments.

[0116] Specifically, the server determines intermediate similar segments in the target video for each video set reference video, and performs segment position comparison for each intermediate similar segment. The server can determine the start time and end time of each intermediate similar segment and perform segment position comparison based on the start time and end time of each intermediate similar segment to obtain a segment comparison result. If the segment comparison result indicates that there is no overlapping relationship, there is no need to process the intermediate similar segments that do not have an overlapping relationship, and they can be retained and used as video set local similar segments in the target video for each video set reference video. If the segment comparison result indicates that there is an overlapping relationship, i.e., if there is segment overlap between the intermediate similar segments, the server determines each intermediate similar segment that has an overlapping relationship and performs segment update for each intermediate similar segment that has an overlapping relationship. For example, various update processes such as deleting, merging, and retaining each intermediate similar segment are performed to obtain video set local similar segments in the target video for each video set reference video. The server determines each intermediate similar segment having an overlapping relationship as a result of the segment comparison, and also determines the respective statistics of each intermediate similar segment having an overlapping relationship and the overlapping duration between each intermediate similar segment. The server performs segment updates for each intermediate similar segment having an overlapping relationship based on the overlapping duration and statistics of each intermediate similar segment having an overlapping relationship, thereby obtaining video set local similar segments in the target video for each video set reference video. Specifically, the server can determine whether merging is necessary based on the length of overlapping duration, and determine whether retention or merging processing is necessary based on the value of the statistics.

[0117] In this embodiment, segment position comparison is performed for intermediate similar segments in the target video with respect to each video set reference video, and segment update is performed for each intermediate similar segment that has an overlapping relationship as a result of the segment comparison. Specifically, segment update is performed for each intermediate similar segment that has an overlapping relationship based on the overlapping duration and statistics of each intermediate similar segment that has an overlapping relationship. Thus, segment update is performed for the overlapping duration and statistics of each intermediate similar segment that has an overlapping relationship. By taking into account the characteristics between each intermediate similar segment, the effect of segment update can be improved, and the accuracy of identifying local similar segments of the video set from the target video can be improved.

[0118] In one embodiment, the step of performing segment position comparison for intermediate similar segments in the target video for each video set reference video to obtain a segment comparison result includes the steps of obtaining a similar segment list consisting of intermediate similar segments in the target video for each video set reference video, sorting each intermediate similar segment in the similar segment list in descending order of statistics and sorting intermediate similar segments with the same statistics in ascending order of start time, and performing segment position comparison for each intermediate similar segment in the similar segment list to obtain a segment comparison result.

[0119] Here, the similar segment list is constructed by sorting the intermediate similar segments in the target video for each video set reference video. In the similar segment list, the intermediate similar segments are sorted in descending order of statistics, and intermediate similar segments with the same statistics are sorted in ascending order of start time. That is, in the similar segment list, the intermediate similar segments are sorted in descending order of statistics, and intermediate similar segments with the same corresponding statistics are sorted in ascending order of start time.

[0120] Specifically, the server obtains a similar segment list consisting of intermediate similar segments in the target video for each video set reference video. The similar segment list can be obtained by the server sorting the intermediate similar segments in advance. Specifically, the intermediate similar segments are first sorted in descending order of statistics, and for intermediate similar segments with the same statistics, the server sorts them in ascending order of start time, thereby obtaining a similar segment list. The server then performs segment position comparison for each intermediate similar segment in the similar segment list to obtain a segment comparison result. In a specific application, the server performs segment position comparison in order from front to back according to the order of the intermediate similar segments in the similar segment list to obtain a segment comparison result.

[0121] Further, the step of performing a segment update for each intermediate similar segment for which an overlapping relationship exists, and obtaining a video set local similar segment for each of the video set reference videos in the target video, includes a step of performing a segment update for a preceding intermediate similar segment by a succeeding intermediate similar segment among each intermediate similar segment for which an overlapping relationship exists, and obtaining a video set local similar segment for each of the video set reference videos in the target video, wherein the preceding intermediate similar segment is positioned before the succeeding intermediate similar segment in the similar segment list.

[0122] In this case, in the similar segment list, a preceding intermediate similar segment is positioned before a succeeding intermediate similar segment. That is, compared to a preceding intermediate similar segment, a succeeding intermediate similar segment is a similar intermediate segment arranged later in the similar segment list among the intermediate similar segments with which there is an overlapping relationship, and compared to a succeeding intermediate similar segment, a preceding intermediate similar segment is a similar intermediate segment arranged earlier in the similar segment list. For example, if the similar segment list includes intermediate similar segment A and intermediate similar segment B and the statistic of intermediate similar segment A is higher than that of intermediate similar segment B, then intermediate similar segment A is positioned before intermediate similar segment B in the similar segment list. In this case, the succeeding intermediate similar segment is intermediate similar segment B, and the preceding intermediate similar segment is intermediate similar segment A.

[0123] Specifically, the server can determine the subsequent similar intermediate segment and the preceding similar intermediate segment for each intermediate similar segment that has an overlapping relationship, and perform segment updates on the preceding intermediate similar segment based on the determined subsequent similar intermediate segment, such as performing various update processes such as deletion, merging, and retention to obtain video set local similar segments in the target video for each video set reference video.

[0124] In this embodiment, based on a similar segment list composed of intermediate similar segments in the target video for each video set reference video, for each intermediate similar segment that has an overlapping relationship, a segment update is performed on the preceding intermediate similar segment using the subsequent intermediate similar segment, thereby ensuring accurate retention of intermediate similar segments with high statistics, improving the effect of segment update and improving the accuracy of identifying local similar segments in the video set from the target video.

[0125] In one embodiment, for intermediate similar segments in the target video with respect to each video set reference video, segment updates are performed for each intermediate similar segment that has an overlapping relationship, and the step of obtaining video set local similar segments for each video set reference video in the target video includes the steps of: performing segment updates for each intermediate similar segment in the target video with respect to each video set reference video that has an overlapping relationship, obtaining updated intermediate similar segments; determining statistics of the updated intermediate similar segments; and, if the statistics of the updated intermediate similar segments exceed a statistics threshold, obtaining video set local similar segments in the target video with respect to each video set reference video based on the updated intermediate similar segments.

[0126] Here, the statistics may include the cumulative number of times the same intermediate similar segment is identified in the intermediate similar segment for each video set reference video identification in the target video. The statistics threshold is used to determine whether the updated intermediate similar segment is a valid video set local similar segment, and the statistics threshold can be set according to actual needs.

[0127] Specifically, the server performs segment update on each intermediate similar segment in the target video that has an overlapping relationship with each video set reference video, thereby obtaining updated intermediate similar segments. The server determines statistics of the updated intermediate similar segments. Specifically, the server performs statistical processing on the updated intermediate similar segments to obtain statistics of the updated intermediate similar segments. The server determines a predetermined statistics threshold, and if the statistics of the updated intermediate similar segment exceed the statistics threshold, the updated intermediate similar segment can be considered a valid video set local similar segment. The server obtains video set local similar segments in the target video for each video set reference video based on the updated intermediate similar segments. For example, the server can use the updated intermediate similar segments as video set local similar segments in the target video for each video set reference video.

[0128] In this embodiment, the validity of the updated intermediate similar segments is determined using a statistical threshold, and after the validity determination, video set local similar segments for each video set reference video in the target video are obtained based on the updated intermediate similar segments, and the validity of the identified video set local similar segments can be ensured.

[0129] In one embodiment, the video identification method further includes obtaining a public video matching the public video type in the target video based on the comprehensive similar segments if the comprehensive similar segments satisfy the determination condition for the public video type.

[0130] Here, the public video type refers to the type of video publicly used for each video, including, but not limited to, opening, closing, and advertisement types. The public video type can be set according to actual needs. The public video type determination condition is used to determine whether the type of the overall similar segment matches the public video type. Specifically, the public video distribution area associated with the public video type is compared with the overall similar segment to determine whether the overall similar segment matches the public video type, thereby determining the type of the overall similar segment. Matching a public video with a public video type means matching a public video type with another public video type. The public video is a video segment whose type has been determined as being overlapping. For example, the public video may be video content that can be overlappingly used in each video, such as an opening, closing, or advertisement.

[0131] Specifically, the server determines a public video type judgment condition, and if the synthetic similar segment satisfies the judgment condition, the server obtains a public video that matches the public video type in the target video based on the synthetic similar segment. For example, the judgment condition for the public video type is that it is in a public video distribution section related to the public video type, and the server determines a time period of the synthetic similar segment and determines whether the time period of the synthetic similar segment is already located in the public video distribution section. If the time period of the synthetic similar segment is in the public video distribution section, the server obtains a public video that matches the public video type based on the synthetic similar segment. In this case, if the public video type is an opening type, the opening in the target video can be obtained based on the synthetic similar segment, and specifically, the synthetic similar segment can be used as the opening of the target video.

[0132] In this embodiment, if the identified comprehensive similar segment satisfies the public video type judgment condition, a public video that matches the public video type in the target video is obtained based on the comprehensive similar segment, thereby identifying a public video that matches the public video type from the target video and improving the identification accuracy when identifying a public video from the target video.

[0133] In one embodiment, when the overall similar segment satisfies the judgment condition of the public video type, the step of obtaining a public video that matches the public video type in the target video based on the overall similar segment includes the step of determining a public video distribution interval to which the public video type of the target video is associated, and when the time period of the overall similar segment is within the public video distribution interval, the step of obtaining a public video that matches the public video type in the target video based on the overall similar segment.

[0134] Here, the public video distribution section is a time distribution section in the target video of the public video belonging to the public video type. For example, if the public video type is an opening type, the related time distribution section may be the first N seconds of the target video, for example, the last 20 seconds of the target video, i.e., the time distribution section is 0s-20s. The time period of the comprehensive similar segment refers to the time span in the target video of the comprehensive similar segment obtained by identification, and can be determined based on the start time and end time of the comprehensive similar segment, and can be the time span directly from the start time to the end time.

[0135] Specifically, the server determines a public video distribution section associated with the public video type of the target video, with different public video types having different public video distribution sections. For example, if the public video type is an opening type, the associated public video distribution section may be the first N seconds of the video, while if the public video type is an ending type, the associated public video distribution section may be the last M seconds of the video. The server determines a time period of the aggregate similar segment. Specifically, the server may determine the time period based on the start time and end time of the aggregate similar segment. If the time period of the aggregate similar segment is within the public video distribution section associated with the public video type, it indicates that the aggregate similar segment is within the time span range corresponding to the public video type. The server obtains a public video matching the public video type of the target video based on the aggregate similar segment. For example, the server may match the aggregate similar segment with the public video type of the target video. If the public video type is an ending type, the server may determine the aggregate similar segment as the ending of the target video.

[0136] In this embodiment, based on the comparison result between the public video distribution section related to the public video type and the time period of the comprehensive similar segment, the public video that matches the public video type in the target video is determined by the comprehensive similar segment, thereby ensuring the accuracy of the public video that matches the public video type from the target video based on the predetermined public video distribution section, and improving the identification accuracy when identifying the public video from the target video.

[0137] In one embodiment, the video identification method further includes determining a start time and an end time of the public video, extracting a non-public video from the target video based on the start time and the end time in response to a video comparison trigger event, and performing a video comparison of the non-public video with the comparison target video.

[0138] Here, the public video is a video segment of a specified type that is used in overlapping. For example, the public video is video content that can be reused in each video, such as an opening, an ending, or an advertisement. The start time of the public video refers to the time when the public video starts, and the end time of the public video refers to the time when the public video ends. The video comparison trigger event is a trigger event for comparing videos, and by comparing the videos, the similarity between the videos can be determined. The non-public video is a video of a segment other than the public video in the target video. The non-public video is not a video segment that is used in overlapping, but is considered to be the main video content of the target video. The comparison target video is a video required for video comparison, and by comparing the non-public video with the comparison target video, the video similarity between the non-public video and the comparison target video can be determined.

[0139] Specifically, the server determines the start time and end time of the public video, and responds to a video comparison trigger event, such as a video comparison event triggered by a user on the terminal, and extracts and obtains a non-public video from the target video based on the start time and end time of the public video. Specifically, the server excludes the public video from the target video based on the start time and end time of the public video, thereby extracting and obtaining a non-public video from the target video. The server obtains a comparison target video and performs a video comparison with the extracted non-public video to obtain a video comparison result. The video comparison result can reflect the content similarity between the comparison target video and the extracted non-public video.

[0140] In this embodiment, based on the start time and end time of the public video, non-public video is extracted from the target video for video comparison with the comparison target video, thereby enabling the non-public video in the target video to be accurately and quickly located, thereby improving the accuracy and processing efficiency of video comparison.

[0141] In one embodiment, the video identification method further includes the steps of determining a skip time point of the public video, playing the target video in response to a video playback event of the target video, and skipping and playing the public video when the playback of the target video reaches the skip time point.

[0142] Here, the skip time refers to the time when a public video needs to be skipped during playback of the target video, i.e., the time when the public video is skipped and not played. The video playback event is a trigger event for playing the target video. Specifically, the server determines a skip time in the public video, and the skip time may be at least one of the start time and the end time in the public video. The server responds to a video playback event for the target video, specifically, triggers a video playback event for the target video on the terminal by the user, plays the target video on the terminal, and skips and plays the public video when the playback of the target video reaches the skip time. That is, the public video is directly skipped and the non-public video in the target video is played. In a specific application, if the public video is an opening, the skip time may be the start time of the public video, that is, when playing the target video, the opening is skipped and the non-public video after the opening is directly played. Furthermore, if the public video is an ending, the skip time may be the end time of the public video, that is, when playing the target video, the ending is skipped and the playback is ended directly, or another video is switched and played.

[0143] In this embodiment, when the target video is being played, if the playback reaches the skip time of the public video, the public video is skipped and played, so that the overlapping public video can be skipped and played during the video playback, thereby improving the efficiency of video playback.

[0144] In one embodiment, the step of performing image matching of video frames between the target video and the video set reference video to obtain video frame pairs further includes the steps of extracting a video frame to be identified from the target video and extracting a video set reference video frame from the video set reference video; extracting video frame features of the video frame to be identified and the video set reference video frame, respectively; and performing feature matching of the video frame features of the video frame to be identified with the video frame features of the video set reference video frame, and obtaining video frame pairs based on the video frame to be identified and the video set reference video frame that have been successfully feature matched.

[0145] Specifically, after obtaining the target video and the video set reference video, the server performs video frame extraction on the target video and the video set reference video, specifically extracting a video frame to be identified from the target video and extracting a video set reference video frame from the video set reference video. The server extracts video frame features of the video frame to be identified and the video set reference video frame, respectively, and performs feature extraction on the video frame to be identified and the video set reference video frame using an image processing model, respectively, to obtain the video frame features of the video frame to be identified and the video set reference video frame. The server performs feature matching on the video frame features of the video set reference video frame with the video frame features of the video set reference video frame, for example, performing feature distance matching, and determining that the video frame to be identified and the video set reference video frame corresponding to a feature distance smaller than a feature distance threshold have been successfully matched. The server obtains video frame pairs based on the video frame to be identified and the video set reference video frame that have been successfully matched.

[0146] In this embodiment, video frames are extracted from the target video and the video set reference video to perform feature matching, and video frame pairs are obtained based on the video frames to be identified and the video set reference video frames that have successfully undergone feature matching.Therefore, similar video segments are identified based on the video frame pairs obtained by image matching, thereby improving the accuracy of similar video segment identification.

[0147] In one embodiment, the step of extracting video frame features of the video frame to be identified and the video frame features of the video set reference video frame respectively includes the step of extracting video frame features of the video frame to be identified and the video frame features of the video set reference video frame respectively using an image processing model.

[0148] Here, the image processing model may be a pre-trained artificial neural network model, such as a convolutional neural network, a residual network, or other various types of network models. Specifically, the server uses the pre-trained image processing model to extract video frame features of the video frame to be identified and the video frame features of the reference video frame of the video set. In specific application, the image processing model may be a pre-trained triple neural network model or a multi-task model.

[0149] Furthermore, the training step of the image processing model includes the steps of obtaining training sample images including classification labels, performing feature extraction and image classification on the training sample images using the image processing model to be trained to obtain sample image features and sample image categories of the training sample images, determining a model loss based on the sample image features, sample image categories, and classification labels, and continuing training after updating the image processing model to be trained based on the model loss, and obtaining the trained image processing model when training is completed.

[0150] Here, the training sample images include classification labels, and the training sample images can be set as a training dataset according to actual needs. The sample image features are image features obtained by performing feature extraction on the training sample images using the image processing model to be trained, and the sample image categories are classification results obtained by performing classification processing on the training sample images using the image processing model to be trained. The model loss updates model parameters in the image processing model to be trained to ensure convergence of the image processing model to be trained and complete model training. Specifically, the server obtains training sample images including classification labels, and performs feature extraction and image classification on the training sample images using the image processing model to be trained, thereby obtaining sample image features and sample image categories output by the image processing model to be trained. The server determines a model loss based on the sample image features, sample image categories, and classification labels. Specifically, the server can determine a triple loss based on the sample image features and a classification loss based on the sample image categories and classification labels. Specifically, the model loss can be a cross-entropy loss, and can be obtained based on the triple loss and the classification loss. The server performs updated continuous training on the image processing model to be trained based on the model loss, and upon completion of training, obtains a trained image processing model, which can extract image features from input image frames and perform image classification processing on the input image frames.

[0151] In this embodiment, the image processing model to be trained is updated and trained based on the model loss determined by the sample image features, the sample image category, and the classification label. The trained image processing model extracts the video frame features of the video frame to be identified and the video frame features of the reference video frame of the video set. The image processing model can fully extract the video frame features of the input video frame, thereby improving the accuracy of video frame matching.

[0152] In one embodiment, the step of identifying a platform-global similar segment in the target video to the platform reference video based on the second matching result obtained by video frame matching between the target video and the platform reference video includes the steps of performing image matching of video frames between the target video and the platform reference video to obtain a video frame pair, where the video frame pair includes a video frame to be identified belonging to the target video and further includes a platform reference video frame image-matched with the video frame to be identified in the platform reference video; determining a time offset of the video frame pair based on the temporal attributes of the video frame to be identified in the video frame pair and the temporal attributes of the video set reference video frame; and screening video frame pairs with matching time offsets, and determining a platform-global similar segment in the target video to the platform reference video based on the temporal attributes of the video frame to be identified in the video frame pair obtained by screening.

[0153] Specifically, the same identification method as for the video set local similar segment can be adopted to identify a platform-global similar segment in the target video to the platform reference video. The server performs image matching of video frames between the target video and the platform reference video, and for the obtained video frame pairs, the server determines the temporal attributes of the video frame to be identified in the video frame pair and the temporal attributes of the platform reference video frame. The server determines the time offset of the video frame pair based on the temporal attributes of the obtained video frame to be identified and the temporal attributes of the platform reference video frame. The server screens each video frame pair based on the time offset and screens video frame pairs with matching time offsets. Based on the screened video frame pairs, the server determines the temporal attributes of the video frame to be identified in the screened video frame pair. Based on the time attributes of the video frame to be identified, the server obtains a platform-global similar segment in the target video to the platform reference video.

[0154] In this embodiment, for the target video and the platform reference video, the time offset of the video frame pair is determined based on the time attributes of the image-matched video frame to be identified and the time attributes of the platform reference video frame, and the platform global similar segment in the target video to the platform reference video is determined based on the time attributes of the video frame to be identified in the video frame pair whose time offsets match after screening. Similar video segments of different durations are flexibly determined based on the image-matched video frame pairs, thereby improving the accuracy of identifying similar video segments in videos.

[0155] The present invention further provides application scenarios for applying the above video identification method, which are specifically applied to the application scenarios as follows:

[0156] When reproducing videos, it is necessary to use relatively pure videos as a material library, and in particular to remove promotional content from the videos that has a beneficial effect on production. For example, when it is necessary to generate a user's compilation video, it is necessary to screen the pure video parts that do not have meaningless content such as user or platform advertisements from the videos previously uploaded by the user as material, and then, through a video smart synthesis method, for example, automatically extract and assemble one video segment with the highest aesthetic evaluation score from each video to generate the user's compilation. In this case, it is very important to clean the opening, ending, and non-main content from the short videos or mini videos uploaded by the user in advance.

[0157] In the case of mini-videos for video users, which are self-recorded, self-produced, or otherwise recorded by individual users and are intended to share their lives, knowledge, practices, skills, and perspectives, the opening and closing sequences may include not only video segments with the user's logo and QR code information, but also the platform's logo. The length of these sequences is 1 to 5 seconds, which is significantly shorter than that of movies and dramas. At the same time, some video creators may randomly change or modify the opening and closing sequences. Furthermore, as platforms focus on different promotional information over time, the platform's opening and closing sequences also change accordingly, resulting in discrepancies in the openings and closing sequences of each video uploaded by users. Furthermore, over time, platform openings and closing sequences may not be correctly identified because new promotional information has been added. How to effectively identify ultra-short openings and closing sequences created by users and, at the same time, adapt to the cleaning of non-main video segments of mini-videos, whose platform openings and closing sequences only stabilize within a certain period of time, is an urgent issue that must be resolved in order to facilitate the secondary production of mini-videos. In addition, when mining the opening and ending of a mini video, it is necessary to consider whether there are openings and endings of the platform logo type. The most direct query method is to compare the target video with the global video on the video platform, that is, to query whether there are overlapping openings and endings between the target mini video and the global video. This requires a lot of time and resources, and is therefore not practical to apply.

[0158] Openings and endings can be information that varies across screens, subtitles, logos, and video themes. It is difficult to uniformly identify specific patterns using machines. Therefore, traditional methods typically involve manually tagging opening and ending information. However, manual tagging requires a large amount of tagging resources each time, resulting in poor processing efficiency. Traditional opening and ending mining methods typically only support mining openings and endings for videos with fixed opening and ending times across multiple videos, which typically involve multiple inputs for a single drama series. This is because they are unable to identify the unique openings and endings of WeMedia's self-produced materials. In reality, the opening and ending times of most videos are not always strictly aligned. Therefore, when different video collection information, opening scenes, etc. are inserted into an opening, it is not possible to strictly guarantee that the opening times are always aligned. Furthermore, conventional opening and ending mining methods can only identify openings with equal duration or endings with equal duration, resulting in inaccurate positioning of openings and endings in videos with unequal durations. When using frame-level video features to identify openings and endings, frame-level video features cannot guarantee successful matching of text-type frame images, such as the main content of a text frame or a title. In reality, regardless of whether the text content is the same or not, the frame-level video features of all text types are similar. Any change in the duration of a text frame can lead to inaccurate positioning of the opening. For example, after a drama is released, a warning is issued that the content is inappropriate. Starting from a certain episode of the drama, a text frame containing the summary content of the video is added to the opening, resulting in a difference in the duration of the text frame between the video of that episode and previous video frames.In addition, many mini-videos do not have corresponding video sets, resulting in no valid videos for mining opening and ending sequences. Furthermore, some mini-videos must be compared with global videos. Global video comparison requires mining a large number of videos, which takes time and is difficult to achieve. Regarding processing methods for building an opening and ending library to mine opening and ending sequences, only those within the library can be queried, and updating the opening and ending library is manual. This makes it difficult to extract openings and endings from a large number of videos. Relying too much on manual labor makes automation impossible, and automatic repetitive processing and maintenance are impossible.

[0159] In light of this, we propose a method for searching and identifying video openings and endings by analyzing the performance of openings and endings in global videos and local videos of the same user account. This method combines frame-level time sequence similarity search in local and global video ranges based on the construction and query of a global common opening and ending library. Specifically, the construction and maintenance of a common opening and ending library improves the detection effectiveness of current openings and endings, and an efficient global video comparison list reduces the number of comparison videos that need to be mined for openings and endings in the global range, thereby achieving mining effectiveness for newly added openings and endings within a limited time. Furthermore, by mining local videos of a user account, we quickly identify irregular opening and ending segments for the user, and finally merge the user's local mining results with the global results to achieve video opening and ending mining. Here, dynamic global mining refers to a method of mining global common openings and endings for global videos that are updated in real time, and performing real-time mining based on the current video query method. In contrast, local identification is a method of mining openings and endings from videos of the same user or series as the query video. The combination of global and local methods allows for a more comprehensive acquisition of openings and endings, improving the accuracy of opening and ending identification.

[0160] The video identification method provided in this embodiment supports the identification of opening and ending segments for any user and platform in a video. It mines a common opening and ending library based on a text OCR (Optical Character Recognition) identification recommendation global matching list, thereby reducing the overall video processing volume and ensuring the effectiveness of common opening and ending mining. It also uses image sequence similarity search to realize cross-searching between two videos, thereby finding overlapping openings and endings. By building a dynamically updated library of common openings and endings, it supports searching the library for openings and endings upon input query, improving response efficiency and supporting the identification of openings and endings for various types of videos. Compared to conventional opening and ending identification methods, the video identification method provided in this embodiment supports the identification of openings and endings of variable lengths and uses video frame similarity sequence search to realize the identification of openings and endings that are not time-aligned or have variable durations. Furthermore, by searching the common opening and ending library and efficiently extracting global videos, openings and endings can be searched and mined, improving the mining ability of common openings and endings. At the same time, it supports the mining of new platform openings and endings, meeting the requirement of dynamically maintaining and identifying common openings and endings due to the dynamic update of platform promotion during application. At the same time, by controlling the global video scope of the search, it is possible to avoid excessive resource consumption for global search of large amounts of data.Furthermore, by maintaining a common opening / ending and keyword library that supports global library search, it not only supports the removal of existing openings and endings, but also supports real-time mining of newly added openings, endings, or keywords. It also provides automatic repair capabilities with simple manual intervention for openings and endings that are missed in the search, further improving the accuracy of video opening and ending identification.

[0161] The video identification method provided in this embodiment can be applied to identifying the opening and closing segments of mini videos, thereby removing the opening and closing segments to obtain the main part of the mini video, and using the mini video for video comparison, etc. In the secondary production of a user's compilation video, as shown in Figure 4, the opening and closing segments are removed from all of a user's uploaded videos, and the main part of the video is retained. Each video is then cut into video segments, one segment per 3 seconds, and all frames of each video segment are aesthetically evaluated and scored. The average score is taken as the aesthetic score of the video for that segment, and the highest aesthetic score for each video among all of the user's videos is obtained. Multiple video segments are stitched together, filter-beautified, and output as the user's compilation video. As shown in FIG. 5, in a scenario where a user compares videos, after identifying the opening and ending of a video uploaded by a user, the main video is retained and the main video is searched for similar time slots in a past video library. If a matching video is found in the past video library, the search engine determines whether the video already exists in the past video library or whether a similar video exists, thereby enabling fast video comparison. As shown in FIG. 6, when a video A is played on a video platform, the opening of the platform introduction screen of the video platform is displayed, specifically, the screen at 2 seconds. As shown in FIG. 7, the video content of the video A is played, specifically, the screen at 20 seconds (including a person) of the video A is played. As shown in FIG. 8, when the playback of the video A ends, the ending of the platform introduction screen of the bio platform is played, specifically, the screen at 1 minute 12 seconds. When editing the video A on the video platform, the opening and ending segments of the platform introduction screen need to be removed in order to retain the main video content.After multiple users upload videos, the platform adds platform logo segments to the videos at the same time period, so that videos with the same logo segment can be found more quickly through global video queries for the same time period, thereby determining that the matched segments have a common ending. As shown in Figure 9, during a first period, video platform A includes text and icon 901 in the opening and closing of its platform introduction screen. As shown in Figure 10, after a certain period of time, during a second period, the opening and closing of video platform A's platform introduction screen include, in addition to text and icon 1001, download promotion information 1002, which may include a download link for the application platform.

[0162] Specifically, in the video identification method provided in this embodiment, as shown in FIG. 11, a query video is a target video for which video identification is to be performed. A user video list for the query video is obtained. Each video in the user video list and the query video belong to the same user account. If the user video list is successfully obtained, openings and endings are mined for each video in the user video list to obtain openings and endings. If the user video list is not obtained, openings and endings are not mined for the user video list. Furthermore, the query video is identified as a common opening and ending. If openings and endings cannot be identified, a global video list in the video platform is obtained. The global video list contains videos extracted from the video platform to which the query video belongs. Openings and endings are mined for the query video based on the global video list to obtain openings and endings. The common opening and ending identification results are merged with the mining results from the user video list to obtain openings and endings, which are then output. Alternatively, the mining results from the global video list are merged with the mining results from the user video list to obtain openings and endings, which are then output. Furthermore, for the mining results of the global video list, common openings and endings are extracted from the mining results, recommended openings and endings corresponding to the extracted common openings and endings are counted and updated, and if the common opening and ending determination condition is met, the extracted common opening and ending is updated to the common opening and ending library after the Tth day.

[0163] Furthermore, for a given query video, other videos of the uploading username are first mined. Here, mining includes searching for similar time periods between video pairs and correcting frame-level OCR keyword queries. If no search results are found in the common opening and ending library, it indicates that the current query video may contain new openings and endings of the platform logo type, which triggers global video mining. Specifically, the identified OCR platform keywords are used to find recent videos containing the same platform keywords from global videos to construct a global video list. A search for similar time periods is performed using the query video and the global list videos. If results are found, this indicates the emergence of a new platform logo type. In this case, the search results are output by merging the video search results under the username. At the same time, the new platform logo type is recommended to the common opening and ending library. If no results are found, it indicates that this video does not have a matching opening and ending in the global library. Furthermore, to ensure the automatic addition of common openings and endings, the new global common openings and endings obtained each time are statistically processed by the opening and ending library to determine whether to recommend updating them to the common opening and ending library.

[0164] As shown in FIG. 12, the video identification method provided in this embodiment includes processes such as global library query, local list mining, global list generation, global list mining, recording new openings and endings in the common opening and ending library, and keyword library maintenance. Specifically, in the global library query, the embedding features of the frame-level images of the query video and the frame-level images of the common opening and ending library can be directly used. Specifically, frame-level images are extracted from the query video and the video in the common opening and ending library, respectively, and the frame-level features of the extracted frame-level images are obtained. A similar time slot search is performed based on the frame-level features, and the matching time slots are regarded as the openings and endings obtained through the search, resulting in identification result 1. Specifically, the query video is matched with multiple openings and endings in the global library to obtain matching time slots, where the longest time slot is the final search result. If no matching time slots for the openings and endings are found, it is determined that the openings and endings in the query video cannot be identified based on the common openings and endings in the global library.

[0165] Global list mining uses the same processing method as local list mining, except for the video list used for search. Frame-level images are obtained from each of the query video and the global list video, and frame-level features of each frame-level image are extracted to perform fixed-segment sequence similarity search processing to obtain identification result 2. Local list mining involves forming two-to-two video pairs between the query video and each video in the user video list. Frame-level images are obtained for each video pair, and frame-level features of the frame-level images are extracted to perform fixed-segment sequence similarity search processing. Similar segments are then generated using video frame images through similar time zone search. After completing all video pair searches, multiple similar segments are obtained, which are then merged to obtain local openings and endings, resulting in identification result 4. Meanwhile, for the frame-level images obtained from the video pairs, frame-level OCR is used to find platform keywords from a keyword library to obtain identification result 3. Identification result 4 is then corrected using identification result 3, and identification result 3 and identification result 4 are merged to obtain the merged result.

[0166] Specifically, for Identification Results 3 and 4, Identification Result 4 is highly reliable opening and ending information obtained by searching two videos. Identification Result 3 is information on whether a frame is invalid, obtained by determining whether the frame contains certain special vocabulary. Therefore, Identification Result 4 is corrected using the information in Identification Result 3. Here, the role of Identification Result 3 is to provide keywords for the opening and ending of the video. For example, an ending may be a promotional screen for a certain video platform, making it invalid for derivative video creations. Therefore, invalid frames near the opening and ending must be removed using this special vocabulary. Specifically, a text search method can be used to remove such characters from the main video. First, the characters to be removed are stored in a keyword library. The OCR obtained by identifying the characters from the input frame image is then queried to determine whether the library keyword exists in the OCR. If the library keyword is found, the frame is considered invalid. The opening and ending times are corrected using the text search results by determining whether all frames are invalid based on whether they contain a match.

[0167] In a specific application, for the opening end time, for example, if the end time of opening [2,18] is 18 seconds, the classification information starting from the opening end time is searched for. If more than 50% of the main screens from the end of the opening to the start of the ending are invalid, the invalid screens are not cleaned. If two or more invalid frames are included within 5 seconds after the end of the opening, i.e., frames 19 to 23, the opening end time is adjusted to the time of the last invalid screen. If there are a certain period of consecutive invalid screens after the end of the opening, the opening end time is adjusted directly to the longest consecutive invalid time. Similarly, for the ending start time, a certain period of time prior to the start time is searched for. If an invalid screen appears, the ending start time is adjusted to the next second after the invalid screen. As shown in Figure 13, for Opening 1, the time of Opening 1 is extended to the end time of the invalid screen containing the identified platform keyword. As shown in Figure 14, for Ending 1, the time of Ending 1 is extended to the start time of the identified invalid screen containing the identified platform keyword.

[0168] Whether querying using a global library, mining using a global list, or mining using a local list, fixed-segment sequence similarity search can be performed based on the frame-level features of frame-level images. Specifically, common opening and closing videos in the global library, global videos in the global list, or user videos in the local list are used as reference videos for the query video to form a video pair with the query video. For frame-level feature extraction, frames are extracted from the video to obtain frame-level images, and frame-level features are extracted for each frame-level image. For example, for a 6-second video at 25 FPS (frames per second), one frame is extracted per second, for a total of six images. Then, for the frame-extracted images, a feature extractor obtains video frame features for each frame, resulting in six video frame features. When a frame extraction method with three frames per second is used, the final opening and closing identification time accuracy is 0.33 seconds. For shorter mini-videos, if higher time accuracy is required, a more dense frame extraction method with 10 frames per second and an accuracy of 0.1 seconds can be used for frame extraction processing. Here, the video frames are extracted by an image feature extractor. The image feature extractor employs the output of a ResNet-101 neural network pooling layer trained on the open-source classification dataset Imagenet to convert each image into a 1x2048 image embedding vector. Imagenet is a large-scale open-source dataset for common object identification. The image feature extractor can also be implemented based on different network structures and different pre-trained model weights.

[0169] Here, image embedding is used to describe the features of image information, including image lower-layer representations, image semantic features, etc. Embedding is not limited to floating-point features, but may also be an image representation consisting of a binary feature vector, i.e., deep hash features. In this embodiment, the embedding features may be binarized deep hash features. The image lower-layer representation is an image embedding from the lower-layer features of deep learning, describing some representation information such as the texture and feature location of the full-view image. The image semantic features are image embeddings from semantic learning, describing the representation of specific designated semantic content parts in the image. For example, when used to describe a dog embedding, the feature of the dog's location in the image is used as the image representation.

[0170] The configuration of the ResNet-101 convolutional neural network (CNN) deep representation module is shown in Table 1 below.

[0171] [Table 1] Furthermore, by performing OCR classification on each frame-extracted image, text information on each image can be identified.

[0172] In embedding-based sequence similarity search processing, when performing video time period matching, for each video pair (i, r) consisting of a query video and a list video, the list video is a video in the global library, global list, or local list, i refers to the query video of the opening and ending waiting to be determined, and r refers to a list video that serves as a reference video. Assuming there are three list videos, for query video i, three sequence similarity search algorithm calculations based on embedding1 and three sequence similarity search algorithm calculations based on embedding2 are required.

[0173] Specifically, the sequence similarity search, also known as the time-slot matching algorithm, processes one pair of videos each time, and the input for each video is its embedding sequence. The threshold in the time-slot matching algorithm can be dynamically adjusted according to the needs of the service or the video being processed. The time-slot matching algorithm specifically pre-sets the distance threshold t0 of the video frame feature embedding as t0 = 0.3. That is, if the Euclidean distance between two embeddings is less than 0.3, it indicates that the two embeddings are from similar frames. The distance threshold can be flexibly set according to actual needs. Frame extraction is performed on the two videos in the video pair to obtain the embedding for each frame. For each frame j in video i, the Euclidean distance between it and each frame embedding in video r is calculated. Frames less than t0 are considered similar frames of j, and a list of similar or matching frames for j, sim-id-list, is obtained. At the same time, the corresponding similar frame time deviation diff-time-list is recorded. For example, for frame j=1, if the similar frame list sim-id-list is [1,2,3], it indicates that it is similar to the 1st, 2nd, and 3rd seconds of the r video, and if the time deviation diff-time-list is [0,1,2], it indicates the distance between the similar frames in sim-id-list and the time represented by the j=1 frame. The default frame extraction is to extract one frame per second, so the frame number is the number of seconds. Therefore, we can obtain the similar frame list SL and the time deviation list TL for all frames in i.

[0174] All frames are traversed and the number of matching frames between video i and video r is counted, i.e., the number of matching frames between video r and video j. If the number of matching frames is less than 1, videos i and r do not have the same video segment, and the opening and ending cannot be mined. Otherwise, the time deviation dt is sorted to obtain the SL list. Specifically, all matching frames in the SL are sorted in ascending order according to diff-time (i.e., dt). If dt is the same, they are sorted in ascending order according to the number of video i. At the same time, the corresponding diff-time list is reconstructed according to this order, i.e., frames with a time difference of 0 are placed first, and frames with a time difference of 1 are placed second. For example, the new SL list might be [10,11], [11,12], [2,4], [3,5], [4,6], [6,9], [7,10].

[0175] To reconstruct data by dt to obtain match-dt-list, the list of all frames in the similar frame list SL of i is reconstructed using the time deviation as the primary key to obtain an ascending list of dt, and similar frames with time deviations of 0, 1, 2, etc. are obtained as match-dt-list:{0:{count,start-id,match-id-list},...}. For example, {2:{3,2,[[2,4],[3,5],[4,6]]}, 3:{2,6,[[6,9],[7,10]]}}, where 2 means a time difference of 2. For example, if the second frame of i and the fourth frame of video vid2 are similar, the time difference between these two frames is 1. count is the number of similar frames in this time deviation. If the second frame of i above and the fourth frame of vid2 are similar, add 1 to count. start-id is the smallest frame id of i in this time difference. For example, if the first frame of i and vid2 are not similar, but the second frame of i and the fourth frame of video vid2 are similar, start-id will be 2.

[0176] Two dt lists in the match-dt-list whose previous and next dts are less than 3 (i.e., matching pairs with a matching deviation of 3s or less) are merged, and the list with a larger dt is merged with the list with a smaller dt. At the same time, the matching for similar frames with a larger dt is updated, and the matching frame list SL is updated. For example, in the above example, the list with a dt of 2 is merged with the list with a dt of 3, and the final result is {2:{5,2,[[2,4],[3,5],[4,6],[6,8],[7,9]]}}. Here, count is the sum of the counts for dt=2 and dt=3, and start-id is the frame where the smallest i video frame is found from the similar frame lists for dt=2 and dt=3. For the list with dt=3, the matched frame numbers are rewritten and merged. For example, [6,9] is rewritten as [6,8] and merged into the similar frame list for dt=2. At the same time, the similar frame pairs with rewritten frame numbers are synchronized and updated in the SL matching frame list in step 5), e.g., [10,11], [11,12], [2,4], [3,5], [4,6], [6,8], [7,9]. As mentioned above, the existence of a merged frame list may disrupt the order of dt or frame IDs, so they must be reordered. Specifically, dt is reordered. That is, the process of reordering dt for the new SL list to obtain an SL list is performed, resulting in a matching frame list sorted in ascending order of dt (ascending order of frame IDs of video i). Data is reconstructed using dt to obtain a match-dt-list. That is, the process of reconstructing data using dt to obtain a match-dt-list is performed again.

[0177] A time duration matching list (match-duration-list) is calculated. Specifically, the time interval between two matching segments is set to be greater than T2 (e.g., 8 seconds, where 1 second is 1 frame, and the difference in frame numbers is 8). For each dt (e.g., dt=2) in match-dt-list, for each frame srcT of video i at dt (e.g., 2 out of 2, 3, 4, 6, and 7 in the example above), if the difference between srcT and the previous srcT is greater than T2 (e.g., if the difference between 2 and the previous srcT is 9, and it is greater than the interval threshold), the previous similar frame pair is merged into one matching segment, new similar frame pair statistics are started from the current srcT, and the similar frames are saved in a temporary list (tmplist). When dt=2 and srcT=2, the similar frames in the previous temporary frame list are saved as matching segments. For example, the similar frames in the previous tmplist=[[10,11],[11,12]] are added to match-duration-list as matching segments. Matching segment information such as [10,11,11,12,1,2,2] is added, where each value is [src-startTime,src-endTime,ref-startTime,ref-endTime,dt,duration,count]. That is, the matching segment stores the start and end frames of video i, the start and end frames of the matching video, the dt of the matching segment, the duration of the matching segment, and the number of matched similar frames. As shown in FIG. 15, the matching segment information includes information such as the start frame time of the target video, the end frame time of the target video, the start frame time of the matching video, and the end frame time of the matching video. The current similar frames are saved in the temporary list as tmplist=[[2,4]]. If the difference between srcT and the previous srcT is less than T2, save the current similar frame in a temporary list tmplist.For example, for dt2, save srcT=3, 4, 6, 7 all in a temporary list, so that we get tmplist=[[2,4],[3,5],[4,6],[6,8],[7,9]]. If we are currently at the last similar frame of this dt (say srcT=7), construct a matching segment from the accumulated similar frames in tmplist and add it to match-duration-list. For example, add [2,7,4,9,2,6,5], where duration is 7-2+1 and count=5 is the count of similar frames, so match-duration-list=[[10,11,11,12,1,2,2],[2,7,4,9,2,6,5]]. The above match-duration-list is sorted in reverse order of the number of similar frames count, for example, match-duration-list=[[2,7,4,9,2,6,5],[10,11,11,12,1,2,2]].

[0178] This process is performed when there are overlapping time periods in match-duration-list. Similar frame calculation traverses all frames of two videos to calculate the distance, and performs operations on those that are similar within a certain threshold range. This can easily result in a frame being similar to multiple frames, resulting in a situation where two matching time periods overlap in match-duration-list. This situation must be handled. Specifically, the duration of the minimum matching segment is set to T3 (e.g., 5, indicating that the shortest matching duration is 5 seconds). For time period i (referring to a time period consisting of src-startTime and src-endTime) in match-duration-list, for time period j=i+1 in match-duration-list, if time period j is included in time period i, j is deleted. As shown in FIG. 16, if the start time of time period i is before the start time of time period j and the end time of time period i is after the end time of time period j, that is, time period j is included in time period i, and j must be deleted. If there is an overlap between i and j and the start time of i is the earliest start time, the start time of j is moved back to the end time position of i, and j is updated. In this case, if the duration of time period j is less than T3, j is deleted; otherwise, the old j is replaced with a new j. As shown in FIG. 17, the start time of time period i is before the start time of time period j, and the end time of time period i is before the end time of time period j, i and j overlap, and the end time of time period i needs to be updated to the end time of time period j. If i and j overlap and the start time of j is the earliest start time, the end time of j is moved back to the start time position of i, and j is updated. In this case, if the duration of time period j is less than T3, j is deleted; otherwise, the old j is replaced with a new j. As shown in FIG. 18, the start time of time period i is after the start time of time period j, and the end time of time period i is after the end time of time period j, i and j overlap, and the start time of time period i needs to be updated to the start time of time period j.Finally, return the matching duration information, e.g., match-duration-list=[[2,7,4,9,2,6,5],[10,11,11,12,1,2,2]], or just the matching segments [[2,7,4,9],[10,11,11,12]].

[0179] To obtain the same matching segments, we perform similarity sequence matching on the query video with the video list, obtain three matching time periods, and then align these three time periods to obtain the same matching segments in the video list based on this embedding. Specifically, suppose video i needs to be mined from videos vid2, vid3, and vid4. Then, we perform the aforementioned video segment matching process on each of the N=3 video pairs [I, vid2], [I, vid3], and [I, vid4] to obtain three matching pieces of information. For example, the first pair of video matching segments returns [[2, 7, 4, 9], [10, 11, 11, 12]]; the second pair of matching segments returns [[2, 7, 4, 9]]; and the third pair returns [[2, 7, 4, 10]]. Count the matching segments, for example, count [2,7,4,9] twice, [2,7,4,10] once, and [10,11,11,12] once. Sort the matching segments in reverse order of count, and if the counts are the same, sort in ascending order of src-startTime, to get match-list=[[2,7,4,9],[2,7,4,10],[10,11,11,12]], count-list=[2,1,1].

[0180] Matching segments that overlap within the match-list are merged. Specifically, the effective overlap rate T4 is set to, for example, 0.5. If the overlapping duration of two time slots exceeds T4, the two counts must be merged for calculation. The effective matching count T5 is set to, for example, 3. If the matching segment count for a segment is greater than T5, the segment cannot be ignored. For time slot i (referring to a time slot consisting of src-startTime and src-endTime) in the match-list, for time slot j = i+1 in the match-list, if time slot i includes time slot j and the duration of j segment is greater than 0.5*i, j is deleted, and the count of i segment is calculated as the count of the original i segment + the count of j segment. If i and j overlap, the overlapping duration is greater than 0.5*i, and the count of j segment is greater than T5, the times of i and j are merged with the longest start and end time, and the count of i segment is calculated as the count of the original i segment + the count of j segment. If the count of the jth segment is less than T5, the jth segment is deleted, and the count of the ith segment is set to the count of the original ith segment + the count of the jth segment. In other words, in this case, the ith and jth segments are not merged, and only the ith segment with the most occurrences is kept. The count of the jth segment is reflected in the count of the new ith segment. If ith and jth segments overlap and the overlap time length is less than 0.5 * the time length of the ith segment, the jth segment is discarded. As shown in Figure 19, if the start time of time period i is before the start time of time period j and the end time of time period i is before the end time of time period j, then ith and jth overlap, and the end time of time period i should be updated to the end time of time period j. On the other hand, if the start time of time period i is after the start time of time period j and the end time of time period i is after the end time of time period j, then ith and jth overlap, and the start time of time period i should be updated to the start time of time period j.

[0181] A new video matching segment match-list (e.g., [[2,7,4,9], [10,11,11,12]]) and count count-list (e.g., [3,1]) are obtained. A valid overlapping occurrence ratio threshold T6 is set. In N pairs of video pair mining, a matching video segment has an overlapping occurrence count x>N*T6, which means it is a valid overlapping segment (e.g., T6=0.5). For the match-list, the valid time period is reserved, and match-list=[[2,7,4,9]] and count=[3] are obtained. Here, match-list is the identification result obtained by performing a fixed segment sequence similarity search on frame-level features and different list videos.

[0182] To generate a global list, global videos from the past week or two weeks are searched for videos with the same OCR keywords, and 10,000 videos are sampled from these to form a global list. Compared to generating a global list directly using all global videos, using videos from the same platform, the same time period, or recent videos reduces the number of videos required for comparison, narrows the scope of updates, and makes it easier to mine new platform openings and endings. If there is no match for the OCR vocabulary in the keyword library, 10,000 videos are randomly sampled from global videos from the past week to form the global list. To efficiently generate a global list, OCR text is pre-extracted from global mini-videos and queried against the keyword library, thereby associating each word in the keyword library with a specific global mini-video. The keyword library contains various keywords, and videos on the video platform are associated with the keywords in the keyword library. Furthermore, the global list and the query videos have the same keywords. At the same time, using 10,000 videos with the same keywords and combining them with 10,000 global random samples can improve generalization performance and keyword identification accuracy. As shown in Figure 12, for a global new video, such as a video newly uploaded by a user on a video platform, frame-level images can be extracted from the global new video, character recognition can be performed on the frame-level images, and keyword queries can be performed using the character recognition results and each keyword in the keyword library to achieve video information collection for the global new video. For example, a relationship between the global new video and the corresponding keyword can be established. A video information collection process can also be performed on each video on the video platform to obtain a global list.

[0183] Regarding the maintenance of the keyword library, as video platforms are constantly emerging, new video platforms may emerge, and the keyword library needs to be dynamically updated and maintained. Keywords of the platform logo segments of the opening and closing sequences that appear on new video platforms can be directly added to the library, thereby realizing the dynamic updating and maintenance of the keyword library. Specifically, during local list mining, the platform keywords of the query video can be obtained, and the obtained platform keywords can be updated in the keyword library.

[0184] Regarding the recording of new openings and endings in the common opening / ending library, recommended openings and endings are generated from anchor point identification result 1 or identification result 2 in list mining and saved in the recommendation library, and the number of appearances N1 and the number of new additions N2 of these openings and endings are recorded. As shown in Figure 20, a single-video common sequence similarity search is performed using frame-level images obtained from the query video to obtain openings and endings, and the number of appearances N1 and the number of new additions N2 of these openings and endings can then be updated. Each time the above video list and single-video mining are performed, the recommendation library is queried for the presence of the opening and ending. If it is determined that the opening and ending are included in the mining results of the openings and endings obtained from the above video list and single-video mining, the number of appearances and the number of new additions N1 and N2 of these openings and endings in the recommendation library are updated. After T days, openings and endings with the highest number of new additions are selected based on the number of new additions and saved in the common opening / ending library.

[0185] Specifically, after highly reliable openings and endings are mined in global list mining, these openings and endings can be used in subsequent video global library query processing. To ensure the validity of the common opening and ending library, a buffer library, i.e., a recommended opening and ending library, can be adopted. This recommended opening and ending library is used to store all openings and endings generated by global list mining, as well as validity information N1 and N2, where N1 is the number of times the opening and ending have appeared, and N2 is the number of times the opening and ending have been newly added. For a given opening and ending, when it is added to the library, N1 is recorded as 1 and N2 is recorded as 0. Each time a query video is input, a query is performed in the recommended opening and ending library. If a match is found, 1 is added to N2 for that opening and ending. After a certain period of time, assuming a time threshold of 7 days, the openings and endings are sorted in descending order based on the number of N2 records. The top 10% of openings and endings with N2 > 100 are selected, and the final recommended openings and endings for this period are recorded in the common opening and ending library. If these openings and endings have been recorded in the common opening and ending library before, all recommended opening and ending library records are updated, i.e., N1 = original N1 + N2, N2 = 0. This starts the next period of statistics. In addition to N1 and N2, time T can also be recorded when adding a video to the library to indicate the number of days the video will be added to the library. Openings and endings whose number of days in the library is a multiple of 7 days are counted daily, and if their N2 record is greater than the specified threshold, they are recorded in the common library. At the same time, the recommended opening and ending library records for all videos that are a multiple of 7 days are updated, i.e., N1 = original N1 + N2, N2 = 0.This will start the next period's statistics. Other threshold determination strategies based on N1, N2, and T can also be adopted to update the common opening and ending library. The time period for updating the recommended opening and ending library to the global opening and ending library can also be adjusted in real time, and will be updated when traffic reaches a certain threshold based on daily video traffic statistics.

[0186] A merged result is generated from classification results 3 and 4, and then merged with classification result 1 or 2. Because both classification results are obtained based on searching multiple video pairs, the resulting matching time periods provide strong opening and ending information, i.e., the time periods are highly reliable as belonging to openings and endings. In this case, the two classification results must be merged to obtain openings and endings that appear multiple times between videos. Specifically, when merging the merged result with classification result 1 or 2, the opening time segments of classification result 1 or 2 are merged, and the maximum time is set as the opening end time, e.g., [2,7], [9,15], [9,13]. After the merged time, [2,15] is output as the opening time period, with 15 as the end time. Similarly, when merging the merged result with the endings of classification result 1 or 2, the minimum time is set as the ending start time, thereby obtaining a comprehensive classification result that includes the openings and endings obtained by comprehensive classification.

[0187] The video identification method provided in this embodiment supports the identification of openings and endings of unequal length. It uses video frame embedding similarity sequence search to identify openings and endings that are not time-aligned or of unequal duration. Furthermore, by combining local and global list embedding mining and user- and platform-dimensional opening and ending identification, the overall detection effect is improved, avoiding the ignorance of platform-dimensional openings and endings in standard mining, thereby enabling more thorough mini-video content cleaning. Furthermore, the mined global openings and endings are stored in a recommended opening and ending library, full network recall statistics, and an official opening and ending library, enabling closed-loop management of opening and ending mining and common openings and endings. In addition to opening and ending identification for mini videos, after limited modifications, the video identification method provided in this embodiment can also be applied to the opening and ending identification process of other types of videos, such as long videos such as movie dramas, for example, it is necessary to limit the video list of global mining for long videos, thereby avoiding comparison with excessive videos, which will cause increased time consumption.

[0188] It should be understood that although the steps in the flowcharts according to the above-described embodiments are displayed sequentially according to the arrows, these steps are not necessarily performed sequentially in the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not limited to a strict order, and these steps may be performed in other orders. Furthermore, at least some of the steps in the flowcharts according to the above-described embodiments may include multiple steps or multiple stages, and these steps or stages may not necessarily be performed simultaneously but may be performed at different times. Furthermore, the order in which these steps or stages are performed does not necessarily have to be sequential but may be alternated with other steps or at least some of the steps or stages in other steps.

[0189] Based on the same inventive idea, the embodiments of the present application further provide a video identification device for realizing the above-mentioned video identification method. The problem-solving embodiments provided by this device are similar to those described in the above-mentioned method, so the specific limitations of one or more of the embodiment examples of the video identification device provided below may refer to the limitations related to the above-mentioned video identification method, and will not be repeated here.

[0190] In one embodiment, as shown in FIG. 21, a video identification device 2100 is provided, the device comprising: a moving image set video acquisition module 2102, a local similar segment identification module 2104, a platform video acquisition module 2106, a global similar segment identification module 2108 and an overall similar segment determination module 2110, wherein: The video set video acquisition module 2102 is configured to acquire a target video and a video set reference video in a video series video set, where the video series video set includes videos belonging to the same series; The locally similar segment identification module 2104 is configured to identify a video set locally similar segment in the target video to the video set reference video based on a first matching result obtained by video frame matching between the target video and the video set reference video; The platform video acquisition module 2106 is configured to acquire a platform reference video from a video platform to which the target video belongs; the global similar segment identification module 2108 is configured to identify a platform global similar segment in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video; The overall similar segment determination module 2110 is configured to determine overall similar segments in the target video to the video set reference video and the platform reference video based on the positions of the video set local similar segments and the platform global similar segments in the target video, respectively.

[0191] In one embodiment, the device further includes a correction update module configured to perform correction updates on video set local similar segments based on correction segments containing correction keywords in the target video to obtain updated video set local similar segments, and the overall similar segment determination module 2110 is further configured to determine overall similar segments in the target video to the video set reference video and the platform reference video based on the positions of the updated video set local similar segments and platform global similar segments, respectively, in the target video.

[0192] In one embodiment, the correction update module includes a correction segment determination module, a timestamp update module, and a similar segment update module, wherein the correction segment determination module is configured to determine a correction segment in the target video including the correction keyword, the timestamp update module is configured to update the timestamp position of a video set locally similar segment in the target video based on the timestamp position of the correction segment in the target video to obtain an updated timestamp position, and the similar segment update module is configured to determine an updated video set locally similar segment in the target video based on the updated timestamp position.

[0193] In one embodiment, the correction segment determination module is further configured to perform character identification on video frames in the target video to obtain character identification results, match the character identification results with the correction keyword to obtain a matching result, and determine a correction segment from the target video that includes the correction keyword based on a video frame associated with the matching result that matches.

[0194] In one embodiment, the platform reference video includes platform public video segments obtained from the public video library of the video platform to which the target video belongs and platform-related videos obtained from the video platform, and the global similar segment identification module 2108 includes a public video matching module, a related video matching module, and a matching result processing module, wherein the public video matching module is configured to perform video frame matching between the target video and the platform public video segments to obtain public video matching results, and if a similar segment cannot be identified based on the public video matching results, the related video matching module is configured to perform video frame matching between the target video and the platform-related video to obtain related video matching results, and the matching result processing module is configured to identify platform global similar segments in the target video to the platform-related video based on the related video matching results.

[0195] In one embodiment, the device further includes a public video update module configured to update the discrimination statistics parameters of the platform global similar segment, obtain updated discrimination statistics parameters, and if the updated discrimination statistics parameters satisfy the platform public judgment criteria, update the platform global similar segment as a platform public video segment in the public video library.

[0196] In one embodiment, the platform video acquisition module 2106 is further configured to acquire platform public video segments from the public video library of the video platform to which the target video belongs, and the global similar segment identification module 2108 is further configured to identify platform global similar segments to the platform public video segments in the target video based on second matching results obtained by video frame matching between the target video and the platform public video segments.

[0197] In one embodiment, the platform video acquisition module 2106 includes a platform determination module, a related video query module, and a video screening module, where the platform determination module is configured to determine the video platform to which the target video belongs and the correction keywords contained in the video frames in the target video, the related video query module is configured to query platform-related videos having a related relationship with the correction keywords in the video platform, and the video screening module is configured to screen from the platform-related videos to obtain a platform reference video according to the reference video screening conditions.

[0198] In one embodiment, the device further includes a related relationship building module configured to perform character identification on video frames in platform videos belonging to a video platform, obtain video keywords, perform matching in a keyword library based on the video keywords, determine target keywords matching the video keywords, and establish related relationships between the platform videos and the target keywords, and the related video query module is further configured to query platform related videos related to the corrected keywords in the video platform based on the related relationships.

[0199] In one embodiment, the overall similar segment determination module 2110 includes a timestamp determination module, a timestamp merging module, and an overall timestamp processing module, wherein the timestamp determination module is configured to determine a first timestamp position of a video set local similar segment in the target video and a second timestamp position of a platform global similar segment in the target video, the timestamp merging module is configured to merge the first timestamp position and the second timestamp position to obtain an overall timestamp position, and the overall timestamp processing module is configured to determine an overall similar segment in the target video to the video set reference video and the platform reference video based on the overall timestamp position.

[0200] In one embodiment, the locally similar segment identification module 2104 includes a video set video frame matching module, a video set offset determination module, and a video set video frame pair processing module, wherein the video set video frame matching module is configured to perform image matching of video frames between a target video and a video set reference video to obtain video frame pairs, the video frame pair including a video frame to be identified belonging to the target video and a video set reference video frame that image matches with the video frame to be identified in the video set reference video, the video set offset determination module is configured to determine a time offset of the video frame pair based on temporal attributes of the video frame to be identified in the video frame pair and temporal attributes of the video set reference video frame, and the video set video frame pair processing module is configured to screen video frame pairs with matching time offsets, and determine a video set locally similar segment in the target video to the video set reference video based on temporal attributes of the video frame to be identified in the video frame pair obtained by screening.

[0201] In one embodiment, the video set video frame pair processing module is further configured to perform numerical matching on the time offsets of each video frame pair, screen video frame pairs whose time offsets numerically match based on the numerical matching result, determine a start time and an end time based on the time attributes of the identified video frame in the video frame pair obtained by screening, and determine a video set local similar segment from the target video to the video set reference video based on the start time and the end time.

[0202] In one embodiment, the moving image set video frame pair processing module is further configured to obtain a video frame pair list consisting of the video frame pairs obtained by screening, in which each video frame pair is sorted in order from smallest to largest time offset value, and video frame pairs with the same time offset are sorted in order from smallest to largest time stamp value of the included video frame to be identified, the time stamp being determined based on the time attribute of the included video frame to be identified, determine a time attribute distance between the time attributes of the video frames to be identified in adjacent video frame pairs in the video frame pair list, determine adjacent video frame pairs whose time attribute distance does not exceed a distance threshold as video frame pairs belonging to the same video segment, and determine a start time and an end time based on the timestamp of the video frames to be identified in the video frame pair belonging to the same video segment.

[0203] In one embodiment, the moving image set video frame pair processing module is further configured to determine a start video frame pair and an end video frame pair from video frame pairs belonging to the same video segment based on timestamps of the identified video frames in the video frame pairs belonging to the same video segment, obtain a start time based on the timestamp of the identified video frame in the start video frame pair, and obtain an end time based on the timestamp of the identified video frame in the end video frame pair.

[0204] In one embodiment, the moving image set video frame pair processing module is further configured to numerically compare the time offsets of each video frame pair respectively, obtain a numerical comparison result, and based on the numerical comparison result, screen video frame pairs from each video frame pair whose numerical difference in time offset is less than a numerical difference threshold, perform offset update for video frame pairs whose numerical difference in time offset is less than the numerical difference threshold, and obtain video frame pairs whose time offsets are numerically matched.

[0205] In one embodiment, there are at least two video set reference videos, and the video set video frame pair processing module is further configured to: screen video frame pairs with matching time offsets; determine intermediate similar segments in the target video to the video set reference video based on the temporal attributes of the identified video frame in the video frame pairs obtained by screening; perform segment update for each intermediate similar segment in the target video that has an overlapping relationship with each video set reference video among the intermediate similar segments to each video set reference video; and obtain video set local similar segments in the target video to each video set reference video.

[0206] In one embodiment, the video set video frame pair processing module is further configured to perform segment updating for each intermediate similar segment in the target video that has an overlapping relationship with each video set reference video among the intermediate similar segments, obtain an updated intermediate similar segment, determine statistics of the updated intermediate similar segment, and if the statistics of the updated intermediate similar segment exceed a statistics threshold, obtain a video set local similar segment in the target video for each video set reference video based on the updated intermediate similar segment.

[0207] In one embodiment, the video set video frame pair processing module is further configured to perform segment position comparison for intermediate similar segments in the target video with respect to each video set reference video, obtain segment comparison results, determine each intermediate similar segment for which an overlap relationship exists as a result of the segment comparison, perform segment update for each intermediate similar segment for which an overlap relationship exists based on the overlap time length and statistics of each intermediate similar segment for which an overlap relationship exists, and obtain video set local similar segments in the target video with respect to each video set reference video.

[0208] In one embodiment, the video set video frame pair processing module is further configured to obtain a similar segment list consisting of intermediate similar segments in the target video for each video set reference video, in which each intermediate similar segment is sorted in order from largest to smallest statistical value, and intermediate similar segments with the same statistical value are sorted in order from earlier to later starting time, and to perform segment position comparison for each intermediate similar segment in the similar segment list to obtain a segment comparison result.

[0209] In one embodiment, the video set video frame matching module is further configured to extract a video frame to be identified from the target video, extract a video set reference video frame from the video set reference video, extract video frame features of the video frame to be identified and the video frame features of the video set reference video frame, respectively, perform feature matching of the video frame features of the video frame to be identified with the video frame features of the video set reference video frame, and obtain a video frame pair based on the video frame to be identified and the video set reference video frame that have been successfully feature matched.

[0210] In one embodiment, the video set video frame matching module is further configured to use an image processing model to extract video frame features of the video frame to be identified and the video frame features of the reference video frame of the video set, respectively, wherein the training step of the image processing model includes the steps of: obtaining training sample images including classification labels; performing feature extraction and image classification on the training sample images using the image processing model to be trained, and obtaining sample image features and sample image categories of the training sample images; determining a model loss based on the sample image features, sample image categories, and classification labels; updating the image processing model to be trained based on the model loss, and then continuing training; and obtaining a trained image processing model when training is completed.

[0211] In one embodiment, the global similar segment identification module 2108 includes a global video frame matching module, a global offset determination module, and a global video frame pair processing module, wherein the global video frame matching module is configured to perform image matching of video frames between the target video and the platform reference video to obtain a video frame pair, the video frame pair including a to-be-identified video frame belonging to the target video and further including a platform reference video frame that image-matches with the to-be-identified video frame in the platform reference video, the global offset determination module is configured to determine a time offset of the video frame pair based on the temporal attributes of the to-be-identified video frame in the video frame pair and the temporal attributes of the video set reference video frame, and the global video frame pair processing module is configured to screen video frame pairs with matching time offsets and determine a platform global similar segment in the target video to the platform reference video based on the temporal attributes of the to-be-identified video frame in the video frame pair obtained by screening.

[0212] In one embodiment, the device further includes a video set identification and update module configured to determine a segment overlap relationship between each video set locally similar segment based on the start time and end time of each video set locally similar segment, perform segment updates for each video set locally similar segment based on the segment overlap relationship, and obtain updated video set locally similar segments in the target video for the video set reference video.

[0213] In one embodiment, the device further includes a public video judgment module configured to obtain a public video matching the public video type in the target video based on the overall similar segment if the overall similar segment meets a judgment condition of the public video type.

[0214] In one embodiment, the public video determination module is further configured to determine a distribution interval of public videos related to the public video type of the target video, and if the time period of the overall similar segment is within the distribution interval of the public video, obtain a public video that matches the public video type in the target video based on the overall similar segment.

[0215] In one embodiment, the device further comprises a video comparison module configured to determine a start time and an end time of the public video, extract the non-public video from the target video based on the start time and the end time in response to a video comparison trigger event, and perform a video comparison of the non-public video with the comparison target video.

[0216] In one embodiment, the device further comprises a video skip module configured to determine a skip time point in the public video, play the target video in response to a video playback event for the target video, and skip and play the public video when playback of the target video reaches the skip time point.

[0217] All or part of each module in the video identification device can be realized by software, hardware, or a combination thereof. Each module can be integrated into a processor in a computer device in the form of hardware or can be independent, or can be stored in a memory in a computer device in the form of software, so that the processor can call the module to perform the corresponding operation.

[0218] In one embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure may be as shown in FIG. 22. The computer device includes a processor, a memory, an input / output (I / O) interface, and a communication interface. The processor, memory, and I / O interface are connected by a system bus, and the communication interface is connected to the system bus by the I / O interface. The processor of the computer device is configured to provide calculation and control functions. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the execution of the operating system and the computer-readable instructions in the non-volatile storage medium. The database of the computer device is configured to store video identification data. The I / O interface of the computer device is configured to allow the processor to exchange information with an external device. The communication interface of the computer device is configured to connect to and communicate with an external terminal via a network. When executed by a processor, the computer-readable instructions cause the processor to implement a video identification method. Those skilled in the art will understand that the structure shown in Figure 22 is only a partial structural block diagram of the technical solution of the present invention, and does not limit the computer device to which the technical solution of the present invention is applied, and that a specific computer device may have more or fewer components than those shown in the drawings, or may combine some components, or may have a different component arrangement.

[0219] In one embodiment, there is further provided a computing device comprising a memory having computer-readable instructions stored thereon, and a processor that executes the computer-readable instructions to implement the steps in each of the method embodiments above.

[0220] In one embodiment, a computer-readable storage medium is provided having stored thereon computer-readable instructions that, when executed by a processor, perform the steps in each of the method embodiments described above.

[0221] In one embodiment, a computer program product is provided that includes computer readable instructions that, when executed by a processor, implement the steps in each of the method embodiments described above.

[0222] In addition, user information (including, but not limited to, user device information, user personal information, etc.) and data (including, but not limited to, data used for analysis, stored data, displayed data, etc.) related to the present invention are information and data approved by the user or fully approved by all parties, and the collection, use, and processing of related data should comply with the relevant laws, regulations, and standards of the relevant countries and regions. Furthermore, with regard to the platform promotion information related to the present invention, users can opt out or conveniently opt out of ad push information.

[0223] It will be obvious to those skilled in the art that all or part of the processes of the methods described above can be achieved by implementing relevant hardware with computer-readable instructions. The computer-readable instructions described above can be stored in a non-volatile computer-readable storage medium, and when the computer-readable instructions are executed, the processes of the method embodiments described above can be implemented. Any reference to a memory, database, or other medium used in the embodiments of the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM), external cache memory, etc. By way of illustration and not limitation, the RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database according to embodiments of the present invention may include at least one of a relational database and a non-relational database. The non-relational database may include, but is not limited to, a distributed database based on blockchain. The processor according to embodiments of the present invention may be, but is not limited to, a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc.

[0224] The technical features in the above examples can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above examples are described. However, as long as there is no contradiction in the combination of these technical features, they should all be considered within the scope of the present specification. The above examples only represent some embodiments of the present invention, and although the description is specific and detailed, it should not be understood as a limitation on the patent scope of the present invention. It should be noted that those skilled in the art can make some modifications and improvements without departing from the concept of the present invention, and all of these should be considered to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined based on the attached claims.

Claims

1. 1. A computer device implemented method for video identification, comprising: determining a target video from a video series movie set, and extracting movie set reference videos from the video series movie set to obtain the target video and the movie set reference videos belonging to the video series movie set, wherein the video series movie set includes videos belonging to the same series; identifying a video set locally similar segment in the target video to the video set reference video based on a first matching result obtained by video frame matching between the target video and the video set reference video; obtaining a platform reference video from a video platform to which the target video belongs; identifying a platform global similar segment in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video; and determining a global similar segment in the target video to the video set reference video and the platform reference video based on the positions of the video set local similar segment and the platform global similar segment in the target video, respectively. Video identification methods.

2. The video identification method includes: Further comprising: performing a correction update on the video set locally similar segments according to the correction segments in the target video including the correction keywords, to obtain updated video set locally similar segments; determining a global similar segment in the target video to the video set reference video and the platform reference video based on the positions of the video set local similar segment and the platform global similar segment in the target video, determining a global similar segment in the target video to the video set reference video and the platform reference video based on the updated positions of the video set local similar segments and the platform global similar segments in the target video, respectively; 2. The video identification method of claim 1.

3. The step of performing a correction update on the video set locally similar segments based on the correction segments including the correction keywords in the target video to obtain updated video set locally similar segments includes: determining a redaction segment in the target video that includes a redaction keyword; updating the timestamp positions of the video set locally similar segments in the target video according to the timestamp positions of the rectification segments in the target video to obtain updated timestamp positions; and determining updated video set local similar segments in the target video based on the updated timestamp positions.

3. The video identification method of claim 2.

4. The step of determining a correction segment in the target video that includes a correction keyword includes: performing text identification on video frames in the target video to obtain text identification results; Matching the text identification result with the corrected keywords to obtain a matching result; and determining a correction segment from the target video that includes the correction keyword based on a video frame associated with a matching result that the match is found.

4. The video identification method according to claim 3, wherein:

5. The platform reference video includes a platform public video segment obtained from a public video library of a video platform to which the target video belongs, and a platform related video obtained from the video platform; identifying a platform global similar segment in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video, performing video frame matching between the target video and the platform public video segment to obtain a public video matching result; If a similar segment cannot be identified based on the public video matching result, performing video frame matching between the target video and the platform-related video to obtain a related video matching result; and identifying platform-global similar segments in the target video to the platform-related video based on the related video matching results.

2. The video identification method of claim 1.

6. After identifying a platform-global similar segment in the target video to the platform-related video based on the related video matching result, the video identification method includes: updating the discrimination statistical parameters of the platform global similar segments to obtain updated discrimination statistical parameters; If the updated discrimination statistical parameters satisfy a platform public determination condition, updating the platform global similar segment as a platform public video segment in the public video library.

6. The video identification method according to claim 5, wherein:

7. The step of obtaining a platform reference video from a video platform to which the target video belongs includes: obtaining a platform-public video segment from a public video library of a video platform to which the target video belongs; identifying a platform global similar segment in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video, and identifying a platform global similar segment in the target video to the platform public video segment based on a second matching result obtained by video frame matching between the target video and the platform public video segment.

2. The video identification method of claim 1.

8. The step of obtaining a platform reference video from a video platform to which the target video belongs includes: determining a video platform to which the target video belongs and correction keywords contained in video frames in the target video; Querying the platform-related videos that have a relational relationship with the corrected keyword in the video platform; and screening the platform-related videos to obtain a platform reference video according to a reference video screening condition.

2. The video identification method of claim 1.

9. The video identification method includes: performing text identification on video frames in platform videos belonging to the video platform to obtain video keywords; matching within a keyword library based on the video keywords to determine target keywords that match the video keywords; establishing an association between the platform video and the target keyword; The step of querying the platform-related videos that have an association relationship with the correction keyword in the video platform includes: querying the video platform for platform-related videos related to the corrective keyword based on the related relationship; 9. The video identification method of claim 8.

10. determining a global similar segment in the target video to the video set reference video and the platform reference video based on the positions of the video set local similar segment and the platform global similar segment in the target video, determining a first timestamp location of the video collection local similar segment in the target video and a second timestamp location of the platform global similar segment in the target video; merging the first timestamp location and the second timestamp location to obtain an overall timestamp location; and determining a comprehensive similar segment in the target video to the video collection reference video and the platform reference video based on the comprehensive timestamp position.

2. The video identification method of claim 1.

11. The step of identifying a video set locally similar segment in the target video to the video set reference video based on a first matching result obtained by video frame matching between the target video and the video set reference video includes: performing image matching of video frames between the target video and the video set reference video to obtain a video frame pair, the video frame pair including a video frame to be identified belonging to the target video, and a video set reference video frame that image-matches the video frame to be identified in the video set reference video; determining a time offset of the video frame pair based on a temporal attribute of the identified video frame and a temporal attribute of a moving image set reference video frame in the video frame pair; screening video frame pairs having matching time offsets, and determining a video set locally similar segment in the target video to the video set reference video based on a temporal attribute of the identified video frame in the screened video frame pairs. Video identification method according to any one of claims 1 to 10, characterized in that

12. The step of screening the video frame pairs having matching time offsets and determining a video set locally similar segment in the target video to the video set reference video based on a temporal attribute of a video frame to be identified in the video frame pairs obtained by screening includes: performing numerical matching on the time offsets of each of the video frame pairs, and screening video frame pairs whose time offsets match numerically based on the numerical matching results; determining a start time and an end time based on temporal attributes of the video frame to be identified in the screened video frame pair; determining a video set local similar segment from the target video to the video set reference video based on the start time and the end time.

12. The video identification method of claim 11.

13. The step of performing numerical matching on the time offsets of each of the video frame pairs, and screening video frame pairs having numerically matching time offsets based on the numerical matching results, includes: Numerically comparing the time offsets of each of the video frame pairs respectively to obtain a numerical comparison result; based on the numerical comparison result, screening each of the video frame pairs for video frame pairs whose numerical difference in time offset is less than a numerical difference threshold; performing offset updates on video frame pairs whose numerical difference in time offset is less than a numerical difference threshold to obtain video frame pairs whose time offsets match numerically.

13. The video identification method of claim 12.

14. The video set reference video is at least two, and the step of screening video frame pairs having matching time offsets and determining a video set local similar segment in the target video to the video set reference video based on a time attribute of the identified video frame in the video frame pair obtained by screening includes: screening video frame pairs with matching time offsets, and determining intermediate similar segments in the target video to the movie set reference video based on temporal attributes of the identified video frames in the screened video frame pairs; and performing segment updating for each intermediate similar segment in the target video that has an overlapping relationship with each of the video set reference videos among the intermediate similar segments in the target video, to obtain a video set local similar segment in the target video that is associated with each of the video set reference videos.

12. The video identification method of claim 11.

15. The step of performing segment updating for each intermediate similar segment in the target video that has an overlapping relationship with each of the motion set reference videos among the intermediate similar segments in the target video to obtain a motion set local similar segment in the target video to each of the motion set reference videos includes: performing segment position comparison for intermediate similar segments in the target video with respect to each of the movie set reference videos to obtain a segment comparison result; determining each intermediate similar segment having an overlapping relationship as a result of the segment comparison; and performing segment updating for each intermediate similar segment having an overlapping relationship based on the overlapping time length and statistics of each intermediate similar segment having an overlapping relationship, thereby obtaining a video set local similar segment in the target video for each of the video set reference videos.

15. The video identification method of claim 14.

16. The step of performing image matching of video frames between the target video and the video set reference video to obtain video frame pairs includes: extracting a video frame to be identified from the target video and extracting a video set reference video frame from the video set reference video; extracting video frame features of the identification target video frame and video frame features of the moving image set reference video frame respectively; and performing feature matching of the video frame features of the identification target video frame with the video frame features of the video set reference video frame, and obtaining a video frame pair based on the identification target video frame and the video set reference video frame that have been successfully feature matched.

12. The video identification method of claim 11.

17. identifying a platform global similar segment in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video, performing image matching of video frames between the target video and the platform-reference video to obtain a video frame pair, the video frame pair including a video frame to be identified belonging to the target video, and a platform-reference video frame that image-matches the video frame to be identified in the platform-reference video; determining a time offset of the video frame pair based on a temporal attribute of the identified video frame and a temporal attribute of a moving image set reference video frame in the video frame pair; screening video frame pairs having matching time offsets, and determining a platform-global similar segment in the target video to the platform reference video based on a temporal attribute of the identified video frame in the screened video frame pairs.

2. The video identification method of claim 1.

18. A video identification device, a video set video acquisition module configured to determine a target video from a video series video set, and extract a video set reference video from the video series video set to acquire the target video and the video set reference video belonging to the video series video set, wherein the video series video set includes videos belonging to the same series; a local similar segment identification module configured to identify a video set locally similar segment in the target video to the video set reference video based on a first matching result obtained by video frame matching between the target video and the video set reference video; a platform video acquisition module configured to acquire a platform reference video from a video platform to which the target video belongs; a global similar segment identification module configured to identify a platform global similar segment in the target video to the platform reference video based on a second matching result obtained by video frame matching between the target video and the platform reference video; an overall similar segment determination module configured to determine overall similar segments in the target video to the video set reference video and the platform reference video based on the positions of the video set local similar segments and the platform global similar segments in the target video; Equipped with Video identification device.

19. A computing device comprising: a memory in which computer-readable instructions are stored; and a processor which, when said computer-readable instructions are executed, performs the steps of the method of any one of claims 1 to 10 and 17.

20. A computer program product causing a processor to carry out the method according to any one of claims 1 to 10 and 17.

Citation Information

Patent Citations

  • Video clip identification method and device, equipment and storage medium

    CN114550070A

  • Moving picture processor, moving picture processing method and moving picture processing program

    JP2005130416A

  • Method, system, and program for topic guidance in video content using sequence pattern mining

    JP2019024192A

  • Video scene classification device and video scene classification method

    US20090257649A1