Method, device, electronic device and storage medium for determining risk level of video

Through the screening of semantic information and content labels of videos and combined with feature extraction methods, the problem of inefficient determination of video infringement risk level in the prior art is solved, and efficient and accurate judgment of infringement risk is achieved.

CN120277238BActive Publication Date: 2025-08-12SHANGHAI CANGUANG VIDEO TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510781440.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-12
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing video infringement risk level determination scheme is inefficient and insufficiently accurate, and cannot effectively deal with infringement of massive videos.

Method used

By obtaining the semantic information and content labels of the video, combining the feature extraction method, the video is divided into multiple levels, and filtering it based on the semantics and content similarity to determine the risk level of the video.

Benefits of technology

It improves the efficiency and accuracy of determining the risk level of video infringement, saves computing resources, and ensures the accuracy of the infringement judgment process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277238B_ABST
    Figure CN120277238B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, electronic device, and storage medium for determining the risk level of a video. The method includes: obtaining a reference video and a set of videos to be compared; obtaining first reference semantic information and first comparative semantic information; determining multiple first videos and multiple second videos in the set of videos to be compared based on the reference semantic information and the comparative semantic information; obtaining a reference content tag and a comparative content tag; determining multiple third videos in the multiple second videos based on the reference content tag and the comparative content tag; obtaining a reference identification feature, a third video, and a comparative identification feature; and determining the risk level information of each video to be compared based on the first reference semantic information, the first comparative semantic information, the reference content tag, the comparative content tag, the reference identification feature, and the comparative identification feature. This application achieves improved efficiency in determining the risk level of a video while ensuring accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a method, device, electronic device, and storage medium for determining the risk level of a video. Background Art

[0002] With the rapid development of digital media technology, short video platforms, film and television production, and other fields generate massive amounts of video content daily. However, this is accompanied by an increasingly severe problem of video copyright infringement. Furthermore, existing video processing technologies offer diverse means of concealing copyright infringement, such as intelligent speed change, facial reenactment, and audio separation and reconstruction. These technological breakthroughs not only significantly enhance the concealment of copyright infringement but also pose unprecedented challenges to traditional video duplication detection mechanisms.

[0003] Current mainstream approaches for determining video infringement risk levels primarily employ two technical approaches: frame-by-frame comparison based on frame-level features and hash matching based on keyframe extraction. The former requires pixel-by-pixel comparison of the target video and the video to be inspected, resulting in high complexity, low efficiency, and inability to handle massive amounts of video. While the latter reduces complexity by extracting feature points, it still suffers from significant accuracy issues.

[0004] Therefore, there is an urgent need for a video infringement risk level determination solution that is both efficient and accurate. Summary of the Invention

[0005] In order to solve the technical problem of low efficiency and accuracy of existing video infringement risk level determination solutions, the present invention provides a video risk level determination method, device, electronic device and storage medium.

[0006] In a first aspect, an embodiment of the present application provides a method for determining a risk level of a video, the method comprising:

[0007] Obtaining a reference video and a set of videos to be compared; the set of videos to be compared includes multiple videos to be compared;

[0008] Acquire first reference semantic information of the reference video and first comparison semantic information of each of the videos to be compared;

[0009] determining, in the set of videos to be compared, a plurality of first-level videos and a plurality of second-level videos based on the first reference semantic information and each of the comparison semantic information;

[0010] Obtaining a reference content tag of the reference video and a comparison content tag of each second-level video;

[0011] determining a plurality of third-level videos from the plurality of second-level videos based on the reference content tag and each of the compared content tags;

[0012] Obtaining a reference identification feature of the reference video and a comparison identification feature of each of the third-level videos and each of the first-level videos;

[0013] Risk level information of each of the videos to be compared is determined based on the first reference semantic information, each of the first comparison semantic information, the reference content tag, each of the comparison content tags, the reference identification feature, and each of the comparison identification features.

[0014] In an optional embodiment, the set of videos to be compared includes a plurality of fourth-level videos, which are videos to be compared in the set of videos to be compared excluding the plurality of first-level videos and the plurality of second-level videos; the plurality of second-level videos includes a plurality of fifth-level videos, which are videos to be compared in the plurality of second-level videos excluding the plurality of third-level videos;

[0015] The determining, based on the first reference semantic information, each of the first comparison semantic information, the reference content tag, each of the comparison content tags, the reference identification feature, and each of the comparison identification features, of the risk level information of each of the to-be-compared videos includes:

[0016] determining risk level information of the plurality of fourth-level videos based on the first reference semantic information and the first comparative semantic information of each of the fourth-level videos;

[0017] determining risk level information of the plurality of fifth-level videos based on the reference content tag and the comparison content tag of each of the fifth-level videos;

[0018] Based on the reference identification feature and each of the compared identification features, risk level information of the plurality of first-level videos and the plurality of third-level videos is determined.

[0019] In an optional embodiment, obtaining the first reference semantic information of the reference video and the first comparison semantic information of each of the videos to be compared includes:

[0020] Acquiring text information and / or audio information of the reference video and text information and / or audio information of each of the videos to be compared;

[0021] Performing semantic recognition processing on the text information and / or sound information of the reference video to obtain first reference semantic information of the reference video;

[0022] Perform semantic recognition processing on the text information and / or sound information of each of the videos to be compared to obtain first comparison semantic information of each of the videos to be compared.

[0023] In an optional embodiment, determining a plurality of first-level videos and a plurality of second-level videos in the set of videos to be compared based on the first reference semantic information and each piece of the first comparison semantic information includes:

[0024] For each video to be compared, execute:

[0025] The currently executed video to be compared is regarded as the current comparison video;

[0026] Determining a first similarity between the current comparison video and the reference video based on the first reference semantic information and the first comparison semantic information of the current comparison video;

[0027] If the first similarity is lower than or equal to a first preset threshold and higher than or equal to a second preset threshold, determining that the current comparison video is a first-level video; or if the first similarity is lower than the first preset threshold, determining that the current comparison video is a second-level video;

[0028] The plurality of first-ranked videos are determined based on each of the first-ranked videos, and the plurality of second-ranked videos are determined based on each of the second-ranked videos.

[0029] In an optional embodiment, determining the risk level information of the plurality of fourth-level videos based on the first reference semantic information and the first comparative semantic information of each fourth-level video includes:

[0030] If the first similarity is higher than the first preset threshold, determining a high risk level as the risk level information of the current comparison video;

[0031] The risk level information of the plurality of fourth-level videos is determined based on the risk level information of each of the current comparison videos.

[0032] In an optional embodiment, obtaining the reference content tag of the reference video and the comparison content tag of each second-level video includes:

[0033] Determining a target correction video from the plurality of fourth-level videos based on a first similarity between each of the fourth-level videos and the reference video; wherein the first similarity corresponding to the target correction video is greater than or equal to the first similarities corresponding to other videos in the plurality of fourth-level videos;

[0034] Obtaining a corrected content label of the target corrected video and a candidate content label of the reference video;

[0035] determining the reference content label based on the revised content label and the candidate content label;

[0036] Obtain a comparative content label for each second-level video.

[0037] In an optional embodiment, obtaining the reference content tag of the reference video and the comparison content tag of each second-level video includes:

[0038] performing a preprocessing operation on the reference video and each of the second-level videos to obtain a preprocessed reference video and a plurality of preprocessed second-level videos; the preprocessing operation comprising at least one of noise reduction, watermark removal, resolution reduction, and frame reduction;

[0039] Based on the preprocessed reference video and the plurality of preprocessed second-level videos, a reference content label of the reference video and a comparative content label of each of the second-level videos are obtained.

[0040] In an optional embodiment, determining a plurality of third-level videos from the plurality of second-level videos based on the reference content tag and each of the comparison content tags includes:

[0041] For each of the second-level videos, perform:

[0042] The second-level video currently being executed is regarded as the current comparison video;

[0043] Determining a second similarity between the current comparison video and the reference video based on the reference content tag and the comparison content tag of the current comparison video;

[0044] If the second similarity is higher than or equal to a third preset threshold, determining that the current comparison video is the third-level video;

[0045] The plurality of third-level videos are determined based on each of the third-level videos.

[0046] In an optional embodiment, determining the risk level information of the plurality of fifth-level videos based on the reference content tag and the comparison content tag of each fifth-level video includes:

[0047] If the second similarity is lower than the third preset threshold, determining a low risk level as the risk level information of the current comparison video;

[0048] The risk level information of the plurality of fifth-level videos is determined based on the risk level information of each of the current comparison videos.

[0049] In an optional embodiment, before obtaining the reference identification feature of the reference video and the comparison identification feature of each of the third-level videos and each of the first-level videos, the method further includes:

[0050] Determining a reference key frame from a plurality of reference frames of the reference video; the reference key frame is used to obtain a reference identification feature of the reference video;

[0051] Performing semantic recognition processing on the reference key frame to obtain second reference semantic information of the reference key frame;

[0052] performing semantic recognition processing on a plurality of video frames of the third-level video to obtain third comparative semantic information corresponding to the plurality of video frames;

[0053] If the second reference semantic information of the reference key frame matches the third comparative semantic information of the target video frame among the multiple video frames, the target video frame is determined to be a key frame of the third-level video; the target video frame is used to obtain comparative identification features of the third-level video.

[0054] In an optional embodiment, obtaining a reference identification feature of the reference video and a comparison identification feature of each of the third-level videos and each of the first-level videos includes:

[0055] Extracting a plurality of global picture features and a plurality of local invariant features of the reference key frame;

[0056] Screening multiple global image features and multiple local invariant features of the reference key frame, and fusing them based on an attention mechanism to obtain the reference identification feature;

[0057] Extracting a plurality of global picture features and a plurality of local invariant features of each target video frame;

[0058] Multiple global picture features and multiple local invariant features of each target video frame are screened, and fused based on the attention mechanism to obtain comparative identification features of each first-level video.

[0059] In an optional embodiment, determining the risk level information of the plurality of first-level videos and the plurality of third-level videos based on the reference identification feature and each of the comparison identification features includes:

[0060] For each of the first-level video and the third-level video, perform:

[0061] The first-level video and the third-level video currently being executed are regarded as current comparison videos;

[0062] Determining a third similarity between the current comparison video and the reference video based on the reference identification feature and the comparison identification feature of the current comparison video;

[0063] If the third similarity is higher than or equal to a fourth preset threshold, a high risk level is determined as the risk level information of the current comparison video; or if the third similarity is lower than the fourth preset threshold, a low risk level is determined as the risk level information of the current comparison video;

[0064] The risk level information of the plurality of first-level videos and the plurality of third-level videos is determined based on the risk level information of each of the current comparison videos.

[0065] In a second aspect, an embodiment of the present application provides a device for determining a risk level of a video, the device comprising:

[0066] A first acquisition module is configured to acquire a reference video and a set of videos to be compared; the set of videos to be compared includes a plurality of videos to be compared;

[0067] A second acquisition module is used to acquire the first reference semantic information of the reference video and the first comparison semantic information of each of the videos to be compared;

[0068] A first determining module, configured to determine, in the set of videos to be compared, a plurality of first-level videos and a plurality of second-level videos based on the first reference semantic information and each of the comparison semantic information;

[0069] a third acquisition module, configured to acquire a reference content tag of the reference video and a comparison content tag of each second-level video;

[0070] a second determining module, configured to determine a plurality of third-level videos from the plurality of second-level videos based on the reference content tag and each of the compared content tags;

[0071] a fourth acquisition module, configured to acquire a reference identification feature of the reference video and a comparison identification feature of each of the third-level videos and each of the first-level videos;

[0072] The third module is used to determine the risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content tag, each of the comparison content tags, the reference identification feature and each of the comparison identification features.

[0073] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the risk level determination method of the video of the first aspect.

[0074] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which at least one instruction or at least one program is stored, and the at least one instruction or at least one program is loaded and executed by a processor to implement the risk level determination method of a video of the first aspect.

[0075] In a fifth aspect, embodiments of the present application provide a computer program product or computer program, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method for determining a video risk level according to the first aspect.

[0076] The method, device, electronic device, and storage medium for determining the risk level of a video provided by the embodiments of the present application have the following technical effects:

[0077] Obtain a reference video and a set of videos to be compared; the set of videos to be compared includes multiple videos to be compared; obtain first reference semantic information of the reference video and first comparative semantic information of each of the videos to be compared; based on the first reference semantic information and each comparative semantic information, determine multiple first-level videos and multiple second-level videos in the set of videos to be compared; obtain a reference content label of the reference video and a comparative content label of each of the second-level videos; based on the reference content label and each of the comparative content labels, determine multiple third-level videos in the multiple second-level videos; obtain a reference identification feature of the reference video and a comparative identification feature of each of the third-level videos and each of the first-level videos; determine the risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparative semantic information, the reference content label, each of the comparative content labels, the reference identification feature and each of the comparative identification features. This application performs preliminary screening through semantic information and then performs secondary screening through content tags, dividing multiple videos to be compared into different levels, screening out a small number of videos for computationally complex feature identification, and determining the risk level information of these videos for subsequent infringement judgments, saving a large amount of computing resources, ensuring accuracy while improving the efficiency of video risk level determination. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0079] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;

[0080] Figure 2 This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 1 ;

[0081] Figure 3 This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 2 ;

[0082] Figure 4 This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 3 ;

[0083] Figure 5 This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 4 ;

[0084] Figure 6 This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 5 ;

[0085] Figure 7 This is a schematic structural diagram of a device for determining the risk level of a video provided in an embodiment of the present application;

[0086] Figure 8 This is a hardware structure block diagram of a server for a method for determining a risk level of a video provided in an embodiment of the present application. DETAILED DESCRIPTION

[0087] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0088] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0089] Figure 1 This is a schematic diagram of an application environment provided by an embodiment of the present application. Figure 1 As shown, the application environment may include a server 01 and a client 02 .

[0090] In some possible embodiments, the client 02 may include, but is not limited to, smartphones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, and the like. It may also be software running on the client, such as an application or applet. Optionally, the operating system running on the client may include, but is not limited to, Android, iOS, Linux, Windows, Unix, and the like.

[0091] In some possible embodiments, server 01 obtains a reference video and a set of videos to be compared; the set of videos to be compared includes multiple videos to be compared; obtains first reference semantic information of the reference video and first comparative semantic information of each of the videos to be compared; based on the first reference semantic information and each comparative semantic information, determines multiple first-level videos and multiple second-level videos in the set of videos to be compared; obtains a reference content tag of the reference video and a comparative content tag of each of the second-level videos; based on the reference content tag and each of the comparative content tags, determines multiple third-level videos in the multiple second-level videos; obtains a reference identification feature of the reference video and a comparative identification feature of each of the third-level videos and each of the first-level videos; determines the risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparative semantic information, the reference content tag, each of the comparative content tags, the reference identification feature and each of the comparative identification features.

[0092] Optionally, server 01 may include an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The operating system running on the server may include, but is not limited to, Android, iOS, Linux, Windows, Unix, etc.

[0093] The following describes a specific embodiment of a method for determining the risk level of a video in this application. Figure 2 This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 1 , this specification provides method operation steps such as embodiments or flow charts, but may include more or fewer operation steps based on routine or non-creative work. The order of steps listed in the embodiments is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or server product is executed, it can be executed in the order shown in the embodiments or the drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment). Specifically, Figure 2 As shown, this may include:

[0094] S201: Acquire a reference video and a set of videos to be compared; the set of videos to be compared includes multiple videos to be compared.

[0095] S202: Acquire first reference semantic information of the reference video and first comparison semantic information of each of the videos to be compared.

[0096] S203: Determine a plurality of first-level videos and a plurality of second-level videos in the set of videos to be compared based on the first reference semantic information and each of the comparison semantic information.

[0097] S204: Obtain a reference content tag of the reference video and a comparison content tag of each second-level video.

[0098] S205: Determine a plurality of third-level videos from the plurality of second-level videos based on the reference content tag and each of the compared content tags.

[0099] S206: Obtain a reference identification feature of the reference video and a comparison identification feature of each of the third-level videos and each of the first-level videos.

[0100] S207: Determine risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content tag, each of the comparison content tags, the reference identification feature, and each of the comparison identification features.

[0101] Figure 3 This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 2 , the method may include:

[0102] S301: Obtain a reference video and a set of videos to be compared.

[0103] In a possible embodiment, the reference video is the original video, and the method for determining the risk level of the video of the present application is to determine the degree of similarity of each video in the set of videos to be compared with the reference video, thereby determining the infringement risk level of each video.

[0104] In a possible embodiment, the set of videos to be compared includes multiple videos to be compared, and the number may be tens of thousands.

[0105] S302: Acquire first reference semantic information of the reference video and first comparison semantic information of each of the videos to be compared.

[0106] In a possible embodiment, the step of obtaining the first reference semantic information of the reference video and the first comparison semantic information of each of the videos to be compared includes:

[0107] S312: Acquire text information and / or sound information of the reference video and text information and / or sound information of each of the videos to be compared.

[0108] S322: Perform semantic recognition processing on the text information and / or sound information of the reference video to obtain first reference semantic information of the reference video.

[0109] S332: Perform semantic recognition processing on the text information and / or sound information of each of the videos to be compared to obtain first comparison semantic information of each of the videos to be compared.

[0110] In the embodiment of the present application, the first reference semantic information refers to the semantic information of the reference video, which is used to represent the direct text information in the reference video and the indirect text information converted from sound.

[0111] Similarly, the first comparison semantic information refers to the semantic information of the video to be compared, and is used to represent the direct text information and the indirect text information converted from sound in the video to be compared.

[0112] In a possible embodiment, the reference video and the video to be compared may include two or one of text information and sound information. Videos without text information or sound information are not considered within the scope of this application.

[0113] In a possible embodiment, the text information may be subtitles on the video, text on the screen, etc.; the sound information may be the voice of the characters in the video, the sound in the background, narration, etc.

[0114] Optionally, the reference video contains only text information, and the video to be compared contains only sound information; the reference video contains only text information, and the video to be compared contains only text information; the reference video contains only sound information, and the video to be compared contains only sound information; the reference video contains only sound information, and the video to be compared contains only text information; the reference video contains both sound and text information, and the video to be compared contains both sound and text information.

[0115] When obtaining text information from a video, text extraction technologies such as optical character recognition technology can be used to extract various types of text in the video, and semantic recognition models can be used to perform semantic recognition on various types of text. When obtaining sound information from a video, speech transcription, sound conversion and other technologies can be used to first convert the speech into text, and then semantic recognition models can be used to perform semantic recognition on various types of text.

[0116] In a possible embodiment, when a video contains both voice information and text information, the voice information and text information can be merged and then the merged information can be semantically recognized using a semantic recognition model. Alternatively, the voice information and text information can be semantically recognized separately using a semantic recognition model and then the two semantic information can be merged.

[0117] S303: Determine a plurality of first-level videos and a plurality of second-level videos in the set of videos to be compared based on the first reference semantic information and each of the comparison semantic information.

[0118] Figure 4 This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 3 ,like Figure 4 As shown, in a possible embodiment, based on the first reference semantic information and each of the first comparison semantic information, determining a plurality of first-level videos and a plurality of second-level videos in the set of videos to be compared includes:

[0119] For each video to be compared, execute:

[0120] S313: The currently executed video to be compared is regarded as the current comparison video.

[0121] S323: Determine a first similarity between the current comparison video and the reference video based on the first reference semantic information and the first comparison semantic information of the current comparison video.

[0122] In a possible embodiment, determining the first similarity between the current comparison video and the reference video is calculating a first similarity between the first reference semantic information and the first comparison semantic information.

[0123] Since the first reference semantic information and the first comparative semantic information extracted by the semantic recognition model are both feature vectors, calculating the first similarity between the first reference semantic information and the first comparative semantic information can calculate the cosine similarity, Euclidean distance, Manhattan distance, etc. between the two feature vectors.

[0124] S333: Determine whether the first similarity is lower than or equal to a first preset threshold. If so, execute S343; if not, execute S373.

[0125] S343: Determine whether the first similarity is higher than or equal to a second preset threshold. If so, execute S353; if not, execute S363.

[0126] S353: Determine that the current comparison video is the first-level video.

[0127] In the embodiment of the present application, the second preset threshold is smaller than the first preset threshold.

[0128] In an embodiment of the present application, if the first similarity corresponding to the current comparison video is lower than or equal to the first preset threshold and higher than or equal to the second preset threshold, the current comparison video is determined to be the first-level video, and the first-level video is a video with medium semantic similarity to the reference video.

[0129] S363: Determine that the current comparison video is the second-level video.

[0130] In an embodiment of the present application, if the first similarity corresponding to the current comparison video is lower than or equal to the first preset threshold, the current comparison video is determined to be a second-level video, and the second-level video is a video with low semantic similarity to the reference video.

[0131] S373: Determine that the current comparison video is the fourth-level video, and determine that a high risk level is the risk level information of the current comparison video.

[0132] S383: Determine the multiple first-level videos based on each of the first-level videos, and determine the multiple second-level videos based on each of the second-level videos.

[0133] In a possible embodiment, the fourth-level video is a video to be compared in the set of videos to be compared except for the plurality of first-level videos and the plurality of second-level videos.

[0134] In an embodiment of the present application, if the first similarity corresponding to the current comparison video is higher than the first preset threshold, it is determined that the current comparison video is a fourth-level video, and the fourth-level video has a high semantic similarity to the reference video. At this time, it can be considered that the sound, text and other features of the fourth-level video are highly similar to the reference video, and the infringement risk is relatively high.

[0135] That is to say, this application first divides a large number of videos to be compared into three categories through the steps of semantic recognition and semantic comparison, namely first-level videos, second-level videos and fourth-level videos. Among them, the semantic similarity of first-level videos is medium and needs further classification; the semantic similarity of second-level videos is low and the infringement risk is low, but in order to ensure accuracy, further classification is also required later; the semantic similarity of fourth-level videos is high and the infringement risk is high.

[0136] Through the above settings, multiple fourth-level videos with high similarity are quickly screened out through text and sound extraction and semantic recognition with fast speed, mature technology and low computational complexity, and high infringement risks are determined. Based on the semantic feature information, the videos to be compared are divided into two different sets of first-level videos and second-level videos, which facilitates subsequent different recognition processing and is more targeted, so as to save computing resources.

[0137] After selecting multiple second-level videos from the videos to be compared, the following steps are performed:

[0138] S304: Obtain a reference content tag of the reference video and a comparison content tag of each second-level video.

[0139] In the embodiment of the present application, the reference content tag refers to the content tag of the reference video, which is used to represent the picture content information of the reference video. The comparison content tag refers to the content tag of the second-level video, which is used to represent the picture content information of the second-level video.

[0140] In this embodiment of the present application, a preset object recognition model is used to identify reference object description information of a reference video, such as sky, house, boy, girl, umbrella, and beach. Furthermore, more specific reference object description information can be obtained, such as blue sky, numerous houses, a crying boy, a laughing girl, a parasol, and a sunny beach. Subsequently, a reference content label for the reference video is determined based on the reference object description information.

[0141] In an embodiment of the present application, the comparison object description information of the second-level video is identified by a preset object recognition model, and the comparison content label of the second-level video is determined based on the comparison object description information.

[0142] In the embodiment of the present application, the reference content tag may include multiple tag information, such as attribute tag information, quantity tag information and scene tag information. The object tag information listed above is only exemplary, and other possible object tag information may be included in the embodiment of the present application.

[0143] In a possible embodiment, the reference content tags of a video to be compared may be relatively simple. A high-risk, i.e., highly similar, fourth-level video may be used to modify and supplement the content tags directly extracted from the reference video. The specific steps include:

[0144] S314: Determine a target correction video from the plurality of fourth-level videos based on the first similarity between each of the fourth-level videos and the reference video.

[0145] In a possible embodiment, the first similarity corresponding to the target revised video is greater than or equal to the first similarities corresponding to other videos in the plurality of fourth-level videos, that is, the target revised video has the highest similarity to the reference video.

[0146] S324: Obtain the corrected content label of the target corrected video and the candidate content label of the reference video.

[0147] S334: Determine the reference content label based on the revised content label and the candidate content label.

[0148] S344: Obtain a comparative content label for each second-level video.

[0149] In a possible embodiment, the revised content tag and the candidate content tag may be merged to directly obtain the reference content tag; or the reference content tag may be obtained by filtering out irrelevant tags after merging the content tags.

[0150] In a possible embodiment, before content label extraction, the reference video and the second-level video may be pre-processed to further reduce the computational complexity of object recognition processing and content label extraction.

[0151] Specifically, the reference video and each second-level video can be preprocessed to obtain a preprocessed reference video and multiple preprocessed second-level videos; based on the preprocessed reference video and the multiple preprocessed second-level videos, the reference content label of the reference video and the comparison content label of each second-level video can be obtained.

[0152] In a possible embodiment, the pre-processing operation includes at least one of noise reduction processing, watermark removal processing, resolution reduction processing, and frame reduction processing.

[0153] Specifically, noise reduction processing refers to removing noise such as picture noise in the video. The video image after noise reduction is clearer, which helps to more accurately identify objects and extract content labels.

[0154] Watermark removal refers to removing watermarks such as logos and text from videos to restore the original content of the video and prevent graphic and text watermarks from affecting object recognition and content label extraction.

[0155] Downscaling involves reducing the video resolution (for example, from 1080p to 720p) to reduce file size and computational complexity. This process reduces the number of pixels in the video, significantly speeding up object recognition and reducing resource consumption.

[0156] Frame rate reduction refers to lowering the video frame rate (for example, from 30fps to 15fps) to reduce the number of frames and computational complexity. There is often a lot of redundant information between adjacent frames in a video. Frame rate reduction can reduce redundant frames while preserving the content of key frames.

[0157] Through the above preprocessing operations, the effect and efficiency of content tagging processing can be significantly optimized.

[0158] S305: Determine a plurality of third-level videos from the plurality of second-level videos based on the reference content tag and each of the compared content tags.

[0159] Figure 5 This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 4 ,like Figure 5 As shown, in a possible embodiment, determining a plurality of third-level videos from the plurality of second-level videos based on the reference content tag and each of the comparison content tags includes:

[0160] For each of the second-level videos, perform:

[0161] S315: The second-level video currently being executed is regarded as the current comparison video.

[0162] S325: Determine a second similarity between the current comparison video and the reference video based on the reference content tag and the comparison content tag of the current comparison video.

[0163] S335: Determine whether the second similarity is higher than or equal to a third preset threshold. If so, execute S345; if not, execute S355.

[0164] S345: Determine that the current comparison video is the third-level video.

[0165] In the embodiment of the present application, if the second similarity between the reference content tag and the comparison content tag is higher than or equal to a third preset threshold, the current comparison video is determined to be the third-level video.

[0166] S355: Determine that the current comparison video is the fifth-level video, and determine a low risk level as the risk level information of the current comparison video.

[0167] In the embodiment of the present application, if the second similarity between the reference content tag and the comparison content tag is lower than a third preset threshold, the current comparison video is determined to be the fifth-level video.

[0168] In a possible embodiment, the plurality of fifth-level videos are videos to be compared among the plurality of second-level videos excluding the plurality of third-level videos.

[0169] That is to say, this application further classifies second-level videos with low semantic similarity through content identification and content label extraction, and divides second-level videos into third-level videos and fifth-level videos based on content similarity. Among them, third-level videos have high content similarity and high infringement risk, and need further classification; fifth-level videos have low content similarity and low infringement risk.

[0170] Through the above settings, multiple second-level videos with medium semantic similarity that were initially screened out through semantic information are further screened through content tags to screen out third-level videos with low semantic similarity but high content tag similarity. Subsequent feature extraction operations are performed to screen out fifth-level videos with low semantic similarity but low content tag similarity, and determine that the infringement risk is low.

[0171] By comparing the similarity of content tags, which is computationally intensive, resource-intensive, and time-consuming, we can further divide the videos to be compared into different categories, ensuring that each video is matched with an appropriate similarity calculation method, which is highly targeted and accurate.

[0172] After selecting multiple third-level videos from the second-level videos, proceed as follows:

[0173] S306: Determine a reference key frame from a plurality of reference frames of the reference video.

[0174] In an embodiment of the present application, the reference key frame is at least one frame that best represents the content features of the reference video and is used to obtain the reference identification features of the reference video.

[0175] Before performing the feature extraction operation, a key frame determination operation is required. By selecting only a few key video frames, on the one hand, the amount of calculation is further reduced, and on the other hand, the interference of irrelevant information is reduced, thereby ensuring the effectiveness of the extracted feature information.

[0176] S307: Perform semantic recognition processing on the reference key frame to obtain second reference semantic information of the reference key frame.

[0177] S308: Perform semantic recognition processing on the multiple video frames of the third-level video to obtain third comparative semantic information corresponding to the multiple video frames.

[0178] S309: If the second reference semantic information of the reference key frame matches the third comparison semantic information of the target video frame among the multiple video frames, determine that the target video frame is a key frame of the third-level video.

[0179] In this embodiment of the present application, the second reference semantic information refers to the semantic information of the reference key frame, which is used to represent the text information of the reference key frame, that is, the most representative text information in the reference video. The third comparative semantic information refers to the semantic information of the target video frame of the third-level video, which is used to represent the text information of the target video frame, that is, the most representative text information in the third-level video.

[0180] In the embodiment of the present application, the target video frame is used to obtain the comparative identification features of the third-level video.

[0181] In a possible embodiment, semantic recognition of video frames is similar to semantic recognition of the entire video. The method of determining whether the second reference semantic information matches the third comparative semantic information can also be adopted by calculating the similarity. If the fourth similarity between the second reference semantic information and the third comparative semantic information is greater than the fifth preset threshold, it is considered that the second reference semantic information of the reference key frame matches the third comparative semantic information of the target video frame.

[0182] S310: Acquire a reference identification feature of the reference video and a comparison identification feature of each of the third-level videos and each of the first-level videos.

[0183] In the embodiment of the present application, the reference identification feature refers to the identification feature of the reference video, which is used to characterize the picture features of the reference key frame of the reference video, and is the feature that is retained after various deformation, cropping, acceleration and other operations; the comparison identification feature refers to the identification feature of the first-level video and the third-level video, which is used to characterize the picture features of the target video frames of the first-level video and the third-level video.

[0184] S3101: Extracting multiple global picture features and multiple local invariant features of the reference key frame.

[0185] In one possible embodiment, global image features may include statistical features of the entire image, such as color distribution, texture features, and overall features. Therefore, when extracting global image features from a reference keyframe, attention is generally paid to the statistical information and overall visual characteristics of the entire image. Global image features may include color distribution, texture features, and overall structure.

[0186] For example, RGB histogram statistics and color moments are used to capture the color distribution characteristics of key frames, reflecting the overall color distribution in the image. Multi-scale Gabor filtering (such as the improved MSAF algorithm) is used to calculate the filter response in eight directions to capture texture information at different scales and directions in the image. Furthermore, principal component analysis (PCA) is used to reduce the dimensionality of high-dimensional features, or pre-trained convolutional neural networks (CNNs) (such as ResNet and EfficientNet) are used to extract high-level semantic features to comprehensively describe the global visual information of the image. These methods can effectively extract global image features of the reference key frames, providing a foundation for subsequent video comparison and analysis.

[0187] In one possible embodiment, when extracting local invariant features from a reference keyframe, attention is typically focused on detailed information in local regions within the image. These features remain unchanged under image transformations (such as rotation, scaling, and brightness changes). For example, SIFT (Scale-Invariant Feature Transform) or ORB (Oriented FAST and Rotated BRIEF) algorithms are used to extract keypoints and their descriptors. These algorithms are robust to rotation, scaling, and brightness changes. Local texture features are extracted using Local Binary Patterns (LBP) or Gabor filtering to describe the texture patterns around the keypoints. Harris corner detection or edge detection algorithms (such as Canny) are used to extract corner and edge information within the image to describe local structural information. These local invariant features can stably reflect the detailed content of the keyframes, providing robust support for video comparison and matching.

[0188] S3102: Filter multiple global image features and multiple local invariant features of the reference key frame, and fuse them based on the attention mechanism to obtain the reference identification feature.

[0189] In a possible embodiment, when screening the global picture features and local invariant features of the reference key frame, the global features (such as color distribution, texture features, CNN high-level semantic features) and local features (such as SIFT, ORB key point descriptors) are first screened through a feature importance evaluation method (such as feature importance scoring based on random forest or XGBoost) to retain features that are highly discriminative and contribute highly to the description of the video content.

[0190] The filtered features are then fused using an attention mechanism. This involves calculating the weight of each feature through an attention network (such as a Transformer or self-attention module), enabling the model to dynamically focus on features that are more important to the task at hand. Ultimately, the weighted global and local features are concatenated or summed to generate robust and discriminative reference features. This fusion approach not only preserves key information but also effectively improves the expressiveness of features.

[0191] S3103: Extracting multiple global picture features and multiple local invariant features of each target video frame.

[0192] S3104: Filter multiple global screen features and multiple local invariant features of each target video frame, and fuse them based on the attention mechanism to obtain comparative identification features of each first-level video.

[0193] In the embodiment of the present application, the feature extraction method for the target video frame is similar to the feature extraction method for the reference key frame.

[0194] S311: Determine risk level information of the plurality of first-level videos and the plurality of third-level videos based on the reference identification feature and each of the comparison identification features.

[0195] Since the semantic similarity of the first-level videos is medium, that is, they have a certain degree of similarity, even if the content similarity is judged by content label extraction and comparison (that is, the second-level screening), it is impossible to accurately judge the similarity between the first-level videos and the reference videos.

[0196] Similarly, since the semantic similarity of the third-level videos is low but the content tag similarity is high, such contradictory results cannot accurately determine the similarity between the third-level videos and the reference videos.

[0197] Therefore, it is necessary to perform precise feature extraction and feature comparison on the first-level video and the third-level video to obtain accurate results.

[0198] Figure 6This is a flow chart of a method for determining the risk level of a video provided in an embodiment of the present application. Figure 5 ,like Figure 6 As shown, in a possible embodiment, determining the risk level information of the plurality of first-level videos and the plurality of third-level videos based on the reference identification feature and each of the comparison identification features includes the following steps:

[0199] For each of the first-level video and the third-level video, perform:

[0200] S3111: The first-level video and the third-level video currently being executed are regarded as current comparison videos.

[0201] S3112: Determine a third similarity between the current comparison video and the reference video based on the reference identification feature and the comparison identification feature of the current comparison video.

[0202] S3113: Determine whether the third similarity is higher than or equal to a fourth preset threshold. If so, execute S3114; if not, execute S3115.

[0203] S3114: Determine the high risk level as the risk level information of the current comparison video.

[0204] In an embodiment of the present application, if the third similarity between the reference identification feature and the comparison identification feature is higher than a fourth preset threshold, the risk level information of the current comparison video is determined to be a high risk level, that is, the video is highly similar to the reference video.

[0205] S3115: Determine the low risk level as the risk level information of the current comparison video.

[0206] In an embodiment of the present application, if the third similarity between the reference identification feature and the comparison identification feature is lower than or equal to the fourth preset threshold, the risk level information of the current comparison video is determined to be a low risk level, that is, the video has low similarity with the reference video.

[0207] By extracting the global image features and local invariant features of the key frames and fusing them together, we obtain unique and robust identification features. Based on the identification features, we determine the infringement risk levels of the first-level and third-level videos that were difficult to judge in the previous steps. This method is highly targeted and accurate.

[0208] The embodiment of the present application also provides a device for determining the risk level of a video. Figure 7 is a structural diagram of a device for determining the risk level of a video provided in an embodiment of the present application, such as Figure 7 As shown, the apparatus 400 includes:

[0209] A first acquisition module 410 is configured to acquire a reference video and a set of videos to be compared; the set of videos to be compared includes a plurality of videos to be compared;

[0210] A second acquisition module 420 is configured to acquire first reference semantic information of the reference video and first comparison semantic information of each of the videos to be compared;

[0211] A first determining module 430 is configured to determine, in the set of videos to be compared, a plurality of first-level videos and a plurality of second-level videos based on the first reference semantic information and each piece of the first comparison semantic information;

[0212] A third acquisition module 440 is configured to acquire a reference content tag of the reference video and a comparison content tag of each second-level video;

[0213] A second determining module 450 is configured to determine a plurality of third-level videos from the plurality of second-level videos based on the reference content tag and each of the comparison content tags;

[0214] A fourth acquisition module 460 is configured to acquire a reference identification feature of the reference video and a comparison identification feature of each of the third-level videos and each of the first-level videos;

[0215] The third determination module 470 is configured to determine risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content tag, each of the comparison content tags, the reference identification feature, and each of the comparison identification features.

[0216] In an optional embodiment, the set of videos to be compared includes multiple fourth-level videos, and the multiple fourth-level videos are the videos to be compared in the set of videos to be compared excluding the multiple first-level videos and the multiple second-level videos; the multiple second-level videos include multiple fifth-level videos, and the multiple fifth-level videos are the videos to be compared in the multiple second-level videos excluding the multiple third-level videos; the third determination module is further used to determine the risk level information of the multiple fourth-level videos based on the first reference semantic information and the first comparison semantic information of each of the fourth-level videos; determine the risk level information of the multiple fifth-level videos based on the reference content tag and the comparison content tag of each of the fifth-level videos; and determine the risk level information of the multiple first-level videos and the multiple third-level videos based on the reference identification feature and each of the comparison identification features.

[0217] In an optional embodiment, the second acquisition module is also used to obtain the text information and / or sound information of the reference video and the text information and / or sound information of each of the videos to be compared; perform semantic recognition processing on the text information and / or sound information of the reference video to obtain the first reference semantic information of the reference video; perform semantic recognition processing on the text information and / or sound information of each of the videos to be compared to obtain the first comparison semantic information of each of the videos to be compared.

[0218] In an optional embodiment, the first determination module is also used to execute for each video to be compared: treating the video to be compared that is currently being executed as the current comparison video; determining the first similarity between the current comparison video and the reference video based on the first reference semantic information and the first comparison semantic information of the current comparison video; if the first similarity is lower than or equal to a first preset threshold and higher than or equal to a second preset threshold, determining that the current comparison video is the first-level video; or; if the first similarity is lower than the first preset threshold, determining that the current comparison video is the second-level video; determining the multiple first-level videos based on each first-level video, and determining the multiple second-level videos based on each second-level video.

[0219] In an optional embodiment, the third determination module is also used to determine the high risk level as the risk level information of the current comparison video if the first similarity is higher than the first preset threshold; and determine the risk level information of the multiple fourth-level videos based on the risk level information of each of the current comparison videos.

[0220] In an optional embodiment, the third acquisition module is further used to determine a target correction video among the multiple fourth-level videos based on the first similarity between each of the fourth-level videos and the reference video; the first similarity corresponding to the target correction video is greater than or equal to the first similarities corresponding to other videos in the multiple fourth-level videos; obtain the correction content label of the target correction video and the candidate content label of the reference video; determine the reference content label based on the correction content label and the candidate content label; and obtain the comparison content label of each second-level video.

[0221] In an optional embodiment, the third acquisition module is further used to perform preprocessing operations on the reference video and each of the second-level videos to obtain a preprocessed reference video and multiple preprocessed second-level videos; the preprocessing operation includes at least one of noise reduction processing, watermark removal processing, resolution reduction processing and frame reduction processing; based on the preprocessed reference video and the multiple preprocessed second-level videos, the reference content label of the reference video and the comparison content label of each of the second-level videos are obtained.

[0222] In an optional embodiment, the second determination module is further used to execute for each second-level video: treat the second-level video currently being executed as the current comparison video; determine the second similarity between the current comparison video and the reference video based on the reference content tag and the comparison content tag of the current comparison video; if the second similarity is higher than or equal to a third preset threshold, determine that the current comparison video is the third-level video; and determine the multiple third-level videos based on each third-level video.

[0223] In an optional embodiment, the third determination module is also used to determine the low risk level as the risk level information of the current comparison video if the second similarity is lower than the third preset threshold; and determine the risk level information of the multiple fifth-level videos based on the risk level information of each of the current comparison videos.

[0224] In an optional embodiment, the method further includes:

[0225] a fourth determining module, configured to determine a reference key frame from a plurality of reference frames of the reference video; the reference key frame being used to obtain a reference identification feature of the reference video;

[0226] a fifth acquisition module, configured to perform semantic recognition processing on the reference key frame to acquire second reference semantic information of the reference key frame;

[0227] a sixth acquisition module, configured to perform semantic recognition processing on a plurality of video frames of the third-level video to acquire third comparative semantic information corresponding to the plurality of video frames;

[0228] The fifth determination module is used to determine that the target video frame is a key frame of the third-level video if the second reference semantic information of the reference key frame matches the third comparative semantic information of the target video frame among the multiple video frames; the target video frame is used to obtain the comparative identification feature of the third-level video.

[0229] In an optional embodiment, the fourth acquisition module is also used to extract multiple global picture features and multiple local invariant features of the reference key frame; screen the multiple global picture features and multiple local invariant features of the reference key frame, and obtain the reference identification features based on the attention mechanism; extract multiple global picture features and multiple local invariant features of each of the target video frames; screen the multiple global picture features and multiple local invariant features of each of the target video frames, and obtain the comparison identification features of each first-level video based on the attention mechanism.

[0230] In an optional embodiment, the third module is further used to execute for each of the first-level video and the third-level video: regard the currently executed first-level video and the second-level video as the current comparison video; determine the third similarity between the current comparison video and the reference video based on the reference identification feature and the comparison identification feature of the current comparison video; if the third similarity is higher than or equal to a third preset threshold, determine the high risk level as the risk level information of the current comparison video; or; if the third similarity is lower than the third preset threshold, determine the low risk level as the risk level information of the current comparison video; determine the risk level information of the multiple first-level videos and the multiple third-level videos based on the risk level information of each of the current comparison videos.

[0231] The device and method embodiments in the embodiments of this application are based on the same application concept.

[0232] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal, a server or a similar computing device. Taking running on a server as an example, Figure 8 This is a hardware structure diagram of a server for a method for determining the risk level of a video provided in an embodiment of the present application. Figure 8As shown, the server 500 may vary significantly depending on its configuration or performance. It may include one or more central processing units (CPUs) 510 (the processor 510 may include, but is not limited to, a processing device such as a microprocessor (MCU) or a programmable logic device (FPGA), a memory 530 for storing data, and one or more storage media 520 (e.g., one or more mass storage devices) for storing application programs 523 or data 522. The memory 530 and storage media 520 may be either transient or persistent storage. The program stored in the storage medium 520 may include one or more modules, each of which may include a series of instruction operations on the server. Furthermore, the CPU 510 may be configured to communicate with the storage medium 520 to execute the series of instruction operations in the storage medium 520 on the server 500. The server 500 may also include one or more power supplies 560, one or more wired or wireless network interfaces 550, one or more input and output interfaces 540, and / or one or more operating systems 521, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0233] The input / output interface 540 can be used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the server 500. In one embodiment, the input / output interface 540 may include a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the input / output interface 540 may be a radio frequency (RF) module for wireless communication with the Internet.

[0234] It can be understood by those skilled in the art that Figure 8 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 8 More or fewer components than shown, or with Figure 8 Different configurations shown.

[0235] An embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the above-mentioned data processing method.

[0236] An embodiment of the present application also provides a computer-readable storage medium, which can be set in a server to store at least one instruction, at least one program, code set or instruction set related to a method for determining the risk level of a video in a method embodiment. The at least one instruction, the at least one program, the code set or instruction set is loaded and executed by the processor to implement the above-mentioned method for determining the risk level of the video.

[0237] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network servers in the computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard drive, a magnetic disk, or an optical disk, among other media capable of storing program code.

[0238] It can be seen from the embodiments of the method, device, electronic device or storage medium for determining the risk level of a video provided by the above-mentioned present application that a reference video and a set of videos to be compared are obtained in the present application; the set of videos to be compared includes multiple videos to be compared; the first reference semantic information of the reference video and the first comparative semantic information of each of the videos to be compared are obtained; based on the reference semantic information and each of the comparative semantic information, multiple first-level videos and multiple second-level videos are determined in the set of videos to be compared; the reference content label of the reference video and the comparative content label of each of the second-level videos are obtained; based on the reference content label and each of the comparative content labels, multiple third-level videos are determined in the multiple second-level videos; the reference identification feature of the reference video and the comparative identification feature of each of the third-level videos and each of the first-level videos are obtained; and the risk level information of each of the videos to be compared is determined based on the first reference semantic information, each of the first comparative semantic information, the reference content label, each of the comparative content labels, the reference identification feature and each of the comparative identification features. This application performs preliminary screening through semantic information and then performs secondary screening through content tags, dividing multiple videos to be compared into different levels, screening out a small number of videos for computationally complex feature identification, and determining the risk level information of these videos for subsequent infringement judgments, saving a large amount of computing resources, ensuring accuracy while improving the efficiency of video risk level determination.

[0239] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0240] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0241] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0242] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for determining the risk level of a video, characterized in that: include: Obtain a reference video and a video set to be compared; The video set to be compared includes multiple videos to be compared; Acquire first reference semantic information of the reference video and first comparison semantic information of each of the videos to be compared; determining, in the set of videos to be compared, a plurality of first-level videos and a plurality of second-level videos based on the first reference semantic information and each piece of the first comparison semantic information; The set of videos to be compared includes a plurality of fourth-level videos, the plurality of fourth-level videos being videos to be compared in the set of videos to be compared excluding the plurality of first-level videos and the plurality of second-level videos; the semantic similarity of the fourth-level videos is greater than the semantic similarity of the first-level videos; and the semantic similarity of the first-level videos is greater than the semantic similarity of the second-level videos; Obtaining a reference content tag of the reference video and a comparison content tag of each second-level video; Based on the reference content tag and each of the comparison content tags, a plurality of third-ranked videos are determined from the plurality of second-ranked videos; the plurality of second-ranked videos include a plurality of fifth-ranked videos, the plurality of fifth-ranked videos being videos to be compared from the plurality of second-ranked videos excluding the plurality of third-ranked videos; and content tag similarity of the third-ranked videos is greater than content tag similarity of the fifth-ranked videos; Obtaining a reference identification feature of the reference video and a comparison identification feature of each of the third-level videos and each of the first-level videos; Determining risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content tag, each of the comparison content tags, the reference identification feature, and each of the comparison identification features; The obtaining of the reference content tag of the reference video and the comparison content tag of each second-level video includes: Determining a target correction video from the plurality of fourth-level videos based on a first similarity between each of the fourth-level videos and the reference video; wherein the first similarity corresponding to the target correction video is greater than or equal to the first similarities corresponding to other videos in the plurality of fourth-level videos; Obtaining a corrected content label of the target corrected video and a candidate content label of the reference video; determining the reference content label based on the revised content label and the candidate content label; Obtain a comparative content label for each second-level video.

2. The method for determining the risk level of a video according to claim 1, wherein: The set of videos to be compared includes a plurality of fourth-level videos, which are videos to be compared in the set of videos to be compared excluding the plurality of first-level videos and the plurality of second-level videos; the plurality of second-level videos includes a plurality of fifth-level videos, which are videos to be compared in the plurality of second-level videos excluding the plurality of third-level videos; The determining, based on the first reference semantic information, each of the first comparison semantic information, the reference content tag, each of the comparison content tags, the reference identification feature, and each of the comparison identification features, of the risk level information of each of the to-be-compared videos includes: determining risk level information of the plurality of fourth-level videos based on the first reference semantic information and the first comparative semantic information of each of the fourth-level videos; determining risk level information of the plurality of fifth-level videos based on the reference content tag and the comparison content tag of each of the fifth-level videos; Based on the reference identification feature and each of the compared identification features, risk level information of the plurality of first-level videos and the plurality of third-level videos is determined.

3. The method for determining the risk level of a video according to claim 1, wherein: The obtaining of the first reference semantic information of the reference video and the first comparison semantic information of each of the videos to be compared includes: Acquiring text information and / or audio information of the reference video and text information and / or audio information of each of the videos to be compared; Performing semantic recognition processing on the text information and / or sound information of the reference video to obtain first reference semantic information of the reference video; Perform semantic recognition processing on the text information and / or sound information of each of the videos to be compared to obtain first comparison semantic information of each of the videos to be compared.

4. The method for determining the risk level of a video according to claim 2, wherein: The determining, based on the first reference semantic information and each of the first comparison semantic information, a plurality of first-level videos and a plurality of second-level videos in the set of videos to be compared includes: For each video to be compared, execute: The currently executed video to be compared is regarded as the current comparison video; Determining a first similarity between the current comparison video and the reference video based on the first reference semantic information and the first comparison semantic information of the current comparison video; If the first similarity is lower than or equal to a first preset threshold and higher than or equal to a second preset threshold, determining that the current comparison video is a first-level video; or if the first similarity is lower than the first preset threshold, determining that the current comparison video is a second-level video; The plurality of first-ranked videos are determined based on each of the first-ranked videos, and the plurality of second-ranked videos are determined based on each of the second-ranked videos.

5. The method for determining the risk level of a video according to claim 4, wherein: The determining, based on the first reference semantic information and the first comparative semantic information of each of the fourth-level videos, risk level information of the plurality of fourth-level videos includes: If the first similarity is higher than the first preset threshold, determining a high risk level as the risk level information of the current comparison video; The risk level information of the plurality of fourth-level videos is determined based on the risk level information of each of the current comparison videos.

6. The method for determining the risk level of a video according to claim 1, wherein: The obtaining of the reference content tag of the reference video and the comparison content tag of each second-level video includes: performing a preprocessing operation on the reference video and each of the second-level videos to obtain a preprocessed reference video and a plurality of preprocessed second-level videos; the preprocessing operation comprising at least one of noise reduction, watermark removal, resolution reduction, and frame reduction; Based on the preprocessed reference video and the plurality of preprocessed second-level videos, a reference content label of the reference video and a comparative content label of each of the second-level videos are obtained.

7. The method for determining the risk level of a video according to claim 2, wherein: The determining a plurality of third-level videos from the plurality of second-level videos based on the reference content tag and each of the compared content tags includes: For each of the second-level videos, perform: The second-level video currently being executed is regarded as the current comparison video; Determining a second similarity between the current comparison video and the reference video based on the reference content tag and the comparison content tag of the current comparison video; If the second similarity is higher than or equal to a third preset threshold, determining that the current comparison video is the third-level video; The plurality of third-level videos are determined based on each of the third-level videos.

8. The method for determining the risk level of a video according to claim 7, wherein: The determining risk level information of the plurality of fifth-level videos based on the reference content tag and the comparison content tag of each fifth-level video includes: If the second similarity is lower than the third preset threshold, determining a low risk level as the risk level information of the current comparison video; The risk level information of the plurality of fifth-level videos is determined based on the risk level information of each of the current comparison videos.

9. The method for determining the risk level of a video according to claim 2, wherein: Before obtaining the reference identification feature of the reference video and the comparison identification feature of each of the third-level videos and each of the first-level videos, the method further includes: Determining a reference key frame from a plurality of reference frames of the reference video; the reference key frame is used to obtain a reference identification feature of the reference video; Performing semantic recognition processing on the reference key frame to obtain second reference semantic information of the reference key frame; performing semantic recognition processing on a plurality of video frames of the third-level video to obtain third comparative semantic information corresponding to the plurality of video frames; If the second reference semantic information of the reference key frame matches the third comparative semantic information of the target video frame among the multiple video frames, the target video frame is determined to be a key frame of the third-level video; the target video frame is used to obtain comparative identification features of the third-level video.

10. The method for determining the risk level of a video according to claim 9, wherein: The obtaining of the reference identification feature of the reference video and the comparison identification feature of each of the third-level videos and each of the first-level videos includes: Extracting a plurality of global picture features and a plurality of local invariant features of the reference key frame; Screening multiple global image features and multiple local invariant features of the reference key frame, and fusing them based on an attention mechanism to obtain the reference identification feature; Extracting a plurality of global picture features and a plurality of local invariant features of each target video frame; Multiple global picture features and multiple local invariant features of each target video frame are screened, and fused based on the attention mechanism to obtain comparative identification features of each first-level video.

11. The method for determining the risk level of a video according to claim 10, wherein: The determining, based on the reference identification feature and each of the comparison identification features, risk level information of the plurality of first-level videos and the plurality of third-level videos includes: For each of the first-level video and the third-level video, perform: The first-level video and the third-level video currently being executed are regarded as current comparison videos; Determining a third similarity between the current comparison video and the reference video based on the reference identification feature and the comparison identification feature of the current comparison video; If the third similarity is higher than or equal to a fourth preset threshold, a high risk level is determined as the risk level information of the current comparison video; or if the third similarity is lower than the fourth preset threshold, a low risk level is determined as the risk level information of the current comparison video; The risk level information of the plurality of first-level videos and the plurality of third-level videos is determined based on the risk level information of each of the current comparison videos.

12. A device for determining the risk level of a video, characterized in that: include: A first acquisition module is used to acquire a reference video and a set of videos to be compared; The video set to be compared includes multiple videos to be compared; A second acquisition module is used to acquire the first reference semantic information of the reference video and the first comparison semantic information of each of the videos to be compared; A first determining module, configured to determine, in the set of videos to be compared, a plurality of first-level videos and a plurality of second-level videos based on the first reference semantic information and each of the comparison semantic information; The set of videos to be compared includes a plurality of fourth-level videos, the plurality of fourth-level videos being videos to be compared in the set of videos to be compared excluding the plurality of first-level videos and the plurality of second-level videos; the semantic similarity of the fourth-level videos is greater than the semantic similarity of the first-level videos; and the semantic similarity of the first-level videos is greater than the semantic similarity of the second-level videos; a third acquisition module, configured to acquire a reference content tag of the reference video and a comparison content tag of each second-level video; a second determining module configured to determine, based on the reference content tag and each of the comparison content tags, a plurality of third-rated videos from the plurality of second-rated videos; the plurality of second-rated videos including a plurality of fifth-rated videos, the plurality of fifth-rated videos being videos to be compared from the plurality of second-rated videos excluding the plurality of third-rated videos; and the content tag similarity of the third-rated videos being greater than the content tag similarity of the fifth-rated videos; a fourth acquisition module, configured to acquire a reference identification feature of the reference video and a comparison identification feature of each of the third-level videos and each of the first-level videos; a third determining module, configured to determine risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content tag, each of the comparison content tags, the reference identification feature, and each of the comparison identification features; The obtaining of the reference content tag of the reference video and the comparison content tag of each second-level video includes: Determining a target correction video from the plurality of fourth-level videos based on a first similarity between each of the fourth-level videos and the reference video; wherein the first similarity corresponding to the target correction video is greater than or equal to the first similarities corresponding to other videos in the plurality of fourth-level videos; Obtaining a corrected content label of the target corrected video and a candidate content label of the reference video; determining the reference content label based on the revised content label and the candidate content label; Obtain a comparative content label for each second-level video.

13. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method for determining the risk level of a video as described in any one of claims 1-11.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the method for determining the risk level of a video according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Video processing method, device and equipment and computer readable storage medium

    CN115049953A