Video risk level determination method and device, electronic equipment and storage medium
Through a multi-level screening method, combined with semantic recognition, content label matching and feature extraction, the problem of inefficient determination of video infringement risk level in the prior art is solved, and efficient and accurate infringement risk assessment is achieved.
Patent Information
- Application Number
- CN202510781440.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing video infringement risk level determination scheme is inefficient and insufficiently accurate, making it difficult to effectively deal with infringement of massive videos.
By obtaining semantic information, content labels and feature information of the reference video and the video to be compared, a multi-level screening method is used, including semantic recognition, content label matching and feature extraction, to determine the risk level of the video.
It improves the efficiency and accuracy of determining the risk level of video infringement, saves computing resources, and ensures the accuracy of infringement judgments.
Smart Images

Figure CN120277238A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of video processing, and in particular, to a method, apparatus, electronic device, and storage medium for determining the risk level of a video. Background Art
[0002] With the rapid development of digital media technology, a huge amount of video content is generated every day in fields such as short video platforms and film and television creation. However, along with this is the increasingly serious problem of video copyright infringement. Moreover, existing video processing technologies provide diverse means of covering up infringement behaviors, such as intelligent speed change, face reenactment technology, audio separation and recombination technology, etc. These technological breakthroughs not only significantly enhance the concealment of infringement behaviors, but also pose an unprecedented challenge to traditional video duplicate checking mechanisms.
[0003] Currently, the mainstream video infringement risk level determination solutions mainly adopt two types of technical paths: the frame-by-frame comparison method based on frame-level features and the hash matching method based on key frame extraction. The former needs to perform pixel-level comparison between the target video and the video to be inspected, with high complexity and low efficiency, and cannot cope with a large number of videos; although the latter reduces the complexity by extracting feature points, there are still significant accuracy problems.
[0004] Therefore, there is an urgent need for a video infringement risk level determination solution that combines efficiency and accuracy. Summary of the Invention
[0005] In order to solve the technical problems of low efficiency and accuracy of existing video infringement risk level determination solutions, the present invention provides a method, apparatus, electronic device, and storage medium for determining the risk level of a video.
[0006] In a first aspect, an embodiment of the present application provides a method for determining the risk level of a video, the method including: Obtain a reference video and a set of videos to be compared; the set of videos to be compared includes a plurality of videos to be compared; Obtain first reference semantic information of the reference video and first comparison semantic information of each of the videos to be compared; Based on the first reference semantic information and each of the comparison semantic information, determine a plurality of first-level videos and a plurality of second-level videos in the set of videos to be compared; Obtain reference content tags of the reference video and comparison content tags of each of the second-level videos; Based on the reference content tags and each of the comparison content tags, determine a plurality of third-level videos in the plurality of second-level videos; Obtain reference identification features of the reference video and comparison identification features of each of the third-level videos and each of the first-level videos; Determine the risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content label, each of the comparison content labels, the reference identification feature, and each of the comparison identification features.
[0007] In an alternative embodiment, the set of videos to be compared includes a plurality of fourth-level videos, and the plurality of fourth-level videos are the videos to be compared in the set of videos to be compared other than the plurality of first-level videos and the plurality of second-level videos; the plurality of second-level videos include a plurality of fifth-level videos, and the plurality of fifth-level videos are the videos to be compared in the plurality of second-level videos other than the plurality of third-level videos; The determining the risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content label, each of the comparison content labels, the reference identification feature, and each of the comparison identification features includes: Determine the risk level information of the plurality of fourth-level videos based on the first reference semantic information and the first comparison semantic information of each of the fourth-level videos; Determine the risk level information of the plurality of fifth-level videos based on the reference content label and the comparison content label of each of the fifth-level videos; Determine the risk level information of the plurality of first-level videos and the plurality of third-level videos based on the reference identification feature and each of the comparison identification features.
[0008] In an alternative embodiment, the obtaining the first reference semantic information of the reference video and the first comparison semantic information of each of the videos to be compared includes: Obtain the text information and / or audio information of the reference video and the text information and / or audio information of each of the videos to be compared; Perform semantic recognition processing on the text information and / or audio information of the reference video to obtain the first reference semantic information of the reference video; Perform semantic recognition processing on the text information and / or audio information of each of the videos to be compared to obtain the first comparison semantic information of each of the videos to be compared.
[0009] In an alternative embodiment, the determining a plurality of first-level videos and a plurality of second-level videos in the set of videos to be compared based on the first reference semantic information and each of the first comparison semantic information includes: Execute for each of the videos to be compared: Regard the currently executing video to be compared as the current comparison video; Determine a first similarity between the current comparison video and the reference video based on the first reference semantic information and the first comparison semantic information of the current comparison video; If the first similarity is lower than or equal to a first preset threshold and higher than or equal to a second preset threshold, determine that the current comparison video is the first-level video; or; if the first similarity is lower than the first preset threshold, determine that the current comparison video is the second-level video; Determine the multiple first-level videos based on each of the first-level videos, and determine the multiple second-level videos based on each of the second-level videos.
[0010] In an alternative embodiment, the determining the risk level information of the multiple fourth-level videos based on the first reference semantic information and the first comparison semantic information of each of the fourth-level videos includes: If the first similarity is higher than the first preset threshold, determine that the high risk level is the risk level information of the current comparison video; Determine the risk level information of the multiple fourth-level videos based on the risk level information of each of the current comparison videos.
[0011] In an alternative embodiment, the obtaining the reference content label of the reference video and the comparison content label of each of the second-level videos includes: Based on the first similarity between each of the fourth-level videos and the reference video, determine a target corrected video among the multiple fourth-level videos; the first similarity corresponding to the target corrected video is greater than or equal to the first similarities corresponding to other videos among the multiple fourth-level videos; Obtain the corrected content label of the target corrected video and the candidate content label of the reference video; Based on the corrected content label and the candidate content label, determine the reference content label; Obtain the comparison content label of each of the second-level videos.
[0012] In an alternative embodiment, the obtaining the reference content label of the reference video and the comparison content label of each of the second-level videos includes: Perform a preprocessing operation on the reference video and each of the second-level videos to obtain a preprocessed reference video and multiple preprocessed second-level videos; the preprocessing operation includes at least one of noise reduction processing, watermark removal processing, resolution reduction processing, and frame rate reduction processing; Based on the preprocessed reference video and the multiple preprocessed second-level videos, obtain the reference content label of the reference video and the comparison content label of each of the second-level videos.
[0013] In an alternative embodiment, determining a plurality of third-level videos from the plurality of second-level videos based on the reference content tag and each of the comparison content tags includes: Performing the following for each of the second-level videos: Regarding the currently executing second-level video as the current comparison video; Determining a second similarity between the current comparison video and the reference video based on the reference content tag and the comparison content tag of the current comparison video; If the second similarity is higher than or equal to a third preset threshold, determining the current comparison video as the third-level video; Determining the plurality of third-level videos based on each of the third-level videos.
[0014] In an alternative embodiment, determining risk level information of the plurality of fifth-level videos based on the reference content tag and the comparison content tag of each of the fifth-level videos includes: If the second similarity is lower than the third preset threshold, determining the low risk level as the risk level information of the current comparison video; Determining the risk level information of the plurality of fifth-level videos based on the risk level information of each of the current comparison videos.
[0015] In an alternative embodiment, before obtaining the reference identification feature of the reference video and the comparison identification features of each of the third-level videos and each of the first-level videos, further included is: Determining reference key frames among a plurality of reference frames of the reference video; the reference key frames are used to obtain the reference identification feature of the reference video; Performing semantic recognition processing on the reference key frames to obtain second reference semantic information of the reference key frames; Performing semantic recognition processing on a plurality of video frames of the third-level video to obtain third comparison semantic information corresponding to the plurality of video frames; If the second reference semantic information of the reference key frames matches the third comparison semantic information of a target video frame among the plurality of video frames, determining the target video frame as the key frame of the third-level video; the target video frame is used to obtain the comparison identification feature of the third-level video.
[0016] In an alternative embodiment, obtaining the reference identification feature of the reference video and the comparison identification features of each of the third-level videos and each of the first-level videos includes: Extracting a plurality of global picture features and a plurality of local invariant features of the reference key frames; Screen the multiple global picture features and multiple local invariant features of the reference key frame, and fuse them based on the attention mechanism to obtain the reference identification feature; Extract the multiple global picture features and multiple local invariant features of each target video frame; Screen the multiple global picture features and multiple local invariant features of each target video frame, and fuse them based on the attention mechanism to obtain the comparison identification feature of each first-level video.
[0017] In an alternative embodiment, determining the risk level information of the multiple first-level videos and the multiple third-level videos based on the reference identification feature and each comparison identification feature includes: Perform the following for each of the first-level videos and the third-level videos: Regard the currently executing first-level video and third-level video as the current comparison video; Based on the reference identification feature and the comparison identification feature of the current comparison video, determine the third similarity between the current comparison video and the reference video; If the third similarity is higher than or equal to a fourth preset threshold, determine that the high risk level is the risk level information of the current comparison video; or; if the third similarity is lower than the fourth preset threshold, determine that the low risk level is the risk level information of the current comparison video; Determine the risk level information of the multiple first-level videos and the multiple third-level videos based on the risk level information of each current comparison video.
[0018] In a second aspect, an embodiment of the present application provides a device for determining the risk level of a video. The device includes: A first acquisition module, configured to acquire a reference video and a set of videos to be compared; the set of videos to be compared includes multiple videos to be compared; A second acquisition module, configured to acquire the first reference semantic information of the reference video and the first comparison semantic information of each video to be compared; A first determination module, configured to determine multiple first-level videos and multiple second-level videos in the set of videos to be compared based on the first reference semantic information and each comparison semantic information; A third acquisition module, configured to acquire the reference content label of the reference video and the comparison content label of each second-level video; A second determination module, configured to determine multiple third-level videos in the multiple second-level videos based on the reference content label and each comparison content label; A fourth acquisition module, configured to acquire the reference identification features of the reference video and the comparison identification features of each of the third-level videos and each of the first-level videos; A third module, configured to determine the risk level information of each to-be-compared video based on the first reference semantic information, each of the first comparison semantic information, the reference content label, each of the comparison content labels, the reference identification features, and each of the comparison identification features.
[0019] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method for determining the risk level of a video in the first aspect.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which at least one instruction or at least one program is stored. The at least one instruction or at least one program is loaded and executed by the processor to implement the method for determining the risk level of a video in the first aspect.
[0021] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for determining the risk level of a video in the first aspect.
[0022] The method, device, electronic device, and storage medium for determining the risk level of a video provided by the embodiments of the present application have the following technical effects: Obtain a reference video and a set of videos to be compared; the set of videos to be compared includes multiple videos to be compared; obtain first reference semantic information of the reference video and first comparison semantic information of each video to be compared; based on the first reference semantic information and each comparison semantic information, determine multiple first-level videos and multiple second-level videos in the set of videos to be compared; obtain a reference content label of the reference video and a comparison content label of each second-level video; based on the reference content label and each comparison content label, determine multiple third-level videos among the multiple second-level videos; obtain a reference identification feature of the reference video and a comparison identification feature of each third-level video and each first-level video; based on the first reference semantic information, each first comparison semantic information, the reference content label, each comparison content label, the reference identification feature, and each comparison identification feature, determine risk level information of each video to be compared. In this application, preliminary screening is performed through semantic information, and then secondary screening is performed through content labels. Multiple videos to be compared are divided into different levels, a small number of videos are selected for computationally complex feature identification, and the risk level information of these videos is determined for subsequent infringement determination, saving a large amount of computing resources and improving the efficiency of determining the risk level of videos while ensuring accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application; Figure 2 is a flowchart of a method for determining the risk level of a video provided by an embodiment of the present application Figure 1 ; Figure 3 is a flowchart of a method for determining the risk level of a video provided by an embodiment of the present application Figure 2 ; Figure 4 is a flowchart of a method for determining the risk level of a video provided by an embodiment of the present application Figure 3 ; Figure 5 is a flowchart of a method for determining the risk level of a video provided by an embodiment of the present application Figure 4 ; Figure 6It is a flowchart illustration of a method for determining the risk level of a video provided by an embodiment of the present application Figure 5 ; Figure 7 It is a schematic structural diagram of a device for determining the risk level of a video provided by an embodiment of the present application; Figure 8 It is a hardware structure block diagram of a server for a method for determining the risk level of a video provided by an embodiment of the present application. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] Figure 1 It is a schematic diagram of an application environment provided by an embodiment of the present application, as Figure 1 shown. The application environment may include a server 01 and a client 02.
[0028] In some possible embodiments, the above-mentioned client 02 may include, but is not limited to, types of clients such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, etc. It may also be software running on the above-mentioned clients, such as application programs, applets, etc. Optionally, the operating systems running on the clients may include, but are not limited to, Android systems, IOS systems, linux, windows, Unix, etc.
[0029] In some possible embodiments, server 01 obtains a reference video and a set of videos to be compared; the set of videos to be compared includes multiple videos to be compared; obtains first reference semantic information of the reference video and first comparison semantic information of each video to be compared; based on the first reference semantic information and each comparison semantic information, determines multiple first-level videos and multiple second-level videos in the set of videos to be compared; obtains a reference content label of the reference video and a comparison content label of each second-level video; based on the reference content label and each comparison content label, determines multiple third-level videos among the multiple second-level videos; obtains a reference identification feature of the reference video and a comparison identification feature of each third-level video and each first-level video; based on the first reference semantic information, each first comparison semantic information, the reference content label, each comparison content label, the reference identification feature and each comparison identification feature, determines risk level information of each video to be compared.
[0030] Optionally, server 01 may include an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The operating system running on the server may include, but is not limited to, Android, IOS, Linux, Windows, Unix, etc.
[0031] The following introduces specific embodiments of a method for determining the risk level of a video in this application. Figure 2 It is a flowchart showing a method for determining the risk level of a video provided by an embodiment of this application. Figure 1 This specification provides method operation steps such as in the embodiment or flowchart, but based on routine or non-creative labor, there may be more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or server product executes, it can be executed in the order shown in the embodiment or the drawing or executed in parallel (for example, in an environment with parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, it may include: S201: Obtain a reference video and a set of videos to be compared; the set of videos to be compared includes multiple videos to be compared.
[0032] S202: Obtain first reference semantic information of the reference video and first comparison semantic information of each video to be compared.
[0033] S203: Determine multiple first - level videos and multiple second - level videos in the set of videos to be compared based on the first reference semantic information and each of the comparison semantic information.
[0034] S204: Obtain the reference content label of the reference video and the comparison content label of each of the second - level videos.
[0035] S205: Determine multiple third - level videos in the multiple second - level videos based on the reference content label and each of the comparison content labels.
[0036] S206: Obtain the reference identification feature of the reference video and the comparison identification features of each of the third - level videos and each of the first - level videos.
[0037] S207: Determine the risk - level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content label, each of the comparison content labels, the reference identification feature, and each of the comparison identification features.
[0038] Figure 3 It is a flowchart illustration of a method for determining the risk level of a video provided by an embodiment of the present application Figure 2 The method may include: S301: Obtain a reference video and a set of videos to be compared.
[0039] In a possible embodiment, the reference video is an original video, and the method for determining the risk level of the video of the present application is to determine the similarity degree of each video in the set of videos to be compared relative to the reference video, so as to determine the infringement risk level of each video.
[0040] In a possible embodiment, the set of videos to be compared includes multiple videos to be compared, and the number may be thousands or tens of thousands.
[0041] S302: Obtain the first reference semantic information of the reference video and the first comparison semantic information of each of the videos to be compared.
[0042] In a possible embodiment, the step of obtaining the first reference semantic information of the reference video and the first comparison semantic information of each of the videos to be compared includes: S312: Obtain the text information and / or audio information of the reference video and the text information and / or audio information of each of the videos to be compared.
[0043] S322: Perform semantic recognition processing on the text information and / or audio information of the reference video to obtain the first reference semantic information of the reference video.
[0044] S332: Perform semantic recognition processing on the text information and / or voice information of each of the videos to be compared, and obtain first comparison semantic information for each of the videos to be compared.
[0045] In an embodiment of the present application, the first reference semantic information refers to the semantic information of the reference video, and is used to represent the direct text information in the reference video and the indirect text information converted from the voice.
[0046] Similarly, the first comparison semantic information refers to the semantic information of the video to be compared, and is used to represent the direct text information in the video to be compared and the indirect text information converted from the voice.
[0047] In a possible embodiment, the reference video and the videos to be compared may include two or one of text information and voice information. Videos without text information or voice information are not within the scope of consideration of the present application.
[0048] In a possible embodiment, the text information may be subtitles on the video, text on the screen, etc.; the voice information may be the speech of the characters in the video, the background sound, the voiceover, etc.
[0049] Optionally, the reference video only has text information, and the video to be compared only has voice information; the reference video only has text information, and the video to be compared only has text information; the reference video only has voice information, and the video to be compared only has voice information; the reference video only has voice information, and the video to be compared only has text information; the reference video has both voice information and text information, and the video to be compared has both voice information and text information.
[0050] When obtaining the text information of the video, various types of text in the video can be extracted by using text extraction technologies such as optical character recognition technology, and semantic recognition is performed on various types of text by using a semantic recognition model. When obtaining the voice information of the video, technologies such as speech transcription and voice conversion can be used to first convert the voice into text, and then semantic recognition is performed on various types of text by using a semantic recognition model.
[0051] In a possible embodiment, when the video has both voice information and text information, the voice information and the text information can be merged and then semantic recognition is performed on the merged information by using a semantic recognition model, or semantic recognition can be respectively performed on the voice information and the text information by using a semantic recognition model, and then the two semantic informations are merged.
[0052] S303: Based on the first reference semantic information and each of the comparison semantic informations, determine multiple first-level videos and multiple second-level videos in the set of videos to be compared.
[0053] Figure 4 It is a flowchart of a method for determining the risk level of a video provided by an embodiment of the present application Figure 3 , as Figure 4 shown. In a possible embodiment, based on the first reference semantic information and each piece of the first comparison semantic information, multiple first-level videos and multiple second-level videos are determined in the set of videos to be compared, including: Perform the following for each video to be compared: S313: Regard the currently executing video to be compared as the current comparison video.
[0054] S323: Based on the first reference semantic information and the first comparison semantic information of the current comparison video, determine the first similarity between the current comparison video and the reference video.
[0055] In a possible embodiment, determining the first similarity between the current comparison video and the reference video is to calculate the first similarity between the first reference semantic information and the first comparison semantic information.
[0056] Since both the first reference semantic information and the first comparison semantic information extracted by the semantic recognition model are feature vectors, calculating the first similarity between the first reference semantic information and the first comparison semantic information can calculate the cosine similarity, Euclidean distance, Manhattan distance, etc. between the two feature vectors.
[0057] S333: Determine whether the first similarity is lower than or equal to the first preset threshold. If so, execute S343; if not, execute S373.
[0058] S343: Determine whether the first similarity is higher than or equal to the second preset threshold. If so, execute S353; if not, execute S363.
[0059] S353: Determine that the current comparison video is the first-level video.
[0060] In the embodiment of the present application, the second preset threshold is less than the first preset threshold.
[0061] In the embodiment of the present application, if the first similarity corresponding to the current comparison video is lower than or equal to the first preset threshold and higher than or equal to the second preset threshold, determine that the current comparison video is the first-level video, and the first-level video is a video with a medium semantic similarity to the reference video.
[0062] S363: Determine that the current comparison video is the second-level video.
[0063] In an embodiment of the present application, if the first similarity corresponding to the current comparison video is lower than or equal to the first preset threshold, it is determined that the current comparison video is a second-level video, and the second-level video is a video with a low semantic similarity to the reference video.
[0064] S373: Determine that the current comparison video is the fourth-level video, and determine that the high-risk level is the risk level information of the current comparison video.
[0065] S383: Determine the multiple first-level videos based on each of the first-level videos, and determine the multiple second-level videos based on each of the second-level videos.
[0066] In a possible embodiment, the fourth-level video is a comparison video to be compared in the set of comparison videos to be compared other than the multiple first-level videos and the multiple second-level videos.
[0067] In an embodiment of the present application, if the first similarity corresponding to the current comparison video is higher than the first preset threshold, it is determined that the current comparison video is a fourth-level video. The fourth-level video has a high semantic similarity to the reference video. At this time, it can be considered that the features such as the sound and text of the fourth-level video are highly similar to the reference video, and the infringement risk is relatively high.
[0068] That is to say, in the present application, through the steps of semantic recognition and semantic comparison, a large number of comparison videos to be compared are divided into three categories, namely the first-level video, the second-level video, and the fourth-level video. Among them, the semantic similarity of the first-level video is medium and needs to be further classified; the semantic similarity of the second-level video is low and the infringement risk is low, but for the sake of accuracy, further classification is also required in the follow-up; the semantic similarity of the fourth-level video is high and the infringement risk is high.
[0069] Through the above settings, multiple fourth-level videos with high similarity are quickly screened out through text and voice extraction and semantic recognition with fast speed, mature technology, and small computational complexity, and a high infringement risk is determined. And based on the semantic feature information, the comparison videos to be compared are divided into two different sets, namely the first-level video and the second-level video, which is convenient for subsequent different recognition processes, has stronger pertinence, and saves computational resources.
[0070] After screening out multiple second-level videos from the comparison videos to be compared, the following steps are performed: S304: Obtain the reference content label of the reference video and the comparison content label of each second-level video.
[0071] In an embodiment of the present application, the reference content label refers to the content label of the reference video and is used to represent the picture content information of the reference video. The comparison content label refers to the content label of the second-level video and is used to represent the picture content information of the second-level video.
[0072] In the embodiments of the present application, the reference object description information of the reference video is recognized by a preset object recognition model. For example, sky, house, boy, girl, umbrella, beach. Further, more specific reference object description information can also be obtained, such as blue sky, numerous houses, crying boys, laughing girls, sunshades and beaches in the sun. Subsequently, based on the reference object description information, the reference content tags of the reference video are determined.
[0073] In the embodiments of the present application, the comparison object description information of the second-level video is recognized by a preset object recognition model, and the comparison content tags of the second-level video are determined based on the comparison object description information.
[0074] In the embodiments of the present application, the reference content tags may include various tag information, such as attribute tag information, quantity tag information, and scene tag information. The object tag information listed above is only exemplary, and other possible object tag information can be included in the embodiments of the present application.
[0075] In a possible embodiment, the reference content tags of a video to be compared may be relatively simple. The content tags directly extracted from the reference video can also be corrected and supplemented by using fourth-level videos with high risks, that is, high similarities. The specific steps include: S314: Based on the first similarity between each of the fourth-level videos and the reference video, determine the target correction video among the multiple fourth-level videos.
[0076] In a possible embodiment, the first similarity corresponding to the target correction video is greater than or equal to the first similarities corresponding to other videos among the multiple fourth-level videos. That is to say, the target correction video has the highest similarity with the reference video.
[0077] S324: Obtain the correction content tags of the target correction video and the candidate content tags of the reference video.
[0078] S334: Based on the correction content tags and the candidate content tags, determine the reference content tags.
[0079] S344: Obtain the comparison content tags of each of the second-level videos.
[0080] In a possible embodiment, the correction content tags and the candidate content tags can be merged to directly obtain the reference content tags; the reference content tags can also be obtained after screening out irrelevant tags after merging the content tags.
[0081] In a possible embodiment, before content label extraction, the reference video and the second-level videos can also be preprocessed first to further reduce the computational load of object recognition processing and content label extraction.
[0082] Specifically, the reference video and each of the second-level videos can be preprocessed to obtain a preprocessed reference video and multiple preprocessed second-level videos; based on the preprocessed reference video and the multiple preprocessed second-level videos, the reference content label of the reference video and the comparison content label of each of the second-level videos are obtained.
[0083] In a possible embodiment, the preprocessing operation includes at least one of noise reduction processing, watermark removal processing, resolution reduction processing, and frame rate reduction processing.
[0084] Specifically, noise reduction processing refers to removing noise such as picture noise in the video. After noise reduction, the video picture is clearer, which helps to perform object recognition more accurately and extract content labels.
[0085] Watermark removal processing refers to removing watermarks such as Logos and text in the video to restore the original content of the video and avoid graphic and text-like watermarks from affecting object recognition and content label extraction.
[0086] Resolution reduction processing refers to reducing the resolution of the video (such as from 1080p to 720p) to reduce the size of the video file and the computational complexity. After reducing the resolution, the number of pixels in the video decreases, and the speed of object recognition is significantly improved, and the consumed resources are significantly reduced.
[0087] Frame rate reduction processing refers to reducing the frame rate of the video (such as from 30fps to 15fps) to reduce the number of frames in the video and the computational complexity. There is usually a large amount of redundant information between adjacent frames in the video. Frame rate reduction processing can reduce redundant frames while retaining the content information of key frames.
[0088] Through the above preprocessing operations, the effect and efficiency of content label processing can be significantly optimized.
[0089] S305: Based on the reference content label and each of the comparison content labels, determine multiple third-level videos from the multiple second-level videos.
[0090] Figure 5 It is a schematic flow of a method for determining the risk level of a video provided by an embodiment of the present application Figure 4 , as Figure 5 shown, in a possible embodiment, based on the reference content label and each of the comparison content labels, determining multiple third-level videos from the multiple second-level videos includes: Execute for each of the said second-level videos: S315: Regard the currently executing second-level video as the current comparison video.
[0091] S325: Determine the second similarity between the current comparison video and the reference video based on the reference content label and the comparison content label of the current comparison video.
[0092] S335: Determine whether the second similarity is higher than or equal to the third preset threshold. If so, execute S345; if not, execute S355.
[0093] S345: Determine that the current comparison video is the third-level video.
[0094] In the embodiments of the present application, if the second similarity between the reference content label and the comparison content label is higher than or equal to the third preset threshold, determine that the current comparison video is the third-level video.
[0095] S355: Determine that the current comparison video is the fifth-level video, and determine that the low-risk level is the risk level information of the current comparison video.
[0096] In the embodiments of the present application, if the second similarity between the reference content label and the comparison content label is lower than the third preset threshold, determine that the current comparison video is the fifth-level video.
[0097] In a possible embodiment, the multiple fifth-level videos are the videos to be compared among the multiple second-level videos except the multiple third-level videos.
[0098] That is to say, through content recognition and content label extraction, the present application further classifies the second-level videos with low semantic similarity, and divides the second-level videos into third-level videos and fifth-level videos according to content similarity. Among them, the third-level videos have high content similarity and high infringement risk, and need to be further classified; the fifth-level videos have low content similarity and low infringement risk.
[0099] Through the above settings, among the multiple second-level videos with medium semantic similarity initially screened by semantic information, further screening is performed through content labels. The third-level videos with low semantic similarity but high content label similarity are screened out for subsequent feature extraction operations. The fifth-level videos with low semantic similarity and low content label similarity are screened out and determined to have a low infringement risk.
[0100] By comparing the content label similarities with a large amount of calculation, more resources occupied, and longer time-consuming, different categories of videos to be compared are further divided, ensuring that each type of video is matched with an appropriate similarity calculation method, with strong pertinence and high accuracy.
[0101] After selecting multiple third - level videos from the above - mentioned second - level videos, the following operations are performed: S306: Determine reference key frames among multiple reference frames of the reference video.
[0102] In the embodiments of the present application, the reference key frame is at least one frame of picture that can best represent the content characteristics of the reference video, and is used to obtain the reference identification characteristics of the reference video.
[0103] Before performing the feature extraction operation, it is also necessary to perform the key frame determination operation. By only selecting several key video frames, on the one hand, the calculation amount is further reduced, and on the other hand, the interference of irrelevant information is reduced, ensuring the effectiveness of the extracted feature information.
[0104] S307: Perform semantic recognition processing on the reference key frame to obtain the second reference semantic information of the reference key frame.
[0105] S308: Perform semantic recognition processing on multiple video frames of the third - level video to obtain the third comparison semantic information corresponding to the multiple video frames.
[0106] S309: If the second reference semantic information of the reference key frame matches the third comparison semantic information of the target video frame among the multiple video frames, determine that the target video frame is the key frame of the third - level video.
[0107] In the embodiments of the present application, the second reference semantic information refers to the semantic information of the reference key frame, which is used to characterize the text information of the reference key frame, that is, the most representative text information in the reference video. The third comparison semantic information refers to the semantic information of the target video frame of the third - level video, which is used to characterize the text information of the target video frame, that is, the most representative text information in the third - level video.
[0108] In the embodiments of the present application, the target video frame is used to obtain the comparison identification characteristics of the third - level video.
[0109] In a possible embodiment, performing semantic recognition on video frames is similar to performing semantic recognition on the entire video. The method of determining whether the second reference semantic information matches the third comparison semantic information can also adopt the method of calculating similarity. If the fourth similarity between the second reference semantic information and the third comparison semantic information is greater than the fifth preset threshold, it is considered that the second reference semantic information of the reference key frame matches the third comparison semantic information of the target video frame.
[0110] S310: Obtain the reference identification characteristics of the reference video and the comparison identification characteristics of each third - level video and each first - level video.
[0111] In the embodiments of the present application, the reference identification feature refers to the identification feature of the reference video, which is used to characterize the frame features of the reference key frames of the reference video and remains after various operations such as deformation, cropping, and acceleration; the comparison identification feature refers to the identification features of the first-level video and the third-level video, which are used to characterize the frame features of the target video frames of the first-level video and the third-level video.
[0112] S3101: Extract multiple global frame features and multiple local invariant features of the reference key frame.
[0113] In a possible embodiment, the global frame features may include statistical features of the entire frame, such as color distribution, texture features, overall features, etc. Therefore, when extracting the global frame features of the reference key frame, the statistical information and overall visual characteristics of the entire frame are usually concerned. The global frame features may include color distribution, texture features, and overall structure, etc.
[0114] For example, the color distribution features of the key frame are captured through RGB histogram statistics and color moments, reflecting the overall distribution of colors in the frame; the multi-scale Gabor filter (such as the improved MSAF algorithm) is used to calculate the filtering responses in 8 directions to capture the texture information at different scales and directions in the frame; at the same time, principal component analysis (PCA) is used to reduce the dimension of the high-dimensional features, or a pre-trained convolutional neural network (CNN) (such as ResNet, EfficientNet) is used to extract high-level semantic features, so as to comprehensively describe the global visual information of the frame. These methods can effectively extract the global frame features of the reference key frame and provide a basis for subsequent video comparison and analysis.
[0115] In a possible embodiment, when extracting the local invariant features of the reference key frame, the detail information of the local area in the frame is usually concerned, and these features remain unchanged under image transformations (such as rotation, scaling, and brightness change). For example, SIFT (Scale-Invariant Feature Transform) or ORB (Oriented FAST and Rotated BRIEF) is used to extract key points and their descriptors, and these algorithms are robust to rotation, scaling, and brightness change; local texture features are extracted through local binary pattern (LBP) or Gabor filter to describe the texture pattern around the key points; at the same time, Harris corner detection or edge detection algorithms (such as Canny) are used to extract corner points and edge information in the frame to describe the local structure information. These local invariant features can stably reflect the detailed content of the key frame and provide robust support for video comparison and matching.
[0116] S3102: Screen multiple global frame features and multiple local invariant features of the reference key frame, and fuse them based on the attention mechanism to obtain the reference identification feature.
[0117] In a possible embodiment, when screening the global frame features and local invariant features of the reference key frame, first, a feature importance evaluation method (such as feature importance scoring based on random forest or XGBoost) is used to screen the global features (such as color distribution, texture features, CNN high-level semantic features) and local features (such as SIFT, ORB key point descriptors), and features with strong discriminability and high contribution to the description of video content are retained.
[0118] Subsequently, the screened features are fused based on the attention mechanism. Specifically, an attention network (such as Transformer or self-attention module) is used to calculate the weight of each feature, enabling the model to dynamically focus on features that are more important for the current task. Finally, the weighted global features and local features are concatenated or summed to generate a robust and discriminative reference identification feature. This fusion method can not only retain key information but also effectively improve the expression ability of features.
[0119] S3103: Extract multiple global frame features and multiple local invariant features of each target video frame.
[0120] S3104: Screen multiple global frame features and multiple local invariant features of each target video frame, and fuse them based on the attention mechanism to obtain the comparison identification feature of each first-level video.
[0121] In the embodiment of the present application, the feature extraction method for the target video frame is similar to that for the reference key frame.
[0122] S311: Based on the reference identification feature and each comparison identification feature, determine the risk level information of the multiple first-level videos and the multiple third-level videos.
[0123] Since the semantic similarity of the first-level videos is medium, that is, they have a certain similarity. Even if the content similarity is judged by means of content label extraction and comparison (that is, the second-level screening), it is impossible to accurately judge the similarity between the first-level videos and the reference video.
[0124] Similarly, since the semantic similarity of the third-level videos is low but the content label similarity is high, such contradictory results also cannot accurately judge the similarity between the third-level videos and the reference video.
[0125] Therefore, it is necessary to perform precise feature extraction and feature comparison on the first-level videos and the third-level videos to obtain accurate results.
[0126] Figure 6 It is a schematic flow chart of a method for determining the risk level of a video provided by an embodiment of the present application. Figure 5 , such as Figure 6 shown, in a possible embodiment, based on the reference identification feature and each of the comparison identification features, determining the risk level information of the multiple first-level videos and the multiple third-level videos includes the following steps: Execute for each of the first-level videos and the third-level videos: S3111: Treat the currently executing first-level video and third-level video as the current comparison video.
[0127] S3112: Based on the reference identification feature and the comparison identification feature of the current comparison video, determine the third similarity between the current comparison video and the reference video.
[0128] S3113: Determine whether the third similarity is higher than or equal to a fourth preset threshold. If so, execute S3114; if not, execute S3115.
[0129] S3114: Determine that the high risk level is the risk level information of the current comparison video.
[0130] In the embodiment of the present application, if the third similarity between the reference identification feature and the comparison identification feature is higher than the fourth preset threshold, determine that the risk level information of the current comparison video is the high risk level, that is, the video is highly similar to the reference video.
[0131] S3115: Determine that the low risk level is the risk level information of the current comparison video.
[0132] In the embodiment of the present application, if the third similarity between the reference identification feature and the comparison identification feature is lower than or equal to the fourth preset threshold, determine that the risk level information of the current comparison video is the low risk level, that is, the video has a low similarity to the reference video.
[0133] By extracting the global picture feature and local invariant feature of the key frame through feature extraction, and fusing them to obtain an identification feature with uniqueness and robustness, the infringement risk level of the first-level video and the third-level video that are difficult to judge in the foregoing steps is based on the identification feature, which has strong pertinence and high accuracy.
[0134] The embodiment of the present application also provides a device for determining the risk level of a video, Figure 7 It is a schematic structural diagram of a device for determining the risk level of a video provided by an embodiment of the present application. As Figure 7 shown, the device 400 includes: The first acquisition module 410 is configured to acquire a reference video and a set of videos to be compared; the set of videos to be compared includes multiple videos to be compared; The second acquisition module 420 is configured to acquire first reference semantic information of the reference video and first comparison semantic information of each of the videos to be compared; The first determination module 430 is configured to determine, based on the first reference semantic information and each of the first comparison semantic information, multiple first-level videos and multiple second-level videos in the set of videos to be compared; The third acquisition module 440 is configured to acquire reference content tags of the reference video and comparison content tags of each of the second-level videos; The second determination module 450 is configured to determine, based on the reference content tags and each of the comparison content tags, multiple third-level videos from the multiple second-level videos; The fourth acquisition module 460 is configured to acquire reference identification features of the reference video and comparison identification features of each of the third-level videos and each of the first-level videos; The third determination module 470 is configured to determine risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content tags, each of the comparison content tags, the reference identification features, and each of the comparison identification features.
[0135] In an alternative embodiment, the set of videos to be compared includes multiple fourth-level videos, and the multiple fourth-level videos are videos to be compared in the set of videos to be compared other than the multiple first-level videos and the multiple second-level videos; the multiple second-level videos include multiple fifth-level videos, and the multiple fifth-level videos are videos to be compared in the multiple second-level videos other than the multiple third-level videos; the third determination module is further configured to determine risk level information of the multiple fourth-level videos based on the first reference semantic information and the first comparison semantic information of each of the fourth-level videos; determine risk level information of the multiple fifth-level videos based on the reference content tags and the comparison content tags of each of the fifth-level videos; and determine risk level information of the multiple first-level videos and the multiple third-level videos based on the reference identification features and each of the comparison identification features.
[0136] In an alternative embodiment, the second acquisition module is further configured to acquire the text information and / or audio information of the reference video and the text information and / or audio information of each of the videos to be compared; perform semantic recognition processing on the text information and / or audio information of the reference video to obtain first reference semantic information of the reference video; perform semantic recognition processing on the text information and / or audio information of each of the videos to be compared to obtain first comparison semantic information of each of the videos to be compared.
[0137] In an alternative embodiment, the first determination module is further configured to perform, for each video to be compared: regard the currently executing video to be compared as the current comparison video; determine a first similarity between the current comparison video and the reference video based on the first reference semantic information and the first comparison semantic information of the current comparison video; if the first similarity is lower than or equal to a first preset threshold and higher than or equal to a second preset threshold, determine that the current comparison video is a first-level video; or; if the first similarity is lower than the first preset threshold, determine that the current comparison video is a second-level video; determine the multiple first-level videos based on each of the first-level videos, and determine the multiple second-level videos based on each of the second-level videos.
[0138] In an alternative embodiment, the third determination module is further configured to, if the first similarity is higher than the first preset threshold, determine that the high-risk level is the risk level information of the current comparison video; determine the risk level information of the multiple fourth-level videos based on the risk level information of each of the current comparison videos.
[0139] In an alternative embodiment, the third acquisition module is further configured to determine a target correction video among the multiple fourth-level videos based on the first similarity between each of the fourth-level videos and the reference video; the first similarity corresponding to the target correction video is greater than or equal to the first similarity corresponding to other videos among the multiple fourth-level videos; acquire the correction content label of the target correction video and the candidate content label of the reference video; determine the reference content label based on the correction content label and the candidate content label; acquire the comparison content label of each of the second-level videos.
[0140] In an alternative embodiment, the third acquisition module is further configured to perform a preprocessing operation on the reference video and each of the second-level videos to obtain a preprocessed reference video and a plurality of preprocessed second-level videos; the preprocessing operation includes at least one of noise reduction processing, watermark removal processing, resolution reduction processing, and frame rate reduction processing; based on the preprocessed reference video and the plurality of preprocessed second-level videos, obtain a reference content label of the reference video and a comparison content label of each of the second-level videos.
[0141] In an alternative embodiment, the second determination module is further configured to perform, for each of the second-level videos: regard the currently executing second-level video as the current comparison video; based on the reference content label and the comparison content label of the current comparison video, determine a second similarity between the current comparison video and the reference video; if the second similarity is higher than or equal to a third preset threshold, determine that the current comparison video is a third-level video; determine the plurality of third-level videos based on each of the third-level videos.
[0142] In an alternative embodiment, the third determination module is further configured to, if the second similarity is lower than the third preset threshold, determine that the low risk level is the risk level information of the current comparison video; determine the risk level information of the plurality of fifth-level videos based on the risk level information of each of the current comparison videos.
[0143] In an alternative embodiment, it further includes: A fourth determination module, configured to determine a reference key frame among a plurality of reference frames of the reference video; the reference key frame is used to obtain a reference identification feature of the reference video; A fifth acquisition module, configured to perform semantic recognition processing on the reference key frame to obtain second reference semantic information of the reference key frame; A sixth acquisition module, configured to perform semantic recognition processing on a plurality of video frames of the third-level video to obtain third comparison semantic information corresponding to the plurality of video frames; A fifth determination module, configured to, if the second reference semantic information of the reference key frame matches the third comparison semantic information of a target video frame among the plurality of video frames, determine that the target video frame is a key frame of the third-level video; the target video frame is used to obtain a comparison identification feature of the third-level video.
[0144] In an alternative embodiment, the fourth acquisition module is further configured to extract a plurality of global picture features and a plurality of local invariant features of the reference key frame; screen the plurality of global picture features and the plurality of local invariant features of the reference key frame, and fuse them based on an attention mechanism to obtain the reference identification feature; extract a plurality of global picture features and a plurality of local invariant features of each target video frame; screen the plurality of global picture features and the plurality of local invariant features of each target video frame, and fuse them based on an attention mechanism to obtain a comparison identification feature of each first-level video.
[0145] In an alternative embodiment, the third module is further configured to perform, for each of the first-level videos and the third-level videos: regard the currently executing first-level video and the second-level video as the current comparison video; based on the reference identification feature and the comparison identification feature of the current comparison video, determine a third similarity between the current comparison video and the reference video; if the third similarity is higher than or equal to a third preset threshold, determine that the high-risk level is the risk level information of the current comparison video; or; if the third similarity is lower than the third preset threshold, determine that the low-risk level is the risk level information of the current comparison video; determine the risk level information of the plurality of first-level videos and the plurality of third-level videos based on the risk level information of each current comparison video.
[0146] The device and method embodiments in this application are based on the same application concept.
[0147] The method embodiments provided in the embodiments of this application can be executed on a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 8 is a hardware structure block diagram of a server for a method for determining the risk level of a video provided in the embodiments of this application. As Figure 8As shown, the server 500 can vary significantly due to different configurations or performances, and may include one or more central processing units (CPUs) 510 (the processor 510 may include, but is not limited to, processing devices such as a microprocessor MCU or a field programmable gate array FPGA), a memory 530 for storing data, and one or more storage media 520 for storing application programs 523 or data 522 (such as one or more mass storage devices). Among them, the memory 530 and the storage media 520 can be transient storage or persistent storage. The programs stored in the storage media 520 may include one or more modules, and each module may include a series of instruction operations in the server. Further, the central processing unit 510 can be set to communicate with the storage media 520 and execute a series of instruction operations in the storage media 520 on the server 500. The server 500 may also include one or more power supplies 560, one or more wired or wireless network interfaces 550, one or more input / output interfaces 540, and / or one or more operating systems 521, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0148] The input / output interface 540 can be used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by the communication provider of the server 500. In one example, the input / output interface 540 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the input / output interface 540 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0149] Those of ordinary skill in the art can understand that Figure 8 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the server 500 may also include more or fewer components than Figure 8 shown, or have a different configuration from Figure 8 shown.
[0150] An embodiment of the present application provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above data processing method.
[0151] An embodiment of the present application further provides a computer-readable storage medium. The storage medium can be disposed in a server to store at least one instruction, at least one program, a code set, or an instruction set related to a method for determining the risk level of a video in a method embodiment. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method for determining the risk level of the above video.
[0152] Optionally, in this embodiment, the above storage medium may be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0153] As can be seen from the embodiments of the method, apparatus, electronic device, or storage medium for determining the risk level of a video provided by the present application, in the present application, a reference video and a set of videos to be compared are obtained; the set of videos to be compared includes multiple videos to be compared; first reference semantic information of the reference video and first comparison semantic information of each video to be compared are obtained; based on the reference semantic information and each comparison semantic information, multiple first-level videos and multiple second-level videos are determined in the set of videos to be compared; reference content tags of the reference video and comparison content tags of each second-level video are obtained; based on the reference content tags and each comparison content tag, multiple third-level videos are determined among the multiple second-level videos; reference identification features of the reference video and comparison identification features of each third-level video and each first-level video are obtained; based on the first reference semantic information, each first comparison semantic information, the reference content tags, each comparison content tag, the reference identification features, and each comparison identification feature, the risk level information of each video to be compared is determined. The present application performs a preliminary screening through semantic information, then performs a secondary screening through content tags, divides multiple videos to be compared into different levels, screens out a small number of videos for computationally complex feature identification, and determines the risk level information of these videos for subsequent infringement determination, saving a large amount of computing resources and improving the efficiency of determining the risk level of a video while ensuring accuracy.
[0154] It should be noted that the above-mentioned sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. Moreover, the above-mentioned specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0155] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments.
[0156] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware or by a program instructing the relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.
[0157] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for determining the risk level of a video, characterized in that, Including: Obtain a reference video and a set of videos to be compared; The set of videos to be compared includes multiple videos to be compared; Obtain the first reference semantic information of the reference video and the first comparison semantic information of each video to be compared; Based on the first reference semantic information and each first comparison semantic information, determine multiple first-level videos and multiple second-level videos in the set of videos to be compared; Obtain the reference content label of the reference video and the comparison content label of each second-level video; Based on the reference content label and each comparison content label, determine multiple third-level videos among the multiple second-level videos; Obtain the reference identification feature of the reference video and the comparison identification feature of each third-level video and each first-level video; Based on the first reference semantic information, each first comparison semantic information, the reference content label, each comparison content label, the reference identification feature, and each comparison identification feature, determine the risk level information of each video to be compared.
2. The method for determining the risk level of a video according to claim 1, wherein The set of videos to be compared includes multiple fourth-level videos, and the multiple fourth-level videos are the videos to be compared in the set of videos to be compared other than the multiple first-level videos and the multiple second-level videos; the multiple second-level videos include multiple fifth-level videos, and the multiple fifth-level videos are the videos to be compared in the multiple second-level videos other than the multiple third-level videos; The determining the risk level information of each video to be compared based on the first reference semantic information, each first comparison semantic information, the reference content label, each comparison content label, the reference identification feature, and each comparison identification feature includes: Based on the first reference semantic information and the first comparison semantic information of each fourth-level video, determine the risk level information of the multiple fourth-level videos; Based on the reference content label and the comparison content label of each fifth-level video, determine the risk level information of the multiple fifth-level videos; Based on the reference identification feature and each comparison identification feature, determine the risk level information of the multiple first-level videos and the multiple third-level videos.
3. The method for determining the risk level of a video according to claim 1, wherein The obtaining the first reference semantic information of the reference video and the first comparison semantic information of each video to be compared includes: Obtain the text information and / or audio information of the reference video and the text information and / or audio information of each video to be compared; Perform semantic recognition processing on the text information and / or audio information of the reference video to obtain the first reference semantic information of the reference video; Perform semantic recognition processing on the text information and / or audio information of each video to be compared to obtain the first comparison semantic information of each video to be compared.
4. The risk level determination method for a video according to claim 2, characterized in that The determining multiple first-level videos and multiple second-level videos in the set of videos to be compared based on the first reference semantic information and each first comparison semantic information includes: Execute for each video to be compared: Take the to-be-compared video currently being executed as the current comparison video; Based on the first reference semantic information and the first comparison semantic information of the current comparison video, determine the first similarity between the current comparison video and the reference video; If the first similarity is lower than or equal to the first preset threshold and higher than or equal to the second preset threshold, determine that the current comparison video is the first-level video; or; if the first similarity is lower than the first preset threshold, determine that the current comparison video is the second-level video; Determine the multiple first-level videos based on each of the first-level videos, and determine the multiple second-level videos based on each of the second-level videos.
5. A method for determining the risk level of a video according to claim 4, characterized in that, The determining the risk level information of the multiple fourth-level videos based on the first reference semantic information and the first comparison semantic information of each of the fourth-level videos includes: If the first similarity is higher than the first preset threshold, determine that the high risk level is the risk level information of the current comparison video; Determine the risk level information of the multiple fourth-level videos based on the risk level information of each of the current comparison videos.
6. The method for determining the risk level of a video according to claim 5, characterized in that The obtaining the reference content label of the reference video and the comparison content label of each of the second-level videos includes: Based on the first similarity between each of the fourth-level videos and the reference video, determine a target correction video among the multiple fourth-level videos; the first similarity corresponding to the target correction video is greater than or equal to the first similarities corresponding to other videos among the multiple fourth-level videos; Obtain the correction content label of the target correction video and the candidate content label of the reference video; Based on the correction content label and the candidate content label, determine the reference content label; Obtain the comparison content label of each of the second-level videos.
7. A method for determining the risk level of a video according to claim 1, characterized in that, The obtaining the reference content label of the reference video and the comparison content label of each of the second-level videos includes: Perform a preprocessing operation on the reference video and each of the second-level videos to obtain a preprocessed reference video and multiple preprocessed second-level videos; the preprocessing operation includes at least one of noise reduction processing, watermark removal processing, resolution reduction processing, and frame rate reduction processing; Based on the preprocessed reference video and the multiple preprocessed second-level videos, obtain the reference content label of the reference video and the comparison content label of each of the second-level videos.
8. The method for determining the risk level of a video according to claim 2, wherein The determining multiple third-level videos among the multiple second-level videos based on the reference content label and each of the comparison content labels includes: Perform on each of the second-level videos: Take the second-level video currently being executed as the current comparison video; Based on the reference content label and the comparison content label of the current comparison video, determine the second similarity between the current comparison video and the reference video; If the second similarity is higher than or equal to the third preset threshold, determine that the current comparison video is the third-level video; Determine the multiple third-level videos based on each of the third-level videos.
9. The method for determining the risk level of a video according to claim 8, characterized in that Determining the risk level information of the multiple fifth-level videos based on the reference content tags and the comparison content tags of each of the fifth-level videos includes: If the second similarity is lower than the third preset threshold, determining that the low risk level is the risk level information of the current comparison video; Determining the risk level information of the multiple fifth-level videos based on the risk level information of each of the current comparison videos.
10. A method for determining the risk level of a video according to claim 2, characterized in that, Before obtaining the reference identification features of the reference video and the comparison identification features of each of the third-level videos and each of the first-level videos, it further includes: Determining reference key frames in multiple reference frames of the reference video; the reference key frames are used to obtain the reference identification features of the reference video; Performing semantic recognition processing on the reference key frames to obtain second reference semantic information of the reference key frames; Performing semantic recognition processing on multiple video frames of the third-level video to obtain third comparison semantic information corresponding to the multiple video frames; If the second reference semantic information of the reference key frame matches the third comparison semantic information of the target video frame among the multiple video frames, determining that the target video frame is the key frame of the third-level video; the target video frame is used to obtain the comparison identification features of the third-level video.
11. A method for determining the risk level of a video according to claim 10, characterized in that, The obtaining the reference identification features of the reference video and the comparison identification features of each of the third-level videos and each of the first-level videos includes: Extracting multiple global picture features and multiple local invariant features of the reference key frame; Screening the multiple global picture features and multiple local invariant features of the reference key frame, and fusing them based on the attention mechanism to obtain the reference identification features; Extracting multiple global picture features and multiple local invariant features of each of the target video frames; Screening the multiple global picture features and multiple local invariant features of each of the target video frames, and fusing them based on the attention mechanism to obtain the comparison identification features of each of the first-level videos.
12. The method for determining the risk level of a video according to claim 11, wherein The determining the risk level information of the multiple first-level videos and the multiple third-level videos based on the reference identification features and each of the comparison identification features includes: Performing the following operations on each of the first-level videos and the third-level videos: Regarding the currently executing first-level video and third-level video as the current comparison video; Determining the third similarity between the current comparison video and the reference video based on the reference identification features and the comparison identification features of the current comparison video; If the third similarity is higher than or equal to the fourth preset threshold, determining that the high risk level is the risk level information of the current comparison video; or if the third similarity is lower than the fourth preset threshold, determining that the low risk level is the risk level information of the current comparison video; Determining the risk level information of the multiple first-level videos and the multiple third-level videos based on the risk level information of each of the current comparison videos.
13. A device for determining the risk level of a video, characterized in that Including: A first obtaining module, configured to obtain a reference video and a set of videos to be compared; The set of videos to be compared includes multiple videos to be compared; A second acquisition module, configured to acquire first reference semantic information of the reference video and first comparison semantic information of each of the videos to be compared; A first determination module, configured to determine a plurality of first-level videos and a plurality of second-level videos in the set of videos to be compared based on the first reference semantic information and each of the comparison semantic information; A third acquisition module, configured to acquire a reference content label of the reference video and a comparison content label of each of the second-level videos; A second determination module, configured to determine a plurality of third-level videos from the plurality of second-level videos based on the reference content label and each of the comparison content labels; A fourth acquisition module, configured to acquire a reference identification feature of the reference video and a comparison identification feature of each of the third-level videos and each of the first-level videos; A third determination module, configured to determine risk level information of each of the videos to be compared based on the first reference semantic information, each of the first comparison semantic information, the reference content label, each of the comparison content labels, the reference identification feature, and each of the comparison identification features.
14. An electronic device, characterized in that, The electronic device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method for determining the risk level of a video according to any one of claims 1-12.
15. A computer-readable storage medium, characterized in that, At least one instruction or at least one program is stored in the computer-readable storage medium. The at least one instruction or the at least one program is loaded and executed by the processor to implement the method for determining the risk level of a video according to any one of claims 1-12.
Citation Information
Patent Citations
Video processing method, device and equipment and computer readable storage medium
CN115049953A
Method and system for similarity contents search
KR1020100047110A
Multimedia content filtering
US20080228928A1
Method And Apparatus For Retrieving Video, Device And Medium
US20210209155A1
Method and system for profiling a reference image and an object-of-interest therewithin
US20230145362A1