Video recognition method and device, electronic equipment and storage medium

By filtering the interaction data of the objects in the video to be identified, the relevance between the video content and the title can be accurately identified, solving the problem of identifying clickbait videos, improving the accuracy and efficiency of video recognition, and enhancing the user experience.

CN116975358BActive Publication Date: 2026-05-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-04-20
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies, clickbait videos often have content that differs significantly from user expectations, resulting in a poor viewing experience and difficulty in accurately identifying the target video.

Method used

By acquiring object interaction data from the video to be identified, filtering relevant interaction data related to the target information, determining the relevance between the video content and the title, and judging whether the video is the target video based on the relevance threshold.

Benefits of technology

It improves the accuracy of target video recognition, reduces data processing volume, and enhances video analysis efficiency and user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116975358B_ABST
    Figure CN116975358B_ABST
Patent Text Reader

Abstract

The application discloses a video recognition method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises: obtaining object interaction data and a video title of a to-be-recognized video; filtering relevant interaction data related to target information from the object interaction data; determining the relevance of the to-be-recognized video and the video title according to the relevant interaction data; and determining whether the to-be-recognized video is a target video according to the comparison result of the relevance and a relevance threshold. In the application, the relevant interaction data can more accurately reflect the relevance of the video content of the to-be-recognized video and the video title, thereby improving the recognition accuracy of the target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet information processing technology, and more specifically, to a video recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] With the development of the internet, more and more new ways of content dissemination have emerged. Taking video content as an example, many frequently pushed video content content suitable for viewing on mobile devices and during short leisure periods has appeared on various new media platforms. When users choose videos they are interested in, they usually filter by title first. To attract users' attention, content providers release some clickbait videos. Clickbait videos attract users by exaggerating the content of the video title and raising users' viewing expectations, but the video content often deviates significantly from users' expectations, resulting in a poor viewing experience. Summary of the Invention

[0003] In view of this, embodiments of this application propose a video recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product that can accurately identify target videos.

[0004] In a first aspect, embodiments of this application provide a video recognition method, the method comprising: acquiring object interaction data and video title of a video to be recognized; filtering relevant interaction data related to target information from the object interaction data, wherein the target information is used to indicate the video content of the video to be recognized, and / or, the target information is used to indicate that the video to be recognized is a target video; determining the relevance between the video to be recognized and the video title based on the relevant interaction data; and determining whether the video to be recognized is a target video based on a comparison result of the relevance and a relevance threshold.

[0005] Secondly, embodiments of this application provide a video recognition device, comprising: an acquisition module for acquiring object interaction data and a video title of a video to be recognized; a filtering module for filtering relevant interaction data related to target information from the object interaction data, wherein the target information is used to indicate the video content of the video to be recognized, and / or, the target information is used to indicate that the video to be recognized is a target video; a relevance determination module for determining the relevance between the video to be recognized and the video title based on the relevant interaction data; and a video determination module for determining whether the video to be recognized is a target video based on a comparison result of the relevance and a relevance threshold.

[0006] Optionally, the filtering module is further configured to acquire video content data of the video to be identified as the target information; filter interactive data related to the video content data from the object interaction data as the relevant interactive data; the relevance determination module is further configured to determine the title relevance probability corresponding to the relevant interactive data based on the video title, wherein the title relevance probability is used to characterize the relevance between the video to be identified and the video title.

[0007] Optionally, the filtering module is further configured to: determine the content relevance probability of the corresponding object interaction data based on the video content data, wherein the content relevance probability is used to characterize the correlation between the object interaction data and the video content data; determine the interaction popularity of the object interaction data, wherein the interaction popularity is used to characterize the interactivity of the object interaction data itself; determine the key information probability of the object interaction data based on the content relevance probability and the interaction popularity; and filter the top N object interaction data with higher key information probabilities from the object interaction data as the relevant interaction data, where N is a positive integer.

[0008] Optionally, the object interaction data includes at least one object interaction data; the filtering module is further configured to obtain the number of likes and replies for each object interaction data; determine the overall interaction volume of the at least one object interaction data based on the number of likes and replies for each object interaction data; perform a weighted summation on the number of likes and replies for each object interaction data to obtain the weighted sum for each object interaction data; calculate the ratio of the weighted sum for each object interaction data to the overall interaction volume to obtain the interaction popularity corresponding to each object interaction data.

[0009] Optionally, the filtering module is further configured to: determine the media features of the video to be identified based on the video content data; obtain the text features of the object interaction data; fuse the text features and the media features to obtain a first fused feature; and analyze the first fused feature using a first correlation analysis model to obtain the content relevance probability corresponding to the object interaction data.

[0010] Optionally, the relevance determination module is further configured to: obtain the key information probability of the relevant interaction data, wherein the key information probability is used to characterize the importance of the relevant interaction data in all object interaction data corresponding to the video to be identified; filter target interaction data from the relevant interaction data based on the relevant interaction data and the key information probability; obtain the title features of the video title and the text features of the target interaction data; fuse the title features and the text features of the target interaction data to obtain a second fused feature; and analyze the second fused feature through a second relevance analysis model to obtain the title relevance probability corresponding to the relevant interaction data.

[0011] Optionally, the video determination module is further configured to determine the video to be identified as the target video when the title relevance probability is less than the relevance threshold.

[0012] Optionally, the filtering module is further configured to obtain target keywords as target information, wherein the target keywords are obtained based on sample interaction data with labeled information, wherein the labeled information is used to characterize that the sample interaction data indicates that the sample video is the target video, and the sample video is the video corresponding to the sample interaction data; and to filter out the interaction data that hits the target keywords from the object interaction data as relevant interaction data; the relevance determination module is further configured to determine the title hit probability corresponding to the relevant interaction data, wherein the title hit probability is used to characterize the relevance between the video to be identified and the video title.

[0013] Optionally, the device further includes: a keyword acquisition module, used to acquire the sample interaction data; extract keywords from the sample interaction data to obtain multiple preliminary keywords; determine the frequency of occurrence of the multiple preliminary keywords in the sample interaction data; and filter target keywords from the multiple preliminary keywords whose frequency of occurrence exceeds a preset frequency.

[0014] Optionally, the relevant interactive data includes at least one relevant interactive data; the relevance determination module is further configured to determine the text features and interaction popularity of each relevant interactive data, wherein the interaction popularity is used to characterize the interactivity of the relevant interactive data itself; analyze the text features of each relevant interactive data through a title hit analysis model to obtain the initial title hit probability of the corresponding relevant interactive data; and obtain the title hit probability based on the initial title hit probability and interaction popularity of each relevant interactive data.

[0015] Optionally, the video determination module is further configured to determine the video to be identified as the target video when the title hit probability exceeds the relevance threshold.

[0016] Optionally, the filtering module is further configured to acquire video content data and target keywords of the video to be identified as the target information, wherein the target keywords are obtained based on sample interaction data with labeled information, and the labeled information is used to characterize that the sample interaction data indicates that the sample video is the target video, and the sample video is the video corresponding to the sample interaction data; filter first interaction data related to the video content data from the object interaction data; filter interaction data that hit the target keywords from the object interaction data as the second interaction data; and use the first interaction data and the second interaction data as related interaction data; the relevance determination module is further configured to determine the title relevance probability corresponding to the first interaction data based on the video title, wherein the title relevance probability is used to characterize the relevance between the video to be identified and the video title; determine the title hit probability corresponding to the second interaction data, wherein the title hit probability is used to characterize the relevance between the video to be identified and the video title; the video determination module is further configured to perform a weighted summation of the title hit probability and the title relevance probability; and determine whether the video to be identified is the target video based on the comparison result of the summation result and the relevance threshold.

[0017] Optionally, the device further includes: a prompt message sending module, configured to send a prompt message or stop distributing the target video if it is determined that the video to be identified is the target video.

[0018] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory; the memory stores a computer program, and the processor executes the computer program to implement the above-described method.

[0019] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.

[0020] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0021] This application provides a video recognition method, apparatus, electronic device, and storage medium. By filtering relevant interaction data related to target information from object interaction data of the video to be recognized, determining the relevance between the video to be recognized and the video title based on the relevant interaction data, and further determining whether the video to be recognized is the target video based on the relevance. Since relevant interaction data can more accurately reflect the relevance between the video content and the video title of the video to be recognized, the recognition accuracy of the target video can be improved. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this application;

[0024] Figure 2 A flowchart of a video recognition method according to an embodiment of this application is shown;

[0025] Figure 3 A flowchart of a video recognition method according to yet another embodiment of this application is shown;

[0026] Figure 4 A flowchart of a filtering method for first interactive data in an embodiment of this application is shown;

[0027] Figure 5 A flowchart of a method for obtaining interaction popularity in an embodiment of this application is shown;

[0028] Figure 6 A flowchart illustrating a method for obtaining content relevance probability in an embodiment of this application is shown;

[0029] Figure 7 A schematic diagram illustrating the process of obtaining content relevance probability in an embodiment of this application is shown;

[0030] Figure 8 A flowchart of a method for obtaining title relevance probability in an embodiment of this application is shown;

[0031] Figure 9 A schematic diagram illustrating the process of obtaining the title-related probability in an embodiment of this application is shown;

[0032] Figure 10 A flowchart of a video recognition method provided in another embodiment of this application is shown;

[0033] Figure 11 A flowchart illustrating a method for obtaining title hit probability in an embodiment of this application is shown;

[0034] Figure 12 This illustration shows a schematic diagram of the process for obtaining the title hit probability in an embodiment of this application;

[0035] Figure 13 A flowchart of a video recognition method provided in another embodiment of this application is shown;

[0036] Figure 14 A schematic diagram illustrating the recognition process of the video to be recognized in an embodiment of this application is shown;

[0037] Figure 15 A flowchart of a video recognition method provided in another embodiment of this application is shown;

[0038] Figure 16 A block diagram of a video recognition device according to one embodiment of this application is shown;

[0039] Figure 17 A structural block diagram of an electronic device for performing a video recognition method according to an embodiment of this application is shown. Detailed Implementation

[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0041] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0043] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this application. For example... Figure 1 As shown, this application scenario includes terminal 101 and server 102, which are connected via wired or wireless network communication. Terminal 101 can be a smartphone, tablet, laptop, desktop computer, smart home appliance, vehicle terminal, aircraft, wearable device terminal, virtual reality device, or other terminal device capable of video playback. This terminal can run a video playback application or other applications that can call the video playback application (such as instant messaging applications, shopping applications, search applications, game applications, forum applications, map and traffic applications, etc.).

[0044] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Server 102 can be used to provide services for the applications running on terminal 101.

[0045] In this process, terminal 101 can send the request of an object to server 102. The object can be a user. Therefore, server 102 can provide corresponding video content based on the request of the object, or server 102 can directly send video content to terminal 101, or server 102 can send customized video content or video content related to the object's interests based on the object information bound to terminal 101.

[0046] Regarding the issues mentioned in the background technology, server 102 can use video recognition methods to identify videos to be sent, recognizing those that are clearly clickbait videos and feeding back videos that better meet the user's needs to terminal 101. The inventors discovered through research that if the relevance between the video's content (such as dialogue or narration added by the author) and the video title is used to determine whether a video is clickbait, even when the video title is obscure, it can accurately determine whether the video to be identified is clickbait.

[0047] Based on this, the inventors have proposed a video recognition method, apparatus, electronic device, and storage medium as provided in this application. The method involves acquiring object interaction data and a video title of a video to be recognized; filtering relevant interaction data related to target information from the object interaction data, where the target information indicates the video content of the video to be recognized, and / or indicates that the video to be recognized is a target video; determining the relevance between the video to be recognized and the video title based on the relevant interaction data; and determining whether the video to be recognized is a target video based on a comparison of the relevance and a relevance threshold. Relevant interaction data can more accurately reflect the relevance between the video content and the video title of the video to be recognized, thereby improving the accuracy of target video recognition.

[0048] Please see Figure 2 , Figure 2 This application illustrates a flowchart of a video recognition method according to an embodiment of the present application. This method can be applied to an electronic device, which can be... Figure 1 The method, which includes the terminal 101 or server 102 shown, comprises:

[0049] S110. Obtain the object interaction data and video title of the video to be identified.

[0050] The video to be identified can refer to any video that needs to be determined to be the target video. For example, it could be a video responding to the needs of the object, a video to be distributed, or a video to be sent to terminal 101 based on the object's identity information, historical data, or object subscription information.

[0051] The target video in this application embodiment can refer to clickbait videos or other videos with identifiable features. The video to be identified can be in formats such as MKV, MP4, or AVI, and this application is not limited to these formats. The video to be identified can be sent to the electronic device by a server (or other multimedia device) connected to the electronic device, or it can be a video obtained by the electronic device from a website via a network address.

[0052] The video title to be identified can be a name or descriptive information added by the video publisher or editor. For example, the movie title added by the publisher of a movie video can be used as the video title, and the descriptive information added by the publisher of a short video to describe the content of the short video can also be used as the video title. Alternatively, the platform can extract the title information based on the descriptive information or video content. For example, the video title for the movie "Kung Fu" is "Kung Fu," and the short video's description is "Teaching you how to learn 300 words in 3 minutes," so the corresponding video title would be "Learn 300 words in 3 minutes."

[0053] In this application, object interaction data may refer to user interaction data, which may be interactive data posted by users in response to the video to be identified, such as bullet comments and other comments. Object interaction data may include at least one of the following: characters, images, emoticons, etc.

[0054] In this application, object interaction data can refer to the text data corresponding to the original object interaction data. When the original object interaction data includes non-text data, the non-text data needs to be converted into the corresponding text data, and the converted result is used as the object interaction data of this application. For example, the original non-text data can be text information converted from the object's voice interaction data. For example, when the object is a user, the object's voice interaction data can be voice comments or voice bullet comments posted by the user on the video.

[0055] When the original object interaction data includes characters, the object interaction data can be derived from the original object interaction data itself. When the original object interaction data includes emoticons, if the emoticons contain characters, the text can be extracted from the emoticons using Optical Character Recognition (OCR, the process of converting text in a file to be recognized into text format through character recognition), and the extracted text can be used as object interaction data. When the original object interaction data includes emoticons, if the emoticons do not contain characters, the emoticon content information (e.g., crying, happy, sad, etc.) can be obtained, and the emoticon content information can be used as object interaction data. When the object interaction data includes images, text can be extracted from the images using optical character recognition technology and used as object interaction data, or object recognition can be performed directly on the images, and the names of the recognized objects can be used as object interaction data.

[0056] In this application, when identifying a specific video, all comments and bullet comments can be obtained as object interaction data. In other embodiments, when identifying a specific video, all comments and bullet comments can be obtained, and then the obtained comments and bullet comments are filtered to select a portion of bullet comments and comments with more characters, fewer repeated characters, or fewer non-characters as object interaction data.

[0057] If a video to be identified has been previously identified using the method of this application, and as time passes and pageviews change, the number of object interaction data for that video increases with object-related actions (such as sending bullet comments and sending comments), then when the number of newly added object interaction data exceeds a preset update count, all object interaction data can be retrieved, and the identification process for the video to be identified can be repeated according to the method of this application. The preset update count can be set based on the length of the video to be identified, the number of views of the video to be identified, or user needs.

[0058] S120. Filter relevant interaction data related to target information from the object interaction data, wherein the target information is used to indicate the video content of the video to be identified, and / or, the target information is used to indicate that the video to be identified is the target video.

[0059] Target information can be used to indicate the video content of the video to be identified, or it can indicate that the video to be identified is a target video. Target information can also refer to both the video content of the video to be identified and the video to be identified as a target video. A target video can refer to a clickbait video.

[0060] It should be noted that when the target information indicates the video content of the video to be identified, the target information and the content of the video to be identified have a very high degree of consistency. The target information can be used as substitute data for the content of the video to be identified. When the target information indicates that the video to be identified is the target video, the probability that the video to be identified is the target video is high, and the probability that the video to be identified is not the target video is low.

[0061] When target information is used to indicate the video content of the video to be identified, the target information can be obtained based on the video description information (e.g., video synopsis), subtitle data, and dialogue data of the video to be identified. When target information is used to indicate that the video to be identified is a target video, the target information can be obtained based on the target video interaction data corresponding to the video identified as the target video. For example, target video interaction data can be bullet screen data such as "This title is too exaggerated," "This title and video are completely irrelevant," and "This title is too fake."

[0062] When the target information is used to indicate the video content of the video to be identified, the filtered relevant interaction data can be object interaction data that is highly relevant to the video content; when the target information is used to indicate that the video to be identified is the target video, the filtered relevant interaction data can be object interaction data that indicates that the video to be identified is the target video; when the target information indicates both the video content of the video to be identified and the video to be identified as the target video, the filtered relevant interaction data includes object interaction data that is highly relevant to the video content and object interaction data that indicates that the video to be identified is the target video.

[0063] S130. Based on the relevant interaction data, determine the correlation between the video to be identified and the video title.

[0064] When all relevant interactive data are object interactive data that are highly relevant to the video content, for any one relevant interactive data, the relevance of that relevant interactive data to the video title is determined as a sub-relevance. Then, the sub-relevances of all relevant interactive data are summarized (the summarization method can be a weighted summation algorithm, etc.) to obtain the relevance summary result. The relevance summary result is used as the relevance between the video to be identified and the video title.

[0065] When all relevant interactive data are object interactive data indicating that the video to be identified is the target video, for any one relevant interactive data, when the relevant interactive data is object interactive data indicating that the video to be identified is the target video, the probability that the relevant interactive data indicates that the video to be identified is the target video is determined, and then the probabilities of each relevant interactive data are summarized (the summarization method can be a weighted sum or other algorithm) to obtain the probability summary result, and the probability summary result is used as the relevance between the video to be identified and the video title.

[0066] When the relevant interactive data includes both object interactive data that is highly relevant to the video content and object interactive data that indicates the video to be identified as the target video, for each relevant interactive data that is highly relevant to the video content, the relevance between the relevant interactive data and the video title is determined as a sub-relevance. Then, all the sub-relevances of the relevant interactive data that are highly relevant to the video content are summarized (the summarization method can be a weighted sum or other algorithm) to obtain the relevance summary result. For any relevant interactive data that indicates the video to be identified as the target video, the probability of the relevant interactive data indicating the video to be identified as the target video is determined. Then, the probabilities of each relevant interactive data indicating the video to be identified as the target video are summarized (the summarization method can be a weighted sum or other algorithm) to obtain the probability summary result. Finally, the relevance summary result and the probability summary result are summarized (using a weighted sum or other algorithm) to obtain the total relevance. This total relevance is used as the relevance between the video to be identified and the video title. The higher the total relevance, the more similar the video content and the video title are, and the lower the probability that the video to be identified is the target video.

[0067] In some implementations, for each relevant interactive data that has a high relevance to the video content, the text features of the relevant interactive data can be determined as the first text feature, and the text features of the video title can be determined as the second text feature. The first and second text features are then fused, and the fused result is analyzed to obtain the relevance between each relevant interactive data and the video title, which is taken as a sub-relevance. Then, the sub-relevances corresponding to all object interactive data are summarized (the summarization method can be a weighted summation algorithm, etc.) to obtain the relevance summary result. The relevance summary result is used as the relevance between the video to be identified and the video title. The higher the relevance summary result, the more similar the relevant interactive data and the video title are, indicating that the video content of the video to be identified is more similar to the video title, and the lower the probability that the video to be identified is the target video.

[0068] It should be noted that, in one possible implementation, a trained language model can be used to obtain the text features of relevant interactive data and video titles. The language model can be, for example, a pre-trained model (Bidirectional Encoder Representation from Transformers, BERT), a Bidirectional Long Short-Term Memory (BiLSTM) network model, a Bidirectional Gated Recurrent Unit (BiGRU) network model, or other neural network models; this embodiment does not limit the specific model used.

[0069] In one possible implementation, the text features of relevant interactive data can be determined by segmenting the relevant interactive data into words to obtain word vectors for each word, then inputting the word vector sequence into a language model, and using the output of the language model as the text features. Similarly, the text features of video titles can be determined by segmenting the video titles into words to obtain word vectors for each word, then inputting the word vector sequence into a language model, and using the output of the language model as the text features.

[0070] Alternatively, the result of fusing the first and second text features can be input into the trained classification model to obtain the sub-relevance of each relevant interaction data with the video title. Then, the sub-relevances of all object interaction data can be weighted and summed to obtain the relevance summary result.

[0071] In other implementations, for each piece of relevant interactive data indicating that the video to be identified is the target video, the first text feature corresponding to the relevant interactive data can be input into the trained title hit analysis model to obtain the probability corresponding to that piece of relevant interactive data output by the title hit analysis model (this probability represents the probability that the relevant interactive data indicates that the video to be identified is the target video). Then, the probabilities corresponding to all relevant interactive data are summed to obtain the probability summary result. The higher the probability summary result, the higher the probability that the video to be identified is the target video. The title hit analysis model can be obtained by training a neural network model (e.g., BiLSTM, BiGRU, BERT, etc.) with training samples. The training samples can include object interaction data as samples and the relevance between the object interaction data as samples and the video title.

[0072] Video recognition based on the selected relevant interaction data can greatly reduce data processing time and improve video analysis efficiency.

[0073] S140. Based on the comparison results of the correlation and the correlation threshold, determine whether the video to be identified is the target video.

[0074] When the relevant interaction data only includes object interaction data that is highly relevant to the video content, the total correlation result obtained by aggregating the sub-correlation of each relevant interaction data is lower than the corresponding first threshold, and the video to be identified is determined to be the target video. When the relevant interaction data only includes object interaction data that indicates that the video to be identified is the target video, the total probability result obtained by aggregating the probabilities of each relevant interaction data reaches the corresponding second threshold, and the video to be identified is determined to be the target video. When the relevant interaction data includes both object interaction data that is highly relevant to the video content and object interaction data that indicates that the video to be identified is the target video, the total correlation result obtained by aggregating the sub-correlation of each object interaction data that is highly relevant to the video content is determined, and the total probability result obtained by aggregating the probabilities of each object interaction data that indicates that the video to be identified is the target video is determined. The total correlation result and the total probability result are then aggregated to obtain the total correlation. If the total correlation does not reach the corresponding third threshold, the video to be identified is determined to be the target video.

[0075] The three thresholds mentioned above are different and can be determined based on the video content of the video to be identified, the amount of data of object interaction data, and the requirements.

[0076] The video recognition method provided in this embodiment acquires object interaction data and video title of the video to be recognized; filters relevant interaction data related to target information from the object interaction data, where the target information indicates the video content of the video to be recognized, and / or indicates that the video to be recognized is a target video; determines the relevance between the video to be recognized and the video title based on the relevant interaction data; and determines whether the video to be recognized is a target video based on the comparison result of the relevance and the relevance threshold. Relevant interaction data can more accurately reflect the relevance between the video content and the video title of the video to be recognized, thereby improving the accuracy of target video recognition. The above method fully utilizes object interaction data, such as comments and bullet comments. User comments and bullet comments are user feedback based on the current video content, which can not only supplement the representation ability of the video content, but also directly reflect the tendency of clickbait titles. By mining information from object interaction data, the modeling and understanding effect of video content can be improved while ensuring computational efficiency. Furthermore, combining object interaction data with direct target video recognition further improves the recognition effect, providing effective support for improving the viewing experience of video platform users. Meanwhile, since the amount of data in object interaction data is generally less than that in video content, the amount of data processing in the video analysis process is reduced, thus improving the efficiency of video analysis.

[0077] Please see Figure 3 , Figure 3 This application illustrates a flowchart of a video recognition method according to another embodiment of the present application. This method can be applied to an electronic device, which can be... Figure 1 The method, which includes the terminal 101 or server 102 shown, comprises:

[0078] S210. Obtain the object interaction data and video title of the video to be identified.

[0079] The description of S210 is the same as that of S110, and will not be repeated here.

[0080] S220. Obtain the video content data of the video to be identified, and use it as the target information.

[0081] In this application, the video content data of the video to be identified may include subtitle data and dialogue data. Subtitle data can be extracted from the video frames using OCR technology, and dialogue data can be obtained by recognizing the dialogue in the video using Automatic Speech Recognition (ASR) technology.

[0082] It is understandable that some video frames in the video to be identified correspond to subtitle data. OCR technology can extract subtitle data from the video frames that carry subtitle data, without needing to extract subtitles from all video frames. At the same time, some subtitles may exist in multiple different video frames. For multiple video frames carrying the same subtitle, only one video frame needs to be extracted for subtitle extraction.

[0083] In some implementations, the video to be identified can have built-in or external subtitle data, and the built-in or external subtitle data can be obtained directly without extracting subtitle data from the video frames.

[0084] A single subtitle or a line of dialogue can be considered as a video content data set, and a video to be identified can correspond to multiple video content data sets.

[0085] S230. Filter the interaction data related to the video content data from the object interaction data, and use it as the relevant interaction data.

[0086] From all object interaction data in the video to be identified, object interaction data that is relevant to the video content data is selected as relevant interaction data. Relevant interaction data has a high similarity to the video content data and can be used as data to indicate the video content of the video to be identified.

[0087] Relevant interaction data can be filtered based on its relevance to the video content, its engagement level, or a combination of both. Engagement level represents the user's attention to and / or acceptance of the data. For example, relevant interaction data might be those with high likes and / or replies. Higher likes and / or replies indicate higher engagement, meaning the data better reflects the specific content of the video being identified, and thus, a higher relevance between the interaction data and the video content.

[0088] S240. Based on the video title, determine the title relevance probability corresponding to the relevant interaction data, wherein the title relevance probability is used to characterize the relevance between the video to be identified and the video title.

[0089] After obtaining the relevant interaction data, the title relevance probability corresponding to the interaction data is determined based on the video title. This title relevance probability reflects the relevance between the video to be identified and the video title. A higher title relevance probability indicates a higher relevance between the interaction data and the video title, a better match between the video content data and the video title, and a greater relevance between the video to be identified and the video title. Conversely, a lower title relevance probability indicates a lower relevance between the interaction data and the video title, a less favorable match between the video content data and the video title, and a less relevant relationship between the video to be identified and the video title.

[0090] In some implementations, the relevant interaction data may include multiple data sets. For each relevant interaction data set, the corresponding title relevance probability is determined based on the video title. Then, the title relevance probabilities of all relevant interaction data sets can be weighted and summed, or the maximum value can be taken to determine the final title relevance probability. This final title relevance probability serves as a characterization of the relevance between the video to be identified and the video title.

[0091] S250. When the title relevance probability is less than the relevance threshold, the video to be identified is determined to be the target video.

[0092] A lower title relevance probability indicates a lower correlation between the relevant interaction data and the video title, a greater mismatch between the video content data and the video title, and a less relevant relationship between the video to be identified and the video title. When the title relevance probability is below the relevance threshold, it indicates a low relevance between the relevant interaction data and the video title, a mismatch between the video content data and the video title, and the video to be identified is the target video, i.e., a clickbait video. The relevance threshold can be determined based on requirements or the video content.

[0093] Similarly, when the title relevance probability reaches the relevance threshold, it indicates that the relevant interaction data is highly correlated with the video title, the video content data matches the video title, and the video to be identified is not the target video, that is, the video to be identified is not a clickbait video.

[0094] In this embodiment, relevant interactive data related to the video content data is selected as the basis for identifying whether the video to be identified is the target video. The relevant interactive data accurately reflects the specific content of the video to be identified, and the correlation of the relevant interactive data can accurately reflect the correlation between the video to be identified and the video title, thereby improving the identification accuracy of the video to be identified.

[0095] At the same time, only the relevant interaction data that has been selected is analyzed, without having to process all the object interaction data. This reduces the amount of object interaction data to be processed, improves data processing efficiency, and thus improves video recognition efficiency.

[0096] Please see Figure 4 , Figure 4 A flowchart illustrating a filtering method for first interactive data in an embodiment of this application is shown. This method can be used to obtain the first interactive data described in the above embodiments. The method can be used with an electronic device, which can be... Figure 1 The method for the terminal 101 or server 102 shown includes:

[0097] S310. Based on the video content data, determine the content relevance probability corresponding to the object interaction data, wherein the content relevance probability is used to characterize the correlation between the object interaction data and the video content data.

[0098] Different object interaction data have varying degrees of correlation with video content data. Based on the video content data, the content relevance probability corresponding to the object interaction data is determined. For a given object interaction data, a higher content relevance probability indicates a higher correlation between the object interaction data and the video content data; conversely, a lower content relevance probability indicates a lower correlation between the object interaction data and the video content data.

[0099] When there are multiple object interaction data, the content relevance probability of each object interaction data can be determined, thereby obtaining the content relevance probability of all object interaction data.

[0100] S320. Determine the interaction heat of the object interaction data, wherein the interaction heat is used to characterize the interactivity of the object interaction data itself.

[0101] Each object's interaction data corresponds to its own interaction popularity. For each object's interaction data, the higher the interaction popularity, the higher the attention and recognition the object's interaction data receives from the object, and the better the object's interaction data matches the video content data. Conversely, the lower the interaction popularity, the lower the attention and recognition the object's interaction data receives from the object, and the less the object's interaction data matches the video content data.

[0102] When there are multiple object interaction data, the interaction popularity of each object interaction data can be determined, thereby obtaining the interaction popularity of all object interaction data.

[0103] S330. Determine the key information probability of the object interaction data based on the content relevance probability and the interaction popularity.

[0104] For each object interaction data point, the key information probability is determined based on its content relevance probability and interaction popularity. This key information probability characterizes the importance of the object interaction data within all object interaction data corresponding to the video to be identified. A higher key information probability indicates greater importance, while a lower probability indicates less importance.

[0105] In the specific implementation of this application, for each object interaction data, the product of the content relevance probability of the object interaction data and the interaction popularity can be calculated, and the product can be used as the key information probability of the object interaction data, thereby obtaining the key information probability of each of all object interaction data.

[0106] S340. Select the top N object interaction data with the highest probability of key information from the object interaction data, and use them as the relevant interaction data, where N is a positive integer.

[0107] Since higher probability of key information indicates more important object interaction data, and lower probability indicates less important object interaction data, the relevant interaction data to be filtered are the top N object interaction data with higher probability of key information. Users can further determine the specific value of N based on actual needs, the length of the video to be identified, and other information. Specifically, it can be the top N object interaction data with a key information probability exceeding a preset key information probability, and the preset key information probability can be set based on needs; this application does not impose any limitations on this.

[0108] In this embodiment, the importance of object interaction data is represented by the probability of key information. Object interaction data with higher importance is selected as relevant interaction data, so that the relevant interaction data is more consistent with the video content of the video to be identified, thereby further improving the accuracy of the correlation between object interaction data and video title.

[0109] Please see Figure 5 , Figure 5 The flowchart illustrates a method for obtaining interaction popularity in an embodiment of this application. This method can be used to obtain the interaction popularity in the above embodiments. The method can be used in an electronic device, which can be... Figure 1 The method for the terminal 101 or server 102 shown includes:

[0110] S410. Obtain the number of likes and replies for each of the object's interaction data.

[0111] In this application, the object interaction data of the video to be identified includes at least one. For each object interaction data, there may be a like button, which can be clicked to like the object interaction data. The number of likes for the object interaction data can be displayed around the like button. The number of likes can be the total number of likes for the object interaction data, or the number of likes for the object interaction data within a fixed period of time (e.g., one month). This application does not limit this.

[0112] Each interactive data point can also include a reply box and a send button. Users can edit a reply in the reply box and then click the corresponding send button to send the reply. Each sent reply is considered a single reply. A single reply can include multiple paragraphs of text, with no limit on the number of characters. For each reply, each time a non-empty reply is edited and the send button is clicked, that message is considered a single reply. The reply count can refer to the total number of replies sent for each interactive data point.

[0113] S420. Determine the overall interaction volume of the at least one object interaction data based on the number of likes and replies for each object interaction data.

[0114] The interaction volume of each object's interaction data can be determined based on the number of likes and replies. Specifically, the interaction volume of each object's interaction data can be obtained by weighted summing of the number of likes and replies. The weights of likes and replies can be set based on requirements, for example, both can be 0.5.

[0115] Then, sum the individual interaction amounts of all object interaction data to obtain the overall interaction amount of all object interaction data.

[0116] S430. The number of likes and the number of replies for each object interaction data are weighted and summed to obtain the weighted sum for each object interaction data.

[0117] Specifically, the number of likes and replies for each object's interaction data is weighted and summed to obtain the weighted sum for each object's interaction data. This weighted sum can also be the interaction volume of the object's interaction data itself.

[0118] S440. Calculate the ratio of the weighted sum of each object interaction data to the overall interaction volume to obtain the interaction popularity corresponding to each object interaction data.

[0119] After determining the weighted sum of each object interaction data point, the weighted sum is compared with the determined overall interaction volume. The ratio is used as the interaction intensity of that object interaction data point. For all object interaction data points, the interaction intensity of each individual object interaction data point is determined. Alternatively, the interaction intensity of that object interaction data point can be the ratio of its own interaction volume to the overall interaction volume.

[0120] Please see Figure 6 , Figure 6 A flowchart illustrating a method for obtaining content relevance probability in an embodiment of this application is shown. This method can be used to obtain the content relevance probability in the above embodiments. The method is applied to an electronic device, which can be... Figure 1 The method for the terminal 101 or server 102 shown includes:

[0121] S510. Determine the media features of the video to be identified based on the video content data.

[0122] The video content data can be input into a trained media feature analysis model to obtain the text features of the corresponding video content data output by the media feature analysis model, which are then used as media features. The media feature analysis model can be a pre-trained model (BERT model), a bidirectional long short-term memory network model (BiLSTM model), a bidirectional gated recurrent unit network model (BiGRU model), or other neural network models; this embodiment does not limit the specific model used.

[0123] For each piece of video content data, the video content data can be input into the media feature analysis model to obtain the media features of the video content data.

[0124] S520. Obtain the text features of the object interaction data.

[0125] The interaction data can be input into a trained interaction data analysis model to obtain the text features of the corresponding interaction data output by the interaction data analysis model, which are then used as media features. The media feature analysis model can be a pre-trained model (BERT model), a bidirectional long short-term memory network model (BiLSTM model), a bidirectional gated recurrent unit network model (BiGRU model), or other neural network models; this embodiment does not limit the specific model used.

[0126] For each object interaction data, the object interaction data can be input into the interaction data analysis model to obtain the text features of the object interaction data.

[0127] S530. The text features and the media features are fused to obtain the first fused feature.

[0128] The media features of each video content data point and the text features of each object interaction data point can be fused separately to obtain the corresponding first fused features. For example, if three video content data points correspond to three media features and three object interaction data points correspond to three text features, the resulting first fused features will include nine first fused features.

[0129] Specifically, the fusion of media features of video content data and text features of object interaction data can be achieved through a bidirectional attention mechanism. Specifically, attention is applied to the media features through the text features of the object interaction data, and attention is applied to the text features of the object interaction data through the media features. The two attention representations are then concatenated as the first fused feature.

[0130] S540. Analyze the first fusion feature using the first correlation analysis model to obtain the content relevance probability of the corresponding object interaction data.

[0131] The first correlation analysis model includes a normalized index (softmax) layer. The softmax layer can perform binary classification on the first fused features to obtain the output probability, which is used as the content relevance probability.

[0132] The first correlation analysis model is pre-trained. One possible training method for the first correlation analysis model is to obtain the textual features of the object interaction data and the media features of the video content data of the target sample video, as well as their correlation. The target sample video includes samples where the object interaction data and video content data are related; their label is "related" and the corresponding category probability is 1. The target sample video also includes samples where the object interaction data and video content data are not related; these samples have no label or are labeled "unrelated" and the corresponding category probability is 0. The first correlation analysis model is trained based on the textual features of the object interaction data and the media features of the video content data of the target sample video, as well as their labels.

[0133] Please see Figure 7 , Figure 7 A schematic diagram illustrating the process of obtaining content relevance probability in an embodiment of this application is shown.

[0134] The video content data is analyzed using a media feature analysis model to obtain corresponding media features. Then, the interactive data (each individual object interaction data) is analyzed using an interactive data analysis model to obtain corresponding text features. The obtained media features and the text features of the object interaction data are then fused (feature fusion can refer to the fusion through the bidirectional attention mechanism mentioned above) to obtain the first fused feature. Finally, the first fused feature is analyzed using a first correlation analysis model to obtain the content relevance probability of the corresponding object interaction data output by the first correlation analysis model.

[0135] In this embodiment, the text features of object interaction data and the media features of video content data are fused together. Then, the fused features are analyzed to obtain the content relevance probability, which represents the correlation between video content and object interaction data. This allows for the determination of key information probabilities with high accuracy based on the content relevance probability.

[0136] Please see Figure 8 , Figure 8 A flowchart illustrating a method for obtaining title relevance probability in an embodiment of this application is shown. This method can be used to obtain the title relevance probability in the above embodiments. The method is applied to an electronic device, which can be... Figure 1 The method for the terminal 101 or server 102 shown includes:

[0137] S610. Obtain the key information probability of the relevant interaction data, wherein the key information probability is used to characterize the importance of the relevant interaction data in all object interaction data corresponding to the video to be identified.

[0138] The method for obtaining the key information probability of relevant interactive data is described in S330 above and will not be repeated here.

[0139] S620. Based on the relevant interaction data and the probability of the key information, filter the target interaction data from the relevant interaction data.

[0140] The relevant interactive data can be selected from multiple sources. Each of the multiple relevant interactive data points corresponds to a probability of its own key information. Based on the multiple relevant interactive data points and their corresponding key information probabilities, a target interactive data point is selected from the relevant interactive data.

[0141] This can be done by determining the text features of the relevant interactive data (the method for determining the text features of the object interactive data is the same as that for determining the text features of the object interactive data, and will not be repeated here), then multiplying the text features of the interactive data with the corresponding key information probabilities to obtain the processed text features, and then performing MaxPooling operation on the processed text features to obtain the final text features. The relevant interactive data corresponding to the final text features is used as the target interactive data.

[0142] S630. Obtain the title features of the video title and the text features of the target interactive data.

[0143] The video title can be input into a video title analysis model to obtain the text features of the corresponding video title output by the model, which are then used as title features. The video title analysis model can be a pre-trained model (BERT model), a bidirectional long short-term memory network model (BiLSTM model), a bidirectional gated recurrent unit network model (BiGRU model), or other neural network models; this embodiment does not limit the specific model used.

[0144] The method for obtaining the text features of the target interaction data is described in S620 and will not be repeated here.

[0145] S640. The title feature and the text feature of the target interaction data are fused to obtain a second fused feature.

[0146] One approach is to fuse the title features and the text features of the target interaction data using a bidirectional attention mechanism to obtain a fused feature, which serves as the second fused feature. Specifically, attention is applied to the title features using the text features of the target interaction data, and attention is also applied to the text features of the target interaction data using the title features. The two attention representations are then concatenated to obtain the second fused feature.

[0147] S650. Analyze the second fusion feature using the second correlation analysis model to obtain the title relevance probability corresponding to the relevant interactive data.

[0148] The first correlation analysis model includes a normalized index (softmax) layer. The softmax layer can perform binary classification on the second fusion feature to obtain the output probability, which is used as the title relevance probability.

[0149] The first correlation analysis model is pre-trained, and its training method is the same as that for the first correlation analysis model, so it will not be repeated here. The training samples for the second correlation analysis model can include new target video samples. These new target video samples include those where user interaction data is related to the video title; their label is "relevant," and their corresponding category probability is 1. The new target video samples also include those where user interaction data is unrelated to the video title; these samples either have no label or are labeled "unrelated," and their corresponding category probability is 0.

[0150] Please see Figure 9 , Figure 9 A schematic diagram illustrating the process of obtaining the title-related probability in an embodiment of this application is shown.

[0151] M relevant interactive data points are selected based on the probability of key information from each relevant interactive data point, where M is an integer greater than 1. Each relevant interactive data point is analyzed using an interactive data analysis model to obtain corresponding text features. These text features are then multiplied by the probability of key information from each relevant interactive data point to obtain processed text features. Target interactive data is then selected from these processed text features, using a method such as MaxPooling.

[0152] The text features of the video title are determined based on the video title analysis model and used as title features. The title features are then fused with the text features of the determined target interactive data (feature fusion can refer to the fusion through the bidirectional attention mechanism mentioned above) to obtain the second fused feature. The second fused feature is then input into the second relevance analysis model to obtain the title relevance probability of the corresponding relevant interactive data output by the second relevance analysis model.

[0153] In this embodiment, the target interactive data most relevant to the video content is selected from the relevant interactive data. The text features of the target interactive data are fused with the title features of the video title. Then, the fused features are analyzed to obtain the title relevance probability, which represents the relevance between the object interactive data and the video title. The title relevance probability has a high accuracy. At the same time, the target interactive data representing the object interactive data has a high relevance to the video content, so the title relevance probability can accurately reflect the relevance between the video content and the video title.

[0154] Please see Figure 10 , Figure 10 A flowchart of a video recognition method provided in another embodiment of this application is shown; the method can be used in an electronic device, which can be... Figure 1 The method for the terminal 101 or server 102 shown includes:

[0155] S710: Obtain object interaction data and video title of the video to be identified.

[0156] The description of S710 is the same as that of S110, and will not be repeated here.

[0157] S720. Obtain target keywords as target information. The target keywords are obtained based on sample interaction data with labeled information. The labeled information is used to characterize that the sample interaction data indicates that the sample video is the target video. The sample video is the video corresponding to the sample interaction data.

[0158] S730. Filter out the interaction data that matches the target keyword from the object interaction data, and use it as relevant interaction data.

[0159] Multiple target keywords can be included, and these target keywords can form a keyword table containing all target keywords. Each object's interaction data is compared with the keyword table to determine if the interaction data matches a specific target keyword in the table. If a match is found, the object's interaction data is considered relevant interaction data.

[0160] A sample video can refer to a video identified as the target video, i.e., a clickbait video. The sample video includes a large amount of object interaction data, among which there are object interaction data annotated with annotation information. This annotated object interaction data serves as sample interaction data. The annotation information indicates that the sample interaction data indicates the corresponding sample video is the target video. The annotation information can be any identifier, including numbers, letters, or a combination of numbers and letters. For example, a comment with the content "This title is really exaggerated" can be annotated with annotation information 1, where annotation information 1 indicates that the comment indicates the video is a clickbait video. Conversely, a comment with the content "This title is too accurate" can be annotated with annotation information 0, where annotation information 0 indicates that the comment indicates the video is not a clickbait video.

[0161] Target keywords are determined based on keywords in the sample interaction data of the sample video. Specifically, the method for determining target keywords includes: acquiring the sample interaction data; extracting keywords from the sample interaction data to obtain multiple preliminary keywords; determining the frequency of occurrence of the multiple preliminary keywords in the sample interaction data; and selecting target keywords from the multiple preliminary keywords whose frequency of occurrence exceeds a preset frequency. The preset frequency can be determined based on the number of multiple preliminary keywords, and this application does not limit it.

[0162] The relevant interactive data that matches the target keywords of the clickbait video's interactive data indicates that the video to be identified is a clickbait video.

[0163] S740. Determine the title hit probability corresponding to the relevant interactive data, wherein the title hit probability is used to characterize the correlation between the video to be identified and the video title.

[0164] The determined title hit probability expresses the probability that the relevant interaction data indicates the video to be identified is the target. The higher the title hit probability, the higher the probability that the relevant interaction data indicates the video to be identified is the target, and the higher the probability that the video to be identified is the target video. Conversely, the lower the title hit probability, the lower the probability that the relevant interaction data indicates the video to be identified is a clickbait video, and the lower the probability that the video to be identified is the target video.

[0165] S750. When the title hit probability exceeds the relevance threshold, the video to be identified is determined to be the target video.

[0166] When the title hit probability exceeds the relevance threshold, the title hit probability is high, indicating that the relevant interaction data indicates a high probability that the video to be identified is the target video, and the video to be identified is the target video. When the title hit probability does not exceed the relevance threshold, the title hit probability is low, indicating that the relevant interaction data indicates a low probability that the video to be identified is the target video, and the video to be identified is not the target video. The relevance threshold value can be set based on requirements, such as 0.3.

[0167] In this embodiment, a portion of the object interaction data is selected as relevant interaction data, and only the relevant interaction data is analyzed. This reduces the amount of data analysis required for object interaction data, improves data analysis efficiency, and thus improves video recognition efficiency.

[0168] Please see Figure 11 , Figure 11 A flowchart illustrating a method for obtaining title hit probability in an embodiment of this application is shown. This method can be used to obtain the title hit probability in the above embodiments. The method can be used in an electronic device, which can be... Figure 1 The method for the terminal 101 or server 102 shown includes:

[0169] S810. Determine the text features and interaction popularity of each of the relevant interactive data, wherein the interaction popularity is used to characterize the interactivity of the relevant interactive data itself.

[0170] The text features of the relevant interactive data are obtained by referring to the method for obtaining the text features of the object's interactive data mentioned above, and will not be repeated here.

[0171] The interaction volume of each relevant interactive data point can be determined based on its number of likes and replies. Specifically, this can be achieved by weighted summing of the likes and replies for each relevant interactive data point. The weights for likes and replies can be set based on requirements, such as both being 0.5. Then, the interaction volumes of all object interactive data points are summed to obtain the overall interaction volume of all relevant interactive data points.

[0172] Calculate the ratio of the weighted sum of each relevant interactive data point to the total interaction volume of the relevant interactive data to obtain the interaction popularity of each relevant interactive data point. After determining the weighted sum of each relevant interactive data point, compare this weighted sum to the total interaction volume of the relevant interactive data; the ratio is used as the interaction popularity of that relevant interactive data point. For all relevant interactive data points, determine the interaction popularity of each individual interactive data point. Alternatively, the interaction popularity of that relevant interactive data point can be determined as the ratio of its own interaction volume to the total interaction volume of the relevant interactive data point.

[0173] S820. Analyze the text features of each relevant interactive data using the title hit analysis model to obtain the initial title hit probability of the corresponding relevant interactive data.

[0174] The text features of each relevant interaction data point are input into the title hit analysis model to obtain the initial title hit probability corresponding to that relevant interaction data, as output by the model. The title hit analysis model can be obtained by training a neural network model, such as a pre-trained model (BERT model), a Bidirectional Long Short-Term Memory network model (BiLSTM model), or a Bidirectional Gated Recurrent Unit network model (BiGRU model). Training samples can include interaction data indicating that a video is a clickbait video and its corresponding class probability 1; training samples can also include interaction data indicating that a video is not a clickbait video and its corresponding class probability 0.

[0175] S830. The title hit probability is obtained based on the initial title hit probability and interaction popularity of each of the relevant interactive data.

[0176] The initial title hit probability and interaction popularity of each relevant interactive data can be multiplied, and then the products of all relevant interactive data can be weighted and summed to obtain the final title hit probability. The weight of the product of each relevant interactive data can be determined according to the interaction popularity of each relevant interactive data, and the weight of the product of each relevant interactive data can also be determined according to different user needs. This application does not impose any restrictions.

[0177] Please see Figure 12 , Figure 12 A schematic diagram illustrating the process of obtaining the title hit probability in an embodiment of this application is shown.

[0178] After filtering relevant interaction data from the object interaction data, each relevant interaction data is input into the interaction data analysis model to obtain the text features corresponding to the relevant interaction data output by the interaction data analysis model. Following the interaction popularity determination method for object interaction data in this application, the interaction popularity of the relevant interaction data is determined. The text features of the relevant interaction data are then input into the title hit analysis model to obtain the initial title hit probability corresponding to the relevant interaction data. The product of the initial title hit probability and the interaction popularity of each determined relevant interaction data is calculated, and the weighted sum of the products of all relevant interaction data is obtained to obtain the final title hit probability.

[0179] Relevant interaction data is the most telling indicator of whether a video is clickbait. This allows the probability of the title matching the determined relevant interaction data to accurately reflect the probability that the video to be identified is clickbait, thus improving the accuracy of the identification results.

[0180] Please see Figure 13 , Figure 13 The flowchart shown is a video recognition method according to another embodiment of this application. The method can be used in an electronic device, which can be... Figure 1 The method for the terminal 101 or server 102 shown includes:

[0181] S910: Obtain object interaction data and video title of the video to be identified.

[0182] S920. Obtain the video content data and target keywords of the video to be identified as the target information. The target keywords are obtained based on the sample interaction data with labeled information. The labeled information is used to characterize that the sample interaction data indicates that the sample video is the target video. The sample video is the video corresponding to the sample interaction data.

[0183] S930. Filter the first interaction data related to the video content data from the object interaction data.

[0184] S940. Select the interaction data that hits the target keyword from the object interaction data and use it as the second interaction count.

[0185] S950, The first interactive data and the second interactive data are used as relevant interactive data.

[0186] S960. Based on the video title, determine the title relevance probability corresponding to the first interactive data, wherein the title relevance probability is used to characterize the relevance between the video to be identified and the video title.

[0187] S970. Determine the title hit probability corresponding to the second interactive data, wherein the title hit probability is used to characterize the correlation between the video to be identified and the video title.

[0188] The descriptions of S910, S920, S930, S940, S960, and S970 are all described in accordance with the descriptions of S220 and S720, respectively. The descriptions of S930, S940, S960, S970, and S970 are all described in accordance with the descriptions of S740. These descriptions will not be repeated here.

[0189] In this embodiment, the relevant interactive data includes two types of data: first interactive data related to the video content data and second interactive data that hits the target keyword.

[0190] S980. The title hit probability and the title relevance probability are weighted and summed.

[0191] S990. Based on the comparison between the summation result and the correlation threshold, determine whether the video to be identified is the target video.

[0192] In this application, the weights of the title hit probability and the title relevance probability can be determined based on the actual video length, video category, and user needs of the video to be identified, and the weights of the title hit probability and the title relevance probability can be 1.

[0193] A higher weighted sum of the title hit probability and the title relevance probability indicates a lower probability that the video to be identified is the target video, and a lower weighted sum of the title hit probability and the title relevance probability indicates a higher probability that the video to be identified is the target video. A relevance threshold can be used to determine whether a video to be identified is the target video.

[0194] If the weighted sum of the title hit probability and the title relevance probability reaches the relevance threshold, the video to be identified is not the target video; otherwise, it is the target video. The relevance threshold can be 0.6, etc.

[0195] It should be noted that the correlation thresholds in S140, S250, S750 and S990 above can be different, and the corresponding correlation thresholds can be determined according to the requirements and different implementation scenarios.

[0196] Please see Figure 14 , Figure 14 A schematic diagram of the recognition process of the video to be recognized in an embodiment of this application is shown.

[0197] Determine the content relevance probability and interaction popularity of the corresponding object interaction data, further determine the key information probability of the object interaction data, filter the first interaction data based on the key information probability of the object interaction data, and then filter the second interaction data that hits the target keyword from the object interaction data.

[0198] Based on the determined key information probability of each first interactive data, target interactive data is filtered, and then the text features of the target interactive data are fused with the title features of the video title. The result of feature fusion is then input into the second relevance analysis model to obtain the title relevance probability.

[0199] Each identified second interactive data point is input into the interactive data analysis model to obtain the text features of the second interactive data. The text features of the second interactive data are then input into the title hit analysis model to obtain the initial title hit probability. The product of the initial title hit probability and the interaction popularity of each second interactive data point is calculated, and the products of each second interactive data point are weighted and summed to obtain the title hit probability.

[0200] Finally, the title hit probability and title relevance probability are weighted and summed to obtain the sum result. The sum result is then compared with the relevance threshold to determine the recognition result of the video to be identified: whether the video to be identified is the target video or not.

[0201] In this embodiment, the video to be identified is identified in two ways, and the two analysis results are combined to make the combined analysis results more accurate.

[0202] To better understand this solution, the video recognition method provided in this application embodiment will be illustrated with examples in a specific application scenario. In this scenario, the target video is a clickbait video.

[0203] 1. Obtain object interaction data from the video to be identified;

[0204] The object interaction data includes multiple data points (which can be multiple bullet comments, multiple comments, or a combination of bullet comments and comments). If the object interaction data of the video to be identified has not changed (no new object interaction data is added), it is not necessary to re-identify the video.

[0205] The amount of object interaction data obtained from the video to be identified is very large, possibly tens of thousands of comments and tens of thousands of bullet comments. At the same time, there is a lot of noise data in the published bullet comments and comments. It is necessary to filter the published comments, bullet comments and other object interaction data to filter out the first interaction data related to the video content data and the second interaction data that indicates the video to be identified as the target video.

[0206] 2. Filtering of the first interactive data;

[0207] First, dialogue in the video to be identified can be recognized using ASR technology, and subtitle data can be extracted from the video using OCR to obtain video content data. The video content data can include text data corresponding to the dialogue and subtitle data. Then, the video content data can be input into the corresponding BERT model for deep representation to obtain media features. The object interaction data can also be input into the corresponding BERT model for deep representation to obtain text features of the object interaction data. The media features and text features of the object interaction data are fused using an attention mechanism, and the fused result is input into the first correlation analysis model to obtain the content relevance probability. The content relevance probability is used to represent the correlation between the video content data and the object interaction data.

[0208] The number of likes and replies for each object's interaction data is counted, and the number of likes and replies for each object's interaction data is weighted and summed (the weight of each like and reply can be 0.5) to obtain the interaction volume of each object's interaction data. The interaction volumes of all object interaction data are summed to obtain the overall interaction volume of the video to be identified. Then, the ratio of the interaction volume of each object's interaction data to the overall interaction volume is used as the interaction popularity of that object's interaction data.

[0209] For a given object interaction data point, the product of its content relevance probability and its interaction popularity is used as the key information probability of that object interaction data point. After determining the key information probabilities of all object interaction data points, the top K object interaction data points with key information probabilities exceeding a preset key information probability and having relatively high key information probabilities are selected as the first interaction data points, where K is a non-zero integer.

[0210] 3. Filtering of the second interactive data;

[0211] The platform has labeled a large amount of object interaction data, indicating whether the object interaction data indicates that the video is a clickbait video. For example, comments such as "This title is simply too exaggerated" and "The title is completely irrelevant" indicate that the video is a clickbait video, while comments such as "The video is really beautiful" do not indicate that the video is a clickbait video. Keyword mining can be performed on this labeled object interaction data to extract high-frequency keywords (e.g., keywords that appear more frequently than a preset frequency) to construct target keywords. A keyword table is then built using these target keywords. This keyword table can be used to initially filter the object interaction data of the video to be identified, and the object interaction data that matches the target keywords in the keyword table is determined as the second interaction data.

[0212] 4. Determine the title relevance probability corresponding to the first interactive data;

[0213] For each first interaction data point, it is input into the corresponding BERT model to obtain the text features of each first interaction data point. The text features of the first interaction data point are multiplied with the corresponding key information probabilities to obtain the processed result. After obtaining the processed results of all first interaction data points, the MaxPooling algorithm is used to select one result from the processed results of all first interaction data points. The first interaction data point corresponding to the selected result is determined as the target interaction data point.

[0214] The video title is input into the corresponding BERT model to obtain the title features, and the text features of the target interaction data are obtained. The title features and the text features of the target interaction data are fused, and the fused result is input into the second correlation analysis model to obtain the corresponding title relevance probability.

[0215] When the title relevance probability is less than the corresponding threshold, the video to be identified can be determined as a clickbait video.

[0216] 5. Determine the title hit probability corresponding to the second interactive data;

[0217] Each second interaction data point is input into the corresponding BERT model to obtain its text features. Following the method described above for determining the interaction popularity of object interaction data, the interaction popularity of each second interaction data point is determined. The text features of the second interaction data are then input into the title hit analysis model to obtain the initial title hit probability corresponding to that second interaction data point. Next, the initial title hit probability of each second interaction data point is multiplied by its interaction popularity, and the products corresponding to all second interaction data points are weighted and summed (the weights of the products corresponding to each second interaction data point can be set according to requirements and are not limited), to obtain the title hit probability.

[0218] When the title relevance probability exceeds the corresponding threshold, the video to be identified can be determined as a clickbait video.

[0219] 6. Combine title hit probability and title relevance probability for video recognition;

[0220] Following the above method, the title relevance probability and title hit probability are obtained. The title relevance probability and title hit probability are then summed with weights (the weights of the title relevance probability and title hit probability can be set according to the video to be identified and user needs) to obtain the summation result. Based on the comparison result of the summation result with the preset relevance threshold, it is determined whether the video to be identified is a clickbait video.

[0221] When the summation result is less than the corresponding threshold, the video to be identified is determined to be a clickbait video; when the summation result is not less than the corresponding threshold, the video to be identified is determined not to be a clickbait video.

[0222] By mining the interactive data of the video to be identified, the first interactive data that can represent the video content is constructed, thereby improving the representation ability of the video content. At the same time, by combining the second interactive data that directly represents the video to be identified as a clickbait video, the recall and accuracy of clickbait video identification are improved. In particular, it can improve the identification effect of clickbait cases that are difficult to identify from video titles and video content alone.

[0223] Please see Figure 15 , Figure 15 This illustration shows a flowchart of a video recognition method provided in another embodiment of the present application. The method can be used in electronic devices and includes:

[0224] S1010: Obtain the object interaction data and video title of the video to be identified.

[0225] S1020. Filter relevant interaction data related to target information from the object interaction data, wherein the target information is used to indicate the video content of the video to be identified, and / or, the target information is used to indicate that the video to be identified is the target video.

[0226] S1030. Based on the relevant interaction data, determine the correlation between the video to be identified and the video title.

[0227] S1040. Based on the comparison results of the correlation and the correlation threshold, determine whether the video to be identified is the target video.

[0228] The descriptions of S1010-S1040 are the same as those of S110-S140, and will not be repeated here.

[0229] S1050. If it is determined that the video to be identified is the target video, then a prompt message is sent, or the distribution of the target video is stopped.

[0230] Once the video to be identified is the target video, which may provide a poor viewing experience, a prompt message can be sent to allow the user to decide whether to continue watching the video. Alternatively, the distribution of the video to be identified can be stopped directly, so that the user will not be able to watch the target video, thereby further improving the user's viewing experience.

[0231] In some possible implementations, the video server may have the video recognition method of this application built in. After the video server obtains the video to be recognized from other platforms or servers, it determines the video to be recognized as the target video according to the method of this application, adds prompt information to the video to be recognized that is determined to be the target video, and sends the video to be recognized with the added prompt information to the video platform of the video server, so that when watching the video to be recognized with the added prompt information through the video platform, the prompt information can be used to determine that the video is the target video.

[0232] In some other possible implementations, after the video server determines that the video to be identified is the target video according to the method of this application, it can directly stop distributing the video to be identified, so that the user will not be able to watch the target video.

[0233] In this embodiment, after determining that the video to be identified is the target video, a prompt message is output or the distribution of the video to be identified is stopped. This can assist the video platform in making distribution decisions, thereby reducing the possibility that the user will watch the target video and improving the user's video viewing experience.

[0234] Please see Figure 16 , Figure 16 The diagram shows a block diagram of a video recognition device according to an embodiment of this application. The device 1100 includes:

[0235] The acquisition module 1110 is used to acquire object interaction data and video title of the video to be identified;

[0236] The filtering module 1120 is used to filter relevant interaction data related to target information from the object interaction data, wherein the target information is used to indicate the video content of the video to be identified, and / or, the target information is used to indicate that the video to be identified is the target video;

[0237] The relevance determination module 1130 is used to determine the relevance between the video to be identified and the video title based on the relevant interaction data.

[0238] The video determination module 1140 is used to determine whether the video to be identified is the target video based on the comparison result of the correlation and the correlation threshold.

[0239] Optionally, the filtering module 1120 is further configured to acquire video content data of the video to be identified as the target information; filter interactive data related to the video content data from the object interaction data as the relevant interactive data; the relevance determination module 1130 is further configured to determine the title relevance probability corresponding to the relevant interactive data based on the video title, wherein the title relevance probability is used to characterize the relevance between the video to be identified and the video title.

[0240] Optionally, the filtering module 1120 is further configured to: determine the content relevance probability of the corresponding object interaction data based on the video content data, wherein the content relevance probability is used to characterize the correlation between the object interaction data and the video content data; determine the interaction popularity of the object interaction data, wherein the interaction popularity is used to characterize the interactivity of the object interaction data itself; determine the key information probability of the object interaction data based on the content relevance probability and the interaction popularity; and filter the top N object interaction data with higher key information probabilities from the object interaction data as the relevant interaction data, where N is a positive integer.

[0241] Optionally, the object interaction data includes at least one object interaction data; the filtering module 1120 is further configured to obtain the number of likes and replies for each object interaction data; determine the overall interaction volume of the at least one object interaction data based on the number of likes and replies for each object interaction data; perform a weighted summation on the number of likes and replies for each object interaction data to obtain the weighted sum for each object interaction data; calculate the ratio of the weighted sum for each object interaction data to the overall interaction volume to obtain the interaction popularity corresponding to each object interaction data.

[0242] Optionally, the filtering module 1120 is further configured to: determine the media features of the video to be identified based on the video content data; obtain the text features of the object interaction data; fuse the text features and the media features to obtain a first fused feature; and analyze the first fused feature through a first correlation analysis model to obtain the content relevance probability corresponding to the object interaction data.

[0243] Optionally, the relevance determination module 1130 is further configured to: obtain the key information probability of the relevant interaction data, wherein the key information probability is used to characterize the importance of the relevant interaction data in all object interaction data corresponding to the video to be identified; filter target interaction data from the relevant interaction data based on the relevant interaction data and the key information probability; obtain the title features of the video title and the text features of the target interaction data; fuse the title features and the text features of the target interaction data to obtain a second fused feature; and analyze the second fused feature through a second relevance analysis model to obtain the title relevance probability corresponding to the relevant interaction data.

[0244] Optionally, the video determination module 1140 is further configured to determine the video to be identified as the target video when the title relevance probability is less than the relevance threshold.

[0245] Optionally, the filtering module 1120 is further configured to obtain target keywords as target information, wherein the target keywords are obtained based on sample interaction data with labeled information, wherein the labeled information is used to characterize that the sample interaction data indicates that the sample video is the target video, and the sample video is the video corresponding to the sample interaction data; and to filter out the interaction data that hits the target keywords from the object interaction data as relevant interaction data; the relevance determination module 1130 is further configured to determine the title hit probability corresponding to the relevant interaction data, wherein the title hit probability is used to characterize the relevance between the video to be identified and the video title.

[0246] Optionally, the device further includes: a keyword acquisition module, used to acquire the sample interaction data; extract keywords from the sample interaction data to obtain multiple preliminary keywords; determine the frequency of occurrence of the multiple preliminary keywords in the sample interaction data; and filter target keywords from the multiple preliminary keywords whose frequency of occurrence exceeds a preset frequency.

[0247] Optionally, the second interactive data includes at least one second interactive data; the relevance determination module 1130 is further configured to determine the text features and interaction popularity of each of the relevant interactive data, wherein the interaction popularity is used to characterize the interactivity of the relevant interactive data itself; analyze the text features of each of the relevant interactive data through a title hit analysis model to obtain the initial title hit probability of the corresponding relevant interactive data; and obtain the title hit probability based on the initial title hit probability and interaction popularity of each of the relevant interactive data.

[0248] Optionally, the video determination module 1140 is further configured to determine the video to be identified as the target video when the title hit probability exceeds the relevance threshold.

[0249] Optionally, the filtering module 1120 is further configured to acquire video content data and target keywords of the video to be identified as the target information, wherein the target keywords are obtained based on sample interaction data with labeled information, and the labeled information is used to characterize that the sample interaction data indicates that the sample video is the target video, and the sample video is the video corresponding to the sample interaction data; filter first interaction data related to the video content data from the object interaction data; filter interaction data that hit the target keywords from the object interaction data as the second interaction data; and use the first interaction data and the second interaction data as related interaction data; the relevance determination module 1130 is further configured to determine the title relevance probability corresponding to the first interaction data based on the video title, wherein the title relevance probability is used to characterize the relevance between the video to be identified and the video title; determine the title hit probability corresponding to the second interaction data, wherein the title hit probability is used to characterize the relevance between the video to be identified and the video title; the video determination module 1140 is further configured to perform a weighted summation of the title hit probability and the title relevance probability; and determine whether the video to be identified is the target video based on the comparison result of the summation result and the relevance threshold.

[0250] Optionally, the device further includes: a prompt message sending module, configured to send a prompt message or stop distributing the target video if it is determined that the video to be identified is the target video.

[0251] Figure 17 A structural block diagram of an electronic device for performing a video recognition method according to an embodiment of this application is shown. The electronic device may be... Figure 1 Terminal 101 or server 102, etc., should be noted that Figure 17 The computer system 1200 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0252] like Figure 17As shown, the computer system 1200 includes a Central Processing Unit (CPU) 1201, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on programs stored in Read-Only Memory (ROM) 1202 or programs loaded from storage portion 1208 into Random Access Memory (RAM) 1203. The RAM 1203 also stores various programs and data required for system operation. The CPU 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An Input / Output (I / O) interface 1205 is also connected to the bus 1204.

[0253] In some embodiments, the following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1210 as needed so that computer programs read from it can be installed into the storage section 1208 as needed.

[0254] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1209, and / or installed from removable medium 1211. When the computer program is executed by central processing unit (CPU) 1201, it performs various functions defined in the system of this application.

[0255] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0256] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0257] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0258] In another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the methods in any of the above embodiments.

[0259] According to one aspect of this application, an electronic device is also provided, comprising: a processor; and a memory storing a computer program, which, when executed by the processor, implements the method in any of the above embodiments.

[0260] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the electronic device to perform the methods of any of the above embodiments.

[0261] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0262] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, and includes several instructions to cause an electronic device to execute the method according to the embodiments of this application.

[0263] It should be noted that in this application, information such as object interaction data, object identity information, and object subscription information requires the authorization of the object. After obtaining the object's authorization for the object interaction data, object identity information, and object subscription information, the above information can be processed, thereby complying with the relevant legal provisions.

[0264] Other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. It should be understood that this application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

[0265] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A video recognition method, characterized in that, The method includes: Obtain object interaction data and video title from the video to be identified; Filter relevant interaction data related to target information from the object interaction data, wherein the target information is used to indicate the video content of the video to be identified, and / or, the target information is used to indicate that the video to be identified is the target video; Based on the relevant interaction data, the correlation between the video to be identified and the video title is determined; Based on the comparison results of the correlation and the correlation threshold, it is determined whether the video to be identified is the target video; The step of filtering relevant interaction data related to the target information from the object interaction data includes: Obtain the video content data of the video to be identified, as the target information; Based on the video content data, determine the content relevance probability corresponding to the object interaction data, wherein the content relevance probability is used to characterize the correlation between the object interaction data and the video content data; Determine the interaction heat of the object interaction data, wherein the interaction heat is used to characterize the interactivity of the object interaction data itself; Based on the content relevance probability and the interaction popularity, determine the probability of key information in the object interaction data; The top N object interaction data with the highest probability of containing key information are selected from the object interaction data and used as the relevant interaction data, where N is a positive integer.

2. The method according to claim 1, characterized in that, The step of determining the relevance between the video to be identified and the video title based on the relevant interaction data includes: Based on the video title, the title relevance probability corresponding to the relevant interaction data is determined, and the title relevance probability is used to characterize the relevance between the video to be identified and the video title.

3. The method according to claim 1, characterized in that, The object interaction data includes at least one object interaction data; determining the interaction popularity of the object interaction data includes: Obtain the number of likes and replies for each of the object's interaction data; The overall interaction volume of the at least one object interaction data is determined based on the number of likes and replies for each of the object interaction data. The number of likes and the number of replies for each object's interaction data are weighted and summed to obtain the weighted sum for each object's interaction data. Calculate the weighted sum of each object interaction data point and its ratio to the overall interaction volume to obtain the interaction popularity corresponding to each object interaction data point.

4. The method according to claim 1, characterized in that, The step of determining the content relevance probability of the corresponding object interaction data based on the video content data includes: Based on the video content data, determine the media characteristics of the video to be identified; Obtain the text features of the object interaction data; The text features and the media features are fused to obtain the first fused feature; The first fusion feature is analyzed using a first correlation analysis model to obtain the content relevance probability of the corresponding object interaction data.

5. The method according to claim 2, characterized in that, The step of determining the title relevance probability corresponding to the relevant interaction data based on the video title includes: The key information probability of the relevant interaction data is obtained, and the key information probability is used to characterize the importance of the relevant interaction data in all object interaction data corresponding to the video to be identified. Based on the relevant interaction data and the probability of the key information, target interaction data is filtered from the relevant interaction data; Obtain the title features of the video title and the text features of the target interactive data; The title features and the text features of the target interactive data are fused to obtain a second fused feature; The second fusion feature is analyzed using a second correlation analysis model to obtain the title relevance probability of the relevant interactive data.

6. The method according to claim 2, characterized in that, The step of determining whether the video to be identified is the target video based on the comparison result of the correlation and the correlation threshold includes: When the title relevance probability is less than the relevance threshold, the video to be identified is determined to be the target video.

7. The method according to claim 1, characterized in that, The step of filtering relevant interaction data related to the target information from the object interaction data includes: The target keywords are obtained as the target information. The target keywords are obtained based on sample interaction data with labeled information. The labeled information is used to characterize that the sample interaction data indicates that the sample video is the target video. The sample video is the video corresponding to the sample interaction data. Interaction data that matches the target keyword is filtered out from the object interaction data and used as relevant interaction data; The step of determining the relevance between the video to be identified and the video title based on the relevant interaction data includes: Determine the title hit probability corresponding to the relevant interactive data, whereby the title hit probability is used to characterize the relevance between the video to be identified and the video title.

8. The method according to claim 7, characterized in that, The method for obtaining the target keywords includes: Obtain the sample interaction data; Keyword extraction was performed on the sample interaction data to obtain several preliminary keywords; Determine the frequency of occurrence of the selected keywords in the sample interaction data; Select target keywords from the multiple initial keywords that appear more frequently than a preset frequency.

9. The method according to claim 7, characterized in that, The relevant interaction data includes at least one relevant interaction data; determining the title hit probability corresponding to the relevant interaction data includes: Determine the text features and interaction popularity of each of the relevant interactive data, wherein the interaction popularity is used to characterize the interactivity of the relevant interactive data itself; The text features of each relevant interactive data are analyzed by the title hit analysis model to obtain the initial title hit probability of the corresponding relevant interactive data; The title hit probability is obtained based on the initial title hit probability and interaction popularity of each relevant interaction data.

10. The method according to claim 7, characterized in that, The step of determining whether the video to be identified is the target video based on the comparison result of the correlation and the correlation threshold includes: When the title hit probability exceeds the relevance threshold, the video to be identified is determined to be the target video.

11. The method according to claim 1, characterized in that, The step of filtering relevant interaction data related to the target information from the object interaction data includes: The video content data and target keywords of the video to be identified are obtained as the target information. The target keywords are obtained based on sample interaction data with labeled information. The labeled information is used to characterize that the sample interaction data indicates that the sample video is the target video. The sample video is the video corresponding to the sample interaction data. Filter the first interaction data related to the video content data from the object interaction data; Interaction data that matches the target keyword is selected from the object interaction data and used as the second interaction data; The first interactive data and the second interactive data are used as relevant interactive data; The step of determining the relevance between the video to be identified and the video title based on the relevant interaction data includes: Based on the video title, determine the title relevance probability corresponding to the first interactive data, wherein the title relevance probability is used to characterize the relevance between the video to be identified and the video title; Determine the title hit probability corresponding to the second interactive data, wherein the title hit probability is used to characterize the correlation between the video to be identified and the video title; The step of determining whether the video to be identified is the target video based on the comparison result of the correlation and the correlation threshold includes: The title hit probability and the title relevance probability are weighted and summed. Based on the comparison between the summation result and the correlation threshold, it is determined whether the video to be identified is the target video.

12. The method according to any one of claims 1 to 11, characterized in that, After determining whether the video to be identified is the target video based on the comparison result of the correlation and the correlation threshold, the method further includes: If the video to be identified is determined to be the target video, a prompt message is sent, or the distribution of the target video is stopped.

13. A video recognition device, characterized in that, The device includes: The acquisition module is used to acquire object interaction data and video title of the video to be identified; A filtering module is used to filter relevant interaction data related to target information from the object interaction data, wherein the target information is used to indicate the video content of the video to be identified, and / or, the target information is used to indicate that the video to be identified is the target video; The relevance determination module is used to determine the relevance between the video to be identified and the video title based on the relevant interaction data. The video determination module is used to determine whether the video to be identified is the target video based on the comparison result of the correlation and the correlation threshold. The filtering module is further configured to: acquire video content data of the video to be identified as the target information; determine the content relevance probability of the corresponding object interaction data based on the video content data, wherein the content relevance probability is used to characterize the correlation between the object interaction data and the video content data; determine the interaction popularity of the object interaction data, wherein the interaction popularity is used to characterize the interactivity of the object interaction data itself; determine the key information probability of the object interaction data based on the content relevance probability and the interaction popularity; and filter the top N object interaction data with higher key information probabilities from the object interaction data as the relevant interaction data, where N is a positive integer.

14. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.

16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.