Hotspot information mining method and device, computer device, and storage medium
By clustering and classifying seed videos on video sharing platforms, a set of trending event videos is identified, solving the problem of misrecommendation caused by relying on external trending sources and achieving more efficient and accurate trending video mining.
Patent Information
- Application Number
- CN202211021681.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2042-08-24
AI Technical Summary
Video sharing platforms rely on trending information from external news websites or search engines, which can lead to the mistaken recommendation of videos related to non-trending events as trending videos, thus reducing the accuracy of identifying videos related to internal trending events.
By acquiring seed videos from the first account, clustering and classification techniques are used to divide the seed videos into multiple video sets based on their text information. The number of seed videos for event categories and advertising categories is counted separately to determine the video sets of hot events and reduce reliance on external hotspot sources.
It improves the accuracy and efficiency of discovering videos related to trending events within video sharing platforms, and reduces the possibility of false recommendations.
Smart Images

Figure CN117033697B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a hot information mining method and device, computer equipment and a computer readable storage medium (referred to as storage medium). BACKGROUND
[0002] With the development of Internet technology, in addition to the traditional self-media public number platform mainly in the form of text and pictures, a video sharing platform that can share videos at any time is also provided. In the video search page or video recommendation page of the video sharing platform, videos related to hot events with high user attention are usually recommended; however, the mining of hot event related videos by the video sharing platform is usually performed by obtaining hot event related information provided by external news websites or external search websites, and then matching the hot event related information with the related information of the videos in the video sharing platform to screen hot event related videos. The mining of hot event related videos by the video sharing platform in the prior art relies too much on the hot event related information provided by external news websites or external search websites, and non-hot event related videos are easily mistaken for hot videos, resulting in that hot event related videos in the video sharing platform cannot be recommended to the video search page or video recommendation page. SUMMARY
[0003] Therefore, it is necessary to provide a hot information mining method, device, computer equipment and storage medium to improve the accuracy of mining hot videos in the video sharing platform.
[0004] In a first aspect, the present application provides a hot information mining method, which comprises:
[0005] obtaining a first account and a seed video of the first account;
[0006] clustering the seed video according to the text information of the seed video to obtain a plurality of video sets;
[0007] first classifying the seed videos under each video set respectively to obtain a first number of seed videos belonging to an event category under each video set;
[0008] second classifying the seed videos under each video set respectively to obtain a second number of seed videos belonging to an advertisement category under each video set;
[0009] determining a hot event video set from the video sets according to the first number and the second number corresponding to each video set.
[0010] In a second aspect, the present application provides a hot information mining device, which comprises:
[0011] The seed video acquisition module is configured to acquire the first account and a seed video of the first account.
[0012] The video clustering module is configured to cluster the seed videos according to text information of the seed videos, to obtain a plurality of video sets.
[0013] The first classification module is configured to respectively perform first classification on the seed videos in each of the video sets, to obtain a first number of seed videos belonging to an event category in each of the video sets.
[0014] The second classification module is configured to respectively perform second classification on the seed videos in each of the video sets, to obtain a second number of seed videos belonging to an advertisement category in each of the video sets.
[0015] The hotspot information acquisition module is configured to determine a hotspot event video set from the video sets according to the first number and the second number corresponding to each of the video sets.
[0016] In a third aspect, the present application further provides a computer device, which comprises:
[0017] one or more processors;
[0018] a memory; and
[0019] one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement the hotspot information mining method.
[0020] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program. The computer program is loaded by a processor to execute steps in the hotspot information mining method.
[0021] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium. The processor executes the computer instructions, so that the computer device executes the method provided in the first aspect.
[0022] The hotspot information mining method, device, computer device and storage medium described above, by obtaining a first account and a seed video of the first account; clustering the seed video according to text information of the seed video to obtain a plurality of video sets; performing first classification on the seed video in each video set respectively to obtain a first number of seed videos belonging to an event category in each video set; performing second classification on the seed video in each video set respectively to obtain a second number of seed videos belonging to an advertisement category in each video set; and determining a hotspot event video set from the video sets according to the first number and the second number corresponding to each video set. By clustering the seed video in the first account, the seed video is divided into video sets with the same described content, and then based on the first number of seed videos of the event category and the second number of seed videos of the advertisement category in the video set, a hotspot event video set with more event category videos and less advertisement category videos is determined from the video set, without relying on external hotspot sources to realize the mining of hotspot event related videos in the video sharing platform, and the mining efficiency and accuracy of the hotspot event related videos are improved. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0024] Figure 1 is a scene diagram of the hotspot information mining method in the embodiments of the present application;
[0025] Figure 2 is a flow diagram of the hotspot information mining method in the embodiments of the present application;
[0026] Figure 3 is a flow diagram of the first account updating step in the embodiments of the present application;
[0027] Figure 4 is a flow diagram of the target video obtaining step in the embodiments of the present application;
[0028] Figure 5 is a flow diagram of the seed video obtaining step in the embodiments of the present application;
[0029] Figure 6 is a flow diagram of the target keyword obtaining step in the embodiments of the present application;
[0030] Figure 7 is a structure diagram of the hotspot information mining device in the embodiments of the present application;
[0031] Figure 8is a structural schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort fall within the protection scope of the present application.
[0033] In the description of the present application, the terms "first", "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0034] In the description of the present application, the word "for example" is used to indicate "as an example, illustration or description". Any embodiment described as "for example" in the present application is not necessarily interpreted as more preferred or more advantageous than other embodiments. The following description is given in order to enable any person skilled in the art to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can realize the present application without using these specific details. In other examples, well-known structures and processes will not be described in detail in order to avoid unnecessary details making the description of the present application obscure. Therefore, the present application is not intended to be limited to the shown embodiments, but is consistent with the broadest scope of principles and features disclosed in the present application.
[0035] It can be understood that the hotspot information mining method of the present embodiment can be executed on a terminal, can be executed on a server, or can be executed by a terminal and a server together. Referring to Figure 1For example, a terminal and a server jointly perform a hotspot information mining method. The hotspot information mining system provided by the embodiments of the present application includes a terminal 110 and a server 120, etc. The terminal 110 and the server 120 are connected through a network, for example, a wired or wireless network, etc. The hotspot information mining device can be integrated in the server. The server 120 can be configured to: acquire a first account and a seed video of the first account; cluster the seed video according to text information of the seed video to obtain a plurality of video sets; perform first classification on the seed video in each video set respectively to acquire a first number of seed videos belonging to an event category in each video set; perform second classification on the seed video in each video set respectively to acquire a second number of seed videos belonging to an advertisement category in each video set; and determine a hotspot event video set from the video sets according to the first number and the second number corresponding to each video set.
[0036] The terminal 110 can receive the hotspot event video set sent by the server 120 and output the hotspot event video set through an output module. Optionally, in some embodiments, the terminal can include a display module configured to display a video search page or a video recommendation page of a video sharing platform. When the hotspot event video set in a current time window is acquired, the seed video in the hotspot event video set is displayed on the video search page or the video recommendation page of the video sharing platform. Based on a triggering operation on a corresponding control of the seed video, the video content of the seed video is played. The terminal 110 can include a mobile phone, a smart television, a tablet computer, a notebook computer, or a personal computer (PC, Personal Computer), etc. A client can also be arranged on the terminal 110. The client can be an application program client or a browser client, etc.
[0037] The process of acquiring the hotspot event video set by the server 120 described above can also be performed by the terminal 110.
[0038] The hotspot information mining method provided by the embodiments of the present application relates to natural language processing (NLP) and data mining in the field of artificial intelligence (AI). The embodiments of the present application can obtain a plurality of video sets by clustering the seed videos according to the text information of the seed videos, and classify the seed videos in each video set, obtain the first quantity of seed videos belonging to the event category and the second quantity of seed videos belonging to the advertisement category in the video set, and finally determine the hotspot event video set from the video set according to the first quantity and the second quantity corresponding to each video set. The embodiments of the present application realize automatic mining of hotspot event related videos, improve the accuracy of hotspot mining results, and also improve the mining efficiency.
[0039] Artificial intelligence is the theory, method, technology and application system for using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision making. Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including hardware and software technologies. Among them, the artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning technology.
[0040] Natural language processing is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science and mathematics. Therefore, the research in this field will involve natural language, i.e. the language used in daily life, so it has a close relationship with the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph and other technologies.
[0041] Data mining refers to a process of searching information hidden in a large amount of data through an algorithm, and is a hot issue in the field of artificial intelligence and database research. Data mining is usually related to computer science, and is achieved through statistics, online analytical processing, information retrieval, machine learning, expert systems (relying on past experience rules), and pattern recognition and many other methods to achieve the above goal. In recent years, data mining has attracted great attention from the information industry, the main reason being that there is a large amount of data that can be widely used, and there is an urgent need to convert these data into useful information and knowledge. The information and knowledge obtained by data mining can be widely used in various application fields, including business management, production control, market analysis, engineering design and scientific exploration.
[0042] The hotspot information mining method of the embodiments of the present application can be applied to various scenes that need to mine hotspot information, for example, when identifying videos related to a hotspot event in a large amount of video information, the hotspot information mining method provided in the embodiments can be used. It should be noted that the order of the following embodiments is not limited as the preferred order of the embodiments.
[0043] Referring to Figure 2 , the embodiments of the present application provide a hotspot information mining method, which is mainly illustrated by taking the server 120 in the above Figure 1 as an example. The method includes steps S210 to S230, which are specifically as follows:
[0044] Step S210, obtaining a first account and a seed video of the first account.
[0045] In this step, the first account refers to an account that publishes multimedia content such as videos in a video sharing platform. Specifically, the first account can be a seed account pre-selected by a person, or a seed account selected from all accounts in the video sharing platform based on hotspot information in a historical time window. It can be understood that the first account refers to an account whose published videos have a large proportion of videos related to a hotspot event. The seed video refers to a video published by the first account on the video sharing platform, for example, the seed video can be a video published by the first account on the video sharing platform within a certain time period.
[0046] Specifically, the server can obtain the first account and extract the seed video of the first account within the current time window. The length of the current time window can be flexibly set, for example, it can be set to 24 hours.
[0047] Step S220, clustering the seed video according to the text information of the seed video to obtain a plurality of video sets.
[0048] The text information is text data used to describe the video content, for example, can be title text, label text, etc. corresponding to the video. After obtaining the seed video, the seed videos with similar video content can be aggregated into a video set according to the text information of the seed video, to obtain multiple different video sets. It can be understood that the clustering of the seed video can be performed by a K-means clustering algorithm, a hierarchical clustering algorithm, etc. and is not limited here.
[0049] For example, in the obtained multiple seed videos, there are seed video A with text information "A place heavy rain", seed video B with text information "A place B area heavy rain", seed video C with text information "A place heavy rain causes many cars to be soaked", seed video D with text information "high school entrance examination score line announced", etc. After clustering the seed videos according to the text information of the seed videos, the seed video A, the seed video B and the seed video C can be classified into one video set, and the seed video D can be classified into a video set alone.
[0050] In an embodiment, the step of clustering the seed videos according to the text information of the seed videos to obtain multiple video sets can specifically include: encoding the text information of each seed video to obtain a text vector of each seed video; calculating the similarity between each seed video based on the text vector corresponding to each seed video; and clustering the seed videos based on the similarity between each seed video to obtain multiple video sets.
[0051] The text information of the seed video can be encoded by a word-to-vector model, or can be encoded by a BERT (Bidirectional Encoder Representation from Transformers) language model, so that the text information of the seed video is converted into a text vector.
[0052] After obtaining the text vectors of the seed videos, the similarity between the text vectors of any two seed videos can be calculated, i.e., the similarity between the text vectors of any two seed videos is calculated, and then the similarity between the text vectors of the two seed videos is determined as the similarity between the seed videos; wherein the calculation method of the similarity between the text vectors can be customized, for example, the similarity between the text vectors of any two seed videos can be calculated by a cosine similarity algorithm, i.e., the cosine value between the included angles of the text vectors of any two seed videos is used as a measure of the difference between the seed videos, the cosine value is close to 1, the included angle tends to 0, indicating that the two text vectors are more similar, i.e., the probability that the two seed videos describe the same event or the same content is greater, the cosine value is close to 0, the included angle tends to 90 degrees, indicating that the two text vectors are less similar, i.e., the probability that the two seed videos describe the same event or the same content is smaller. At this time, the seed videos with a similarity greater than a preset similarity threshold are classified into a video set, and all seed videos are divided into multiple video sets, and the seed videos in each video set describe similar content.
[0053] Further, after obtaining the similarity between any two seed videos, any seed video can be taken as an independent video set, the two video sets (i.e., seed videos) with the smallest similarity can be merged into a new video set, and a new text vector corresponding to the video set can be calculated based on the text vectors of the seed videos in the new video set; then, the similarity between any two video sets is calculated based on the text vectors corresponding to the video sets, so that the two video sets with the smallest similarity are merged into a new video set. The merging of seed videos is repeated according to the above process until the similarity between any video sets is not greater than a preset similarity threshold. The above hierarchical clustering is used to cluster the seed videos based on the text vectors of the seed videos.
[0054] Step S230, respectively classifying the seed videos in each video set to obtain the first number of seed videos belonging to the event category in each video set.
[0055] Step S240, respectively classifying the seed videos in each video set to obtain the second number of seed videos belonging to the advertisement category in each video set.
[0056] The video published by the first account accounts for a large proportion of the videos related to the hot event, but there are still seed videos that are non-event category videos or advertisement category videos. The event category video refers to a video related to a news event, the non-event category video refers to a video not related to a news event, for example, a vlog (video blog) video, a science popularization video, a food video, etc. The advertisement category video refers to a video whose content is marketing a product, and the non-advertisement category video refers to a video whose content has no marketing-related content.
[0057] After obtaining the plurality of video sets, for any video set in which the number of seed videos is greater than 1, the seed videos in the video set are first classified and second classified to determine whether the seed videos in the video set to be processed belong to the event category video and the advertisement category video, and then the first number of seed videos belonging to the event category in the video set and the second number of seed videos belonging to the advertisement category in the video set are counted.
[0058] The first classification or the second classification of the seed video is specifically a video classification of the seed video, which can be performed by a first classifier to predict whether the seed video belongs to the event category video and by a second classifier to predict whether the seed video belongs to the advertisement category video. It can be understood that the first classifier and the second classifier can be a classifier based on a support vector machine or other machine learning algorithm, or a neural network classifier, which is not limited here.
[0059] Specifically, for the seed video in the video set, the text information of the seed video can be encoded to obtain a text vector of the seed video, and then the text vector of the seed video is input into the first classifier and the second classifier to determine whether the seed video belongs to the event category video through the first classifier and whether the seed video belongs to the advertisement category video through the second classifier. It should be noted that the first classifier and the second classifier are trained by a plurality of training data with labels. The training data of the embodiment includes a plurality of training texts, and the label refers to the category information represented by the training text. For example, the label in the training data for the first classifier is used to identify whether the training text is an event category, and the label in the training data for the second classifier is used to identify whether the training text is an advertisement category. The classification model can be trained by other devices and provided to the server, or the server can also train it by itself.
[0060] In step S250, the hot event video set is determined from the video sets according to the first number and the second number corresponding to each video set.
[0061] After obtaining the first quantity of seed videos belonging to the event category and the second quantity of seed videos belonging to the advertisement category in each video set, the video set in which the first quantity is greater than the preset quantity threshold and the second quantity is less than the preset quantity threshold can be determined as the hot event video set. The preset quantity threshold can be set according to actual conditions. For example, the preset quantity threshold can be set as half of the total quantity of seed videos in the corresponding video set. For example, if the total quantity of seed videos in the video set is 10, when the first quantity of seed videos belonging to the event category is greater than 5 and the second quantity of seed videos belonging to the advertisement category is less than 5, the video set can be determined as the hot event video set.
[0062] By screening the video set in which the first quantity of seed videos belonging to the event category is greater than the preset quantity threshold and the second quantity of seed videos belonging to the advertisement category is less than the preset quantity threshold, it is ensured that the screened video set is related to news events and has no advertisement marketing, and the hot event related video is accurately mined.
[0063] It can be understood that after obtaining the hot event video set, the server can push the hot event video set to the terminal, so that the terminal recommends the videos in the hot event video set to the user on the video search page or the video recommendation page of the video sharing platform.
[0064] Further, in an embodiment, after the step of determining the hot event video set from the video set according to the first quantity and the second quantity corresponding to each video set, a hot video can be screened from the hot event video set; and then, a text label of a hot event is extracted based on the text information of the hot video.
[0065] The hot event refers to the event described by all seed videos in the hot event video set, the hot video refers to the video used to represent the event content described by all seed videos in the hot event video set, and the text label of the hot event refers to the text information used to describe the hot event. For example, the hot event video set includes seed video A with the text information "rainstorm in A", seed video B with the text information "rainstorm occurs in B area of A", and seed video C with the text information "rainstorm in A causes many cars to be flooded", the hot video in the hot event video set is seed video A, and the text information "rainstorm in A" of the seed video can be determined as the hot event, and the corresponding text label can be "rainstorm in A", "A", and "rainstorm".
[0066] Specifically, the hotspot video can be selected from the hotspot event video set, specifically, the hotspot video can be selected based on the click quantity of each seed video in the hotspot event video set, the seed video with the largest click quantity can be selected as the hotspot video; or the hotspot video can be selected based on the text length of the title text of each seed video in the hotspot event video set, the seed video with the shortest text length can be selected as the hotspot video; or it can be judged whether the title text of each seed video in the hotspot event video set contains a key entity word, the seed video containing the key entity word in the title text can be determined as the hotspot video; further, each seed video can be scored according to the click quantity of each seed video in the hotspot event video set, the text length of the title text of each seed video, and whether the title text of each seed video contains a key entity word, and then the seed video with the highest score can be determined as the hotspot video in the hotspot event video set. There are various ways to select the hotspot video from the hotspot event video set, which are not limited here.
[0067] After the hotspot video in the hotspot event video set is obtained, the text information of the hotspot video can be determined as the text label of the hotspot event, or the key word or key sentence can be selected from the text information of the hotspot video to determine the text label of the hotspot event. It can be understood that the server can select the video related to the hotspot event from all the videos in the video sharing platform based on the text label of the hotspot event to expand the video quantity in the hotspot event video set.
[0068] In the above hotspot information mining method, the first account and the seed video of the first account are obtained; the seed videos are clustered according to the text information of the seed videos, and a plurality of video sets are obtained; the seed videos in each video set are respectively classified in a first classification, and the first quantity of seed videos belonging to the event category in each video set is obtained; the seed videos in each video set are respectively classified in a second classification, and the second quantity of seed videos belonging to the advertisement category in each video set is obtained; and the hotspot event video set is determined from the video set according to the first quantity and the second quantity corresponding to each video set. By clustering the seed videos in the first account, the seed videos are divided into video sets with the same described content, and then the hotspot event video set with more event category videos and less advertisement category videos is determined from the video set based on the first quantity of event category seed videos and the second quantity of advertisement category seed videos in the video set, without relying on external hotspot sources to realize the mining of hotspot event related videos in the video sharing platform, improving the efficiency and accuracy of hotspot event related video mining.
[0069] In the hot information mining method, the seed video published by the first account greatly influences the accuracy of subsequent hot event and hot event related video mining. In order to improve the accuracy of hot event related video mining, other hot event related accounts can be selected from the video sharing platform as seed accounts based on the obtained hot event video set to supplement or replace the first account. Specifically, in one embodiment, as shown in Figure 3 After the step of determining the hot event video set from the video set according to the first number and the second number corresponding to each video set, the method further includes:
[0070] Step S310, filtering hot videos from the hot event video set, and obtaining target videos from the video library according to the text information of the hot videos.
[0071] Step S320, obtaining a second account corresponding to the target video, and obtaining all videos published by the second account and the total number of the all videos.
[0072] Step S330, obtaining a third number of videos belonging to the hot event from the all videos of the second account.
[0073] Step S340, if the ratio between the third number and the total number of the all videos is greater than a preset ratio, updating the first account according to the second account.
[0074] The hot video refers to a video representing the event content described by all seed videos in the hot event video set. The specific acquisition method can refer to the above embodiment, and will not be repeated again. The video library includes all videos published by the video sharing platform within the current time window. Specifically, after the hot video is determined, the text information of the hot video is matched with the text information of each video in the video library to obtain the target video with a matching success.
[0075] Further, in one embodiment, as shown in Figure 4 The step of filtering the target video from the video library according to the text information of the hot video includes:
[0076] Step S410, obtaining a first label text of the hot video according to the text information of the hot video.
[0077] Step S420, obtaining the text information and a second label text of the original video in the video library.
[0078] Step S430, obtaining label similarity feature information between the hot video and the original video based on the first label text and the second label text.
[0079] In step S440, the text information of the hot video and the text information of the original video are spliced to obtain first spliced text, and first text similarity feature information is obtained based on the first spliced text.
[0080] In step S450, the text information of the hot video and the second label text of the original video are spliced to obtain second spliced text, and second text similarity feature information is obtained based on the second spliced text.
[0081] In step S460, based on the label similarity feature information, the first text similarity feature information, and the second text similarity feature information, a matching result between the original video and the hot video is identified.
[0082] In step S470, a target video is obtained based on the matching result of the original video.
[0083] The label text includes but is not limited to domain label text, key label text, regional label text, and person name label text of the video; the domain label text is used to identify the event domain of the video content, such as the sports domain and the entertainment domain; the key label text is used to identify the keywords of the video content in the corresponding event domain, such as the keywords "Lakers" and "NBA" in the sports domain; the regional label text is used to identify the regional information in the video content, such as Los Angeles and Guizhou; and the person name label text is used to identify the person name information in the video content, such as "James". Specifically, after obtaining the hot video and all original videos in the video library, for any one video, label text corresponding to the video is obtained based on the text information of the video.
[0084] After obtaining the first label text of the hot video and the second label text of each original video, for any one original video, label similarity feature information is extracted based on the corresponding label texts of the original video and the hot video. Specifically, the number of intersections, the number of unions, and the ratio between the number of intersections and the number of unions between the first label text of the hot video and the second label text of the original video can be obtained as the label similarity feature information between the hot video and the original video; further, the number of intersections, the number of unions, and the ratio between the number of intersections and the number of unions between the first label text of the hot video and the second label text of the original video in the domain label text, the key label text, the regional label text, and the person name label text can be obtained as the label similarity feature information between the hot video and the original video.
[0085] Meanwhile, the text information of the hot video and the text information of the original video are spliced to obtain a first spliced text, and the text information of the hot video and the second label text of the original video are spliced to obtain a second spliced text, and then the first spliced text and the second spliced text are encoded to obtain first text similarity feature information corresponding to the first spliced text and second text similarity feature information corresponding to the second spliced text. It can be understood that the first spliced text and the second spliced text can be converted into vector representation by the BERT language model to obtain the first text similarity feature information and the second text similarity feature information.
[0086] Finally, the label similarity feature information, the first text similarity feature information, and the second text similarity feature information are input into a pre-trained video matching model, and a matching result between the original video and the hot video is output by the video matching model. The video matching model can be a pre-trained neural network model. Specifically, the video matching model is trained by a plurality of training data with labels. The training data of the embodiment includes label similarity feature information, first text similarity feature information, and second text similarity feature information corresponding to a plurality of video pairs. The label is used to represent the matching category (including matching category and non-matching category) of each video pair in the training data. The video matching model can be trained by other devices and provided to the server, or the server can also train it by itself.
[0087] The total video refers to all videos published by the second account on the video sharing platform within a certain time period. For example, it can be all videos published by the second account since the account was successfully registered, or all videos published by the second account in the past month. After obtaining the target video matching the hot video, the second account publishing the target video is obtained, and the total video published by the second account is obtained. The text information of each video in the total video of the second account is matched with the text label of the historical hot event to obtain a third number of videos belonging to the hot event in the total video of the second account. When the ratio of the third number to the total number of the total video of the second account is greater than a preset ratio, it can be considered that the video published by the second account has a large proportion of videos related to the hot event. At this time, the second account can be determined as a seed account, and the second account is added to the first account as a seed account.
[0088] Based on the hot event related video, other hot event related videos are further mined from the video sharing platform as seed accounts to supplement or replace the first account, so that the seed account in the hot event related video mining process is updated in real time, and the accuracy of the hot event related video set mining is improved.
[0089] Further, in one embodiment, the first account can also be periodically cleaned up, specifically, for any one first account, the total videos published by the first account are obtained, and the text information of each video in the total videos of the first account is matched with the text labels of historical hot events to obtain a fourth quantity of videos in the total videos of the first account that belong to hot event corresponding videos, and when the ratio of the fourth quantity to the total quantity of the total videos of the first account is less than or equal to a preset ratio, the first account is deleted, thereby eliminating accounts whose published videos have a too low proportion of hot event related videos.
[0090] In addition to the videos published by the first account as seed videos, the mining of hot event related videos based on the text information of the original videos in the video library can also be performed, specifically, in one embodiment, as shown in Figure 5 Before the step of clustering the seed videos according to the text information of the seed videos to obtain a plurality of video sets, the method further includes:
[0091] Step S510, extracting video keywords from the text information of all original videos in the video library; wherein the video keywords include one-word keywords;
[0092] Step S520, counting a first word frequency of each video keyword in a current time window and a second word frequency of each video keyword in a historical time window;
[0093] Step S530, screening target keywords from the video keywords according to the word frequency ratio between the first word frequency and the second word frequency of each video keyword;
[0094] Step S540, screening candidate videos from the video library according to the target keywords, updating the seed videos of the first account based on the candidate videos to obtain updated seed videos.
[0095] Wherein, the extraction of video keywords can specifically be the tokenization processing of the text information of each original video in the video library, so that the text information is cut into individual words, and the video keywords are obtained. One-word keywords refer to one word obtained after tokenization processing. For example, taking the text information "A place heavy rain causes many cars to be soaked" as an example, the keywords "A place", "heavy rain", "car", and "soaked" are one-word keywords extracted from the text information.
[0096] Wherein, the current time window refers to the time window in which the hot event related videos are to be mined; the historical time window refers to a time window selected from the current time window and moving forward, corresponding to the current time window. For example, a historical time window of more than three times the length of the current time window can be selected as the historical time window; for example, if the current time window is 24 hours, the historical time window can be selected as 3 days (i.e. 72 hours) before the current time window.
[0097] wherein the word frequency refers to the number of times that the video keyword appears in the text information of all original videos within a certain time window. In order to facilitate the region, the number of times that the video keyword appears in the text information of the original video within the current time window can be recorded as the first word frequency, and the average number of times that the video keyword appears in the text information of the historical video within the historical time window can be recorded as the second word frequency. It can be understood that the historical video herein refers to the video whose publishing time is within the historical time window, and the average number of times refers to the average number of times in the unit time length of the window length of the current time window. Taking the current time window of 24 hours and the historical time window of 72 hours before the current time window as an example, at this time, the total number of times that the video keyword appears in the text information of the historical video in the unit time length of 24 hours is 120, and the average number of times (i.e. the second word frequency) is 40.
[0098] Specifically, after obtaining the video keyword, for any video keyword, the first word frequency of the video keyword within the current time window and the second word frequency of the video keyword within the historical time window can be counted, and the word frequency ratio between the first word frequency and the second word frequency corresponding to the video keyword can be calculated, and then the video keyword with a word frequency ratio greater than a preset word frequency ratio threshold is determined as the target keyword. For example, the word frequency ratio of any video keyword can be calculated by the following formula (1):
[0099]
[0100] wherein C(W i , T j ) represents the number of times that the video keyword W i appears in the text information of all original videos within the current time window T j , and C(W i , T1,..., T j-1 ) represents the number of times that the video keyword W i appears in the text information of all original videos within the historical time window of T1 to T j-1 .
[0101] It can be understood that if an event has a burst propagation, the word frequency of the keyword corresponding to the text label of the event within the current time window will increase. Since the event does not occur massively in the historical window time, the word frequency of the keyword corresponding to the text label of the event within the historical time window is very low. By calculating the word frequency of the video keyword within the current time window and the historical time window, the video keyword with the burst propagation attribute can be mined, and the video with the burst propagation tendency can be mined, thereby improving the accuracy of the hot video and the hot event corresponding to the hot video.
[0102] Further, the ratio of the first word frequency and the second word frequency of the video keyword is often related to the word frequency size of the video keyword. For example, assuming that the first word frequency of video keyword A is 40, the second word frequency is 110, and the ratio of the first word frequency and the second word frequency is 2.7, and the first word frequency of video keyword B is 6, the second word frequency is 20, and the ratio of the first word frequency and the second word frequency is 3.3, the ratio of the first word frequency and the second word frequency of video keyword B is greater than that of video keyword A, but video keyword A is the target keyword with increased word frequency and explosive spread.
[0103] Therefore, according to the ratio of the first word frequency and the second word frequency of each video keyword, the target keyword is selected from the video keywords. Specifically, the ratio of the first word frequency and the second word frequency of the video keyword is calculated first, then the ratio of the first word frequency and the second word frequency of the video keyword is smoothed by using the Bayesian smoothing algorithm to obtain the smoothed ratio of the first word frequency and the second word frequency, and finally, the target keyword is selected from the video keywords according to the smoothed ratio of the first word frequency and the second word frequency of each video keyword. For example, the ratio of the first word frequency and the second word frequency of any video keyword can be calculated by the following formula (2):
[0104]
[0105] wherein, represents the ratio of the first word frequency and the second word frequency of video keyword W i in the current time window T j and the word frequency in the historical time window. avg represents the average value of the ratio of the first word frequency and the second word frequency of all video keywords in the current time window T j and the word frequency in the historical time window, C(W i , T j ) represents the number of times that video keyword W j appears in the text information of all original videos in the current time window T i , and C avg represents the average value of the number of times that all video keywords appear in the text information of all original videos in the current time window T j .
[0106] By smoothing the ratio of the first word frequency and the second word frequency of the video keyword, the fluctuation of the ratio of the first word frequency and the second word frequency of the video keyword with low word frequency is prevented, so that the error in the selection of the target keyword is reduced, and the accuracy of the mining of the video corresponding to the hot event is improved.
[0107] The candidate video refers to a video whose content is a potential hot event. After the target keyword is obtained, the target keyword can be matched with text information of any original video in the video library to screen the original video whose text information contains the target keyword as the candidate video. After the candidate video is obtained, the candidate video can be used as a supplementary video of the seed video, and the candidate video and the seed video of the first account are used as the seed video processed for subsequent hot information mining. Alternatively, the candidate video can be directly used as the seed video, the candidate video is clustered through the text information of the candidate video, and the candidate video in each video set after clustering is classified in the first and second classifications to obtain the first number of candidate videos belonging to the event category in each video set and the second number of candidate videos belonging to the advertisement category in each video set. Finally, the hot event video set is determined from the video set according to the first number and the second number corresponding to each video set, and hot information mining is realized.
[0108] Further, a single keyword often has no burst trend, but the combination of two keywords has a burst trend. For example, the number of keywords “Beijing” or “snowstorm” included in the text information of the video is relatively stable, but the keyword “Beijing snowstorm” has a burst trend. Therefore, in an embodiment, the video keyword further includes a multi-keyword, and the multi-keyword includes at least two one-keywords; the candidate video is screened from the video library according to the target keyword, as shown in Figure 6 Further, a single keyword often has no burst trend, but the combination of two keywords has a burst trend. For example, the number of keywords “Beijing” or “snowstorm” included in the text information of the video is relatively stable, but the keyword “Beijing snowstorm” has a burst trend. Therefore, in an embodiment, the video keyword further includes a multi-keyword, and the multi-keyword includes at least two one-keywords; the candidate video is screened from the video library according to the target keyword, as shown in
[0109] Step S610, the first one-word frequency of the one-keyword in the multi-keyword in the current time window and the second one-word frequency in the historical time window are counted;
[0110] Step S620, the first multi-word frequency of the multi-keyword in the current time window and the second multi-word frequency in the historical time window are counted;
[0111] Step S630, the target keyword is screened from the multi-keyword according to the first one-word frequency, the second one-word frequency, the first multi-word frequency and the second multi-word frequency corresponding to each multi-keyword.
[0112] The multi-keyword refers to a two-keyword or more than two-keyword. Still taking the text information “A place rainstorm causes many cars to be soaked” as an example, “A place rainstorm” is a multi-keyword constructed by the two one-keywords “A place” and “rainstorm”.
[0113] Specifically, for any pair of multi-keywords, the number of occurrences of each unigram keyword in the multi-keywords in the text information of all original videos in the current time window (i.e., the first unigram frequency) and the average number of occurrences in the text information of all historical videos in the historical time window (i.e., the second unigram frequency) are calculated, respectively, and the number of occurrences of each unigram keyword in the multi-keywords together in the text information of all original videos in the current time window (i.e., the first multi-keyword frequency) and the average number of occurrences together in the text information of all historical videos in the historical time window (i.e., the second multi-keyword frequency) are calculated; then, according to the first unigram frequency, the second unigram frequency, the first multi-keyword frequency and the second multi-keyword frequency of the multi-keywords, it can be judged whether the multi-keywords are target keywords of explosive spread.
[0114] In one embodiment, the first unigram frequency, the second unigram frequency, the first multi-keyword frequency and the second multi-keyword frequency of the multi-keywords can be input into a pre-trained linear regression model, and the probability value of the multi-keywords being target keywords of explosive spread is predicted by the linear regression model, if the probability value is greater than a preset probability threshold, the multi-keywords are determined as target keywords. Specifically, taking an example that the multi-keywords include two unigram keywords, the linear regression model can be shown in the following formula (3):
[0115] z = a1x1 + a2x2 + b1y1 + b2y2 + c1x1y1 + c2x2y2 + d (3)
[0116] wherein, Z represents the probability value of the multi-keywords being target keywords, x1 represents the first unigram frequency of the first unigram keyword in the multi-keywords, x2 represents the ratio between the first unigram frequency and the second unigram frequency of the first unigram keyword in the multi-keywords, y1 represents the first unigram frequency of the second unigram keyword in the multi-keywords, y2 represents the ratio between the first unigram frequency and the second unigram frequency of the second unigram keyword in the multi-keywords, x1y1 represents the first multi-keyword frequency of the multi-keywords, x2y2 represents the ratio between the first multi-keyword frequency and the second unigram frequency of the multi-keywords; a1 and a2 are weight parameters of x1 and x2 respectively, b1 and b2 are weight parameters of y1 and y2 respectively, and d represents a bias value.
[0117] Further, the linear regression model can also introduce more feature data related to the multi-element keyword, for example, the frequency ratio of the uni-element keyword or the multi-element keyword at different time granularities can be introduced; for example, the current time window is selected as the current day (24 hours), the historical time window can be selected as the day before the current time window as the first historical time window, the week before the current time window as the second historical time window, and the month before the current time window as the third historical time window, and then the second uni-element frequency of the uni-element keyword in the first historical time window, the second uni-element frequency in the second historical time window, and the second uni-element frequency in the third historical time window are obtained, and the second multi-element frequency of the multi-element keyword in the first historical time window, the second multi-element frequency in the second historical time window, and the second multi-element frequency in the third historical time window are obtained, and then (for example, the multi-element keyword includes two uni-element keywords) the following pre-trained linear regression formula (4) can be used to determine whether the multi-element keyword is a target keyword for explosive propagation:
[0118] z=a1x1+a2x2+a3x3+a4x4+b1y1+b2y2+b3y3+b4y4+c1x1y1+c2x2y2+c3x3y3+c4x4y4+d (4)
[0119] Wherein, Z represents a probability value of the multi-keyword being the target keyword, x1 represents a first unigram frequency of a first unigram keyword in the multi-keyword, x2 represents a ratio between the first unigram frequency corresponding to the first unigram keyword and a second unigram frequency of a first historical time window, x3 represents a ratio between the first unigram frequency corresponding to the first unigram keyword and a second unigram frequency of a second historical time window, x4 represents a ratio between the first unigram frequency corresponding to the first unigram keyword and a second unigram frequency of a third historical time window; y1 represents a first unigram frequency of a second unigram keyword in the multi-keyword, y2 represents a ratio between the first unigram frequency corresponding to the second unigram keyword and the second unigram frequency of the first historical time window, y3 represents a ratio between the first unigram frequency corresponding to the second unigram keyword and the second unigram frequency of the second historical time window, y4 represents a ratio between the first unigram frequency corresponding to the second unigram keyword and the second unigram frequency of the third historical time window; x1y1 represents a first multi-keyword frequency of the multi-keyword, x2y2 represents a ratio between the first multi-keyword frequency of the multi-keyword and a second multi-keyword frequency of the first historical time window; x3y3 represents a ratio between the first multi-keyword frequency of the multi-keyword and a second multi-keyword frequency of the second historical time window; x4y4 represents a ratio between the first multi-keyword frequency of the multi-keyword and a second multi-keyword frequency of the third historical time window; a1, a2, a3, a4 are weight parameters of x1, x2, x3, x4 respectively, b1, b2, b3, b4 are weight parameters of y1, y2, y3, y4 respectively, and d represents a bias value.
[0120] In addition, the linear regression model can further introduce, as feature data, a ratio between the first multi-keyword frequency of the multi-keyword in the current time window and the first unigram frequency of each unigram keyword in the multi-keyword, to predict whether the multi-keyword is the target keyword of the explosive propagation.
[0121] After the target keyword in the multi-keyword is acquired, the multi-keyword as the target keyword and the unigram keyword as the target keyword can be collectively taken as a final target keyword, and matched with text information of any original video in the video library, to screen an original video whose text information contains the target keyword as a candidate video.
[0122] By screening the target keyword of the explosive propagation in the text information corresponding to the original video in the video library, and then performing hot event related video mining based on the target keyword, the completeness of the hot event related video mining is improved, and the omission of the hot event related video is reduced.
[0123] In order to better implement the hot information mining method provided in the embodiments of the present application, on the basis of the hot information mining method provided in the embodiments of the present application, a hot information mining device is further provided in the embodiments of the present application, likeFigure 7 As shown, the hotspot information mining device 700 comprises:
[0124] A seed video acquisition module 710 is configured to acquire a first account and a seed video of the first account.
[0125] A video clustering module 720 is configured to cluster the seed video according to the text information of the seed video, to obtain a plurality of video sets.
[0126] A first classification module 730 is configured to respectively perform first classification on the seed video in each video set, to obtain a first number of seed videos belonging to an event category in each video set.
[0127] A second classification module 740 is configured to respectively perform second classification on the seed video in each video set, to obtain a second number of seed videos belonging to an advertisement category in each video set.
[0128] A hotspot information acquisition module 750 is configured to determine a hotspot event video set from the video set according to the first number and the second number corresponding to each video set.
[0129] In some embodiments of the present application, the hotspot information acquisition module 750 is further configured to filter a hotspot video from the hotspot event video set, and extract a text label of the hotspot event based on the text information of the hotspot video.
[0130] In some embodiments of the present application, the hotspot information acquisition module 750 is further configured to filter a hotspot video from the hotspot event video set, and acquire a target video from a video library according to the text information of the hotspot video; acquire a second account corresponding to the target video, and acquire all videos published by the second account and a total number of the all videos; acquire a third number of videos belonging to the hotspot event from the all videos of the second account; if a ratio between the third number and the total number of the all videos is greater than a preset ratio, update the first account according to the second account.
[0131] In some embodiments of the present application, the hotspot information obtaining module 750 is further configured to: obtain a first label text of the hotspot video according to the text information of the hotspot video; obtain text information and a second label text of the original video in the video library; obtain label similarity feature information between the hotspot video and the original video based on the first label text and the second label text; splice the text information of the hotspot video and the text information of the original video to obtain a first spliced text, and obtain first text similarity feature information based on the first spliced text; splice the text information of the hotspot video and the second label text of the original video to obtain a second spliced text, and obtain second text similarity feature information based on the second spliced text; identify a matching result between the original video and the hotspot video based on the label similarity feature information, the first text similarity feature information, and the second text similarity feature information; and obtain the target video based on the matching result of the original video.
[0132] In some embodiments of the present application, the seed video obtaining module 710 is further configured to: extract video keywords from the text information of all original videos in the video library; wherein the video keywords include one-element keywords; count a first word frequency of each video keyword in a current time window and a second word frequency in a historical time window; select target keywords from the video keywords according to the word frequency ratio between the first word frequency and the second word frequency of each video keyword; and select candidate videos from the video library according to the target keywords, and update the seed video of the first account based on the candidate videos to obtain an updated seed video.
[0133] In some embodiments of the present application, the video keywords further include multi-element keywords, and the multi-element keywords include at least two one-element keywords; the seed video obtaining module 710 is further configured to: count a first one-element word frequency of a one-element keyword in the multi-element keyword in the current time window and a second one-element word frequency in the historical time window; count a first multi-element word frequency of the multi-element keyword in the current time window and a second multi-element word frequency in the historical time window; and select target keywords from the multi-element keywords according to the first one-element word frequency, the second one-element word frequency, the first multi-element word frequency, and the second multi-element word frequency corresponding to each multi-element keyword.
[0134] In some embodiments of the present application, the video clustering module 720 is configured to: encode the text information of each seed video respectively to obtain a text vector of each seed video; calculate the similarity between each seed video based on the text vector corresponding to each seed video; and cluster the seed videos based on the similarity between each seed video to obtain a plurality of video sets.
[0135] Specific limitations regarding the hotspot information mining device can be found in the limitations of the hotspot information mining method described above, and will not be repeated here. Each module in the aforementioned hotspot information mining device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0136] In some embodiments of this application, the hotspot information mining device 700 can be implemented as a computer program, and the computer program can be implemented in, for example... Figure 8 The computer device shown is running the program. The computer device's memory can store the various program modules that make up the hotspot information mining device 700, for example, Figure 7 The diagram shows a seed video acquisition module 710, a video clustering module 720, a first classification module 730, a second classification module 740, and a hotspot information acquisition module 750. The computer program comprised of these modules causes the processor to execute the steps in the hotspot information mining methods of the various embodiments of this application described in this specification.
[0137] For example, Figure 8 The computer device shown can be used as follows Figure 7 The seed video acquisition module 710 in the hotspot information mining device 700 shown executes step S210. The computer device can execute step S220 via the video clustering module 720. The computer device can execute step S230 via the first classification module 730. The computer device can execute step S240 via the second classification module 740. The computer device can execute step S250 via the hotspot information acquisition module 750. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with external computer devices via a network connection. When the computer program is executed by the processor, it implements a hotspot information mining method.
[0138] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0139] In some embodiments of the present application, a computer device is provided, comprising one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to perform the steps of the hotspot information mining method described above by the processor. The steps of the hotspot information mining method described above can be the steps of the hotspot information mining method in each of the embodiments described above.
[0140] In some embodiments of the present application, a computer readable storage medium is provided, storing a computer program, which is loaded by a processor to make the processor perform the steps of the hotspot information mining method described above. The steps of the hotspot information mining method described above can be the steps of the hotspot information mining method in each of the embodiments described above.
[0141] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments can be included. Any reference to memory, storage, database or other medium used in each embodiment of the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0142] Each technical feature of the above embodiments can be combined arbitrarily, and in order to make the description concise, not all possible combinations of each technical feature in the above embodiments are described, however, as long as the combination of these technical features does not exist contradictory, it should be considered as the scope of the present application.
[0143] The above provides a detailed description of the hotspot information mining method, device, computer device and storage medium provided by the embodiments of the present application. The principle and implementation mode of the present application are described by applying specific examples, and the above description of the embodiments is only used to help understand the method and core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed, and the above description should not be understood as the limitation of the present application.
Claims
1. A method for mining hotspot information, characterized in that, The method comprises: acquiring a first account and a seed video of the first account; clustering the seed video according to text information of the seed video to obtain a plurality of video sets; performing first classification on the seed video in each video set to obtain a first quantity of seed videos belonging to an event category in each video set; performing second classification on the seed video in each video set to obtain a second quantity of seed videos belonging to an advertisement category in each video set; determining a hot event video set from the video sets according to the first quantity and the second quantity corresponding to each video set.
2. The method of claim 1, wherein, After the step of determining a hot event video set from the video sets according to the first quantity and the second quantity corresponding to each video set, the method further comprises: screening a hot video from the hot event video set; extracting a text label of a hot event based on text information of the hot video.
3. The method of claim 1, wherein, After the step of determining a hot event video set from the video sets according to the first quantity and the second quantity corresponding to each video set, the method further comprises: screening a hot video from the hot event video set, and acquiring a target video from a video library according to text information of the hot video; acquiring a second account corresponding to the target video, and acquiring all videos published by the second account and a total quantity of the all videos; acquiring a third quantity of videos belonging to a hot event from the all videos of the second account; if a ratio between the third quantity and the total quantity of the all videos is greater than a preset ratio, updating the first account according to the second account.
4. The method of claim 3, wherein, The step of acquiring a target video from a video library according to text information of the hot video comprises: acquiring a first label text of the hot video according to the text information of the hot video; acquiring text information and a second label text of an original video in the video library; acquiring label similarity feature information between the hot video and the original video based on the first label text and the second label text; splicing the text information of the hot video and the text information of the original video to obtain a first spliced text, and acquiring first text similarity feature information based on the first spliced text; splicing the text information of the hot video and the second label text of the original video to obtain a second spliced text, and acquiring second text similarity feature information based on the second spliced text; identifying a matching result between the original video and the hot video based on the label similarity feature information, the first text similarity feature information, and the second text similarity feature information; acquiring a target video based on the matching result of the original video.
5. The method of claim 1, wherein, Before the step of clustering the seed video according to text information of the seed video to obtain a plurality of video sets, the method further comprises: extracting a video keyword from text information of all original videos in a video library; wherein the video keyword comprises a one-element keyword; statistically acquiring a first word frequency of each video keyword in a current time window and a second word frequency of each video keyword in a historical time window; screen target keywords from the video keywords according to a keyword frequency ratio between the first keyword frequency and the second keyword frequency of each of the video keywords; screen candidate videos from the video library according to the target keywords, update the seed videos of the first account based on the candidate videos, and obtain updated seed videos.
6. The method of claim 5, wherein, The video keywords further include multi-element keywords, and the multi-element keywords include at least two one-element keywords. Before the step of screening candidate videos from the video library according to the target keywords, updating the seed videos of the first account based on the candidate videos, and obtaining updated seed videos, the method further includes: statistically obtaining a first one-element keyword frequency of the one-element keywords in a current time window and a second one-element keyword frequency of the one-element keywords in a historical time window; statistically obtaining a first multi-element keyword frequency of the multi-element keywords in the current time window and a second multi-element keyword frequency of the multi-element keywords in the historical time window; screen target keywords from the multi-element keywords according to the first one-element keyword frequency, the second one-element keyword frequency, the first multi-element keyword frequency, and the second multi-element keyword frequency of each of the multi-element keywords.
7. The method according to any one of claims 1 to 6, characterized in that, The step of clustering the seed videos according to the text information of the seed videos to obtain a plurality of video sets includes: respectively encode the text information of each of the seed videos to obtain a text vector of each of the seed videos; calculate the similarity between each of the seed videos based on the text vector corresponding to each of the seed videos; cluster the seed videos based on the similarity between each of the seed videos to obtain a plurality of video sets.
8. A hot spot information mining apparatus characterized by comprising: The device includes: a seed video acquisition module configured to acquire a first account and seed videos of the first account; a video clustering module configured to cluster the seed videos according to text information of the seed videos to obtain a plurality of video sets; a first classification module configured to respectively perform first classification on the seed videos in each of the video sets to obtain a first number of seed videos belonging to an event category in each of the video sets; a second classification module configured to respectively perform second classification on the seed videos in each of the video sets to obtain a second number of seed videos belonging to an advertisement category in each of the video sets; a hot spot information acquisition module configured to determine a hot spot event video set from the video sets according to the first number and the second number corresponding to each of the video sets.
9. A computer device, comprising: The computer device includes: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement the hot spot information mining method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps of the hot spot information mining method of any one of claims 1 to 7.
11. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions are executed by a processor to implement the steps of the hot spot information mining method of any one of claims 1 to 7.
Citation Information
Patent Citations
News event mining method and device, computer equipment and storage medium
CN108170773A