Method, device and equipment for monitoring hot events and storage medium
Patent Information
- Application Number
- CN202311506806.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-11-13
AI Technical Summary
[0003]现有的技术方案,是基于资讯文章聚类相似度的方式发现热点,等聚类的文章数据达到一定阈值即认为是相关热点事件,此方案缺乏时效性以及误判其他类型文章为热点事件;另外基于用户行为数据捕捉热点,虽然时效性较好,但是无法及时获取相应的资讯文章
[0016]本公开通过从消费端和生产端两方面联合,解决了时效性不及时的问题,并且基于聚类算法,提升了热点事件的识别准确性。
Smart Images

Figure CN117493488B_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the field of data monitoring technology, and in particular relates to a method, apparatus, equipment, and storage medium for monitoring hot events. Background Technology
[0002] Trending events are a type of highly valuable event data, therefore, they need to be monitored accurately and promptly.
[0003] Existing technical solutions identify hot topics based on the similarity of news articles clustered together. Once the clustered articles reach a certain threshold, they are considered relevant hot topics. However, this approach lacks timeliness and may misidentify other types of articles as hot topics. In addition, while capturing hot topics based on user behavior data has better timeliness, it cannot obtain the relevant news articles in a timely manner.
[0004] Therefore, existing technical solutions suffer from poor accuracy and timeliness. Summary of the Invention
[0005] This disclosure provides a method, apparatus, device, and storage medium for monitoring hot events.
[0006] According to a first aspect of this disclosure, a method for acquiring consumer-side data and production-side data is provided; The acquired consumer data is preprocessed, and the sequence data at each time point is determined based on the query terms in the preprocessed consumer data. The sequence data at each time point is then divided according to a preset time window to determine the time series window for the consumer. The acquired production data is preprocessed, and the sequence data for each time point is determined based on the keywords in the preprocessed production data and the number of articles containing the keywords in each time period. The time series window for the production end is determined based on the sequence data for each time point according to the preset time window. Based on the time series windows of the consumer end, the mean and standard deviation of each time series window are determined, and the slope between the predicted point and the adjacent time point is calculated. A preset hotspot identification algorithm is used to determine the hotspot data of the consumer end. Based on the time series windows of the production end, the mean and standard deviation of each time series window are determined, and the slope between the prediction point and the adjacent time point is calculated. A preset hotspot identification algorithm is used to determine the hotspot data of the production end. Based on the hot search terms and hot list in the consumer hot data, the first pre-set clustering algorithm is used to cluster them to obtain consumer hot term event clusters; Based on hot keywords in production-side hot data, when the number of clustered articles with the same event in the production-side hot data reaches a preset threshold, a production-side hot keyword event cluster is obtained. By associating consumer-side hot topic event clusters with production-side hot topic event clusters, hot topics can be identified.
[0007] In some implementations of the first aspect, the consumer-side data includes first consumer-side data and second consumer-side data; The first consumer-side data is determined based on a user search query log dataset within a preset time window; The second consumer-side data is determined based on the hot topic ranking data of the target webpage crawled by web crawlers; The production-side data is determined based on a dataset of news articles crawled in real time by web crawlers.
[0008] In some implementations of the first aspect, the acquired consumer data is preprocessed, including: Filter query terms with a length greater than the first threshold, query terms with a length less than the second threshold, non-event query terms, and links to non-news media from the first consumer data. Determine the search volume of the query terms in the filtered first consumer data at each target time point and the total search volume in the target time period. Based on the total search volume in the target time period, further filter the filtered first consumer data for query terms whose total search volume is less than the third threshold to obtain the final filtered first consumer data. The second consumer data is filtered and deduplicated to obtain the filtered second consumer data. The preprocessing of the acquired production data includes: Filter out news articles lacking titles, news articles lacking body text, news titles containing questions, question marks, exclamation marks, and clickbait news data from the production data; It also extracts entities and keywords from the filtered production data based on a pre-set information extraction algorithm, and counts the number of news articles containing entities or keywords in each time period.
[0009] In some implementations of the first aspect, the method also includes: The pre-processed production data is clustered using the second preset clustering algorithm to obtain a coarse clustering result. In the coarse clustering results, the corresponding articles are indexed sequentially based on their keywords in a preset information index library to obtain articles whose relevance meets the preset filtering conditions, and text pairs are constructed. After preprocessing each text pair, they are sequentially input into a preset event similarity model to determine the similarity between any two articles and whether they belong to the same event. If two articles belong to the same event, they will be tagged with the same event and stored in the event database; If the events do not belong to the same event and there is no corresponding event for the article in the preset information database, then create a new event tag for the article and store it in the event database.
[0010] In some implementations of the first aspect, the hot query terms and hot list based on consumer hot data are clustered using a first pre-set clustering algorithm to obtain consumer hot term event clusters, including: Using the first pre-built clustering algorithm, the textual and semantic similarity between the hot query terms in the consumer hot data and the hot list is judged to obtain consumer hot term event clusters.
[0011] In some implementations of the first aspect, the preset hotspot identification algorithm satisfies the formula: in, y To predict search volume or number of articles at a specific point in time, avg The average value over the time window. Here, is the influence coefficient, and std is the standard deviation of the time window. ratio To predict the slope of adjacent points, threshold It is the threshold of the slope. y i This refers to the number of searches or articles at the current point in time. y i-1 This refers to the search volume or number of articles at the previous point in time. min This represents the maximum number of low-ranking search results or articles. max This represents the minimum value for high search volume or article volume. others To remove as well as Other situations.
[0012] In some implementations of the first aspect, the method also includes: The article quality assessment criteria for production-side data are based on the article's publishing site, author, title, relevance between title and content, and the combination of text and images. The quality of articles produced is determined based on the article quality judgment criteria.
[0013] According to a second aspect of this disclosure, a hotspot event monitoring device is provided, the device comprising: The acquisition module is used to acquire data from both the consumer and production ends. The processing module is used to preprocess the acquired consumer data and determine the sequence data at each time point based on the query terms in the preprocessed consumer data. Based on the sequence data at each time point, the time series window of the consumer is determined according to the preset time window. The processing module is also used to preprocess the acquired production data, and determine the sequence data at each time point based on the keywords in the preprocessed production data and the number of articles containing the keywords in each time period. Based on the sequence data at each time point, the time series window of the production end is determined by dividing it according to a preset time window. The processing module is also used to determine the average value and standard deviation of each time series window based on the time series window of the consumer end, calculate the slope between the prediction point and the adjacent time point, and use a preset hotspot identification algorithm to determine the hotspot data of the consumer end. The processing module is also used to determine the average value and standard deviation of each time series window based on the time series window of the production end, and to calculate the slope between the prediction point and the adjacent time point, and to use a preset hotspot identification algorithm to determine the hotspot data of the production end. The processing module is also used to perform clustering based on the hot query terms and hot list in the consumer hot data using a first preset clustering algorithm to obtain consumer hot term event clusters; The processing module is also used to obtain a production-side hot keyword event cluster based on hot keywords in the production-side hot data, when the number of clustered articles with the same event in the production-side hot data reaches a preset threshold. The processing module is also used to associate the consumer-side hot word event clusters and the production-side hot word event clusters to determine hot events.
[0014] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the program to implement the method described above.
[0015] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to a first aspect of this disclosure.
[0016] This disclosure addresses the issue of untimely information delivery by combining efforts from both the consumer and production sides, and improves the accuracy of identifying trending events based on clustering algorithms.
[0017] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0018] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A flowchart illustrating a method for monitoring hot events according to an embodiment of the present disclosure is shown. Figure 2 A block diagram of a hotspot event monitoring device according to an embodiment of the present disclosure is shown; Figure 3 A block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0020] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0021] Trending events are a type of highly valuable event data, therefore, they need to be monitored accurately and promptly.
[0022] Currently, major search engines recommend relevant trending topics in the search box and display trending content / events, related encyclopedia entries / videos at the top of the search results page. Therefore, the ability to quickly discover relevant trending events and user-focused events from massive amounts of news information / events and user behavior data, thereby guiding users to read about trending events, is a crucial factor affecting user experience. At the same time, from the perspective of content production, it is essential to quickly capture, track, operate, and produce trending events, and build a trending event production and consumption system.
[0023] Therefore, it is currently urgent to address the challenges of ensuring the timeliness, accuracy, and comprehensiveness of information related to trending events. Timeliness issues: Discovering trending events requires timely monitoring of all points on the website. Events need to be discovered promptly from both the search consumption end and the information production end, and corresponding content needs to be extracted in order to effectively help recommend and distribute trending events. For example, the incident of a passenger being trapped in the subway in City A.
[0024] Accuracy issues: The discovery of trending topics and content relies on algorithms to identify / mine trending events. This evaluation is based on the event itself, and it is necessary to determine whether the discovered / mined trending event is indeed an event.
[0025] Therefore, existing technical solutions suffer from poor accuracy and timeliness.
[0026] This invention, based on data from both the consumer search engine and the information production engine, effectively improves the accuracy, timeliness, and comprehensiveness of hot topic discovery through algorithmic capabilities.
[0027] Figure 1 A flowchart illustrating a method for monitoring hotspot events according to an embodiment of this disclosure is shown, such as... Figure 1 As shown, the hotspot event monitoring method 100 may include: S101, acquires consumer and production data; S102, preprocess the acquired consumer data, and determine the sequence data for each time point based on the query terms in the preprocessed consumer data. Then, divide the sequence data for each time point according to a preset time window to determine the time series window for the consumer. S103, preprocess the acquired production data, and determine the sequence data for each time point based on the keywords in the preprocessed production data and the number of articles containing the keywords in each time period. Based on the sequence data for each time point, divide it according to the preset time window to determine the time series window of the production end. S104: Based on the time series window of the consumer end, determine the mean and standard deviation of each time series window, calculate the slope between the prediction point and the adjacent time point, and use the preset hotspot identification algorithm to determine the hotspot data of the consumer end. S105, based on the time series window of the production end, determines the mean and standard deviation of each time series window, and calculates the slope between the prediction point and the adjacent time point, and uses a preset hotspot identification algorithm to determine the hotspot data of the production end; S106, Based on the hot query terms and hot list in the consumer hot data, the first pre-set clustering algorithm is used to cluster them to obtain consumer hot term event clusters; S107. Based on the hot keywords in the production hot data, when the number of clustered articles of the same event in the production hot data reaches a preset threshold, the production hot keyword event cluster is obtained. S108: Associate the consumer-side hot topic event clusters and the production-side hot topic event clusters to identify hot topics.
[0028] In the S101-S108 process, the problem of untimely delivery was solved by combining the consumer and production sides, and the accuracy of identifying hot events was improved based on clustering algorithms.
[0029] In some embodiments, consumer data may include first consumer data and second consumer data; The first consumer-side data is determined based on the user search query log dataset within a preset time window, which is search click log data. This data may include search results corresponding to search terms and user click titles / links. The second consumer-side data is determined based on the hot topic ranking data of the target webpage crawled by web crawlers; The production-side data is determined based on a dataset of news articles crawled in real time by web crawlers, which is the information database.
[0030] In the above embodiments, the comprehensiveness of the data is ensured by collecting and crawling consumer-end data and production-end data.
[0031] In some embodiments, preprocessing the acquired consumer data may include: Filter query terms with a length greater than the first threshold, query terms with a length less than the second threshold, non-event query terms, and links to non-news media from the first consumer data. Determine the search volume of the query terms in the filtered first consumer data at each target time point and the total search volume in the target time period. Based on the total search volume in the target time period, further filter the filtered first consumer data for query terms with a total search volume less than a third threshold to obtain the final filtered first consumer data. Each time point can be set to a day, an hour, or a shorter period according to actual needs. Filter out query terms with low total search volume as needed. At the same time, the number of news media links clicked at each time point can also be counted, for example, the top 3 can be selected. The second consumer data is filtered and deduplicated to obtain filtered second consumer data, which can be considered as hot query terms / events; The preprocessing of the acquired production data includes: Filter out news articles lacking titles, news articles lacking body text, news titles containing questions, question marks, exclamation marks, and clickbait news data from the production data; It also extracts entities, i.e. "proper nouns", and keywords from the filtered production data based on a pre-set information extraction algorithm, and counts the number of news articles containing entities or keywords in each time period.
[0032] In the above embodiments, by filtering production-side data and consumption-side data, the accuracy of the data source is ensured, thereby making subsequent calculations more accurate and timely.
[0033] In some embodiments, S102, the acquired consumer data is preprocessed, and the sequence data at each time point is determined based on the query terms in the preprocessed consumer data. The sequence data at each time point is divided according to a preset time window. In the process of determining the time series window of the consumer, for example, the time window is counted by days, so seven days can be used as a time window. For example, 2023.7.1~2023.7.6 is a time window, 2023.7.2~2023.7.7, and so on.
[0034] In some embodiments, the method may further include: The pre-configured clustering algorithm is used to cluster the pre-processed production data to obtain a rough clustering result. This step is to alleviate the server pressure and process the data in batches. In the coarse clustering results, the corresponding articles are indexed sequentially based on their keywords in a preset information index library to obtain articles whose relevance meets the preset filtering conditions, and text pairs are constructed. After preprocessing each text pair, they are sequentially input into a preset event similarity model to determine the similarity between any two articles and whether they belong to the same event. If two articles belong to the same event, they will be tagged with the same event and stored in the event database; If the events do not belong to the same event and there is no corresponding event for the article in the preset information database, then create a new event tag for the article and store it in the event database; Finally, the clustered events can be stored in an event library for easy retrieval later.
[0035] In the above embodiments, text pairs are constructed by coarse clustering and it is determined that two articles belong to the same event, which ensures the calculation speed and thus the timeliness.
[0036] Specifically, the second pre-set clustering algorithm is mainly a coarse-grained clustering algorithm, which is based on the k-Nearest Neighbor (kNN) algorithm in machine learning. It trains the coarse-grained clustering algorithm on the article title and body content. The number of coarse-grained clusters mainly depends on the number of multi-threaded processes in the streaming process. This is mainly to alleviate the latency caused by single-pass.
[0037] In addition, the main features of the algorithm in the preset event similarity model are: extracting key information (entities, keywords) from the article and extracting the most important sentences related to the title, such as 5 sentences; Based on the TextRank algorithm, the extracted article entities and words are calculated to obtain key entities, keywords and key numbers. Based on this key information, the cosine similarity is calculated by constructing a one-hot encoding matrix to obtain the key information similarity score, which is defined as the text structure similarity score. The similarity algorithm model is trained based on the pre-trained model of Bidirectional Encoder Representations from Transformers (BERT). By adding key sentences to the titles of two articles, a classification model is used to determine whether the two articles are similar events, which is defined as semantic similarity score. While semantic similarity can determine whether two events are similar at the semantic level, it cannot determine their textual similarity. For example, if two articles about heavy rain in location A and location B have high semantic similarity, it would be easy to misclassify them as similar events if semantic similarity were used for judgment. Therefore, it is necessary to judge similarity by textual structure similarity. Furthermore, it is highly sensitive to numbers appearing in the same time title. Therefore, key numbers must be consistent (for example, "Floods in one location caused 3 injuries" and "Floods in another location caused 5 injuries with the number continuing to rise" are not the same event). Thus, high similarity of these key information is necessary to improve the accuracy of similar events.
[0038] In some embodiments, S106, during the process of clustering consumer hot topic event clusters using a first pre-set clustering algorithm based on hot topic query terms and hot topic lists in consumer hot topic data, the first pre-set clustering algorithm clusters similar hot topic query terms for hot topic query terms, determines whether they are similar hot topic terms, and thus forms hot topic event clusters.
[0039] It should also be noted that the first pre-set clustering algorithm judges text similarity and semantic similarity. Text similarity includes the text similarity calculated from key entities, keywords, and sensitive numbers, while the semantic similarity score is the semantic similarity score obtained by training the BERT model classification algorithm. The combination of these two is the first pre-set clustering algorithm.
[0040] In some embodiments, the clustering of hot query terms and hot topic lists based on consumer hot topic data using a first pre-set clustering algorithm yields consumer hot topic event clusters, including: Using the first pre-built clustering algorithm, the textual and semantic similarity between the hot query terms in the consumer hot data and the hot list is judged to obtain consumer hot term event clusters.
[0041] In the above embodiments, text similarity and semantic similarity are used to determine the cluster of hot topic events on the consumer side, thereby accurately obtaining the cluster of hot topic events on the consumer side.
[0042] In some embodiments, the preset hotspot identification algorithm satisfies the formula: in, y To predict search volume or number of articles at a specific point in time, avg The average value over the time window. Here, is the influence coefficient, and std is the standard deviation of the time window. and std Combined, determine whether it exceeds a certain upper limit of the average. ratio To predict the slope of adjacent points, threshold This is the threshold for the slope, which varies slightly depending on the data source. A default value of 0.5 can be used. y i This refers to the number of searches or articles at the current point in time. y i-1 This refers to the search volume or number of articles at the previous point in time. min This represents the maximum number of low-ranking search results or articles. max This represents the minimum value for high search volume or article volume. others To remove as well as Other situations.
[0043] For example, if the search volume is min=50 and max=250, the slope ratio >= threshold can be used to comprehensively determine whether it is a hotspot.
[0044] In addition, after determining the predicted points, in order to continue using the sliding time window for hotspot identification and prevent the peak values of previous hotspot time points from causing inaccurate judgment of subsequent hotspots, it is necessary to perform a smoothing operation on the predicted points.
[0045] The smoothing formula is: in, yi_new This represents the amount of data after smoothing the predicted points. y i This represents the amount of data before smoothing. y i-1 This represents the amount of data at the time point preceding the prediction point. influence This represents the impact factor, with a default value of 0.5.
[0046] Since hotspot identification algorithms can identify whether user query terms are hot topics, but cannot accurately determine whether they are event-related query terms, further filtering is required. The filtering rules can be based on a series of links / titles clicked by the user based on the user query terms, such as filtering to get the top 5, filtering based on manually selected news sites to filter out user query terms not included in these news sites, and filtering out clicked low-quality links through article quality algorithms. In this way, the remaining user query terms can be considered a batch of high-quality hot topic event clusters.
[0047] In some embodiments, the method may further include: using the article's publishing site, author, article title, relevance between title and content, and richness of images and text as criteria for judging the quality of the production-end data; The quality of articles produced is determined based on the article quality judgment criteria.
[0048] In other words, the quality of articles in production-side data can be judged by factors such as the publishing site of the news article (whether it is authoritative), the author (whether it is an author in the field), the article title (whether it is clickbait), whether the title is relevant to the content, and whether it is rich in pictures and text. Then, based on the evaluation criteria for the quality of articles, these features are transformed into algorithmic features, and a high-quality algorithm is trained to judge the quality of articles produced. Article quality algorithm features can specifically include: whether the publishing site is authoritative (e.g., The Paper, People's Daily, etc.), whether the author is an author in the field (whether the author's field matches the topic of the published article), whether the article title is clickbait, whether the title is relevant to the body text, whether it contains images, etc. By fitting these features, an article quality scoring model is trained based on machine learning algorithms, and this model is used to determine whether the article is of high quality.
[0049] Because the algorithm also includes the feature of authority, it not only judges the quality of an article, but also whether the article is authoritative, etc.
[0050] In some embodiments, S108, the process of associating consumer-side hot word event clusters and production-side hot word event clusters to determine hot events may include: By searching the event database for the tags of the news articles that contain clicks in the hot keyword event cluster, you can associate them with the hot keyword cluster and the hot event cluster. If no corresponding event tag is found, the indexing capability is used to retrieve the most relevant event. If it is not found, it means that the corresponding event has not yet been entered, but hot query term event clusters have been discovered / mined. Once the corresponding event is entered, it can be associated with the event. If a hot topic is identified on the production side but cannot be associated with a cluster of hot keywords, it is necessary to analyze the hot topic on the production side, extract the corresponding hot keywords as a cluster of hot keywords, and fill the gap in the timely discovery of hot keyword clusters on the consumer side.
[0051] This invention, by combining the consumer and production ends, can solve both the problems of untimely delivery and incomplete coverage of hot events. At the same time, it improves the accuracy of hot event identification and enhances the clustering accuracy of event clusters based on event clustering algorithms.
[0052] Furthermore, existing clustering algorithms assess similarity based on the original text or semantics, but lack some key elements in the event domain. For example, semantic similarity cannot distinguish between entities or keywords with similar semantics (e.g., place names may have similar spatial vectors semantically, but they are not the same event). Therefore, this invention, based on the original semantic similarity, not only adds text similarity such as key entities / keywords, but also adds sensitive information (such as numbers) in specific domains to strictly limit event similarity, thereby improving the accuracy of event clustering.
[0053] Furthermore, the present invention associates hot word event clusters with hot events. The hot word event clusters obtained by the hot spot identification algorithm are associated with the event library obtained by the clustering / similar event algorithm. Hot events are obtained by associating the event library with clicked links and indexing capabilities.
[0054] In summary, this invention, based on features such as key text information, semantic information, and sensitive information in specific domains, determines an event similarity algorithm. It detects trending events through both consumer and production ends. The consumer end, relying on user behavior, cannot always promptly identify trending events. Therefore, it needs to crawl information data from the production end and utilize aggregation, similarity judgment, and trending event identification capabilities to discover trending events, thus improving timeliness. Meanwhile, the consumer end helps improve the comprehensiveness and accuracy of trending events. The combination of both improves overall timeliness, coverage, and accuracy.
[0055] The above is an introduction to the method embodiments. The following describes the present disclosure further through device embodiments.
[0056] Figure 2 A block diagram of a hotspot event monitoring device according to an embodiment of the present disclosure is shown.
[0057] like Figure 2 As shown, the hotspot event monitoring device 200 includes: Module 201 is used to acquire consumer-side data and production-side data; The processing module 202 is used to preprocess the acquired consumer data, and determine the sequence data at each time point based on the query terms in the preprocessed consumer data. Based on the sequence data at each time point, the time series window of the consumer is determined by dividing it according to a preset time window. The processing module 202 is also used to preprocess the acquired production data, and determine the sequence data at each time point based on the keywords in the preprocessed production data and the number of articles containing the keywords in each time period, and divide the sequence data at each time point according to a preset time window to determine the time series window of the production end. The processing module 202 is also used to determine the average value and standard deviation of each time series window based on the time series window of the consumer end, calculate the slope between the prediction point and the adjacent time point, and use a preset hotspot identification algorithm to determine the hotspot data of the consumer end. The processing module 202 is also used to determine the average value and standard deviation of each time series window based on the time series window of the production end, and to calculate the slope between the prediction point and the adjacent time point, and to use a preset hotspot identification algorithm to determine the hotspot data of the production end. The processing module 202 is further configured to perform clustering based on the hot query terms and hot list in the consumer hot data using a first preset clustering algorithm to obtain consumer hot term event clusters; The processing module 202 is also used to obtain a production-end hot keyword event cluster based on hot keywords in the production-end hot data, when the number of clustered articles with the same event in the production-end hot data reaches a preset threshold. The processing module 202 is also used to associate the consumer-side hot word event cluster and the production-side hot word event cluster to determine hot events.
[0058] In some embodiments, the consumer data includes first consumer data and second consumer data; The first consumer-side data is determined based on a user search query log dataset within a preset time window; The second consumer-side data is determined based on the hot topic ranking data of the target webpage crawled by web crawlers; The production-side data is determined based on a dataset of news articles crawled in real time by web crawlers.
[0059] In some embodiments, the processing module 202 is further configured to filter query terms with a length greater than a first threshold, query terms with a length less than a second threshold, non-event query terms, and links to non-news media in the first consumer data. Determine the search volume of the query terms in the filtered first consumer data at each target time point and the total search volume in the target time period. Based on the total search volume in the target time period, further filter the filtered first consumer data for query terms whose total search volume is less than the third threshold to obtain the final filtered first consumer data. The second consumer data is filtered and deduplicated to obtain the filtered second consumer data. The processing module 202 is also used to filter out news articles that lack titles, news articles that lack body text, news titles containing questions, question marks, exclamation marks, and clickbait news data from the production data. It also extracts entities and keywords from the filtered production data based on a pre-set information extraction algorithm, and counts the number of news articles containing entities or keywords in each time period.
[0060] In some embodiments, the processing module 202 is further configured to use a second preset clustering algorithm to cluster the preprocessed production data to obtain a coarse clustering result; In the coarse clustering results, the corresponding articles are indexed sequentially based on their keywords in a preset information index library to obtain articles whose relevance meets the preset filtering conditions, and text pairs are constructed. After preprocessing each text pair, they are sequentially input into a preset event similarity model to determine the similarity between any two articles and whether they belong to the same event. If two articles belong to the same event, they will be tagged with the same event and stored in the event database; If the events do not belong to the same event and there is no corresponding event for the article in the preset information database, then create a new event tag for the article and store it in the event database.
[0061] In some embodiments, the processing module 202 is further configured to use a first preset clustering algorithm to determine the text similarity and semantic similarity between hot query terms in consumer hot data and hot list, and obtain consumer hot term event clusters.
[0062] In some embodiments, the preset hotspot identification algorithm satisfies the formula: in, y To predict search volume or number of articles at a specific point in time, Here, is the influence coefficient, and std is the standard deviation of the time window. avg The average value over the time window. ratio To predict the slope of adjacent points, threshold It is the threshold of the slope. y i This refers to the number of searches or articles at the current point in time. y i-1 This refers to the search volume or number of articles at the previous point in time. others To remove as well as Other situations.
[0063] In some embodiments, the processing module 202 is further configured to use the article's publishing site, author, article title, relevance between title and content, and richness of images and text as criteria for judging the quality of the production-end data. The quality of articles produced is determined based on the article quality judgment criteria.
[0064] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this disclosure.
[0065] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0066] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0067] Figure 3A block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0068] Device 300 includes a computing unit 301, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 302 or a computer program loaded from storage unit 308 into random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via bus 304. Input / output (I / O) interface 305 is also connected to bus 304.
[0069] Multiple components in device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of monitors, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0070] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of method 100 described above may be performed.
[0071] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0072] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0073] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0074] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0075] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0076] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0077] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0078] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for monitoring trending events, characterized in that, The method includes: Acquire consumer and production data; The acquired consumer data is preprocessed, and the sequence data at each time point is determined based on the query terms in the preprocessed consumer data. The sequence data at each time point is then divided according to a preset time window to determine the time series window for the consumer. The acquired production data is preprocessed, and the sequence data for each time point is determined based on the keywords in the preprocessed production data and the number of articles containing the keywords in each time period. The time series window for the production end is determined based on the sequence data for each time point according to the preset time window. Based on the time series windows of the consumer end, the mean and standard deviation of each time series window are determined, and the slope between the predicted point and the adjacent time point is calculated. A preset hotspot identification algorithm is used to determine the hotspot data of the consumer end. Based on the time series windows of the production end, the mean and standard deviation of each time series window are determined, and the slope between the prediction point and the adjacent time point is calculated. A preset hotspot identification algorithm is used to determine the hotspot data of the production end. Based on the hot search terms and hot list in the consumer hot data, the first pre-set clustering algorithm is used to cluster them to obtain consumer hot term event clusters; Based on hot keywords in production-side hot data, when the number of clustered articles with the same event in the production-side hot data reaches a preset threshold, a production-side hot keyword event cluster is obtained. By associating consumer-side hot topic event clusters with production-side hot topic event clusters, hot topics can be identified.
2. The method according to claim 1, characterized in that, The consumer-side data includes first consumer-side data and second consumer-side data; The first consumer-side data is determined based on a user search query log dataset within a preset time window; The second consumer-side data is determined based on the hot topic ranking data of the target webpage crawled by web crawlers; The production-side data is determined based on a dataset of news articles crawled in real time by web crawlers.
3. The method according to claim 2, characterized in that, The acquired consumer data is preprocessed, including: Filter query terms with a length greater than the first threshold, query terms with a length less than the second threshold, non-event query terms, and links to non-news media from the first consumer data. Determine the search volume of the query terms in the filtered first consumer data at each target time point and the total search volume in the target time period. Based on the total search volume in the target time period, further filter the filtered first consumer data for query terms whose total search volume is less than the third threshold to obtain the final filtered first consumer data. The second consumer data is filtered and deduplicated to obtain the filtered second consumer data. The preprocessing of the acquired production data includes: Filter out news articles lacking titles, news articles lacking body text, news titles containing questions, question marks, exclamation marks, and clickbait news data from the production data; It also extracts entities and keywords from the filtered production data based on a pre-set information extraction algorithm, and counts the number of news articles containing entities or keywords in each time period.
4. The method according to claim 1, characterized in that, The method further includes: The pre-processed production data is clustered using the second preset clustering algorithm to obtain a coarse clustering result. In the coarse clustering results, the corresponding articles are indexed sequentially based on their keywords in a preset information index library to obtain articles whose relevance meets the preset filtering conditions, and text pairs are constructed. After preprocessing each text pair, they are sequentially input into a preset event similarity model to determine the similarity between any two articles and whether they belong to the same event. If two articles belong to the same event, they will be tagged with the same event and stored in the event database; If the events do not belong to the same event and there is no corresponding event for the article in the preset information database, then create a new event tag for the article and store it in the event database.
5. The method according to claim 1, characterized in that, The hot query terms and hot list based on consumer hot data are clustered using a first pre-set clustering algorithm to obtain consumer hot term event clusters, including: Using the first pre-built clustering algorithm, the textual and semantic similarity between the hot query terms in the consumer hot data and the hot list is judged to obtain consumer hot term event clusters.
6. The method according to claim 1, characterized in that, The preset hotspot identification algorithm satisfies the formula: in, y To predict search volume or number of articles at a specific point in time, avg The average value over the time window. Here, is the influence coefficient, and std is the standard deviation of the time window. ratio To predict the slope of adjacent points, threshold It is the threshold of the slope. y i This refers to the number of searches or articles at the current point in time. y i-1 This refers to the search volume or number of articles at the previous point in time. min This represents the maximum number of low-ranking search results or articles. max This represents the minimum value for high search volume or article volume. others To remove as well as Other situations.
7. The method according to claim 1, characterized in that, The method further includes: The article quality assessment criteria for production-side data are based on the article's publishing site, author, title, relevance between title and content, and the combination of text and images. The quality of articles produced is determined based on the article quality judgment criteria.
8. A monitoring device for hotspot events, characterized in that, The device includes: The acquisition module is used to acquire data from both the consumer and production ends. The processing module is used to preprocess the acquired consumer data and determine the sequence data at each time point based on the query terms in the preprocessed consumer data. Based on the sequence data at each time point, the time series window of the consumer is determined according to the preset time window. The processing module is also used to preprocess the acquired production data, and determine the sequence data at each time point based on the keywords in the preprocessed production data and the number of articles containing the keywords in each time period. Based on the sequence data at each time point, the time series window of the production end is determined by dividing it according to a preset time window. The processing module is also used to determine the average value and standard deviation of each time series window based on the time series window of the consumer end, calculate the slope between the prediction point and the adjacent time point, and use a preset hotspot identification algorithm to determine the hotspot data of the consumer end. The processing module is also used to determine the average value and standard deviation of each time series window based on the time series window of the production end, and to calculate the slope between the prediction point and the adjacent time point, and to use a preset hotspot identification algorithm to determine the hotspot data of the production end. The processing module is also used to perform clustering based on the hot query terms and hot list in the consumer hot data using a first preset clustering algorithm to obtain consumer hot term event clusters; The processing module is also used to obtain a production-side hot keyword event cluster based on hot keywords in the production-side hot data, when the number of clustered articles with the same event in the production-side hot data reaches a preset threshold. The processing module is also used to associate the consumer-side hot word event clusters and the production-side hot word event clusters to determine hot events.
9. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory is characterized in that it stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Content popularity prediction method and device based on artificial intelligence and computer equipment
CN111339404A
Hot spot searching method and device, terminal and storage medium
CN112307304A