An information intelligent aggregation method, device and storage medium

By automatically updating the keyword set and using the TF-IDF algorithm, the problem of keyword lag caused by manual maintenance in existing technologies has been solved, realizing intelligent aggregation of information channels and timely updates of trending news.

CN114398544BActive Publication Date: 2026-02-24SHANGHAI JUJUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111675792.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2026-02-24
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The rationality of the keyword set in the existing news channel relies on manual maintenance, which leads to delays in keyword addition and timely updates, causing the channel to miss the latest relevant information.

Method used

By creating an initial set of official keywords and a set of candidate keywords with weighted values, periodically acquiring information, calculating attention levels, automatically updating the keyword set, using word segmentation and TF-IDF algorithms to extract core keywords, setting thresholds to filter keywords, and achieving automatic aggregation and updating.

Benefits of technology

It achieves the ability to generate objective keyword sets in a timely and accurate manner without manual maintenance, and automatically updates trending news, thereby improving the accuracy and timeliness of information aggregation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114398544B_ABST
    Figure CN114398544B_ABST
Patent Text Reader

Abstract

The application provides an information intelligent aggregation method, device and storage medium, wherein the information intelligent aggregation method comprises the following steps: creating an official keyword set based on a required information type, setting an initial weight value for each keyword, setting a total initial weight value of all keywords in the official keyword set as a constant, setting a candidate keyword set, and the candidate keyword set is initially empty, periodically acquiring information to form an information set, screening and extracting an information subset by using the official keyword set, calculating information attention degrees of each information in the information subset according to a total amount of attention of the information subset, processing each information of the information subset, re-distributing keyword weights, and obtaining a new official keyword set and a new candidate keyword set. No manual maintenance is needed, manpower is saved, and an objective keyword set is formed. The keyword set can be automatically updated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of news aggregation, specifically to a method, device, and storage medium for intelligent information aggregation. Background Technology

[0002] Currently, news channels primarily rely on manual selection of relevant topics. Some channels maintain a set of keywords manually, using computer programs to filter news containing those keywords. The drawback is that the validity and completeness of this keyword set depend on the maintainer's skill level, making it highly random and unreliable. Furthermore, a topic constantly evolves, generating new keywords. Adding these new keywords manually often results in a significant delay, causing the channel to frequently miss the latest relevant information. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the purpose of this invention is to provide an intelligent information aggregation method. The technical solution of this invention is as follows:

[0004] A method for intelligent information aggregation includes the following steps:

[0005] Create a formal keyword set based on the required information type, set an initial weight value for each keyword, set the total initial weight value of all keywords in the formal keyword set to a constant, set a candidate keyword set, and the candidate keyword set is initially an empty set;

[0006] Information is periodically acquired to form an information set. A subset of information is then extracted using a set of formal keywords. Based on the total attention received by each information subset, the information attention level of each information in the subset is calculated.

[0007] The information subset is processed to obtain a new set of keywords and a new set of candidate keywords.

[0008] Based on the above technical solution and as a preferred embodiment of the above technical solution: the information subset selection steps include:

[0009] Each piece of information in the information subset is segmented into words, and the core keywords of each piece of information are extracted. The core words of each piece of information are combined into a core keyword set, and a single piece of information attention is assigned to the core keywords in each piece of information in the information subset.

[0010] For each article with a single core keyword, a single core keyword attention score is calculated for each core keyword in the core keyword set.

[0011] The core keyword set, the formal keyword set, and the candidate keyword set are combined to form a new keyword set, and a weight value is assigned to each keyword in the new keyword set;

[0012] Set a first threshold and a second threshold, where the first threshold is greater than the second threshold. Compare the weight value of each keyword in the new keyword set with the first threshold and the second threshold to obtain a new official keyword set, a new candidate keyword set, and a recycled keyword set.

[0013] For each additional period of information, a subset filtering operation is performed to replace the initial key set, the new candidate keyword set, and the recycled keyword set.

[0014] Based on the above technical solution and as a preferred embodiment of the above technical solution: when the weight value of the keywords in the new keyword set is not less than the first threshold, the set of keywords in the new keyword set that are not less than the first threshold is used to replace the formal keyword set.

[0015] Based on the above technical solution and as a preferred solution: when the weight value of a keyword in the new keyword set is not less than the second threshold but less than the first threshold, the keywords in the new keyword set that are not less than the second threshold but less than the first threshold are combined and incorporated into the candidate keyword set.

[0016] Based on the above technical solution and as a preferred solution: when the weight value of a keyword in the new keyword set is less than the second threshold, the keywords in the new keyword set that are less than the second threshold are combined into a recycled keyword set.

[0017] Based on the above technical solution and as a preferred embodiment of the above technical solution: the method for allocating the weight value of each keyword in the new keyword set includes:

[0018] When filtering keywords from the information set acquired in the first cycle, the weight value of each keyword in the new keyword set is the sum of the weight values ​​of that keyword in the formal keyword set and the core keyword set;

[0019] When filtering keywords from information sets acquired beyond the first cycle, the weight value of each keyword in the new keyword set is the sum of the weight values ​​of that keyword in the new official keyword set, the new candidate keyword set, and the core keyword set in the current cycle.

[0020] Based on the above technical solution and as a preferred solution: the attention of each core keyword in the information subset is the sum of the attention of that core keyword in each piece of information.

[0021] Based on the above technical solution and as a preferred embodiment of the above technical solution: the method for calculating the weight value of keywords in the new keyword set further includes weight correction, the weight correction method including the following steps:

[0022] Set the correction percentage; within the first cycle:

[0023] Based on the correction percentage, the weight value of each keyword in the official keyword set is divided into a first corrected weight value and a corrected keyword weight value. Each corrected keyword weight value replaces the weight value of that keyword in the official keyword set.

[0024] The first corrected weight value is obtained by summing the first corrected weight values ​​of each keyword in the formal keyword set. A percentage is then assigned to each core keyword in the core keyword set. The sum of the percentages of all core keywords in the core keyword set is a constant. The first corrected weight value is then distributed according to the percentage of each keyword in the core keyword set, and this distribution is used as the weight value of that core keyword. The more attention a core keyword receives, the larger the percentage it is assigned.

[0025] The weight of each keyword in the new keyword set is the sum of the corrected keyword weight in the formal keyword set and the core keyword weight in the core keyword set.

[0026] In each cycle beyond the first cycle:

[0027] Based on the correction percentage, the weight value of each keyword in the new official keyword set is divided into a first corrected weight value and a corrected keyword weight value. Each corrected keyword weight value in the new official keyword set replaces the weight value of that keyword in the new official keyword set in this period.

[0028] Each candidate keyword in the new candidate keyword set is divided into a second corrected weight value and a corrected candidate keyword weight value. Each corrected candidate keyword weight value of each keyword in the new candidate keyword set replaces the weight value of that candidate keyword in the new candidate keyword set in the current period.

[0029] The first corrected weight value is obtained by summing the first corrected weight values ​​of each keyword in the new official keyword set, and the second corrected weight value is obtained by summing the second corrected weight values ​​of each candidate keyword in the new candidate keyword set. A percentage is assigned to each core keyword in the core keyword set, and the sum of the percentages of all core keywords in the core keyword set is a constant. The sum of the first corrected weight value and the second corrected weight value is distributed according to the percentage of each keyword in the core keyword set, which is used as the weight value of that core keyword. The more attention a core keyword receives, the larger the percentage it is assigned.

[0030] The weight of each keyword in the new keyword set is the sum of the corrected keyword weight in the new official keyword set, the corrected candidate keyword weight in the new candidate keyword set, and the core keyword weight in the core keyword set.

[0031] Based on the above technical solution and as a preferred solution: the weight values ​​of all keywords in the recovered keyword set are summed up and distributed to the new official keyword set and the new candidate keyword set in an equal weight ratio. The total weight of all new keywords in the new official keyword set and the total weight of all candidate keywords in the new candidate keywords set are equal to the total initial weight value.

[0032] An electronic device includes: at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the information intelligent aggregation method described above.

[0033] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the information intelligent aggregation method described above.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] It requires no manual maintenance, saving manpower, and can generate an objective set of keywords in a timely and accurate manner.

[0036] The novel intelligent news aggregation of this invention can automatically update trending news. It only requires the establishment of a formal keyword set, and the intelligent aggregation method of this invention can continuously update it.

[0037] Based on the weight value of each new keyword in the final official keyword set, we can know the objective level of attention given to keywords in a given topic. Attached Figure Description

[0038] The following detailed description of non-limiting embodiments is provided with reference to the accompanying drawings.

[0039] Figure 1 This is a schematic flowchart of the present invention. Detailed Implementation

[0040] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0041] To better illustrate the present invention, the following description is in conjunction with the accompanying drawings. Figure 1 The present invention will be described in detail.

[0042] Example 1:

[0043] A method for intelligent information aggregation includes the following steps:

[0044] A formal keyword set is created based on the required information type. An initial weight value is set for each keyword. The total initial weight value of all keywords in the formal keyword set is set to a constant. A candidate keyword set is set, and the candidate keyword set is initially empty. It should be noted that in this embodiment, the formal keyword set is set as K0, the candidate keyword set is set as K1, the initial weight value is set as A, and the total initial weight value is a constant Wc. Algorithm example: K0 = {a, b, c}, the initial weight value set of the formal keywords A: {a: 0.2, b: 0.3, c: 0.5}, Wc = 1, K1 = {}.

[0045] Information is periodically acquired to form an information set. A subset of information is extracted using a set of relevant keywords. The attention level of each piece of information within the subset is calculated based on its total attention. In this embodiment, the period value can be set according to requirements, such as one day, two days, or other time periods. The acquisition method involves using a web crawler to obtain online information in real time, temporarily storing the acquired information, and proceeding to the next step after the set period value is reached. It should be noted that any method capable of obtaining online information in real time or periodically can be used in this invention. Using the above algorithm as an example, let the information set be T, and the information subset be S. During use, each word in K0 is used to search T. If an article contains any keyword in K0, it can be added to S. In this embodiment, the total attention is set as a constant Ac (equivalently set to 1). The attention set A for each piece of information in the information subset S is calculated according to a certain algorithm: A1, A2, ..., An, where A1 + A2 + ... + An = Ac. It is worth mentioning that in this embodiment, the attention level of each piece of information in the information subset is the number of clicks within the period.

[0046] In this embodiment, the attention level of each core keyword in the information subset is the sum of the attention levels of the keyword in each individual news article.

[0047] The information subset is processed to obtain a new keyword set and a new candidate keyword set. For each additional period of information, the information subset filtering operation is performed, replacing the initial keyword set, the new candidate keyword set, and the recycled keyword set.

[0048] In this specific example, the steps for filtering the information subset include: segmenting each piece of information in the information subset into words, extracting the core keywords of each piece of information, forming a core keyword set from the core words of each piece of information, and assigning a single-article attention level to the core keywords in each piece of information in the information subset; the core keyword set of each piece of information is set as Ci = {C1, C2, ..., Cn}, the attention level of each word in each piece of information in that piece of information is set as TFi, and the total number of words in the information segmentation is set as m. This embodiment's algorithm example segments each piece of information in the information subset S into words and extracts the core keyword set. Specifically, after segmenting a news item from a subset of news items, we obtain a complete set of segmented words for that news item, T = {T1, T2, ..., Tm}. The attention value (TFi) of each segmented word within that news item is equal to the number of times that segmented word (Ti) appears in that news item / the total number of words in that news item (m). IDFi = log(total number of news items n / number of news items containing Ti), TFIDFi = TFi × IDFi. We then sort {T1, T2, ..., Tm} according to their TFIDF values ​​from largest to smallest and select the top 10 words as the core keyword set for that news item. It should be noted that the segmentation of each news item is limited according to its own circumstances, but this is not specified here.

[0049] For each core keyword, a single news article's attention level is calculated based on the attention level of each core keyword in the core keyword set. Specifically, the attention level of Ci is proportionally allocated according to the previously calculated TFIDF value. In this embodiment, a news article has 10 core keywords, and the attention level Fi of keyword Ti = the attention level of the news article A × TFIDFi / (TFIDF1 + TFIDF2 + ... + TFIDF10).

[0050] A new keyword set is formed by combining the core keyword set, the formal keyword set, and the candidate keyword set, and a weight value is assigned to each keyword in the new keyword set. The attention scores of all core keywords appearing in each news item within the news subset are summed up according to the core keywords to obtain a core keyword set C for this period, where the attention score of each core keyword is the sum of the attention scores it receives in each news item. In a specific example of this embodiment, suppose a word segment Ti appears only in news items C1 and C2, indicating that it is a core keyword in C1 and C2. Suppose its attention score in C1 is F1 and its attention score in C2 is F2, then the sum of its attention scores in this round of calculation is equal to the sum of F1 and F2. A Map structure is used to represent the attention scores of the keyword set, where words represent keys and attention scores represent values, such as {T1: F1, T2: F2, ..., Tm: Fm}. Suppose we have two weighted maps: {a: 0.1, b: 0.2, c: 0.3} and {a: 0.2, b: 0.1, d: 0.4}. The new map we get by combining them is {a: 0.2 + 0.1, b: 0.2 + 0.2, c: 0.3, d: 0.4}, which is equivalent to {a: 0.3, b: 0.4, c: 0.3, d: 0.4}.

[0051] In a specific example of this embodiment, the method for allocating the weight value of each keyword in the new keyword set includes: when filtering keywords from the information set acquired in the first period, the weight value of each keyword in the new keyword set is the sum of the weight values ​​of that keyword in the formal keyword set and the core keyword set; when filtering keywords from the information set acquired beyond the first period, the weight value of each keyword in the new keyword set is the sum of the weight values ​​of that keyword in the new formal keyword set, the new candidate keyword set, and the core keyword set in the current period in the previous period. It is worth mentioning that in this embodiment, the method for calculating the weight value of keywords in the new keyword set also includes weight correction, which includes the following steps:

[0052] Set the correction percentage; within the first cycle:

[0053] Based on the correction percentage, the weight value of each keyword in the official keyword set is divided into a first corrected weight value and a corrected keyword weight value. Each corrected keyword weight value replaces the weight value of that keyword in the official keyword set.

[0054] The first adjusted weight value is obtained by summing the first adjusted weight values ​​of each keyword in the formal keyword set. A percentage is then assigned to each core keyword in the core keyword set. The sum of the percentages of all core keywords in the core keyword set is a constant. The first adjusted weight value is then distributed according to the percentage of each keyword in the core keyword set, and this distribution serves as the weight value of that core keyword. The more attention a core keyword receives, the larger the percentage it is assigned. It should be noted that, in the specific example of this embodiment, the formula for calculating the weight value of each core keyword in the core keyword set is (attention of that core keyword / total attention of the sum of attention of each core keyword in the core keyword set) × the first adjusted weight value.

[0055] The weight of each keyword in the new keyword set is the sum of the corrected keyword weight in the formal keyword set and the core keyword weight in the core keyword set.

[0056] In each cycle beyond the first cycle:

[0057] Based on the correction percentage, the weight value of each keyword in the new official keyword set is divided into a first corrected weight value and a corrected keyword weight value. Each corrected keyword weight value in the new official keyword set replaces the weight value of that keyword in the new official keyword set in this period.

[0058] Each candidate keyword in the new candidate keyword set is divided into a second corrected weight value and a corrected candidate keyword weight value. Each corrected candidate keyword weight value of each keyword in the new candidate keyword set replaces the weight value of that candidate keyword in the new candidate keyword set in the current period.

[0059] The first corrected weight value is obtained by summing the first corrected weight values ​​of each keyword in the new official keyword set, and the second corrected weight value is obtained by summing the second corrected weight values ​​of each candidate keyword in the new candidate keyword set. A percentage is assigned to each core keyword in the core keyword set, and the sum of the percentages of all core keywords in the core keyword set is a constant. The sum of the first corrected weight value and the second corrected weight value is distributed according to the percentage of each keyword in the core keyword set, which serves as the weight value of that core keyword. The more attention a core keyword receives, the larger the percentage it is assigned. It should be noted that, in the specific example of this embodiment, the formula for calculating the weight value of each core keyword in the core keyword set is (attention of the core keyword / total attention of the sum of attention of each core keyword in the core keyword set) × (first corrected weight value + second corrected weight value).

[0060] The weight of each keyword in the new keyword set is the sum of the corrected keyword weight in the new official keyword set, the corrected candidate keyword weight in the new candidate keyword set, and the core keyword weight in the core keyword set.

[0061] In this example, the correction percentage is set as R, assuming R = 10%. The attention map of C is assumed to be {T1: F1, T2: F2, ..., Tm: Fm}, and the weight value set of the formal keyword set K0 is {t1: W1, t2: W2, ..., tn: Wn}. 10% of the weight of each keyword in K0 is taken, and the first corrected weight value is W = (W1 + W2 + ... + Wn) × 10%, which is then allocated to C to obtain {T1: w1, T2: w2, ..., Tm: wm}, where wi = W × Fi / (F1 + F2 + ... + Fm). After allocation, the corrected keyword weight value set in the new formal keyword set K0' is {t1: W1 × 90%, t2: W2 × 90%, ..., tn: Wn × 90%}.

[0062] The core keyword weight value of each core keyword in the core keyword set is calculated by combining the attention level of each core keyword in the core keyword set. Specifically, when filtering keywords from the information acquired in the first period, the set of keyword weight values ​​in the new keyword set is the sum of the corrected keyword weight value of the keyword in the formal keyword set and the core keyword weight value in the core keyword set.

[0063] In a specific example of this embodiment, during the information subset filtering process, a first threshold and a second threshold are set, wherein the first threshold is greater than the second threshold. The weight value of each keyword in the new keyword set is compared with the first threshold and the second threshold to obtain a new formal keyword set, a new candidate keyword set, and a recycled keyword set. When the weight value of a keyword in the new keyword set is not less than the first threshold, the keywords in the new keyword set that are not less than the first threshold are combined to form a new formal keyword set to replace the formal keyword set. When the weight value of a keyword in the new keyword set is not less than the second threshold but less than the first threshold, the keywords in the new keyword set that are not less than the second threshold but less than the first threshold are combined and merged into the candidate keyword set. When the weight value of a keyword in the new keyword set is less than the second threshold, the keywords in the new keyword set that are less than the second threshold are combined to form a recycled keyword set.

[0064] The weights of all keywords in the collected keyword set are summed up and then distributed to the new official keyword set and the new candidate keyword set in an equal weight ratio. After the distribution, the total weight of all new keywords in the new official keyword set and the total weight of all candidate keywords in the new candidate keyword set are equal to the total initial weight value.

[0065] When processing the information set obtained in the second cycle, a new set of official keywords replaces the set of official keywords, and a new set of candidate keywords replaces the set of candidate keywords. This process is repeated. Assume the weight value set of the new official keyword set K0' obtained after processing the information in the first cycle is {t1': W1', t2': W2', ..., tn': Wn'}, and the weight value set of the new candidate keyword K1' is {Te1: We1, Te2: We2, ..., Tek: Wek}. The correction coefficient is set to R, assuming R = 10%. Assume the attention map of C is {T1: F1, T2: F2, ..., Tm: F...}. The set of weight values ​​for the new official keyword set K0' is {t1': W1', t2': W2', ..., tn': Wn'}, and the set of weight values ​​for the new candidate keyword K1' is {Te1: We1, Te2: We2, ..., Tek: Wek}. 10% of the weight of each keyword in K0 is taken out, and the sum of the first and second corrected weight values ​​is W = (W1' + W2' + ... + Wn' + We1 + We2 + ... + Wek) × 10%, which is then allocated to C to obtain {T1: w1, T2: w2, ..., Tm: wm}, where wi = W × Fi / (F1 + F2 + ... + Fm). After allocation, the set of corrected keyword weight values ​​in the new official keyword set K0' is {t1': W1'×90%, t2': W2'×90%, ..., tn': Wn'×90%}, and the set of corrected candidate keyword weight values ​​in the new candidate keyword set K1' is {Te1: We1×90%, Te2: We2×90%, ..., Ten: Wen×90%}. The weight value of each keyword in the new keyword set is the sum of its corrected keyword weight value in the new official keyword set, its corrected candidate keyword weight value in the new candidate keyword set, and its core keyword weight value in the core keyword set.

[0066] When processing the information set obtained in the third period, repeat the above steps. The weight value of each keyword in the new keyword set is the sum of the corrected keyword weight value in the new official keyword set in the second period, the corrected candidate keyword weight value in the new candidate keyword set, and the core keyword weight value in the core keyword set.

[0067] Based on the required cycle, the information set obtained in each cycle is processed according to the above information and processing process. The keywords in the new official keyword set are changed and the weight value of the keyword is changed to realize intelligent news aggregation and automatically update popular content.

[0068] Example 2:

[0069] An electronic device includes: at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the information intelligent aggregation method described above.

[0070] Example 3:

[0071] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the above-described intelligent information aggregation method.

[0072] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above. Those skilled in the art can make various changes or modifications within the scope of the claims, as well as arbitrary combinations of the above embodiments, without affecting the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. An information intelligent aggregation method, characterized in that, The method comprises the following steps: Creating a formal keyword set based on the required information type, setting an initial weight value for each keyword in the formal keyword set, setting a constant total initial weight value for all keywords in the formal keyword set, setting a candidate keyword set, and the initial state of the candidate keyword set is an empty set; Periodically obtaining information to form an information set, screening and extracting an information subset by using the formal keyword set, and calculating the information attention degree of each information in the information subset according to the total attention amount of the information subset; Processing the information subset to obtain a new keyword set and a new candidate keyword set; The weight value of each keyword in the new keyword set is further subjected to weight correction before being allocated, and the weight correction method comprises the following steps: Setting a correction percentage; in the first period: According to the correction percentage, the weight value of each keyword in the formal keyword set is divided into a first correction weight value and a correction keyword weight value, and each correction keyword weight value of each keyword in the formal keyword set replaces the weight value of the keyword in the formal keyword set; The first correction weight total value is obtained by superimposing and summing the first correction weight value of each keyword in the formal keyword set, a percentage is allocated to each core keyword in the core keyword set, the sum of the percentages of all core keywords in the core keyword set is a constant, and the first correction weight total value is allocated to the weight value of each keyword in the core keyword set according to the percentage of each keyword in the core keyword set, and the percentage allocated to the core keyword with more attention degree is greater; The weight value of each keyword in the new keyword set is obtained by adding the correction keyword weight value of the keyword in the formal keyword set and the core keyword weight value in the core keyword set; In each period exceeding the first period: According to the correction percentage, the weight value of each keyword in the new formal keyword set is divided into a first correction weight value and a correction keyword weight value, and each correction keyword weight value of each keyword in the new formal keyword set replaces the weight value of the keyword in the new formal keyword set in the current period; Each candidate keyword in the new candidate keyword set is divided into a second correction weight value and a correction candidate keyword weight value, and each correction candidate keyword weight value of each keyword in the new candidate keyword set replaces the weight value of the candidate keyword in the new candidate keyword set in the current period; The first correction weight total value is obtained by superimposing and summing the first correction weight value of each keyword in the new formal keyword set, the second correction weight total value is obtained by superimposing and summing the second correction weight value of each candidate keyword in the new candidate keyword set, a percentage is allocated to each core keyword in the core keyword set, the sum of the percentages of all core keywords in the core keyword set is a constant, and the sum of the first correction weight total value and the second correction weight total value is allocated to the weight value of each keyword in the core keyword set according to the percentage of each keyword in the core keyword set, and the percentage allocated to the core keyword with more attention degree is greater. The weight of each keyword in the new keyword set is the sum of the corrected keyword weight in the new official keyword set, the corrected candidate keyword weight in the new candidate keyword set, and the core keyword weight in the core keyword set.

2. The method of claim 1, wherein, The processing steps for information subsets include: Each piece of information in the information subset is segmented into words, and the core keywords of each piece of information are extracted. The core words of each piece of information are combined into a core keyword set, and a single piece of information attention is assigned to the core keywords in each piece of information in the information subset. For each article with a single core keyword, a single core keyword attention score is calculated for each core keyword in the core keyword set. The core keyword set, the formal keyword set, and the candidate keyword set are combined to form a new keyword set, and a weight value is assigned to each keyword in the new keyword set; Set a first threshold and a second threshold, where the first threshold is greater than the second threshold. Compare the weight value of each keyword in the new keyword set with the first threshold and the second threshold to obtain a new official keyword set, a new candidate keyword set, and a recycled keyword set. For each additional period of information, a subset of information is filtered, and a new set of official keywords, a new set of candidate keywords, and a set of recycled keywords are added.

3. The method of claim 2, wherein, When the weight value of a keyword in the new keyword set is not less than the first threshold, the set of keywords in the new keyword set that are not less than the first threshold is used to replace the original keyword set.

4. The method of claim 2, wherein, When the weight value of a keyword in the new keyword set is not less than the second threshold but less than the first threshold, the keywords in the new keyword set that are not less than the second threshold but less than the first threshold are combined and incorporated into the candidate keyword set.

5. The method of claim 2, wherein, When the weight value of a keyword in the new keyword set is less than the second threshold, the keywords in the new keyword set that are less than the second threshold are combined into a recycled keyword set.

6. The method of claim 2, wherein, The weighting methods for each keyword in the new keyword set include: When filtering keywords from the information set acquired in the first cycle, the weight value of each keyword in the new keyword set is the sum of the weight values ​​of that keyword in the formal keyword set and the core keyword set; When filtering keywords from information sets acquired beyond the first cycle, the weight of each keyword in the new keyword set is the sum of its weight in the previous cycle in the new official keyword set, the new candidate keyword set, and the core keyword set in the current cycle.

7. The method of claim 2, wherein, The attention level of each core keyword in the news subset is the sum of the attention levels of that core keyword in each individual news article.

8. The method of claim 2, wherein, The weights of all keywords in the collected keyword set are summed up and then distributed to the new official keyword set and the new candidate keyword set in an equal weight ratio. After the distribution, the total weight of all new keywords in the new official keyword set and the total weight of all candidate keywords in the new candidate keyword set are equal to the total initial weight value.

9. An electronic device, comprising: include: At least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform an information intelligent aggregation method as described in any one of claims 1-8.

10. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to cause the computer to execute the information intelligent aggregation method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Method and device for acquiring hotspot information

    CN104424278A

  • Corpus keyword automatic extraction algorithm based on data mining

    CN110377724A