Article collection updating method, device, equipment and storage medium

By calculating the reverse file frequency and net editing frequency of candidate phrases for weighted scores, determining the target phrase and updating the article set, the problem of inaccurate judgment of emergencies in the encyclopedia editing history is solved, and the relevance and importance of emergencies are improved.

CN116340324BActive Publication Date: 2025-08-26CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111544569.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-08-26
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

In the history of encyclopedia editing, the TextRank algorithm cannot accurately judge the burst level of phrases, resulting in misjudgment of common phrases. When emergencies overlap in large articles, the importance of emergencies is lost and the correlation is low.

Method used

By calculating the reverse file frequency and net edit frequency of candidate phrases, weighted score calculations are performed, target phrases are determined, and the initial article collection is updated using target phrases to improve the accuracy of burst levels.

Benefits of technology

The accuracy of judgment of the burst level of phrases is improved and the relevance of the target article collection is ensured that the target phrases have a higher burst level in the target article collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116340324B_ABST
    Figure CN116340324B_ABST
Patent Text Reader

Abstract

The present application discloses an article collection updating method, apparatus, device, and storage medium. The method includes: determining at least one candidate phrase from an initial article collection; determining, based on the initial article collection, the reverse file frequency and net edit frequency of each of the at least one candidate phrases; wherein the reverse file frequency is used to characterize the category distinguishing ability of the candidate phrase, and the net edit frequency is used to characterize the number of edits of the candidate phrase; calculating a weighted score for each candidate phrase using the reverse file frequency and net edit frequency of each candidate phrase to determine a weighted burst score for each candidate phrase; determining a target phrase from the at least one candidate phrase based on the weighted burst score; and updating the initial article collection using the target phrase to obtain a target article collection. This method can improve the accuracy of determining the burst level of a phrase and improve the burst level of the target phrase in the target article collection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of public opinion analysis, and in particular to a method, apparatus, device, and storage medium for updating an article collection. Background Art

[0002] Encyclopedia websites, such as Baidu Baike and Wikipedia, are internet encyclopedias. Take Wikipedia, for example, one of the most famous encyclopedias on the internet. All edits to each Wikipedia article, along with related information, are stored in the Wikipedia edit history. Edits are often triggered by real-world events. Editors learn of breaking news from news articles, social media, or other sources and select new content to add to Wikipedia articles. Therefore, compared to news articles, the Wikipedia edit history provides a more summarized and organized view of how an event occurred and evolved.

[0003] Many researchers have studied and proposed various technical implementations to solve the problem of detecting sudden events and key phrases in encyclopedia edit histories. Currently, the main methods include the following: PageRank, TextRank, and DivGraphPointer.

[0004] Among them, PageRank is a page importance ranking algorithm proposed when building an early search system prototype. Currently, many important link analysis algorithms are derived based on the PageRank algorithm. The TextRank algorithm is inspired by the page importance ranking algorithm PageRank. The PageRank algorithm calculates the importance of each page based on the link relationship between pages on the Internet. TextRank regards phrases as nodes in the key phrase graph. The importance of each phrase is calculated based on the co-occurrence relationship of the phrases. The DivGraphPointer method is a graph-based key phrase extraction algorithm used to extract key phrases from documents. It combines traditional graph-based ranking methods and neural network-based methods. Some studies also represent unstructured documents with contextual similarity information in the form of graph databases.

[0005] However, the current solution still has some defects. For example, when using TextRank to determine the burst level of a phrase, some common phrases may have a very high burst level, resulting in misjudgment. In addition, in a large collection of articles, many burst events may overlap in the same time period, causing each burst event to lose its importance. The correlation between the burst phrase and the article collection is low, resulting in a low burst level for the burst phrase. Summary of the Invention

[0006] The present application provides an article collection updating method, apparatus, device and storage medium, which can improve the accuracy of judging the burst level of phrases and improve the burst level of target phrases in a target article collection.

[0007] The technical solution of this application is achieved as follows:

[0008] In a first aspect, an embodiment of the present application provides a method for updating an article collection, the method comprising:

[0009] determining at least one candidate phrase from the initial set of articles;

[0010] Determining, based on the initial set of articles, an inverse document frequency and a net edit frequency of each of the at least one candidate phrase; wherein the inverse document frequency is used to characterize the category distinguishing ability of the candidate phrase, and the net edit frequency is used to characterize the number of times the candidate phrase has been edited;

[0011] performing weighted score calculations on corresponding candidate phrases using the respective reverse document frequencies and the respective net edit frequencies of the at least one candidate phrase to determine a weighted burst score for each of the at least one candidate phrases;

[0012] determining a target phrase from the at least one candidate phrase according to the weighted burst score;

[0013] The initial article set is updated using the target phrase to obtain a target article set.

[0014] In a second aspect, an embodiment of the present application provides an updating device, which includes a determining unit, a calculating unit, and an updating unit, wherein:

[0015] The determining unit is configured to determine at least one candidate phrase from an initial set of articles; and determine, based on the initial set of articles, an inverse document frequency and a net edit frequency of each of the at least one candidate phrase; wherein the inverse document frequency is used to represent the category distinguishing ability of the candidate phrase, and the net edit frequency is used to represent the number of times the candidate phrase has been edited;

[0016] The calculation unit is configured to calculate a weighted score for each candidate phrase using the reverse document frequency and the net edit frequency of each candidate phrase to determine a weighted burst score for each candidate phrase; and determine a target phrase from the at least one candidate phrase based on the weighted burst score;

[0017] The updating unit is configured to update the initial article set using the target phrase to obtain a target article set.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device including a memory and a processor, wherein:

[0019] The memory is used to store a computer program that can be run on the processor;

[0020] The processor is configured to execute the article collection updating method as described in the first aspect when running the computer program.

[0021] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program, which, when executed by at least one processor, implements the article collection updating method as described in the first aspect.

[0022] An article collection updating method, apparatus, device, and storage medium provided by embodiments of the present application include: determining at least one candidate phrase from an initial article collection; determining, based on the initial article collection, a reverse file frequency and a net edit frequency of each of the at least one candidate phrases; wherein the reverse file frequency is used to characterize the category distinguishing ability of the candidate phrase, and the net edit frequency is used to characterize the number of edits of the candidate phrase; calculating a weighted score for each candidate phrase using the reverse file frequency and the net edit frequency of each of the at least one candidate phrases to determine a weighted burst score for each of the at least one candidate phrases; determining a target phrase from the at least one candidate phrase based on the weighted burst score; and updating the initial article collection using the target phrase to obtain a target article collection. In this way, for a given initial article set, when determining the weighted burst score of the candidate phrases therein, the reverse file frequency and net editing frequency of the candidate phrases in the initial article set are utilized, and the influence of the phrase's category discrimination ability and the number of historical edits on the weighted burst score of the phrase is fully considered, so that a weighted burst score that accurately measures the burst level of the phrase can be obtained. Finally, a target phrase is selected from the candidate phrases, and the initial article set is updated with the target phrase, so that in the final target article set, the target phrase has a higher burst level, and the target article set is more relevant to the target phrase and the event represented by the target phrase, thereby improving the judgment accuracy of the phrase burst level and the burst level of the target phrase in the target article set. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A flowchart of a method for updating an article collection provided in an embodiment of the present application;

[0024] Figure 2 A schematic diagram of an editing record of a Wikipedia article provided in an embodiment of the present application;

[0025] Figure 3 A phrase node diagram provided in an embodiment of the present application;

[0026] Figure 4 Another phrase node schematic diagram provided in an embodiment of the present application;

[0027] Figure 5 A flowchart of another article collection updating method provided in an embodiment of the present application;

[0028] Figure 6 A phrase node diagram of a target phrase provided in an embodiment of the present application;

[0029] Figure 7 A phrase node diagram of another target phrase provided in an embodiment of the present application;

[0030] Figure 8 A schematic diagram of changes in the weighted burst score of a target phrase provided in an embodiment of the present application;

[0031] Figure 9 A schematic diagram of a comparison of a target phrase and its adjacent phrases provided in an embodiment of the present application;

[0032] Figure 10 A schematic diagram of the system architecture of an article collection updating system provided in an embodiment of the present application;

[0033] Figure 11 A schematic diagram of the structure of an updating device provided in an embodiment of the present application;

[0034] Figure 12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0035] Figure 13 A schematic diagram of the structure of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It should be understood that the specific embodiments described herein are only used to explain the related applications and are not intended to limit the applications. It should also be noted that for ease of description, only the portions relevant to the related applications are shown in the drawings.

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0038] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0039] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0040] An analysis of current solutions revealed that while TextRank can identify key phrases with intensive editing activity throughout the entire editing history, it does not consider the temporal trends of phrase edits. Common phrases that are frequently edited in most articles often have high TextRank scores (generally, a higher TextRank score indicates a more bursty phrase), yet these common phrases are not necessarily bursty. Furthermore, in a large collection of articles, many bursts can overlap within the same time period, causing each burst to lose its significance.

[0041] Based on this, an embodiment of the present application provides an article collection updating method, the basic idea of ​​which is: determining at least one candidate phrase from an initial article collection; determining the reverse file frequency and net edit frequency of each of the at least one candidate phrases based on the initial article collection; wherein the reverse file frequency is used to characterize the category distinguishing ability of the candidate phrase, and the net edit frequency is used to characterize the number of edits of the candidate phrase; using the reverse file frequency and net edit frequency of each of the at least one candidate phrases to calculate weighted scores for the corresponding candidate phrases, and determining a weighted burst score of each of the at least one candidate phrases; determining a target phrase from the at least one candidate phrase based on the weighted burst score; and using the target phrase to update the initial article collection to obtain a target article collection. In this way, for a given initial article set, when determining the weighted burst score of the candidate phrases therein, the reverse file frequency and net editing frequency of the candidate phrases in the initial article set are utilized, and the influence of the phrase's category discrimination ability and the number of historical edits on the weighted burst score of the phrase is fully considered, so that a weighted burst score that accurately measures the burst level of the phrase can be obtained. Finally, a target phrase is selected from the candidate phrases, and the initial article set is updated with the target phrase, so that in the final target article set, the target phrase has a higher burst level, and the target article set is more relevant to the target phrase and the event represented by the target phrase, thereby improving the judgment accuracy of the phrase burst level and the burst level of the target phrase in the target article set.

[0042] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0043] In one embodiment of the present application, see Figure 1 , which shows a flow chart of a method for updating an article collection provided by an embodiment of the present application. Figure 1 As shown, the method may include:

[0044] S101: Determine at least one candidate phrase from an initial article set.

[0045] It should be noted that the article collection updating method provided in the embodiments of the present application can be applied to an updating device for updating an article collection, or an electronic device incorporating the updating device. The electronic device may be, for example, a computer, a smartphone, a tablet computer, a laptop computer, a PDA, a navigation device, a server, etc., and the embodiments of the present application do not specifically limit this.

[0046] It should also be noted that the initial article set refers to a collection of articles related to a certain topic. At least one phrase can be extracted from these articles; further, some phrases in this at least one phrase (usually referring to some phrases with high burstiness in the initial article set) constitute the theme of the initial article set.

[0047] In some embodiments, the initial article set includes at least one initial article, and determining at least one candidate phrase from the initial article set may include:

[0048] Determine the target period;

[0049] Obtaining the edit increment of at least one initial article during a target period; wherein the edit increment represents the text increment of an article before and after editing;

[0050] Part-of-speech segmentation is performed on the edit increment, and at least one candidate phrase is determined from the result of the part-of-speech segmentation.

[0051] It should be noted that if a phrase is frequently edited during a specific period of time, then it is very likely that an emergency event related to this phrase has occurred during this period of time, and this phrase can be regarded as an emergency phrase during this period of time.

[0052] The embodiment of the present application can determine a sudden target phrase from an initial article set during a target period. Before determining the target phrase, it is necessary to first determine a plurality of candidate phrases, and then determine the target phrase from the candidate phrases.

[0053] It should also be noted that when determining candidate phrases, the embodiment of the present application can first determine a target time period; wherein, the target time period can be a specific week, a day, a month, etc., and the embodiment of the present application does not make specific limitations on this. In the subsequent description, the target time period is mainly used as the target week for illustrative purposes.

[0054] It should also be noted that in this embodiment of the present application, the initial article set is composed of a series of edit increments; where an edit increment is used to represent the text increments before and after an edit. Taking Wikipedia as an example, the content between adjacent edits in Wikipedia articles is more closely related to the sudden emergence of real-world events. Since adding new content to an article is more common than deleting content, the edit increment in this embodiment of the present application primarily refers to the sentences newly added to the article between two adjacent edits.

[0055] For example, for an article, the zeroth edit means that the article has not appeared yet; after the first edit, the article consists of sentences A, B, and C; then the edit increments between the first edit and the zeroth edit are A, B, C; after the second edit, the article consists of sentences A, B, C, D, and E, then the edit increments between the second edit and the first edit are D and E. For example, see Figure 2 , which shows a schematic diagram of an editing record of a Wikipedia article provided in an embodiment of the present application.

[0056] That is, each initial article in the initial article set is composed of several edit increments, that is, the initial article set is composed of a series of edit increments. When it is necessary to analyze the target time period to determine the sudden events in the initial article set during the target time period (sudden events can be represented by phrases with a high level of suddenness), the embodiment of the present application can only obtain the edit increments of the initial articles in the initial article set during the target time period, and perform part-of-speech classification on the edit increments in the target time period, and determine at least one candidate phrase from the results of the part-of-speech classification.

[0057] Specifically, the Apache Open Natural Language Processing (Apache Open NLP) part-of-speech tagging (POS) can be used to segment each sentence in the edit increment into chunks, where chunks containing noun tags are phrases. This allows for the identification of at least one candidate phrase from the edit increments during the target period.

[0058] In addition, when determining candidate phrases, you can also not specify the target time period first, but instead perform part-of-speech classification on all edit increments in the initial article set, and extract at least one candidate phrase from all edit increments. When subsequently determining the target phrase, specify a time period when the target phrase suddenly occurs as the target time period. The embodiments of the present application do not make specific limitations on this.

[0059] S102: Determine the inverse document frequency and net edit frequency of at least one candidate phrase according to the initial article set.

[0060] Among them, the inverse document frequency is used to characterize the category discrimination ability of the candidate phrase, and the net edit frequency is used to characterize the number of edits of the candidate phrase.

[0061] S103: Calculate a weighted score for the corresponding candidate phrase using the reverse document frequency and the net edit frequency of the at least one candidate phrase to determine a weighted burst score for the at least one candidate phrase.

[0062] S104: Determine a target phrase from at least one candidate phrase according to the weighted burst score.

[0063] It should be noted that when determining the target phrase, we first determine the reverse document frequency and net editing frequency of each candidate phrase in the initial article set, and further determine the weighted burst score of each candidate phrase. Based on the weighted burst score, we determine the target phrase from the candidate phrases.

[0064] Among them, Inverse Document Frequency (IDF) can characterize the category distinction ability of a phrase and is a measure of the general importance of a phrase; Net Editing Frequency (NF) can characterize the number of times a phrase has been edited and is used to measure the historical editing activities of each phrase.

[0065] Regarding determining the reverse document frequency, in some embodiments, determining the reverse document frequency of each of at least one candidate phrase based on the initial set of articles includes:

[0066] Obtaining the total number of articles in the initial article set and the number of first articles in the initial article set; wherein the first article refers to an article that contains a first candidate phrase in an edit increment within a target time period, and the first candidate phrase is any one of the at least one candidate phrase;

[0067] A logarithmic operation is performed on the ratio of the total number to the number of the first article to determine the inverse document frequency of the first candidate phrase.

[0068] It should be noted that the inverse document frequency of a candidate phrase can be determined by the total number of articles in the initial article set and the number of articles (ie, first articles) in the initial article set that contain the candidate phrase in the edit increment within the target period.

[0069] Specifically, for any candidate phrase (referred to as the first candidate phrase), the method for determining its inverse document frequency can refer to formula (1):

[0070]

[0071] Among them, V i represents the first candidate phrase; idf(V i ) represents the inverse document frequency of the first candidate phrase; L represents the total number of articles in the initial article set; DF(V i ) represents the number of articles containing the target phrase in the editing increment of the target period in the initial article set, that is, the number of first articles.

[0072] That is, the inverse document frequency is determined by performing a logarithmic operation on the ratio of the total number of articles in the initial article set to the number of the first article. The main idea of ​​the inverse document frequency is: if there is a phrase V in the initial article set, i The fewer articles, that is, The larger the idf(V i ) is larger, the phrase V i Has good ability to distinguish categories; that is, common phrases often DF (V i ) is larger and the frequency of reverse files is smaller.

[0073] Regarding determining the net edit frequency, in some embodiments, determining the net edit frequency of each of at least one candidate phrase based on the initial set of articles includes:

[0074] Determining initial net edit frequencies of the first candidate phrase in a plurality of time periods; wherein the plurality of time periods include a target time period and all time periods before the target time period;

[0075] An exponential moving average calculation is performed on the initial net edit frequencies of the first candidate phrase in a plurality of time periods to determine the net edit frequency of the first candidate phrase.

[0076] It should be noted that the net editing frequency of the candidate phrase can be determined by the initial net editing frequency of the candidate phrase corresponding to the target time period and the initial net editing frequency of the candidate phrase corresponding to the time period before the target time period. By performing exponential moving average calculation on these several initial net editing frequencies, the net editing frequency of the candidate phrase can be determined.

[0077] Specifically, for the first candidate phrase, the net editing frequency can be determined by referring to formula (2):

[0078]

[0079] Among them, V i Indicates the first candidate phrase; NF(V i ) represents the net editing frequency of the first candidate phrase; w represents the target period (taking the target week as an example); NF w (V i ) represents the initial net editing frequency of the first candidate phrase in week w; NF w-1 (V i ) represents the initial net editing frequency of the first candidate phrase in the w-1th week, and so on; η is the damping coefficient, where 0≤η<1, which can usually be 0.6.

[0080] Furthermore, regarding the initial net edit frequencies of the candidate phrases in each time period, in some embodiments, determining the initial net edit frequencies of the first candidate phrase in each of the multiple time periods may include:

[0081] Obtain the number of edits for each article in the initial article set during the first period and the frequency of occurrence of the first candidate phrase in each edit increment of each article;

[0082] determining an initial net edit frequency of the first candidate phrase in the first time period based on the number of edits of each article and the frequency of occurrence of the first candidate phrase in each edit increment of each article;

[0083] The first time period is any one of several time periods.

[0084] It should be noted that the initial net edit frequency of a candidate phrase in a certain period can be determined by the number of edits of each article in the initial article set in the period, and the number of occurrences of the candidate phrase in each edit increment in the initial article set in the period.

[0085] Specifically, for the first candidate phrase, the method for determining its initial net edit frequency in the target time period can refer to Formula (3). The method for determining the net edit frequency of the first candidate phrase in any time period can refer to Formula (3), and it is only necessary to replace the corresponding data with the data of the time period to be calculated.

[0086]

[0087] Among them, V i Indicates the first candidate phrase; NF w (V i) represents the initial net edit frequency of the first candidate phrase in the target period w; rev(a,w) represents the number of edits of the first candidate phrase in the a-th article in the initial article set in the target period w, where 1≤a≤L, L represents the total number of articles in the initial article set; RTFD j,a (V i ) represents the absolute value of the difference in the occurrence frequency of the first candidate phrase between the j-th edit and the j-1-th edit of the a-th article in the target period w, where the occurrence frequency of the first candidate phrase is the number of times the first candidate phrase appears.

[0088] Formula (3) represents the average difference in the frequency of occurrence of the first candidate phrase in all edits of each article in the initial article set during the target period w.

[0089] Through formula (3), the initial net edit frequencies of the first candidate phrase in the target period and each period before the target period can be calculated. Then, the exponential moving mean calculation (also called exponential moving average calculation) of the initial net edit frequencies of each period is performed using formula (2) to obtain the net edit frequency of the first candidate phrase.

[0090] The main idea of ​​net edit frequency is: the greater the average difference in the frequency of a phrase appearing in all edits of each article in the target article set in each period, the more sudden the event is with the increase in the number of edits, the more likely it is to be a common phrase and better reflect the characteristics of the sudden event; conversely, if the difference in the frequency of the phrase appearing in each edit of all articles in all periods is not large, the more likely it is to be a common phrase.

[0091] Furthermore, after obtaining the reverse file frequency and net edit frequency of each candidate phrase, in some embodiments, calculating a weighted score for the corresponding candidate phrase using the reverse file frequency and net edit frequency of at least one candidate phrase to determine a weighted burst score for each candidate phrase may include:

[0092] In the initial article set, determining weighted burst scores of a plurality of first adjacent phrases and a first edge weight between the first candidate phrase and each first adjacent phrase; wherein the first adjacent phrase has an adjacent relationship with the first candidate phrase;

[0093] Determining a second edge weight between each first adjacent phrase and a plurality of second adjacent phrases; wherein the plurality of second adjacent phrases and the first adjacent phrase all have an adjacent relationship;

[0094] performing multiplication calculation on the inverse document frequency of the first candidate phrase and the net edit frequency of the first candidate phrase to determine a burst weight of the first candidate phrase;

[0095] The weighted burst scores of the first adjacent phrases, the first edge weight between the first candidate phrase and each first adjacent phrase, and the second edge weight between each first adjacent phrase and the plurality of second adjacent phrases are weightedly calculated according to the burst weight to obtain the weighted burst score of the first candidate phrase.

[0096] It should be noted that in TextRank, the co-occurrence relationship between phrases is used to construct an edge-weighted graph. Since each phrase is called a node in the edge-weighted graph, the edge-weighted graph is also called a phrase node graph. For example, see Figure 3 , which shows a phrase node schematic diagram provided in an embodiment of the present application (i.e., an example of a phrase node diagram).

[0097] like Figure 3 As shown, for the first candidate phrase V in the initial article set i , V j For V i The first adjacent phrase with an adjacency relationship, V k1 、V k2 、V k3 and V k4 For V j The second adjacent phrase with an adjacency relationship; there is a weighted edge between every two phrases with an adjacency relationship.

[0098] Among them, the judgment of the phrase adjacency relationship can be achieved in the following way: for each article in the initial article set, in a preset sliding window, if two phrases appear together in the preset sliding window, it means that the two phrases have a co-occurrence relationship, that is, there is an adjacency relationship between the two phrases; and there is a weighted edge between the two phrases, and the edge weight is equal to the count of the number of co-occurrences of the two phrases during the sliding process of the preset sliding window in the initial article set.

[0099] In addition, in order to further highlight the burst characteristics of phrases in time periods, when constructing a phrase node graph of candidate phrases in the initial article set, the embodiment of the present application can also construct a phrase node graph only for the editing increments of the target time period, so that when the burst score / weighted burst score is subsequently calculated, it can also be calculated only for the editing increments of the phrases in the target time period.

[0100] The method of using TextRank to calculate the burst score of candidate phrases can be shown in formula (4):

[0101]

[0102] Among them, V i Represents the first candidate phrase; WS(V i) represents the burst score of the first candidate phrase (also called TextRank score); d is the damping factor, which is usually set to 0.85; V j ∈In(Vi) indicates pointing to V i The node set of w is the set of first adjacent phrases that have an adjacent relationship with the first candidate phrase; ji Indicates V i and V j The edge weight between V k ∈Out(V j ) represents V j The node set pointed to is the set of second adjacent phrases other than the first candidate phrase that have an adjacent relationship with the first adjacent phrase; jk Indicates V j and V k The edge weight between WS(V j ) represents the first adjacent phrase V j The burst score.

[0103] It should be noted that TextRank can obtain the burst score of each candidate phrase through iterative calculation. Before the algorithm starts, the burst score of a certain phrase needs to be given in advance.

[0104] Because TextRank doesn't consider the temporal trends of phrase edits, some common phrases that are frequently edited in most articles often have large burst scores. This embodiment of the present application improves on traditional TextRank by adding a burst weight to the phrase when calculating its burst score to reflect the temporal trends in the edit history. This improved TextRank algorithm is called TextRank_nfidf.

[0105] It should also be noted that the burst weight of the candidate phrase is determined by multiplying the aforementioned calculated inverse document frequency and net edit frequency.

[0106] Specifically, the burst weight of the candidate phrase can be determined by referring to formula (5):

[0107] W(V i ) nfidf =nf(V i )×idf(V i ) (5)

[0108] Among them, V i represents the first candidate phrase; W(V i ) nfidf represents the burst weight of the first candidate phrase; nf(V i ) represents the net editing frequency of the first candidate phrase; idf(V i) represents the inverse document frequency of the first candidate phrase.

[0109] The weighted burst scores of the first adjacent phrases, the first edge weight between the first candidate phrase and each first adjacent phrase, and the second edge weight between each first adjacent phrase and the plurality of second adjacent phrases are weightedly calculated according to the burst weight to obtain the weighted burst score of the first candidate phrase.

[0110] When using TextRank_nfidf to calculate the burst score of the candidate phrase, for example, see Figure 4 , which shows another phrase node schematic diagram provided by an embodiment of the present application. Figure 4 As shown, Figure 3 The difference is that for each node in the phrase node graph, a burst weight W(V) is added nfidf , where V represents any node in the phrase node graph, i.e. a phrase.

[0111] In addition, in order to further highlight the sudden characteristics of phrases in a time period, when constructing a phrase node graph of candidate phrases in an initial article set, the embodiment of the present application may also construct a phrase node graph only for the editing increments of the target time period.

[0112] The method of using TextRank_nfidf to calculate the weighted burst score of the candidate phrase can be referred to as shown in formula (6):

[0113]

[0114] Among them, V i Represents the first candidate phrase; WS(V i ) w represents the weighted burst score of the first candidate phrase (also called TextRank_nfidf score); d is the damping factor; W(V i ) nfidf represents the burst weight of the first candidate phrase; V j ∈In(Vi) indicates pointing to V i The node set of w, that is, the first adjacent phrase; ji Indicates V i and V j The edge weight between V k ∈Out(V j ) represents V j The node set pointed to is the second adjacent phrase, and the second adjacent phrase does not include the target phrase; jk Indicates V j and V k The edge weight between WS(V j ) w Represents the first adjacent phrase Vj The weighted burst score of .

[0115] It should be noted that TextRank_nfidf can obtain the burst score of each candidate phrase through iterative calculation. Before the algorithm starts, the weighted burst score of a certain phrase needs to be given in advance, so as to finally obtain the weighted burst score of each candidate phrase.

[0116] Combining the calculation methods of reverse file frequency and net editing frequency, it can be seen that the nf(V i ) is smaller, idf(V i ) is smaller, W(V i ) nfidf The smaller , the smaller the burst weight of the phrase is, and the burst weight of the burst phrase is larger, so that the burst phrase can have a larger weighted burst score.

[0117] That is, TextRank_nfidf adds a burst weight (also called node weight in phrase node graph) W(V i ) nfidf To reduce the burst level of common phrases, it takes into account the time trend of the phrases being edited, so that the calculated weighted burst score can more accurately measure the burstiness of the phrases in the article set.

[0118] After obtaining the weighted burst score of at least one candidate phrase, a candidate phrase is selected from the candidate phrases as a target phrase according to the weighted burst score.

[0119] Specifically, the obtained weighted burst scores may be sorted from high to low, and a candidate phrase may be determined as the target phrase from the candidate phrases corresponding to the top weighted burst scores.

[0120] S105: Using the target phrases, the initial article set is updated to obtain a target article set.

[0121] It should be noted that the aforementioned steps enable the selection of a target phrase with a high burstiness level from the candidate phrases in the initial article set. However, since the initial article set may not necessarily contain all articles related to the target phrase, and not all articles in the initial article set may be related to the target phrase, after determining the target phrase, the embodiment of the present application further uses the target phrase to update the initial article set, so that the articles included in the final target article set are all more relevant to the target phrase, thereby making the target phrase have a clearer and more specific burstiness in the target article set.

[0122] This embodiment provides an article collection updating method, comprising: determining at least one candidate phrase from an initial article collection; determining, based on the initial article collection, a reverse file frequency and a net edit frequency of each of the at least one candidate phrases; wherein the reverse file frequency is used to characterize the category distinguishing ability of the candidate phrase, and the net edit frequency is used to characterize the number of edits of the candidate phrase; calculating a weighted score for each candidate phrase using the reverse file frequency and the net edit frequency of each candidate phrase to determine a weighted burst score for each of the at least one candidate phrases; determining a target phrase from the at least one candidate phrase based on the weighted burst score; and updating the initial article collection using the target phrase to obtain a target article collection. In this way, for a given initial article set, when determining the weighted burst score of the candidate phrases therein, the reverse file frequency and net editing frequency of the candidate phrases in the initial article set are utilized, and the impact of the phrase's category discrimination ability and the number of historical edits on the weighted burst score of the phrase is fully considered, so that a weighted burst score that accurately measures the burst level of the phrase can be obtained. Finally, a target phrase is selected from the candidate phrases, and the target phrase is used to update the initial article set, so that in the final target article set, the target phrase has a higher burst level, and the target article set is more relevant to the target phrase and the event represented by the target phrase.

[0123] In another embodiment of the present application, see Figure 5 , which shows a flow chart of another article collection updating method provided by an embodiment of the present application. Figure 5 As shown, the method may include:

[0124] S201. Obtain the inbound and outbound articles of each article in the initial article set.

[0125] S202: Select at least one article that includes a target phrase in an edit increment during a target period from the inbound and outbound articles of each article.

[0126] S203. Compose a candidate article set based on the at least one selected article.

[0127] It should be noted that the meaning of the initial article set is consistent with the above embodiment and will not be repeated here. Among them, the link-in article quotes the article in the initial article set, and the link-out article is cited by the article in the initial article set.

[0128] According to statistics, the number of inbound and outbound articles in the initial article set is huge. Therefore, the embodiment of the present application further filters the inbound and outbound articles, that is, filters out at least one article that includes the target phrase in the edit increment of the target period to form a candidate article set.

[0129] S204: Update the initial article set using the candidate article set to obtain the target article set.

[0130] It should be noted that in this embodiment of the present application, given an initial article set, a target phrase, and a target time period, the target phrase bursts in the initial article set during the target time period. Because the initial article set may not include all articles related to the target phrase, this embodiment of the present application updates the initial article set using a candidate article set to obtain a target article set, ensuring that the target phrase bursts more clearly and strongly in the target article set.

[0131] In some embodiments, updating the initial article set using the candidate article set to obtain the target article set may include:

[0132] Select an article from the candidate article set and add it to the initial article set;

[0133] The adjacency weight score of each article in the initial article set is determined by increasing the adjacency node weight strategy; the adjacency weight score is used to represent the probability of the article being retained;

[0134] Selecting the lowest adjacency weight score from the determined adjacency weight scores, and deleting the article corresponding to the lowest adjacency weight score from the initial article set to obtain an updated article set;

[0135] When the updated article set meets the preset conditions, the updated article set is determined as the target article set;

[0136] If the updated article set does not meet the preset conditions, the updated article set is used as the initial article set, and the step of selecting an article from the candidate article set and adding it to the initial article set is continued until the updated article set meets the preset conditions to obtain the target article set.

[0137] It is understandable that in a large collection of articles, many bursts (a burst may include one or more phrases) may overlap in the same time period, thus making each burst lose its importance. The present embodiment can obtain an article collection where a target phrase has a clear and strong burst level; if the article collection is as small as possible, then potential topics related to the target phrase can be extracted from the article collection, thereby obtaining an explanation of the target phrase.

[0138] Therefore, in the embodiment of the present application, when updating the initial article set, each time an article is added to the set, another article is deleted to obtain an updated article set, that is, the number of article sets remains unchanged. During the article set update process, an article is first selected from the candidate article set (this article can be called a candidate article) and added to the initial article set. That is, the initial article set at this time includes not only the original initial article but also the newly added article.

[0139] Then, using the Increase Adjacent Node Weight (IANW or IDNW) strategy, we determine the adjacency weight score for each article in the initial article set. This adjacency weight score represents the probability that the article will be retained, that is, the probability that the article will not be deleted. The article with the lowest adjacency weight score is then deleted from the initial article set, thereby completing an update of the initial article set to obtain the updated article. It is understood that the deleted article may be an original article that originally existed in the initial article set, or an article that was added to the initial article set.

[0140] Repeat the process of adding an article from the candidate article set to the initial article set, calculating the adjacency weight score and deleting an article until the updated article set meets the preset conditions, and the latest updated article set is determined as the target article set; otherwise, the updated article set is used as the initial article set, and the aforementioned update steps are repeated until the updated article set meets the preset conditions and the target article set is obtained.

[0141] The preset condition may be that each article in the candidate article set has been added to the initial article set; or that the operation time reaches a preset threshold, wherein the preset threshold may be set to 120 minutes.

[0142] Furthermore, in combination with the aforementioned embodiment, the present application determines the burst level of the target phrase on the initial article set by calculating the weighted burst score of the target phrase. According to formula (6), the weighted burst score of a phrase is positively correlated with the weighted burst score of its adjacent phrases, positively correlated with the edge weights between its adjacent phrases, positively correlated with the burst weight of the phrase, and negatively correlated with the edge weights between the adjacent nodes of the phrase and the adjacent nodes of the adjacent nodes (excluding the phrase itself).

[0143] For example, see Figure 6 , which shows a phrase node diagram of a target phrase provided by an embodiment of the present application. Figure 6 As shown in , with the target phrase p as the center, phrase a, phrase b and phrase c are the adjacent phrases of the target phrase p; phrase e and phrase f are the adjacent phrases of phrase b; phrase d and phrase g are the adjacent nodes of phrase a; phrase h is the adjacent node of phrase c. Figure 6 In

[15] , edge weight can be defined as the probability of transferring from one phrase node to another phrase node after visiting one phrase node. That is, if the edge weight between two phrases is larger, it means that the co-occurrence probability of the two phrases is larger. Then, when visiting one of the nodes (such as when sliding through a preset sliding window), the probability of the other phrase appearing at the same time will also be large.

[0144] To increase the weighted burst score of the target phrase in the article set, we need to increase the edge weights between the target phrase p and its adjacent phrases. This means increasing the occurrence of phrase pairs ap, bp, and cp and reducing the occurrence of phrase pairs ad, ag, be, bf, and ch. That is, the occurrence of phrase pairs ap, bp, and cp increases the likelihood of an article being selected, while the occurrence of phrase pairs ad, ag, be, bf, and ch decreases the likelihood of an article being selected.

[0145] That is, the embodiment of the present application can improve the weighted burst score of the target phrase by increasing the weighted burst weight of the adjacent phrases of the target phrase, as well as the edge weight between the target phrase and the adjacent phrases of the target phrase. The method of determining the adjacent weight score of the article can be called the increase adjacent node weight (IDNW) strategy.

[0146] Regarding the IDNW strategy, in some embodiments, determining the adjacency weight score of each article in the initial article set using the IDNW strategy may include:

[0147] In the second article, determining weighted burst scores of the plurality of third adjacent phrases and co-occurrence frequencies of the target phrase and the plurality of third adjacent phrases;

[0148] determining the co-occurrence frequency of each third adjacent phrase with a plurality of fourth adjacent phrases;

[0149] Determining weighted burst scores and occurrence frequencies of a preset number of phrases other than the target phrase and the third adjacent phrase;

[0150] Obtaining an adjacency weight score for the second article by calculating based on the weighted burst scores of the plurality of third adjacent phrases, the co-occurrence frequencies of the target phrase and the plurality of third adjacent phrases, and the weighted burst scores and occurrence frequencies of a preset number of phrases other than the target phrase and the third adjacent phrases;

[0151] The second article is any article in the initial article set, and in the second article, the plurality of third adjacent phrases all have an adjacent relationship with the target phrase, and the plurality of fourth adjacent phrases all have an adjacent relationship with the third adjacent phrase.

[0152] It should be noted that the second article represents any article in the initial article set after the addition of an article. The calculation method of the weighted burst score of each phrase in the second article in the second article is still based on formula (6). Since formula (6) calculates the weighted burst score of a phrase in an article set, when calculating the weighted burst score of a phrase in an article, it is necessary to adaptively change certain parameters and simultaneously construct a phrase node graph of the target phrase in the second article. The phrase node graph is the same as the aforementioned Figure 4 Similar, except that the phrase nodes only include phrases from the second article, rather than from the entire initial set of articles.

[0153] Specifically, when calculating the weighted burst score of the target phrase in the second article, V i Indicates the target phrase; WS(V i ) w represents the weighted burst score of the target phrase in the second article; d is the damping factor; W(V i ) nfidf V represents the burst weight of the target phrase in the initial article set (the initial article set here refers to the initial article set after adding one article); j ∈In(Vi) means in the second article, pointing to V i The node set of w ji Indicates V i and V j The edge weight between V k ∈Out(V j ) means that in the second article, V j The node set pointed to; w jk Indicates V j and V k The edge weight between WS(V j ) w Indicates the phrase V adjacent to the target phrase in the second article j The weighted burst score of .

[0154] The method for calculating the adjacency weight score of the second article is shown in formula (7):

[0155]

[0156] Where IDNW(A,w,p) represents the adjacency weight score of the second article A in the target time period w and when the target phrase is p; in the numerator, V i ∈In(p) represents all phrases in the second article that have a co-occurrence relationship (adjacent relationship) with the target phrase p, that is, the third adjacent phrase; TextRank_nfidf(V i) represents the weighted burst weight of the target phrase in the second article (i.e. the aforementioned WS(V i ) w , also known as TextRank_nfidf score); PF(V i ,p) represents the co-occurrence frequency of the target phrase and the third adjacent phrase in the second article, that is, the co-occurrence number of the target phrase and the third adjacent phrase in the second article during the sliding process using the preset sliding window.

[0157] In the denominator, V j ∈Out(V i )&V j ≠p means that in the second article, there is a phrase that is adjacent to the third phrase and is not the target phrase, that is, the fourth adjacent phrase; PF(V j ,V i ) represents the co-occurrence frequency of the third adjacent phrase and the fourth adjacent phrase in the second article; V n ∈Top_non_adjacent(p) indicates that in the second article, in addition to the target phrase and the third adjacent phrase, a preset number of other phrases with weighted burst weights ranked high from high to low can be called a preset number of non-adjacent phrases; TextRank_nfidf(V n ) represents the weighted burst score of the non-adjacent phrase in the second article; PF(V n ) represents the frequency of occurrence of the non-adjacent phrase in the second article, that is, the number of occurrences; represents the co-occurrence phrase pairs (V) between each fourth adjacent phrase (excluding the target phrase) that has a co-occurrence relationship with the third adjacent phrase in the second article. j V i , the sum of the co-occurrence frequencies of phrase pairs; It represents the sum of the products of the weighted burst weights of a preset number of non-adjacent phrases with the highest weighted burst scores and the occurrence frequencies of the non-adjacent phrases in the second article that are not adjacent to the target phrase.

[0158] In this way, the adjacency weight score of each article in the initial article set is obtained, and the article with the lowest adjacency weight score is deleted, and another article is added to the initial article set to continue updating the article set.

[0159] For example, see Figure 7 , which shows a phrase node diagram of another target phrase provided by an embodiment of the present application. Figure 7 As shown, it shows the phrase node graph obtained after one (or multiple) article set updates, which is different from Figure 6In contrast, the phrases that are not adjacent to the target phrase p in the phrase node graph and have high weighted burst scores should be as small as possible, and more phrases (phrases a, b, c) that are adjacent to the target phrase p should be retained.

[0160] In this way, the initial article set is updated to the target article set through the aforementioned IDNW strategy.

[0161] It should also be noted that after obtaining the target article set, the embodiment of the present application can also evaluate the update effect of the article set. Therefore, in some embodiments, the method can also include:

[0162] Determining an initial weighted burst score of the target phrase in the initial article set and a target weighted burst score of the target phrase in the target article set;

[0163] determining an initial adjacent phrase of the target phrase in the initial article set and a target adjacent phrase of the target phrase in the target article set;

[0164] Determining a first value of the target phrase based on the initial weighted burst score and the target weighted burst score; and determining a second value of the target phrase based on the initial adjacent phrase and the target adjacent phrase;

[0165] The first value is used to indicate the degree of improvement of the burst level of the target phrase, and the second value is used to indicate the degree of improvement of the context richness of the target phrase.

[0166] It should be noted that in the embodiment of the present application, according to the aforementioned formula (6), the weighted burst score of a phrase in an article set can be calculated, and the burst level of the phrase in the article set can be evaluated based on the weighted burst score. The higher the weighted burst score of the phrase, the higher the burst level of the phrase in the article set. The weighted burst score of the target phrase in the initial article set is called the initial weighted burst score, and the weighted burst score of the target phrase in the target article set is called the target weighted burst score.

[0167] In the embodiment of the present application, the target article set obtained by updating using the IDNW strategy can be evaluated based on the initial weighted burst score of the target phrase in the initial article set and the target weighted burst score of the target phrase in the target article set to evaluate the degree of improvement in the burst level of the target phrase relative to the initial article set, thereby evaluating the update effect of the article set.

[0168] Specifically, the method for determining the degree of improvement of the burst level of the target phrase can refer to formula (8):

[0169]

[0170] Where R1 represents the degree of improvement of the burst level of the target phrase, i.e. the first value; WS(Vi ) wt represents the target weighted burst score of the target phrase in the target article set; WS(V i ) w represents the weighted burst score of the target phrase in the initial article set.

[0171] The larger the first value is, the higher the improvement degree of the burst level of the target phrase is, which means that the updated target article set is more relevant to the target phrase and the event represented by the target phrase.

[0172] In addition, the embodiment of the present application can also intuitively compare the degree of improvement of the burstiness level of the target phrase in the target article set relative to the initial article set through the weighted burstiness score curves of the target phrases in the initial article set and the target article set.

[0173] For example, see Figure 8 , which shows a schematic diagram of the change of the weighted burst score of a target phrase provided by an embodiment of the present application, such as Figure 8 As shown in the figure, it shows the changes in the burst level (weighted burst score) of the target phrase in the initial article set, the target article set, the article set selected by the greedy strategy, and the article set selected by other strategies (such as the TPO strategy), respectively. The horizontal axis represents time (weeks) and the vertical axis represents the burst level of the target phrase in the article set (expressed by the weighted burst score).

[0174] from Figure 8 As can be seen in the figure, in the initial article set, the weighted burst score of the target phrase shows weak peaks at weeks 60 and 90, indicating that events related to the target phrase occurred during these two periods. The weighted burst score of the target phrase in the target article set shows stronger peaks at weeks 60 and 90, more clearly capturing the burstiness of the target phrase and events related to it in the target article set.

[0175] Furthermore, the embodiment of the present application can also obtain adjacent phrases of the target phrase in the initial article set and the target article set respectively. In the target article set, the number of adjacent phrases of the target phrase is greater, and the target phrase has a higher context richness.

[0176] Specifically, the method for determining the degree of improvement of the context richness of the target phrase can refer to formula (9):

[0177] R2=Q2-Q1 (9)

[0178] Among them, R2 represents the degree of improvement of the context richness of the target phrase, that is, the second value; Q2 represents the number of target adjacent phrases of the target phrase in the target article set; Q1 represents the number of initial adjacent phrases of the target phrase in the initial article set.

[0179] The larger the second value is, the higher the degree of improvement in the context richness of the target phrase is, indicating that in the updated target article set, the target phrase has more correct adjacent phrases and a richer context.

[0180] For example, see Figure 9 , which shows a comparison diagram of a target phrase and its adjacent phrases provided by an embodiment of the present application. Figure 9 As shown, (A) represents a schematic diagram of the target phrase and its adjacent phrases in the initial article set, and (B) represents a schematic diagram of the target phrase and its adjacent phrases in the target article set.

[0181] from Figure 9 It can be seen that in the initial article set, the target phrase has three adjacent phrases (also called adjacent nodes): A, B, and C, while in the target article set, the target phrase has five adjacent phrases: A, B, C, D, and E. Obviously, three is less than five, and the second value is 2. This shows that in the target article set, the number of adjacent phrases of the target phrase increases, and the richness of the relevant event context of the target phrase increases.

[0182] The embodiment of the present application provides an article collection updating method. The method selects articles containing a target phrase from the inbound and outbound articles of an initial article to form a candidate article collection. Then, one article is selected from the candidate article collection one at a time to be added to the initial article collection. The adjacency weight score of each article is calculated. The article with the lowest adjacency weight score is deleted to complete the updating of the initial article collection. This ensures that all articles in the updated target article collection are more relevant to the target phrase, and the number of articles included in the target article collection is consistent with the number of articles included in the initial article collection. This ensures that the target phrase has clear and strong burstiness in a smaller article collection. In addition, in the process of updating the article collection, the influence of the adjacent phrases of the target phrase and the edge weights between the two on the adjacency weight scores of the articles, as well as the influence of phrases not connected to the target phrase on the adjacency weight scores of the articles, can be fully considered. Thus, articles more relevant to the target phrase can be retained in the target article collection, thereby improving the burstiness of the target phrase in the target article collection.

[0183] In another embodiment of the present application, see Figure 10 , which shows a schematic diagram of the system architecture of an article collection updating system provided by an embodiment of the present application. The article collection updating method provided by an embodiment of the present application can be based on Figure 10 The system architecture shown in the figure is implemented. Figure 10 As shown, the system architecture may mainly include: an input part 301 , a key phrase extraction model part 302 , and an article update model part 303 .

[0184] The input section 301 is primarily used to input a given initial article set (S), a target phrase (p), and a target week (w). This means that during the target week, the target phrase will appear suddenly in the initial article set. For example, the initial article set can be a collection of articles on Wikipedia named after a certain event, with the total number of articles in the initial article set being |S|. The target phrase is a phrase with a high weighted suddenness score (also known as TextRank_nfidf score) in the initial article set. The target week is a week in which the target phrase appears suddenly.

[0185] Specifically, for a given initial set of articles, the weighted burst score of each phrase in the initial set of articles in each time period (in this embodiment of the application, the time period interval is weekly) can be obtained by calculation through TextRank_nfidf, so that it can be determined which phrases are bursty in each time period, and then a burst phrase in a certain time period (that is, a phrase with a higher weighted burst score in the time period) is selected as the target week and target phrase.

[0186] The key phrase extraction model part 302 is mainly used to use TextRank_nfidf to construct a phrase node graph including burst weights and edge weights based on the initial article set, target phrases and target weeks.

[0187] The article update model part 303 is mainly used to collect candidate articles through the internal links and external link article pages of the initial article set to obtain a candidate article set, and select articles with high adjacency weight scores from the candidate articles to output the target article set. In order to avoid too many nodes, the embodiment of the present application also keeps the number of article sets unchanged.

[0188] for Figure 10 The process architecture shown in the figure takes as input the initial article set S, target phrase p and target week w; the output is the target article set S'.

[0189] The following will describe in detail the workflows of the key phrase extraction model part 302 and the article update model part 303 .

[0190] (1) Key phrase extraction model part 302

[0191] It should be noted that the workflow of the key phrase extraction model part 302 mainly includes:

[0192] S3021. Edit incremental extraction.

[0193] S3022, part-of-speech tagging.

[0194] S3023. Constructing a phrase node graph.

[0195] TextRank is a commonly used method for extracting key phrases (a key phrase is one or more phrases with high burstiness within a collection of articles). It draws inspiration from PageRank's webpage scoring system. TextRank constructs an edge-weighted graph using the co-occurrence relationships between phrases. Given an initial set of articles, the initial set consists of a series of edit increments; the textual differences between two consecutive edits of an article are called edit increments.

[0196] The embodiment of the present application may use the part-of-speech tagger of Apache Open NLP to divide each sentence in the edit increment into blocks, where blocks containing noun tags are considered phrases.

[0197] The embodiment of the present application can construct a phrase node graph for phrases in the initial article set. Specifically, for each article in the initial article set, a preset sliding window can be used for sliding, and the sliding size of the preset sliding window can be M words (or letters); wherein the specific value of M can be set based on actual conditions.

[0198] The schematic diagram of the phrase node graph can refer to the above Figure 3 ,like Figure 3 As shown, if the phrase V i and V j If V appears in the preset sliding window of the maximum M words in the text at the same time, i and V j is a pair of co-occurring phrases, indicating that V i and V j With an adjacency relationship, each node in the phrase node graph is a phrase (in the embodiment of the present application, a node in the phrase node graph represents a phrase in the article collection, so sometimes a phrase is also referred to as a node or phrase node). i To node V j There is a belt weight w ji The edge (V i ,V j ), edge weight w ji During the sliding process of the preset sliding window, V i and V j Count of co-occurrences.

[0199] The function for calculating the TextRank score refers to the above formula (4). Using this function, the TextRank algorithm can Figure 4 The phrase nodes in are scored, and the scoring result is the burst score of the phrase.

[0200] TextRank_nfidf reduces the scores of common phrases by adding a burst weight to the original TextRank. The phrase node graph with the added burst weight can refer to the above Figure 4 The TextRank_nfidf algorithm considers the time trend of phrases being edited and uses the TextRank_nfidf algorithm to Figure 4 When scoring the phrase nodes in , the scoring formula can refer to the above formula (6), and the scoring result is the weighted burst score of the phrase.

[0201] Here, the definition of the burst weight refers to the aforementioned formula (5), the definition of the reverse file frequency refers to the aforementioned formula (1), and the definition of the net edit frequency refers to the aforementioned formula (2).

[0202] The reverse document frequency (RDF) measures the general importance of a phrase. In a collection of articles, the RDF of a phrase can be calculated by dividing the total number of articles in the collection by the number of articles containing the phrase, and then taking the logarithm of the resulting quotient. The key idea behind RDF is that the fewer documents in a collection that contain a phrase, the greater its RDF, indicating that the phrase has good category-discriminating power. Frequently used phrases tend to have lower RDFs and smaller burst weights.

[0203] Net edit frequency measures the historical editing activity of each phrase. A phrase's net edit frequency is determined by its initial net edit frequency during the target week and the period preceding it. The initial net edit frequency is defined using Equation (3) above. It represents the average difference in the frequency of a phrase's occurrence across all edits of each article in the collection during a certain period. The net edit frequency for the target week and each week preceding it is then smoothed using an exponential moving average to obtain the net edit frequency for the phrase.

[0204] The main idea of ​​net edit frequency is: the greater the average difference in the frequency of a phrase appearing in all edits of each article in the article collection each week, the more sudden the event, the more likely it is to be a common phrase. On the contrary, if the frequency of a phrase appearing in each edit of all articles in all time periods is not much different, the more likely it is to be a common phrase.

[0205] In summary, the smaller the net edit frequency of common phrases, the smaller the reverse file frequency, the smaller the product of the net edit frequency and the reverse file frequency, the smaller the burst weight of the phrase, and the larger the section burst weight of the burst phrase.

[0206] (2) Article Update Model Part 303

[0207] It should be noted that the workflow of the article update model part 303 mainly includes:

[0208] S3031. Collection of candidate articles.

[0209] S3032. Article selection.

[0210] S3033. Update article collection.

[0211] In this embodiment of the present application, assuming an initial set of articles, a target phrase, and a target week, the target phrase is bursty in the initial set of articles in the target week. The goal of this embodiment of the present application is to add or delete articles from the initial set of articles so that the burstiness of the target phrase becomes clearer and stronger in the target week. The weighted burstiness score calculated by the TextRank_nfidf algorithm is used here to measure the burstiness of the phrase.

[0212] The candidate article set is determined from the inbound and outbound articles of the initial articles in the initial article set. Statistics show that the number of inbound and outbound articles may be in the tens of thousands. Therefore, a preliminary screening is performed to identify articles containing the target phrase within the target week's edit increment as candidate articles. These candidate articles form the candidate article set (C).

[0213] At this time, there is a candidate article set and an initial article set. The goal of the embodiment of the present application is to select |S| articles from the candidate article set and the initial article set to form a target article set S' so that in the target week, the weighted burst score of the target phrase in the target article set is maximized, and the total number of articles in the target article set is consistent with the total number of articles in the initial article set, that is, to improve the burst level of the target phrase in the article set, and the number of articles in the article set remains unchanged.

[0214] For the initial article set, target phrase, candidate article set and target week, when determining the target article set that makes the target phrase have the maximum burst level, we can try to add all combinations of articles from k = 1 to k = n = |C| (n = |C| represents the number of articles in the candidate article set), and their combinations are Thus, the number of combinations is exponential in magnitude of n, which is very time-consuming in actual calculations. Therefore, the present invention employs a certain search strategy to find the best combination strategy within a reasonable computation time.

[0215] The embodiment of the present application proposes an increased adjacent node weight (IDNW) strategy, which increases the burst weight of the target phrase by increasing the edge weights of the target phrase and the adjacent phrases of the target phrase in the phrase node graph.

[0216] In formula (6), the weighted burst score of a target phrase is positively correlated with the weighted burst score of the phrase adjacent to the target phrase (denoted as the first adjacent phrase), positively correlated with the edge weight between the target phrase and the first adjacent phrase, positively correlated with the burst weight of the target phrase, and negatively correlated with the edge weight between the first adjacent phrase and the second adjacent phrase, where the second adjacent phrase is the adjacent phrase of the first adjacent phrase and does not include the target phrase.

[0217] exist Figure 6 In TextRank_nfidf, the edge weight is defined as the probability of transitioning from one phrase node to another. To maximize the burst weight of the target phrase, we should try to increase the edge weight between the target phrase and the first adjacent phrase, that is, increase the occurrence of phrase pairs ap, bp, and cp, and reduce the occurrence of phrase pairs ad, ag, be, bf, and ch. Therefore, the occurrence of phrase pairs ap, bp, and cp will increase the probability of an article being selected, while the phrase pairs ad, ag, be, bf, and ch will decrease the probability of an article being selected.

[0218] The embodiment of the present application updates the phrase node graph (eg, Figure 6 and Figure 7 ) and the weighted burst score of the phrase.

[0219] The key points of improving the weighted burst score of the target phrase through the IDNW strategy are as follows:

[0220] 1. Reduce the burst weights of phrases other than the target phrase. In the phrase node graph of TextRank_nfidf, the burst weight of a phrase is given by the product of the inverse document frequency and the net edit frequency of the target weekly phrase. Replacing articles in the initial article set can change the frequency of phrase edits in the edit increment of the initial article set. If the net edit frequency of phrases other than the target phrase in the initial article set is reduced by article replacement, the burst weights of phrases other than the target phrase will become relatively low, and the burst weight of the target phrase will become relatively high. Since the weighted burst score of a phrase includes the weights of its nearby phrases, it is effective to reduce the burst weight of phrases far away from the target phrase.

[0221] 2. Reduce the edge weight between the first and second adjacent phrases, that is, reduce the edge weights of phrases that are not adjacent to the target phrase. The edge weights on the phrase node graph are derived from the co-occurrence frequency of the two phrases in the edit increment. The transition probability between phrases is derived from the edge weights. By increasing the edge weight between the target phrase and the first adjacent phrase and reducing the edge weight between the first and second adjacent phrases, the transition probability from the target phrase to the first adjacent node increases, that is, by increasing the edge weight between the target phrase and the first adjacent phrase.

[0222] To maximize the weighted burst score of the target phrase, candidate articles containing the target phrase are extracted from the inbound and outbound links of the articles in the initial article set to obtain a candidate article set. Given an initial article set, a target phrase, a candidate article set, and a target week, the process of generating an updated article set using the IDNW strategy is as follows:

[0223] (1) Extract a candidate article (A) from the candidate article set and add the candidate article to the initial article set (add one article).

[0224] (2) Calculate the adjacency weight score of each article (including candidate articles) in the initial article set and delete the article with the lowest adjacency weight score (delete one article).

[0225] Specifically, this solution performs the following process (in this process, the evolution of the phrase node graph can refer to Figure 6 and Figure 7 ):

[0226] A candidate article is selected from the candidate article set and added to the initial article set, and the adjacency weight score of each article in the initial article set is calculated respectively.

[0227] Increase the occurrence of phrase pairs ap, bp, cp. Since the weighted burst score of the target phrase is positively correlated with the weighted burst score of the first adjacent phrase and the edge weight between the target phrase and the first adjacent phrase, the embodiment of the present application calculates the weighted burst scores of a, b, and c, and calculates the co-occurrence frequencies of the co-occurring phrase pairs (a, p), (b, p), and (c, p) in the article, and makes the product of these two parts as large as possible (that is, the TextRank_nfidf(V) in formula (7) is positively correlated with the weighted burst score of the first adjacent phrase and the edge weight between the target phrase and the first adjacent phrase). i )PF(V i ,p) as large as possible).

[0228] Reduce the edge weights between the first adjacent phrases a, b, c and the second adjacent phrases d, e, f, g, h of the first adjacent phrases a, b, c. Figure 6, since the weighted burst score of the target phrase is negatively correlated with the edge weight between the first adjacent phrase and the second adjacent phrase.,This embodiment of the application calculates the phrase pair frequencies of the co-occurring phrase pairs (a, d), (a, g), (b, e), (b, f), and (c, h) in an article.

[0229] Reduce the occurrence of phrases that are not adjacent to the target node. The embodiment of the present application sets a threshold value n, selects n phrases with high weighted burst scores from the article except the target phrase and the adjacent phrases of the target phrase, and minimizes their occurrence. Since the weighted burst score of a phrase that is not adjacent to the target phrase is positively correlated with the burst weight of the node, if the net editing frequency of phrases other than the target phrase is reduced by article replacement, the burst weight of phrases other than the target phrase becomes relatively low, and the burst weight of the target phrase becomes relatively high. The embodiment of the present application calculates the weighted burst scores of the n phrases that are not adjacent to the target phrase and the phrase frequency of the non-adjacent phrase in the phrase node graph, and makes the product of these two parts as small as possible (that is, TextRank_nfidf(V) in formula (7) n )PF(V n ) as small as possible).

[0230] Finally, based on the calculated adjacency weight score of each article, the article with the lowest adjacency weight score is deleted.

[0231] (3) Go to step (1) until all candidate articles are traversed.

[0232] (4) Output the updated target article set S'.

[0233] This embodiment of the present application uses the IDNW strategy when updating the article set. To complete the experiment within a reasonable time, this embodiment of the present application also sets a time limit of 120 minutes. That is, if the time exceeds 120 minutes and the calculation of the articles in the candidate article set has not been completed, the calculation will not be continued, and the current initial article set will be determined as the target article set.

[0234] Furthermore, the embodiment of the present application can also evaluate the update effect of the article collection, mainly including the difference in the burst level of the target phrase and the diversity of the adjacent phrases of the target phrase.

[0235] With respect to the burst level, the embodiment of the present application can compare the degree to which the IDNW strategy improves the burst level of the target phrase by comparing the weighted burst score of the target phrase in the target article set with the incremental percentage of the weighted burst score of the target phrase in the initial article set. The specific determination method can refer to the aforementioned formula (8). The higher the incremental percentage of the weighted burst score (i.e., the first value), the higher the degree to which the IDNW strategy improves the burst level of the target phrase, and the more relevant the target article set is to the target phrase and the event represented by the target phrase.

[0236] In addition, you can also refer to Figure 8 , intuitively compare the degree to which the IDNW strategy improves the burst level of the target phrase. Figure 8 As can be seen in Figure 3, in burst weeks 60 and 90, the target article set has a higher burst level for the target phrase than the initial article set and the article set updated using other methods. This shows that the IDNW strategy can find the target article set related to the target phrase within a reasonable time, significantly improving the burst level of the target phrase in the target week.

[0237] Regarding the diversity of adjacent phrases, this scheme can show the extent to which the IDNW strategy improves the richness of the relevant event context of the target phrase by comparing the number of adjacent phrases of the target phrase in the initial article set and the target article set.

[0238] like Figure 9 As shown, compared with the initial article set, in the target article set, the number of adjacent phrases of the target phrase increases, and the richness of the relevant event context of the target phrase increases.

[0239] The diversity of the target phrase's neighboring phrases indicates the diversity of the bursts involving the target phrase. If the target phrase has a large number of edges with other phrases, it indicates that the target phrase appears in many different contexts. Diversity can be expressed by the number of neighboring phrases of the target phrase. When updating the article set, it is desirable to maximize the diversity of the target phrase. As can be seen, in the target article set, the number of neighboring phrases of the target phrase increases significantly around the target week.

[0240] In summary, the embodiment of the present application proposes a method for updating an article set based on the IDNW strategy. In the IDNW strategy, phrase pairs are used instead of phrases to search for articles related to emergencies. Not only is the judgment made based on the frequency of the target phrase, but also phrase pairs related to the target phrase of the emergency are considered. Moreover, the running speed is much faster than that of the general greedy search algorithm.

[0241] That is to say, the embodiment of the present application improves the traditional TextRank, and adds the burst weight of the phrase (inverse file frequency × net editing frequency) to TextRank to reflect the time trend in the encyclopedia editing history. The performance of TextRank_nfidf is better than TextRank, and the score of common phrases is reduced, and the burst phrases will have a higher weighted burst score. Moreover, TextRank_nfidf can find actively edited phrases from a relatively small set of articles, which is conducive to explaining the background of the emergency, locating the subset of articles where the emergency occurred, and tracking the evolution of the editing of the emergency. The purpose of the embodiment of the present application is to obtain a set of articles related to the emergency or target phrases of interest, and to update the set of articles related to the emergency or target phrases of interest by exploring the co-occurrence, connection and category structure of phrases in the article editing history, combined with the burst pattern. And TextRank_nfidf is applied to the article set to extract the top-ranked phrases as the theme of the article set, and the update effect of the article set is evaluated by comparing the peak level of the burst phrases and the top-ranked phrases before and after the update.

[0242] TextRank_nfidf can be used to calculate the weighted burst scores of burst phrases in the article set and its edit increments to extract burst phrases and their burst periods. However, the article set may not include all articles related to the target phrase of the sudden event. To make the article set contain more articles related to the target phrase, the embodiment of the present application adds more articles related to the target phrase and deletes irrelevant articles.

[0243] Given a target phrase, target week, and initial article set, by adjusting the article set to maximize the weighted burst score of the target phrase and make it more relevant to the sudden event, the present embodiment can find articles that explain the sudden event and significantly maximize the burst level.

[0244] In addition, TextRank_nfidf is applied to the article collection before and after the update to extract the phrases with the highest weighted burst score as the topics of the article collection before and after the update. By comparing the top-ranked phrases before and after the update, the embodiment of the present application can also evaluate the effect of the article collection update.

[0245] In short, this embodiment provides a method for updating an article collection, which is implemented based on the aforementioned article collection updating system. The specific implementation of the aforementioned embodiment is elaborated in detail through the above embodiment. It can be seen that the embodiment of the present application mainly extracts the target phrase and its burst period from the edit increment of a set of article collections by calculating the weighted burst score of the phrase. The embodiment of the present application can also obtain a phrase node graph, which includes many burst events in the initial article collection. However, the initial article collection may not include all articles related to a certain burst event. In order to obtain more articles related to the target phrase of a certain burst event, the embodiment of the present application adds articles related to the target phrase of the burst event, deletes irrelevant articles, and finally obtains an article collection that better explains the burst event.

[0246] In another embodiment of the present application, see Figure 11 , which shows a schematic diagram of the composition structure of an updating device 40 provided in an embodiment of the present application. Figure 11 As shown, the updating device 40 may include a determining unit 401, a calculating unit 402 and an updating unit 403, wherein:

[0247] The determining unit 401 is configured to determine at least one candidate phrase from the initial article set; and determine, based on the initial article set, an inverse document frequency and a net edit frequency of each of the at least one candidate phrase; wherein the inverse document frequency is used to represent the category distinguishing ability of the candidate phrase, and the net edit frequency is used to represent the number of times the candidate phrase has been edited;

[0248] The calculation unit 402 is configured to calculate a weighted score for each candidate phrase using the reverse document frequency and the net edit frequency of each candidate phrase to determine a weighted burst score for each candidate phrase; and determine a target phrase from the at least one candidate phrase based on the weighted burst score.

[0249] The updating unit 403 is configured to update the initial article set using the target phrase to obtain a target article set.

[0250] In some embodiments, the initial article set includes at least one initial article, and the determination unit 401 is specifically configured to determine a target time period; and obtain an editing increment of at least one initial article in the target time period; wherein the editing increment represents the text increment of an article before and after editing; and perform part-of-speech classification on the editing increment, and determine at least one candidate phrase from the result of the part-of-speech classification.

[0251] In some embodiments, the determination unit 401 is further specifically configured to obtain the total number of articles included in the initial article set and the number of first articles in the initial article set; wherein the first article refers to an article that contains a first candidate phrase in the edit increment within the target time period, and the first candidate phrase is any one of the at least one candidate phrase; and determine the initial net editing frequency of the first candidate phrase in each of several time periods; wherein the several time periods include the target time period and all time periods before the target time period; and perform a logarithmic operation on the ratio of the total number to the number of first articles to determine the inverse file frequency of the first candidate phrase; and perform an exponential moving average calculation on the initial net editing frequency of the first candidate phrase in each of the several time periods to determine the net editing frequency of the first candidate phrase.

[0252] In some embodiments, the determination unit 401 is further specifically configured to obtain the number of edits of each article in the initial article set within the first time period and the frequency of occurrence of the first candidate phrase in each edit increment of each article; and determine the initial net edit frequency of the first candidate phrase in the first time period based on the number of edits of each article and the frequency of occurrence of the first candidate phrase in each edit increment of each article; wherein the first time period is any one of a plurality of time periods.

[0253] In some embodiments, the calculation unit 402 is specifically configured to determine, in the initial article set, weighted burst scores of several first adjacent phrases and a first edge weight between the first candidate phrase and each first adjacent phrase; wherein the first adjacent phrase has an adjacent relationship with the first candidate phrase; and determine the second edge weight between each first adjacent phrase and several second adjacent phrases; wherein the several second adjacent phrases all have an adjacent relationship with the first adjacent phrase; and multiply the reverse file frequency of the first candidate phrase and the net edit frequency of the first candidate phrase to determine the burst weight of the first candidate phrase; and perform weighted calculation on the weighted burst scores of the several first adjacent phrases, the first edge weight between the first candidate phrase and each first adjacent phrase, and the second edge weight between each first adjacent phrase and the several second adjacent phrases according to the burst weight to obtain a weighted burst score of the first candidate phrase.

[0254] In some embodiments, the updating unit 402 is specifically configured to obtain the inbound and outbound articles of each article in the initial article set; select at least one article in the edit increment of the target time period that includes the target phrase from the inbound and outbound articles of each article; form a candidate article set based on the selected at least one article; and update the initial article set using the candidate article set to obtain the target article set.

[0255] In some embodiments, the updating unit 402 is further specifically configured to select an article from the candidate article set and add it to the initial article set; and use the IDNW strategy to determine the adjacency weight score of each article in the initial article set; wherein the adjacency weight score is used to characterize the probability of the article being retained; and select the lowest adjacency weight score from the determined adjacency weight scores, and delete the article corresponding to the lowest adjacency weight score from the initial article set to obtain an updated article set; and if the updated article set meets the preset conditions, determine the updated article set as the target article set; and if the updated article set does not meet the preset conditions, use the updated article set as the initial article set, and continue to perform the step of selecting an article from the candidate article set and adding it to the initial article set until the updated article set meets the preset conditions to obtain the target article set.

[0256] In some embodiments, the updating unit 402 is further specifically configured to determine, in the second article, the weighted burst scores of several third adjacent phrases and the co-occurrence frequency of the target phrase and the several third adjacent phrases; and determine the co-occurrence frequency of each of the third adjacent phrases with several fourth adjacent phrases; and determine the weighted burst scores and occurrence frequencies of a preset number of phrases other than the target phrase and the third adjacent phrases; and calculate based on the weighted burst scores of the several third adjacent phrases, the co-occurrence frequency of the target phrase and the several third adjacent phrases, and the weighted burst scores and occurrence frequencies of a preset number of phrases other than the target phrase and the third adjacent phrases to obtain an adjacency weight score of the second article; wherein the second article is any article in the initial article set, and in the second article, the several third adjacent phrases all have an adjacency relationship with the target phrase, and the several fourth adjacent phrases all have an adjacency relationship with the third adjacent phrase.

[0257] In some embodiments, the determination unit 401 is further configured to determine an initial weighted burst score of the target phrase in the initial article set and a target weighted burst score of the target phrase in the target article set; and determine an initial adjacent phrase of the target phrase in the initial article set and a target adjacent phrase of the target phrase in the target article set; and determine a first value of the target phrase based on the initial weighted burst score and the target weighted burst score; and determine a second value of the target phrase based on the initial adjacent phrase and the target adjacent phrase; wherein the first value is used to indicate a degree of improvement in the burst level of the target phrase, and the second value is used to indicate a degree of improvement in the context richness of the target phrase.

[0258] It is understood that in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular system. Furthermore, the various components in this embodiment can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The aforementioned integrated units can be implemented in the form of hardware or software functional modules.

[0259] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0260] Therefore, this embodiment provides a computer storage medium storing a computer program. When the computer program is executed by at least one processor, the steps of the article collection updating method described in any one of the aforementioned embodiments are implemented.

[0261] Based on the above-mentioned composition of an updating device 40 and computer storage medium, see Figure 12 , which shows a schematic diagram of the structure of an electronic device 50 provided in an embodiment of the present application. Figure 12 As shown, it may include: a communication interface 501, a memory 502 and a processor 503; each component is coupled together via a bus system 504. It is understood that the bus system 504 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 504 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 12 Various buses are labeled as bus system 504. Among them, the communication interface 501 is used to receive and send signals in the process of sending and receiving information between other external network elements;

[0262] Memory 502, used to store computer programs that can be run on processor 503;

[0263] The processor 503 is configured to, when running the computer program, execute:

[0264] determining at least one candidate phrase from the initial set of articles;

[0265] Determining, based on the initial set of articles, an inverse document frequency and a net edit frequency of at least one candidate phrase; wherein the inverse document frequency is used to characterize the category distinguishing ability of the candidate phrase, and the net edit frequency is used to characterize the number of times the candidate phrase has been edited;

[0266] Calculating a weighted score for each candidate phrase using the inverse document frequency and the net edit frequency of each candidate phrase to determine a weighted burst score for each candidate phrase;

[0267] determining a target phrase from at least one candidate phrase according to the weighted burst score;

[0268] The target phrase is used to update the initial article set to obtain the target article set.

[0269] It is understood that the memory 502 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DRRAM). The memory 502 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0270] The processor 503 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 503 or by software instructions. The processor 503 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 502, and the processor 503 reads the information in the memory 502 and, in conjunction with its hardware, completes the steps of the above method.

[0271] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof.

[0272] For software implementation, the techniques described herein can be implemented by modules (e.g., procedures, functions, etc.) that perform the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0273] Optionally, as another embodiment, the processor 503 is further configured to execute the method described in any one of the aforementioned embodiments when running the computer program.

[0274] See also Figure 13 , which shows a schematic diagram of the composition structure of another electronic device 50 provided in an embodiment of the present application. Figure 13 As shown, the electronic device 50 at least includes the updating device 40 according to any one of the aforementioned embodiments.

[0275] For the electronic device 50, for a given initial article set, when determining the weighted burst score of a candidate phrase therein, the reverse file frequency and net editing frequency of the candidate phrase in the initial article set are utilized, and the influence of the phrase's category distinction ability and the number of historical edits on the weighted burst score of the phrase is fully considered, so that a weighted burst score that accurately measures the burst level of the phrase can be obtained. Finally, a target phrase is selected from the candidate phrases, and the initial article set is updated using the target phrase, so that in the final target article set, the target phrase has a higher burst level, and the target article set is more relevant to the target phrase and the event represented by the target phrase, thereby improving the judgment accuracy of the phrase burst level and the burst level of the target phrase in the target article set.

[0276] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.

[0277] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0278] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0279] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0280] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0281] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0282] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for updating an article collection, characterized in that: The method comprises: determining at least one candidate phrase from the initial set of articles; Determining, based on the initial set of articles, an inverse document frequency and a net edit frequency of each of the at least one candidate phrase; wherein the inverse document frequency is used to characterize the category distinguishing ability of the candidate phrase, and the net edit frequency is used to characterize the number of times the candidate phrase has been edited; performing weighted score calculations on corresponding candidate phrases using the respective reverse document frequencies and the respective net edit frequencies of the at least one candidate phrase to determine a weighted burst score for each of the at least one candidate phrases; determining a target phrase from the at least one candidate phrase according to the weighted burst score; Updating the initial article set using the target phrase to obtain a target article set; The step of calculating a weighted score for a corresponding candidate phrase using the reverse file frequency and the net edit frequency of each candidate phrase to determine a weighted burst score of each candidate phrase includes: Determining, in the initial article set, weighted burst scores of a plurality of first adjacent phrases and a first edge weight between a first candidate phrase and each of the first adjacent phrases; wherein the first adjacent phrase has an adjacent relationship with the first candidate phrase, and the first candidate phrase is any one of the at least one candidate phrase; Determining a second edge weight between each of the first adjacent phrases and a plurality of second adjacent phrases; wherein the plurality of second adjacent phrases all have an adjacent relationship with the first adjacent phrase; performing multiplication calculation on the inverse document frequency of the first candidate phrase and the net edit frequency of the first candidate phrase to determine a burst weight of the first candidate phrase; The weighted burst scores of the first candidate phrases, the first edge weights between the first candidate phrase and each of the first adjacent phrases, and the second edge weights between each of the first adjacent phrases and the plurality of second adjacent phrases are weightedly calculated according to the burst weights to obtain a weighted burst score of the first candidate phrase.

2. The method according to claim 1, characterized in that The initial article set includes at least one initial article, and determining at least one candidate phrase from the initial article set includes: Determine the target period; Obtaining an edit increment of the at least one initial article during the target time period; wherein the edit increment represents a text increment of an article before and after editing; Part-of-speech classification is performed on the edit increment, and the at least one candidate phrase is determined from the result of the part-of-speech classification.

3. The method according to claim 2, characterized in that Determining the reverse document frequency and net edit frequency of each of the at least one candidate phrase based on the initial set of articles includes: Obtaining the total number of articles in the initial article set and the number of first articles in the initial article set; wherein the first article refers to an article that contains the first candidate phrase in an edit increment within the target time period; determining an initial net edit frequency of each of the first candidate phrases in a plurality of time periods; wherein the plurality of time periods include the target time period and all time periods before the target time period; performing a logarithmic operation on a ratio of the total number to the number of the first articles to determine an inverse document frequency of the first candidate phrase; An exponential moving average calculation is performed on the initial net edit frequencies of the first candidate phrase in the plurality of time periods to determine the net edit frequency of the first candidate phrase.

4. The method according to claim 3, characterized in that The determining of the initial net edit frequencies of the first candidate phrases in a plurality of time periods includes: Obtaining the number of edits of each article in the initial article set within a first time period and the frequency of occurrence of the first candidate phrase in each edit increment of each article; determining an initial net edit frequency of the first candidate phrase in the first time period based on the number of edits of each article and the frequency of occurrence of the first candidate phrase in each edit increment of each article; The first time period is any one of the multiple time periods.

5. The method according to claim 1, characterized in that The updating of the initial article set by using the target phrase to obtain the target article set includes: Obtaining the inbound and outbound links of each article in the initial article set; Selecting at least one article that includes the target phrase in the editing increment of the target period from the inbound articles and outbound articles of each article; forming a candidate article set based on the at least one selected article; The initial article set is updated using the candidate article set to obtain the target article set.

6. The method according to claim 5, characterized in that The updating of the initial article set by using the candidate article set to obtain the target article set includes: Selecting an article from the candidate article set and adding it to the initial article set; Determining an adjacency weight score for each article in the initial article set using an IDNW strategy; wherein the adjacency weight score is used to represent the probability of the article being retained; Selecting a lowest adjacency weight score from the determined adjacency weight scores, and deleting the article corresponding to the lowest adjacency weight score from the initial article set to obtain an updated article set; In the case where the updated article set meets a preset condition, determining the updated article set as the target article set; If the updated article set does not meet the preset condition, the updated article set is used as the initial article set, and the step of selecting an article from the candidate article set and adding it to the initial article set is continued until the updated article set meets the preset condition, so as to obtain the target article set.

7. The method according to claim 6, characterized in that Determining the adjacency weight score of each article in the initial article set using the IDNW strategy includes: In the second article, determining weighted burst scores of a plurality of third adjacent phrases and co-occurrence frequencies of the target phrase and the plurality of third adjacent phrases; determining a co-occurrence frequency between each of the third adjacent phrases and a plurality of fourth adjacent phrases; determining weighted burst scores and occurrence frequencies of a preset number of phrases other than the target phrase and the third adjacent phrase; Obtaining an adjacency weight score for the second article by calculating based on the weighted burst scores of the plurality of third adjacent phrases, the co-occurrence frequencies of the target phrase and the plurality of third adjacent phrases, and the weighted burst scores and occurrence frequencies of the preset number of phrases other than the target phrase and the third adjacent phrases; The second article is any article in the initial article set, and in the second article, the plurality of third adjacent phrases all have an adjacent relationship with the target phrase, and the plurality of fourth adjacent phrases all have an adjacent relationship with the third adjacent phrase.

8. The method according to any one of claims 1 to 7, characterized in that The method further comprises: determining an initial weighted burst score of the target phrase in the initial article set and a target weighted burst score of the target phrase in the target article set; Determining initial adjacent phrases of the target phrase in the initial article set and target adjacent phrases of the target phrase in the target article set; determining a first value of the target phrase based on the initial weighted burst score and the target weighted burst score; and determining a second value of the target phrase based on the initial adjacent phrase and the target adjacent phrase; The first value is used to indicate the degree to which the burst level of the target phrase is improved, and the second value is used to indicate the degree to which the context richness of the target phrase is improved.

9. An updating device, characterized in that: The updating device includes a determining unit, a calculating unit and an updating unit, wherein: The determining unit is configured to determine at least one candidate phrase from an initial set of articles; and determine, based on the initial set of articles, an inverse document frequency and a net edit frequency of each of the at least one candidate phrase; wherein the inverse document frequency is used to represent the category distinguishing ability of the candidate phrase, and the net edit frequency is used to represent the number of times the candidate phrase has been edited; The calculation unit is configured to calculate a weighted score for each candidate phrase using the reverse document frequency and the net edit frequency of each candidate phrase to determine a weighted burst score for each candidate phrase; and determine a target phrase from the at least one candidate phrase based on the weighted burst score; The updating unit is configured to update the initial article set using the target phrase to obtain a target article set; The method of calculating a weighted score for each candidate phrase using the inverse document frequency and the net edit frequency of each candidate phrase to determine a weighted burst score for each candidate phrase includes: determining, in the initial article collection, weighted burst scores of a plurality of first adjacent phrases and a first edge weight between a first candidate phrase and each of the first adjacent phrases; wherein the first adjacent phrase has an adjacent relationship with the first candidate phrase, and the first candidate phrase is any candidate phrase from the at least one candidate phrase; and determining second edge weights between each of the first adjacent phrases and a plurality of second adjacent phrases; wherein the plurality of second adjacent phrases all have an adjacent relationship with the first adjacent phrase; and multiplying the inverse document frequency of the first candidate phrase by the net edit frequency of the first candidate phrase to determine a burst weight of the first candidate phrase; and performing weighted calculation on the weighted burst scores of the plurality of first adjacent phrases, the first edge weight between the first candidate phrase and each of the first adjacent phrases, and the second edge weight between each of the first adjacent phrases and a plurality of second adjacent phrases according to the burst weights to obtain a weighted burst score for the first candidate phrase.

10. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein: The memory is used to store a computer program that can be run on the processor; The processor is configured to execute the article collection updating method according to any one of claims 1 to 8 when running the computer program.

11. A computer storage medium, characterized in that The computer storage medium stores a computer program, and when the computer program is executed by at least one processor, the article collection updating method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Information pushing method and apparatus

    CN108241667A

  • Method for detecting microblog emergencies

    CN110543590A