Event title extraction method and system and computer terminal
The word graph is constructed through the BERTopic theme model and the improved cluster search algorithm, and the event titles in the government hotline text are automatically extracted, solving the problems of low statistical efficiency and lagging response of hot events, and achieving fast and accurate event title extraction.
Patent Information
- Application Number
- CN202510764295.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the number of people's appeal texts accepted by government hotlines has increased dramatically, resulting in low statistical efficiency of hot events, strong subjectivity of manual interpretation and difficulty in dealing with real-time data flow, resulting in lagging response.
The event title extraction method is adopted, the theme clustering is carried out through the BERTopic theme model, word graph is constructed, node and edge weights are calculated, and the optimal path is searched using the improved cluster search algorithm, and event titles are automatically extracted.
It improves the efficiency of hot-spot events extraction, reduces the subjectivity of manual interpretation, and can quickly respond to hot-spot events, making it easier for relevant staff to handle it in a timely manner.
Smart Images

Figure CN120277213A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information processing, and particularly relates to an event title extraction method, system and computer terminal. Background Art
[0002] With the increasing demand for social governance, the data of government service hotlines has expanded rapidly, and the number of texts of people's demands accepted by government service hotlines has increased sharply. How to efficiently count hot events has become a key challenge in improving service quality and management efficiency.
[0003] For multi-document analysis, topic clustering technology is currently often used to classify massive texts, but the classification results only stay at the level of "topic sets". In fact, an event is a concise summary of the core content of a topic. Staff still need to manually interpret the content of topic clusters, resulting in low extraction efficiency. In addition, manually interpreting and generating event topic sentences is highly subjective and difficult to handle real-time data streams, leading to a lag in the response to hot events. Summary of the Invention
[0004] The present invention provides an event title extraction method, system and computer terminal to solve at least one aspect of the problems in the prior art.
[0005] In a first aspect, the present invention provides an event title extraction method, which includes: Step 101: Obtain target text data, where the target text data includes multiple target texts; Step 102: Perform topic clustering on all the target texts included in the target text data to obtain several clusters; Step 103: According to a preset screening rule, screen the texts of each cluster to obtain the screened clusters; Step 104: For each screened cluster, gather the texts within the cluster to form their respective candidate sets; Step 105: Extract the event title of each candidate set to obtain all the event titles corresponding to all the target texts included in the target text data.
[0006] Further, in step 105, the method for extracting the event title of each candidate set all includes: Construct a word graph corresponding to all the texts in the candidate set; Calculate the weight of each node in the word graph to obtain the weights of each node in the word graph; Calculate the weight of each edge in the word graph to obtain the weights of each edge in the word graph; Based on the obtained node weights and edge weights, search for the optimal path in the word graph to obtain the target path; In the direction from the start end to the end end along the path, the word segments on each node in the target path are sequentially concatenated to obtain the event title of the candidate set.
[0007] Further, construct word graphs corresponding to all texts in the candidate set, specifically including: Step 110: Generate a word graph containing a start node and an end node, and then add an edge with a direction from the start node to the end node between the start node and the end node; Step 120: For the first text in the candidate set , use a word segmenter to perform word segmentation to obtain all word segments of the text and the part-of-speech of each word segment; Step 130: According to a preset stop word list, remove the stop words in all word segments of the text to obtain all word segments after removing the stop words of the text and their part-of-speech; Step 140: Disconnect the edge between the start node and the end node, and then, according to the order of all word segments after removing the stop words of the text and their part-of-speech in the text , insert new nodes sequentially in the disconnected place in the direction from the start node to the end node. After each word segment is inserted, record the word frequency of the node word segment as 1 on the corresponding inserted node; then add edges between adjacent nodes in the word graph, and the direction of the edge is from the previous node to the next node among adjacent nodes; Step 150: Traverse each text in the remaining texts in the candidate set , , where n is the total number of texts in the candidate set. For each text traversed, perform the following steps: Step 1501: For the text traversed, use a word segmenter to perform word segmentation to obtain all word segments of the text and their part-of-speech; Step 1502: According to a preset stop word list, remove the stop words in all word segments of the obtained text to obtain all word segments after removing the stop words of the text and their part-of-speech; Step 1503: Traverse each word segment after removing the stop words of the text one by one in the order in which they appear in the text , and for each word segment traversed, perform the following steps: Check whether there is a target node in the entire word graph. The target node is a node whose name and part-of-speech of the node word segment are the same as those of the currently traversed word segment: If it exists, determine whether merging the currently traversed token with the target node will create a loop. If it is determined that a loop will be created, create a new node in the word graph, record the currently traversed token and its part of speech on the newly created node, and record the word frequency of the token in the current node as 1 in the newly created node. If it is determined that no loop will be created, merge the currently traversed token with the target node and increment the value of the word frequency recorded in the target node by 1; If it does not exist, create a new node in the word graph, record the currently traversed token and its part of speech on the newly created node, and record the word frequency of the token in the current node as 1 in the newly created node; Step 1504, after step 1503 is executed: According to the order in which the tokens appear in the text add in-edges and out-edges to each node of the tokens after removing stop words in the text inserted into the word graph; Step 160, the word graph obtained after step 150 is completed is the word graph corresponding to all texts in the constructed candidate set.
[0008] Furthermore, the method for determining whether merging the currently traversed token with the target node will create a loop includes: Obtain the previous token of the currently traversed token in the token sequence. The token sequence is a sequence obtained by sorting the tokens after removing stop words in the text in the order in which they appear in the text ; Determine whether the previous token is on the path where the out-edge of the target node is located: If so, it is determined that a loop will be created; If not, determine whether the previous token is on the path between the adjacent node in the in-edge direction of the target node in the word graph and the starting node: If so, it is determined that a loop will be created; If not, it is determined that no loop will be created.
[0009] Furthermore, the searching for the optimal path in the word graph based on the obtained node weights and edge weights to obtain the target path includes: Based on the obtained node weights and edge weights, use an improved beam search algorithm to search for the optimal path in the word graph to obtain the target path; The improved beam search algorithm is improved in that it improves the calculation of the score of the searched path for ranking the searched paths to obtain the optimal path, specifically including: In the search of the beam search algorithm before improvement, each time the end node is reached in the path search, an additional score is added to the score of the current path statistically calculated by the beam search algorithm before improvement based on the node weights and edge weights of the current path, so as to adjust the final score of the current path for the beam search algorithm to calculate the ranking of the current path and search for the optimal path in the word graph; The calculation formula of the additional score is:
[0010] Among them, represents the additional score of the current path, is the path length of the current path, L1 is the preset path length, is a real coefficient, represents the word frequency recorded on the last node in the current path in the word graph.
[0011] Furthermore, in step 102, theme clustering is performed on all target texts included in the target text data, and the text-topic probability of each text under each cluster is also obtained; In step 103, the preset screening rules include: Delete texts in the cluster whose text content length is less than the preset text length threshold L2; The value range of L2 is 5 characters to 10 characters; Delete texts in the cluster whose text-topic probability is less than or equal to the preset probability threshold P; For texts that belong to multiple clusters at the same time, only keep them in the cluster with the highest text-topic probability; For each cluster, delete redundant texts in the cluster.
[0012] Furthermore, the method for calculating the node weight of each node in the word graph all includes: Calculate the term frequency-inverse document frequency of each node in the word graph, and the specific calculation formula is: , , Among them, represents the th node in the word graph, , , represents the th text in the candidate set, represents the node 's word frequency, represents the node 's inverse document frequency, and N represents the total number of texts in the candidate set; Calculate each node Information entropy , and the specific calculation formula is as follows: , where ; Calculate the semantic relevance of each node in the word graph in the context , and the specific calculation formula is as follows: , where represents the semantic relevance of the word segmentation on node in the text of the candidate set , , represents the cosine similarity represents the word vector of the text where the word segmentation on node is located is the word vector of the text , is the sentence weight , represents the number of times the statistical sentence appears in the candidate set; Calculate the overall similarity of each node in the word graph with other nodes in the word graph , and the specific formula is as follows: , where represents the normalization factor whose value is the total number of nodes in the current word graph represents node the cosine similarity between the word segmentation on it and the word segmentation on node in the word graph; Calculate the weight of each node in the word graph , and the calculation formula is as follows:
[0013] where is the weight coefficient .
[0014] Furthermore, the calculation formula for the edge weight of each edge in the word graph is:
[0015] where represents the score of the edge between two nodes in the word graph The word segmentations on the nodes in the representative word graph and the word segmentations on the nodes in the representative word graph and the word segmentations on the nodes the probability values on a pre-trained Chinese bigram language model The word segmentations on the nodes in the representative word graph the occurrence probability in the text within the candidate set The word segmentations on the nodes in the representative word graph the occurrence probability in the text within the candidate set represent the edges the number of adjacent occurrences in the text within the candidate set of the two nodes on the edges and the number of adjacent occurrences in the text within the candidate set
[0016] Second, the present invention provides an event title extraction system, which includes: A text acquisition module for acquiring target text data, where the target text data includes multiple target texts; A clustering module for performing topic clustering on all the target texts included in the target text data by using the BERTopic topic model to obtain several clusters; A screening module for screening the texts of each cluster according to a preset screening rule to obtain the screened clusters; A candidate module for, for each screened cluster, collecting the texts within the cluster to form their respective candidate sets; An event title extraction module for correspondingly extracting the event titles of each candidate set to obtain all the event titles corresponding to all the target texts included in the target text data.
[0017] Third, the present invention provides a computer terminal, which includes: A processor; A memory for storing the execution instructions of the processor; wherein, the processor is configured to execute the methods described in the above aspects.
[0018] It can be seen from the above technical solutions that the present invention has the following advantages: After clustering to obtain clusters, the present invention screens each cluster, then obtains the candidate set corresponding to each screened cluster, and then extracts the event title for each candidate set, and then obtains the event titles of all topics extracted from the target text data, thereby avoiding leaving the clustering result only at the level of "topic set" for manual interpretation of the topic cluster content, which helps to improve the efficiency of extracting event titles to a certain extent.
[0019] The present invention does not require manual interpretation of the topic cluster content, avoiding the generation of event topic sentences by manual interpretation, overcoming the problems of strong subjectivity in manual interpretation and difficulty in dealing with real-time data streams, helping to quickly extract all hot events in the target text data, and facilitating timely response by relevant staff. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic flow chart of the method provided by an embodiment of the present invention.
[0022] Figure 2 It is a schematic diagram of the system provided by an embodiment of the present invention.
[0023] Figure 3 It is a schematic diagram of the computer terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The event title extraction method provided by the present invention can realize the mining of hot events in the target text of the target text data, thus helping to meet the high requirements for text data processing in social governance, and further providing certain support for the solution of public demands.
[0025] In order to facilitate a clear description of the technical solutions of the embodiments of the present invention, some terms and technologies involved in the embodiments of the present invention are briefly introduced below: BERTopic: It is a topic modeling technology based on BERT, which combines the semantic understanding ability of BERT and the non-parametric clustering algorithm HDBSCAN, and can better capture the complex semantic structure of the text.
[0026] The specific execution steps of the event title extraction method will be described in detail below. For the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details.
[0027] It should be understood that when used in the specification of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0028] The statements such as "an embodiment" or "some embodiments" described in the present invention mean that the specific features, structures or characteristics described in the embodiment are included in one or more embodiments of the present invention. Thus, the statements such as "in some embodiments" and "in other some embodiments" that appear in different places in the present invention do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0030] Figure 1 It is a schematic flow chart of a method for extracting event titles provided for the embodiments of the present invention. Among them, Figure 1 The execution subject can be an event title extraction system. The event title extraction method provided by the embodiments of the present invention is executed by a computer terminal. Correspondingly, the event title extraction system runs in the computer terminal. According to different requirements, the order of the steps in this flow chart can be changed, and some can be omitted.
[0031] Please refer to Figure 1 , and the method includes the following steps 101 to step 105.
[0032] Step 101: Obtain target text data, where the target text data includes multiple target texts.
[0033] It can be understood that the target text data is the data to be extracted for event titles.
[0034] Specifically, the target text data can be the text data of the public's demands accepted by the government service hotline.
[0035] Step 102: Perform topic clustering on all the target texts included in the target text data to obtain several clusters.
[0036] Before performing topic clustering on all the target texts included in the target text data, preprocess the target text data first. The preprocessing includes data cleaning. The data cleaning includes removing irrelevant characters, such as removing non-Chinese characters, garbled codes, HTML tags and other noises in the text.
[0037] Step 103: According to the preset screening rules, screen the texts of each cluster to obtain the screened clusters of each type.
[0038] Step 104: For each screened cluster, gather the texts within the cluster to form their respective candidate sets.
[0039] Step 105: Extract the event titles of each candidate set to obtain all the event titles corresponding to all the target texts included in the target text data.
[0040] Exemplarily, in step 102, the BERTopic topic clustering model can be used to perform topic clustering on all the target texts included in the target text data.
[0041] Optionally, when using the BERTopic topic clustering model to perform topic clustering on the target text data, the text-topic probability of each text under each clustering cluster can also be obtained in the clustering result.
[0042] It can be understood that the text-topic probability of the text under the cluster is the probability that the text belongs to the cluster.
[0043] Step 105: Extract the event titles of each candidate set to obtain all the event titles corresponding to all the target texts included in the target text data.
[0044] In the present invention, first perform topic clustering on all the target texts included in the target text data, then screen the texts of each cluster, then for each screened cluster, gather the texts within the cluster to form their respective candidate sets, and finally extract the event titles of each candidate set to obtain all the event titles corresponding to all the target texts included in the target text data.
[0045] In some embodiments, the method for extracting the event title of each candidate set in step 105 all includes: Construct a word graph corresponding to all the texts in the candidate set; Calculate the weights of each node in the word graph to obtain the weights of each node in the word graph; Calculate the weights of each edge in the word graph to obtain the weights of each edge in the word graph; Based on the obtained node weights and edge weights, search for the optimal path in the word graph to obtain the target path; According to the direction of the path from the start end to the end end, splice the word segments on each node in the target path in sequence to obtain the event title of the candidate set.
[0046] Understandably, the present application provides a new method for extracting event titles. The event titles extracted by this method take into account all the text extractions in the candidate set, and the extraction results are focused.
[0047] In some other embodiments, the present application constructs word graphs corresponding to all the text in the candidate set, specifically including: Step 110: Generate a word graph containing a start node and an end node, and then add an edge with a direction from the start node to the end node between the start node and the end node; Step 120: For the first text in the candidate set , use a word segmentation tool to perform word segmentation to obtain all the word segments of the text and the part-of-speech of each word segment; Step 130: According to a preset stop word list, remove the stop words in all the word segments of the text to obtain all the word segments and their part-of-speech after removing the stop words of the text ; Step 140: Disconnect the edge between the start node and the end node, and then, according to the order of the word segments after removing the stop words of the text in the text , insert new nodes one by one in the direction from the start node to the end node at the disconnection point. After each word segment is inserted, record the word frequency of the node word segment as 1 on the corresponding inserted node; then add edges between adjacent nodes in the word graph, and the direction of the edge is from the previous node to the next node among adjacent nodes; Step 150: Traverse each text in the remaining texts in the candidate set , where n is the total number of texts in the candidate set. For each text traversed, execute steps 1501 to 1503: Step 1501: For the text traversed, use a word segmentation tool to perform word segmentation to obtain all the word segments of the text and their part-of-speech; Step 1502: According to a preset stop word list, remove the stop words in all the word segments of the obtained text to obtain all the word segments and their part-of-speech after removing the stop words of the text ; Step 1503: Traverse each word segment after removing the stop words of the text one by one in the order in which they appear in the text , and for each word segment traversed, execute the following steps: Search for the target node in the entire word graph. The target node is the node whose name and part of speech of the node segmentation are the same as those of the currently traversed segmentation: If it exists, determine whether combining the currently traversed segmentation with the target node will create a loop. If the determination is yes, create a new node in the word graph, record the currently traversed segmentation and its part of speech on the newly created node, and record the word frequency of the segmentation in the current node as 1 in the newly created node. If the determination is no, combine the currently traversed segmentation with the target node and increment the value of the word frequency recorded in the target node by 1; If it does not exist, create a new node in the word graph, record the currently traversed segmentation and its part of speech on the newly created node, and record the word frequency of the segmentation in the current node as 1 in the newly created node; Step 1504. After step 1503 is executed: According to the order in which the segmentations appear in the text in the text, add incoming and outgoing edges to each node of the segmentations after removing stop words in the text inserted into the word graph; Step 160. The word graph obtained after step 150 is executed is the word graph corresponding to all texts in the constructed candidate set.
[0048] In the process of generating the word graph in this application, the word frequency TF (i.e., the number of occurrences) of each node segmentation in the text of the candidate set is counted in the word graph. Thus, while completing the construction of the word graph, the word frequency TF of each node segmentation in all texts in the candidate set is also counted.
[0049] The node segmentation is the segmentation on the node.
[0050] It can be understood that the word graph in this application is a directed graph, and the incoming and outgoing edges are the incoming edge and the outgoing edge.
[0051] In a directed graph: Adjacent nodes: The two nodes (i.e., two vertices) at both ends of an edge are called adjacent nodes; Incoming edge of a node: It refers to the edge with the current node as the end point; Outgoing edge of a node: It refers to the edge with the current node as the starting point.
[0052] The word graph introduces two auxiliary nodes: the start node (STARTNODE) and the end node (ENDNODE). STARTNODE is used to identify the starting point of the word graph, and it points to the first word node of each input text. ENDNODE is used to identify the end point of the word graph, and the last word node of each input text points to this node. These two special nodes do not belong to the actual content of the input text, but are used to ensure the integrity of the word graph structure and clear path identification.
[0053] The stop word list contains pre-set words with no specific meaning, such as adverbs, modal particles, conjunctions, etc. For example, Table 1 shows an exemplary stop word list.
[0054] Table 1: Stop Word List
[0055] In specific implementation, those skilled in the art can set the stop words in the stop word list according to actual needs.
[0056] In other embodiments, the method for the present application to determine whether merging the currently traversed word segment with the target node will generate a loop includes: Obtain the previous word segment of the currently traversed word segment in the word segment sequence, where the word segment sequence is a sequence obtained by sorting the word segments after removing stop words from the text according to the order in which the word segments appear in the text in the text ; Determine whether the previous word segment is on the path where the out-edge of the target node is located: If so, it is determined that a loop will be generated; If not, determine whether the previous word segment is on the path between the adjacent node in the in-edge direction of the target node in the word graph and the starting node: If so, it is determined that a loop will be generated; If not, it is determined that no loop will be generated.
[0057] The present application can actively avoid the generation of loops in the word graph construction stage, which not only helps to fundamentally solve the loop problem, but also helps to retain the possible optimal path. This not only helps to improve the robustness and accuracy of the model, but also helps to ensure that the finally generated topic sentence (i.e., event title) can completely and accurately convey the core information of the input sentence.
[0058] Optionally, the above-mentioned searching for the optimal path in the word graph based on the obtained node weights and edge weights to obtain the target path includes: Based on the obtained node weights and edge weights, use an improved beam search algorithm to search for the optimal path in the word graph to obtain the target path.
[0059] In one embodiment, the improvement of the above-mentioned improved beam search algorithm lies in improving the beam search algorithm to calculate the scores of the searched paths for ranking the searched paths to obtain the optimal path, specifically including: During the search of the beam search algorithm before improvement, each time the end node is reached in the path search, an additional score is added to the score of the current path statistically calculated by the beam search algorithm before improvement based on the node weights and edge weights of the current path, so as to adjust the final score of the current path for the beam search algorithm to calculate the ranking of the current path and search for the optimal path in the word graph; The calculation formula of the additional score is as follows:
[0060] wherein, represents the additional score of the current path, is the path length of the current path, L1 is a preset path length, is a real number coefficient, represents the word frequency recorded on the last node in the current path in the word graph.
[0061] In some embodiments, the value range of the above L1 is 5 nodes to 7 nodes, and the value range of the above is .
[0062] In this embodiment, L1 takes 5 nodes, .
[0063] After the word graph is constructed, there are many connected paths in the word graph. Based on the weights of the nodes and edges in the word graph, the present application uses an improved beam search algorithm to search for the path with the highest score (i.e., the total path weight) in the word graph as the optimal path, and then generates an event topic sentence (i.e., an event title) of the candidate set corresponding to the word graph (each candidate set corresponds to a topic) based on this path.
[0064] In some other embodiments, in step 103, the preset screening rules include the following Rule 1 to Rule 4.
[0065] Rule 1: Delete the text in the cluster whose text content length is less than the preset text length threshold L2.
[0066] Exemplarily, the value range of L2 is 5 characters to 10 characters.
[0067] Rule 1 is used to control the length of the text in the cluster. The text with a content length shorter than 5 to 10 characters contains less content, and most of the texts with such a length are incomplete sentences with low information density. The present invention discards them, which helps to remove interference and improve the extraction accuracy to a certain extent.
[0068] Rule 2: Delete the text-cluster topic probability (less than or equal to) the text with the preset probability threshold P.
[0069] The use of Rule 2 helps the present application filter out irrelevant or low-correlation documents and screen out texts with a high degree of relevance to the theme, thereby helping to reduce the impact of interference information, control the quantity and quality of the sources of texts in the candidate set below, and ensure that the event titles extracted based on the candidate set are more focused within the corresponding theme range.
[0070] Rule 3: For texts that belong to multiple clusters simultaneously, only keep them in the cluster with the highest text-topic probability.
[0071] Rule 3 corresponds to the screening of multi-topic documents. When a certain text belongs to multiple themes, it is preferentially classified into the category with the highest text-topic probability.
[0072] For example, there is a target text c, which belongs to cluster a and cluster b simultaneously. Among them, the probability of belonging to cluster a (i.e., the text-topic probability) is 0.96, and the probability of belonging to cluster b (the text-topic probability) is 0.75. Then, delete the text c in cluster b and keep the text c in cluster a.
[0073] Rule 4: For each cluster, delete redundant texts within the cluster.
[0074] Among the clusters obtained by clustering, there are often some texts that are similar, such as having basically the same meaning but different expressions. In specific implementation, a text similarity measurement method can be used to evaluate the similarity between texts in the cluster, and for texts whose similarity reaches the preset similarity threshold, redundant deletion is performed. This helps to avoid information redundancy and improve information coverage.
[0075] In some specific embodiments, the method for calculating the node weight of each node in the word graph of the present application all includes: Calculating the term frequency-inverse document frequency of each node in the word graph, and the specific calculation formula is: , , where represents the th node in the word graph, , , represents the th text in the candidate set, represents the term frequency of node , represents the inverse document frequency of node , and N represents the total number of texts in the candidate set; Calculating the information entropy of each node in the word graph , the specific calculation formula is: , where ; Calculate the semantic relevance of each node in the word graph in the context, the specific calculation formula is: , the specific calculation formula is: , where, represents the semantic relevance of the word segmentation on node in the text of the candidate set, in the semantic relevance, , represents the cosine similarity, represents the word vector of the text where the word segmentation on node in the word graph is located, is the word vector of the text , is the sentence weight, , represents the statistical sentence the number of times it appears in the candidate set; Calculate the overall similarity of each node in the word graph with other nodes in the word graph , the specific formula is: , where, represents the normalization factor, the value is the total number of nodes in the current word graph, represents the node the cosine similarity between the word segmentation on and the word segmentation on node in the word graph; Calculate the weight of each node in the word graph the weight , the calculation formula is:
[0076] where, is the weight coefficient, .
[0077] In this application, the calculation of node weights takes into account the term frequency and inverse document frequency (TF-IDF) of nodes, selects high-frequency words that can better represent the theme, and reduces the interference of common words, which helps to ensure that the selected words are unique. In the calculation of node weights in this application, for word segments with low term frequency and IDF, the semantic association between the node word segments and the theme is considered. In the calculation of node weights in this application, information entropy is used to consider the word segments that cover diverse texts in the candidate set text. In the calculation of node weights in this application, it is also considered to extract words with high similarity in the same theme as core words through clustering similarity to reduce the interference of outlier words.
[0078] In some other specific embodiments, the calculation formula for the edge weight of each edge in the word graph is:
[0079] Wherein, represents the score of the edge between two nodes in the word graph, represents the probability value of the word segment on the node in the word graph and the word segment on the node on the pre-trained Chinese bigram language model, represents the occurrence probability of the word segment on the node in the candidate set text, represents the occurrence probability of the word segment on the node in the candidate set text, represents the number of adjacent occurrences of the word segments on the two nodes on the edge and in the candidate set text.
[0080] Optionally, the method for obtaining the above-mentioned trained Chinese bigram language model includes: Construct a training set; Build a pre-trained Chinese bigram language model; Use the training set to train the network architecture to obtain a trained Chinese bigram language model.
[0081] Specifically, the method for obtaining the above-mentioned trained Chinese bigram language model includes: 1. Construct a training set.
[0082] Step 1, data collection: Scrape approximately 20,000 news texts from news websites such as China News Service and Toutiao as general corpus.
[0083] When scraping texts, ensure that the scraped news texts cover as many topics and fields as possible.
[0084] Collect historical government texts as domain corpus. The historical government texts are Step 2: Data cleaning: Perform data cleaning on the general corpus, and collect the cleaned general corpus to form a general corpus library.
[0085] Perform data cleaning on the domain corpus, and collect the cleaned domain corpus to form a domain corpus library.
[0086] Data cleaning includes removing irrelevant characters, including removing non-Chinese characters, garbled codes, HTML tags and other noises in the text.
[0087] Step 3: Sentence normalization: Unify the number formats and date formats (such as "2023 → YYYY") in the texts of the general corpus library and the domain corpus library to obtain the general corpus library and the domain corpus library with normalized sentences.
[0088] Step 4: Duplicate removal: For each of the general corpus library and the domain corpus library with normalized sentences, remove the completely duplicate sentences in the corpus library to avoid data deviation.
[0089] So far, the general corpus library and the domain corpus library obtained after duplicate removal form a training corpus.
[0090] Step 5: Word segmentation and part-of-speech tagging: For each text in the general corpus of the training corpus and each text in the domain corpus, use a word segmentation tool to perform word segmentation, obtain all the word segments of each text in the general corpus of the training corpus and each text in the domain corpus, and retain the part-of-speech tags of each word segment. The retained form of the word segment and its part-of-speech tag is: word segment / part-of-speech. For example, "road / N", "collapse / V", etc.
[0091] Step 6: Stop word removal: For the word segments of each text in the general corpus of the training corpus and the word segments of each text in the domain corpus of the training corpus, remove the stop words (such as "de", "le", etc.) according to the preset stop word list, obtain all the word segments after stop word removal of each text in the general corpus of the training corpus and the domain corpus of the training corpus, and retain the part-of-speech of these word segments.
[0092] For each text in the general corpus of the training corpus, all the word segments after removing stop words and their part-of-speech tags are used as a general corpus sample.
[0093] For each text in the domain corpus of the training corpus, all the word segments after removing stop words and their part-of-speech tags are used as a domain corpus sample.
[0094] Collect the domain corpus samples from various domains to obtain a domain corpus training set.
[0095] Collect the general corpus samples to obtain a general corpus training set.
[0096] The domain corpus training set and the general corpus training set constitute the training set.
[0097] II. Build the network architecture of the Chinese bigram language model.
[0098] Load the base model: Load the pre-trained Chinese bigram language model of an open-source tool (such as KenLM).
[0099] Configure incremental training: Set the domain data weight to 0.7 and the general data implicit weight to 0.3 to balance domain adaptability and generality.
[0100] III. Model training 1. Incremental training: Use the KenLM tool to fine-tune the pre-trained Chinese bigram model with the data in the general corpus training set.
[0101] Enable the Kneser-Ney smoothing algorithm to automatically handle low-frequency word combinations (such as "manhole cover / N damaged / V").
[0102] 2. Parameter optimization: Optimize the model parameters through the domain corpus training set to strengthen the probability of domain word pairs (such as "road → collapse").
[0103] Based on the KenLM tool, use the data in the general corpus training set to perform incremental training on the pre-trained Chinese bigram model. During each training process, use the Kneser-Ney smoothing algorithm to optimize the model parameters until the model converges or reaches the pre-set number of training times. Then, use the domain corpus training set to optimize the parameters of the model trained with the general corpus training set to obtain the trained Chinese bigram language model.
[0104] As Figure 2 shown, the present invention provides an event title extraction system, which includes: A text acquisition module 201, configured to acquire target text data, where the target text data includes a plurality of target texts; The clustering module 202 is configured to perform topic clustering on all target texts included in the target text data by using the BERTopic topic model, so as to obtain a plurality of clusters; The screening module 203 is configured to screen the texts of each cluster according to a preset screening rule, so as to obtain the screened clusters of each type; The candidate module 204 is configured to, for each screened cluster, collect the texts within the cluster to form their respective corresponding candidate sets; The event title extraction module 205 is configured to correspondingly extract the event titles of each candidate set, so as to obtain all the event titles corresponding to all the target texts included in the target text data.
[0105] Figure 3 FIG. 10 is a schematic structural diagram of a computer terminal 300 provided by an embodiment of the present invention. The terminal 300 can be used to execute the event title extraction method provided by the embodiment of the present invention.
[0106] Among them, the terminal 300 may include: a processor 310, a memory 320, and a communication unit 330. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0107] Among them, the memory 320 can be used to store the execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the terminal 300 can execute some or all of the steps in the above method embodiments.
[0108] The processor 310 is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 320, and by invoking data stored in the memory, it executes various functions of the electronic terminal and / or processes data. The processor may be composed of an integrated circuit (IC), for example, it may be composed of a single packaged IC, or it may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 310 may only include a central processing unit (CPU). In the embodiments of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.
[0109] The communication unit 330 is used to establish a communication channel so that the storage terminal can communicate with other terminals. It receives user data sent by other terminals or sends user data to other terminals.
[0110] It should be noted that the same or similar parts in the embodiments of this specification can be referred to each other.
[0111] In addition, it should be noted that the word segmentation tools described in this specification all use the jieba word segmentation tool.
[0112] In specific implementation, the BERTopic topic clustering model is used to perform topic clustering on the target text data. In addition to obtaining all clustering clusters (each cluster corresponds to a topic, that is, corresponds to a topic), it can also obtain and output the number of documents under each clustering cluster and the relevant topic words of the topics corresponding to each clustering cluster, which is convenient for understanding the main content of each topic or topic.
[0113] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for extracting event titles, characterized in that The method includes: Step 101, obtain target text data, where the target text data includes multiple target texts; Step 102, perform topic clustering on all the target texts included in the target text data to obtain several clusters; Step 103, according to a preset screening rule, screen the texts of each cluster to obtain the screened clusters; Step 104, for each screened cluster, gather the texts within the cluster to form their respective candidate sets; Step 105, extract the event titles of each candidate set to obtain all the event titles corresponding to all the target texts included in the target text data.
2. The method according to claim 1, wherein In Step 105, the method for extracting the event titles of each candidate set all includes: Construct a word graph corresponding to all the texts in the candidate set; Calculate the weights of each node in the word graph to obtain the weights of each node in the word graph; Calculate the weights of each edge in the word graph to obtain the weights of each edge in the word graph; Based on the obtained node weights and edge weights, search for the optimal path in the word graph to obtain the target path; In the direction from the start end to the end end of the path, splice the word segments on each node in the target path in sequence to obtain the event title of the candidate set.
3. The method according to claim 2, characterized in that, Constructing a word graph corresponding to all the texts in the candidate set specifically includes: Step 110, generate a word graph containing a start node and an end node, and then add an edge with a direction from the start node to the end node between the start node and the end node; Step 120: Segment the first text in the candidate set using a word segmentation tool to obtain all the word segments of the text and the part-of-speech of each word segment; Step 130: Remove stop words from all the word segments of the text according to a preset stop word list, and obtain all the word segments and their parts of speech after stop word removal of the text Step 140: Disconnect the edge between the starting node and the ending node, and then all the word segments after stop word removal and their part-of-speech of the text shall, according to the sequence of the word segments in the text, be inserted by newly creating nodes in the direction from the starting node to the ending node at the disconnection point in turn. After each word segment is inserted, record the word frequency of the node word segment as 1 on the corresponding inserted node respectively; then add edges between adjacent nodes in the word graph, and the direction of the edge is from the previous node to the next node among adjacent nodes; Step 150: Traverse each piece of text in the remaining text in the candidate set , , where n is the total number of texts in the candidate set, and for each piece of text traversed , the following steps are performed: Step 1501: Segment the traversed text , and use a word segmentation tool to segment the text to obtain all the word segments and their part-of-speech tags of the text Step 1502, according to a preset stop word list, remove the stop words in all the word segments of the obtained text to obtain all the word segments of the text after removing the stop words and their word natures Step 1503, according to the order in which the word segments appear in the text traverse the text word segments after removing stop words one by one in the order of their appearance and for each traversed word segment, perform the following steps: Search in the entire word graph to check if there is a target node, where the target node is a node whose name and part of speech of the node word segment are the same as those of the currently traversed word segment: If it exists, determine whether merging the currently traversed word segment with the target node will generate a loop. If it is determined to be yes, create a new node in the word graph, record the currently traversed word segment and its part of speech on the newly created node, and record the word frequency of the word segment in the current node as 1 in the newly created node. If it is determined to be no, merge the currently traversed word segment with the target node and add 1 to the value of the word frequency recorded in the target node; If it does not exist, create a new node in the word graph, record the currently traversed word segment and its part of speech on the newly created node, and record the word frequency of the word segment in the current node as 1 in the newly created node; Step 1504, after Step 1503 is executed: According to the order in which the word segments appear in the text add incoming and outgoing edges to each node of the word segments after stop word removal in the word graph for the text inserted into the word graph; Step 160, the word graph obtained after Step 150 is executed is the word graph corresponding to all the texts in the constructed candidate set.
4. The method according to claim 3, wherein The method for determining whether merging the currently traversed word segment with the target node will generate a loop includes: Obtain the previous token of the currently traversed token in the token sequence, where the token sequence is a sequence obtained by sorting each token after stop word removal of the text according to the order in which the tokens appear in the text in the text ; Determine whether the previous word segment is on the path where the out-edge of the target node is located: If it is, then it is determined that a loop will be generated; If not, determine whether the previous word segment is on the path between the adjacent node in the in-edge direction of the target node in the word graph and the start node: If it is, then it is determined that a loop will be generated; If not, then it is determined that no loop will be generated.
5. The method according to claim 2, wherein The searching for the optimal path in the word graph based on the obtained node weights and edge weights to obtain the target path includes: Based on the obtained node weights and edge weights, use an improved beam search algorithm to search for the optimal path in the word graph to obtain the target path; Improved beam search algorithm, the improvement lies in improving the beam search algorithm to calculate the scores of the searched paths for ranking the searched paths to obtain the optimal path, specifically including: In the search of the beam search algorithm before improvement, each time the end node is reached in the path search, an additional score is added to the score of the current path statistically calculated by the beam search algorithm before improvement based on the node weights and edge weights of the current path, so as to adjust the final score of the current path for the beam search algorithm to calculate the ranking of the current path to search for the optimal path in the word graph; The calculation formula of the additional score is: Among them, represents the additional score of the current path, is the path length of the current path, L1 is a preset path length, is a real number coefficient, represents the word frequency recorded on the last node in the current path in the word graph.
6. The method according to claim 1, wherein In step 102, perform topic clustering on all target texts included in the target text data, and also obtain the text-topic probability of each text under each cluster; In step 103, the preset screening rules include: Delete texts in the cluster with a text content length less than the preset text length threshold L2; the value range of L2 is 5 characters to 10 characters; Delete texts in the cluster with a text-topic probability less than or equal to the preset probability threshold P; For texts that belong to multiple clusters at the same time, only keep them in the cluster with the highest text-topic probability; For each cluster, delete redundant texts within the cluster.
7. The method according to claim 2, characterized in that, The methods for calculating the node weights of each node in the word graph all include: Calculate the term frequency-inverse document frequency of each node in the word graph, and the specific calculation formula is: , , Among them, represents the th node in the word graph, , , represents the th text in the candidate set, represents the word frequency of the node, represents the inverse document frequency of the node, and N represents the total number of texts in the candidate set; Calculate the information entropy of each node in the word graph , , and the specific calculation formula is as follows: , Among them ; Calculate each node in the word graph Semantic relevance in context , and the specific calculation formula is as follows: , Among them, represents the semantic relevance of the word segmentation on the node in the candidate set text ; , represents the cosine similarity, represents the word vector of the text where the word segmentation on the node is located, is the text ; is the sentence weight, , represents the statistical number of times the sentence appears in the candidate set; Calculate the overall similarity of each node in the word graph with other nodes in the word graph , and the specific formula is as follows: , Among them, represents a normalization factor, and its value is the total number of nodes in the current word graph, represents a node and the cosine similarity between the word segmentation on the node and the word segmentation on the node within the word graph; Calculate the weight of each node in the word graph The weight of , and the calculation formula is as follows: Among them, is the weight coefficient, .
8. The method according to claim 7, wherein The calculation formula of the edge weight of each edge in the word graph is: Among them, represents the score of the edge between two nodes in the word graph . represents the probability value of the word segmentation on the node in the word graph on the pre-trained Chinese bigram language model and the word segmentation on the node in the word graph . represents the occurrence probability of the word segmentation on the node in the text within the candidate set . represents the occurrence probability of the word segmentation on the node in the text within the candidate set . represents the number of adjacent occurrences of the word segmentations on the two nodes on the edge in the text within the candidate set , .
9. An event title extraction system, characterized in that, The system includes: A text acquisition module for acquiring target text data, where the target text data includes multiple target texts; A clustering module for performing topic clustering on all target texts included in the target text data using the BERTopic topic model to obtain several clusters; A screening module for screening texts for each cluster according to the preset screening rules to obtain the screened clusters; A candidate module for, for each screened cluster, collecting the texts within the cluster to form their respective candidate sets; An event title extraction module for correspondingly extracting the event titles of each candidate set to obtain all the event titles corresponding to all the target texts included in the target text data.
10. A computer terminal, characterized in that, Including: A processor; A memory for storing the execution instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Hot topic extraction method and device, terminal equipment and storage medium
CN111460153A
Method and device for generating conference summary based on conference records and storage medium
CN112765344A
Automatically optimized and updated theme library construction method and hot event real-time updating method
CN113934910A
Hot event mining method and system
CN119293743A
Topic bridging determination using topical graphs
US20180129752A1