Event Clustering / Context Construction Method and Related Devices, Equipment, and Storage Media

By using keyword subgraphs and graph neural network models in event clustering, the problem of low event clustering accuracy is solved, and higher clustering accuracy and event context construction is achieved to meet users' clear understanding of event development.

CN114357159BActive Publication Date: 2025-07-11IFLYTEK CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111509493.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-07-11
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

In the prior art, the accuracy of event clustering is low, resulting in the text of the same event being clustered incorrectly, making it difficult to quickly and accurately summarize the events of interest to readers.

Method used

By obtaining the structural characteristics and semantic characteristics of the candidate text, a keyword subgraph is formed, and the keywords are divided into communities based on the keyword subgraph, and candidate texts of the same event are clustered in the community. The graph neural network model is used for semantic vector representation and spectral clustering, and event nodes and story trees are constructed.

Benefits of technology

It improves the accuracy and recall rate of event clustering, can better identify the text of the same event, form a clear context of event development, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114357159B_ABST
    Figure CN114357159B_ABST
Patent Text Reader

Abstract

The present application discloses an event clustering / context construction method and its related devices, equipment, and storage media. Among them, the event clustering method includes: obtaining candidate texts; extracting keywords of the candidate texts based on the structural features and semantic features of the words in the candidate texts to form a keyword subgraph for each candidate text; dividing the keywords into several communities based on the keyword subgraph, and clustering the candidate texts into the communities according to the keywords of each candidate text respectively; in each community, clustering the candidate texts describing the same event into the same event node based on the keyword subgraph. The above solution can improve the accuracy of event clustering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of information processing, and in particular, to an event clustering / context construction method and related devices, equipment, and storage media thereof. Background Art

[0002] In today's era of information explosion, a vast amount of news reports and other texts emerge every day. At the same time, these texts contain a large amount of redundant or overlapping information and may cover different topics, making it increasingly difficult for ordinary readers to digest a large amount of text. Therefore, more and more scholars are committed to researching how to quickly and accurately summarize events that readers are interested in from the vast amount of text.

[0003] Currently, there is a technology that clusters texts including the same keyword together to perform event clustering on a vast amount of text. For example, for two texts, "Company A was awarded the first license in Region B" and "Latest progress of an event of Company A", they may be simply clustered together just because they both include the keyword "Company A". In fact, the two texts describe two different events respectively, resulting in low accuracy of event clustering. In view of this, how to improve the accuracy of event clustering has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by the present application is to provide an event clustering / context construction method and related devices, equipment, and storage media thereof, which can improve the accuracy of event clustering.

[0005] To solve the above technical problem, a first aspect of the present application provides an event clustering method, including: obtaining candidate texts; respectively extracting keywords of the candidate texts based on the structural features and semantic features of the words in the candidate texts to form a keyword subgraph for each candidate text; dividing the keywords into several communities based on the keyword subgraph, and clustering the candidate texts into the communities according to the keywords of each candidate text; in each community, clustering the candidate texts describing the same event into the same event node based on the keyword subgraph.

[0006] To solve the above technical problem, a second aspect of the present application provides an event context construction method, including: after obtaining event nodes by using the event clustering method in the first aspect, the method further includes: performing structured display on the event nodes to construct several story trees.

[0007] To solve the above technical problems, a third aspect of the present application provides an event clustering device, including: a candidate text acquisition module, a formation module, a first clustering module, and a second clustering module; the candidate text acquisition module is configured to acquire candidate texts; the formation module is configured to respectively extract keywords of the candidate texts based on the structural features and semantic features of the words in the candidate texts, and form a keyword subgraph for each candidate text; the first clustering module is configured to divide the keywords into several communities based on the keyword subgraph, and cluster the candidate texts into the communities according to the keywords of each candidate text respectively; the second clustering module is configured to, in each community, cluster the candidate texts describing the same event into the same event node based on the keyword subgraph.

[0008] To solve the above technical problems, a fourth aspect of the present application provides an event context construction device, including: an event node acquisition module and a construction module; the event node acquisition module is configured to acquire event nodes; the construction module is configured to perform a structured display on the event nodes and construct several story trees.

[0009] To solve the above technical problems, a fifth aspect of the present application provides an electronic device, which includes a memory and a processor coupled to each other, and the processor is configured to execute program instructions stored in the memory to implement the event clustering method in the first aspect above, or implement the event context construction method in the second aspect above.

[0010] To solve the above technical problems, a sixth aspect of the present application provides a computer-readable storage medium, on which program instructions that can be run by a processor are stored, and the program instructions are configured to implement the event clustering method in the first aspect above, or implement the event context construction method in the second aspect above.

[0011] In the above solution, after acquiring the candidate texts, keywords of the candidate texts are respectively extracted based on the structural features and semantic features of the words in the candidate texts, a keyword subgraph is formed for each candidate text, the keywords are divided into several communities based on the keyword subgraph, and the candidate texts are clustered into the communities according to the keywords of each candidate text respectively; then, in each community, the candidate texts describing the same event are clustered into the same event node based on the keyword subgraph. Based on this, the keyword subgraph takes into account both structural features and semantic features, the interpretability of the keywords is better, and the connection between the keywords of the same topic and the same event node is closer. Therefore, community clustering and event node clustering are performed based on the keyword subgraph, which can improve the accuracy of event clustering. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a schematic flowchart of an embodiment of the event clustering method of the present application;

[0013] Figure 2 It is a schematic flowchart of another embodiment of the event clustering method of the present application;

[0014] Figure 3 It is a schematic diagram of community division in the event clustering method of the present application;

[0015] Figure 4 It is a schematic flowchart of an embodiment of the event context construction method of the present application;

[0016] Figure 5 It is a schematic framework diagram of an embodiment of the event clustering device of the present application;

[0017] Figure 6 It is a schematic framework diagram of an embodiment of the event context construction method of the present application;

[0018] Figure 7 It is a schematic framework diagram of an embodiment of the electronic device of the present application;

[0019] Figure 8 It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0020] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0021] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0022] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "plurality" in this article means two or more than two.

[0023] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the event clustering method of the present application.

[0024] Specifically, the following steps may be included:

[0025] Step S11: Obtain candidate texts.

[0026] The candidate text is the text that needs to be subjected to event clustering and can be texts in various forms such as news reports. Based on the candidate text, information such as the topic and event of the candidate text can be determined. Different candidate texts may belong to the same or similar topics or may describe the same or similar events.

[0027] The method for obtaining the candidate text is not specifically limited. For example, the original text can be directly used as the candidate text, or the candidate text can be screened from the original text according to keywords, etc. In an open embodiment, in order to perform event clustering according to the actual needs of the user, when obtaining the candidate text, the user input requirements are received; the named entity recognition model is used to perform semantic parsing on the user input requirements to obtain index keywords; the index keywords are used to screen out the text containing at least one index keyword in the text collection as the candidate text. After receiving the user input requirements, when performing semantic parsing on the user input requirements, it may include various semantic parsing implementation methods such as using the named entity recognition model to perform semantic parsing on the user input requirements or using an algorithm for text recognition to perform semantic parsing on the user input requirements. The index keywords can be entity words, trigger words, etc. of the candidate text. For example, if the candidate text is "the latest progress of a certain event", the index keywords can be "a certain", "event", and "latest progress". In order to have higher query efficiency when screening the candidate text, the texts in the text collection can be stored in the form of an inverted index. Screening the candidate text using the user input requirements and supporting user search, the event clustering or subsequent event context construction implemented based on the candidate text that fits the user input requirements is what the user really cares about, improving the user experience.

[0028] Step S12: Extract the keyword of each candidate text based on the structural feature and semantic feature of the words in the candidate text to form a keyword subgraph of each candidate text.

[0029] The candidate text includes several text components such as words, and each word includes structural features and semantic features. The structural features include word frequency, word co-occurrence information, etc. Among them, the word frequency represents the number of times a word appears in the candidate text, and the word co-occurrence information indicates the situation where several words appear together in the candidate text. Specifically, the conditional probability in the following text can be referred to. The semantic feature is the meaning of the word. For example, the meaning of the word "Beijing" is a place name. The keywords of the candidate text often represent the specific content of the candidate text to a certain extent. Therefore, when extracting keywords from the candidate text, the keywords of the candidate text are extracted based on the structural features and semantic features of the words in the candidate text, so as to consider both the structural features and semantic features of the words at the same time, increase the interpretability of the keywords for the candidate text, and each candidate text independently forms a keyword subgraph, treating the candidate text as a whole. Compared with the existing method of mixing all the keywords of all candidate texts together, when each candidate text independently forms a keyword subgraph, the relevance between the keywords and the candidate text is stronger. In other disclosed embodiments, it is also possible to extract the keywords of the candidate text only based on the structural features or semantic features of the words in the candidate text to form a keyword subgraph for each candidate text.

[0030] The extraction method of keywords is not specifically limited and can be any keyword extraction method. In an open embodiment, when extracting keywords of candidate texts based on the structural and semantic features of words in the candidate texts and forming a keyword subgraph for each candidate text, first, extract keywords of the candidate texts using a keyword extraction method based on statistical characteristics, and use a named entity recognition model to extract keywords of the candidate texts; then, use the keywords as nodes of the keyword subgraph, and connect the keywords that meet the co-occurrence condition with edges, and discard the keywords that do not meet the co-occurrence condition to form a keyword subgraph. The keyword extraction method based on statistical characteristics can be a keyword extraction method based on term frequency-inverse document frequency, a keyword extraction method based on text ranking, etc., which can mainly reflect the structural characteristics of words and some semantic characteristics, and is an unsupervised way to extract keywords. When using a named entity recognition model to extract keywords of the candidate texts, the keywords can reflect the semantic characteristics of words. Specifically, the keywords can be entity words such as person names, place names, and organization names, which is a supervised way to extract keywords. Since the keywords extracted by the keyword extraction method based on statistical characteristics are not comprehensive enough, especially the semantic characteristics are not comprehensive enough, the semantic characteristics of the keywords are combined by using a named entity recognition model, making the keywords more comprehensive, more representative, and more explanatory. In addition, for keyword extraction, a method combining supervised and unsupervised learning is designed, which utilizes the structural and semantic characteristics of keywords. And compared with the existing semi-supervised extraction method, the supervised and unsupervised combination method is used to extract keywords, and the extracted keywords have better interpretability. For the model training of the named entity recognition model, when annotating the text for supervised training, only entity words can be annotated, which reduces the annotation difficulty of keywords. And for the actual use of the named entity recognition model for keyword extraction, a strategy combining unsupervised learning and supervised learning is adopted. When the named entity recognition model processes a large number of candidate texts, it has the characteristics of simple unsupervised structure, high operation efficiency, rich semantics, and high accuracy of supervised learning, improving the keyword extraction effect and enhancing the interpretability of keywords and the recall rate of keyword extraction.

[0031] Each candidate text corresponds to a keyword subgraph. Each node of the keyword subgraph is a keyword of the candidate text, and the edge indicates that the two keywords at the endpoints meet the co-occurrence condition. The co-occurrence condition is that the number of candidate texts containing both keywords exceeds a preset number of texts, and the quotient of the number of candidate texts containing both keywords and the sum of the number of candidate texts containing each of the two keywords respectively exceeds a preset conditional probability. The quotient of the number of candidate texts containing both keywords and the sum of the number of candidate texts containing each of the two keywords respectively is also called the conditional probability, and its calculation formula is as follows:

[0032]

[0033] where Pr{w i |w j} represents the conditional probability, DF i,j represents the number of texts that contain both the keyword w i and w j among all candidate texts, DF j is the number of texts that contain w j among all candidate texts, DF i is the number of texts that contain w i among all candidate texts.

[0034] Since keywords among candidate texts related in content may have the characteristic of a high co-occurrence frequency, when constructing the keyword subgraph of each candidate text, keywords that do not meet the co-occurrence condition will not form the edges of the keyword subgraph, that is, keywords that do not meet the co-occurrence condition are discarded, which can improve the quality of keywords, avoid the influence of this part of keywords that are not representative of candidate texts on subsequent clustering, and at the same time ensure that the keywords in the keyword subgraph are closely related and can better reflect their corresponding candidate texts.

[0035] Step S13: Divide the keywords into several communities based on the keyword subgraph, and cluster the candidate texts into the communities according to the keywords of each candidate text respectively.

[0036] Since keywords can reflect the content and topics of candidate texts, if there are many identical keywords between two candidate texts, their topics are very likely to be the same or similar. After classifying identical, similar, or highly relevant keywords into the same community, at this time, the keywords in the same community represent the same topic. Then, according to the keywords of the candidate texts, the candidate texts are clustered into the community. In this way, candidate texts with similar content and the same topic will be clustered into the same community, while candidate texts with greatly different content and different topics will be clustered into different communities. To achieve text clustering of the same topic based on keywords, after obtaining the keyword subgraph, the keywords are divided into several communities on the basis of the keyword subgraph. Compared with directly mixing the keywords of candidate texts to divide communities, in this solution, due to the small distance and close connection between the keywords of the keyword subgraph of the same candidate text, the community division is more in line with the similarity and relevance between the keywords. Keywords in the same community often better represent the same topic. Therefore, when dividing the keywords into several communities based on the keyword subgraph, because the distance between the keywords of the keyword subgraph is small and they are closely connected, after dividing the community on this basis, the keywords in the community are also closely connected and can better represent the same topic. Then, according to the keywords of each candidate text, the candidate texts are respectively clustered into the community, so that candidate texts with similar content and the same topic will be clustered into the same community, while candidate texts with greatly different content and different topics will be clustered into different communities. Through the above method, candidate texts belonging to the same topic are clustered into the same community, while candidate texts of different topics belong to different communities respectively. Therefore, using the keyword subgraph to complete the division of keywords in each community further improves the screening accuracy of candidate texts related to the user input requirements.

[0037] Step S14: In each community, based on the keyword subgraph, cluster candidate texts describing the same event into the same event node.

[0038] Under the same topic, corresponding to different development stages, there may be several events. Therefore, candidate texts with the same topic, the same or similar content, and describing the same event are clustered into the same event node, while candidate texts with the same topic but different content are clustered into different event nodes in the same community.

[0039] After candidate texts belonging to the same topic are clustered into the same community, the candidate texts in the same community are already relatively similar. On this basis, since the keyword subgraph corresponds one-to-one with the candidate texts, by comparing different keyword subgraphs, in each community, based on the keyword subgraph, cluster candidate texts describing the same event into the same event node to achieve event clustering.

[0040] In the above solution, after obtaining the candidate texts, keywords of the candidate texts are extracted respectively based on the structural features and semantic features of the words in the candidate texts, a keyword sub-graph of each candidate text is formed, the keywords are divided into several communities based on the keyword sub-graph, and the candidate texts are clustered into the communities respectively according to the keywords of each candidate text; then in each community, the candidate texts describing the same event are clustered into the same event node based on the keyword sub-graph. Based on this, the keyword sub-graph takes into account both the structural features and semantic features, the interpretability of the keywords is better, and the connection between the keywords of the same topic and the same event node is closer. Therefore, performing community clustering and event node clustering based on the keyword sub-graph can improve the accuracy of event clustering.

[0041] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another embodiment of the event clustering method of the present application.

[0042] Specifically, it may include the following steps:

[0043] Step S21: Obtain candidate texts;

[0044] Step S22: Extract keywords of the candidate texts respectively based on the structural features and semantic features of the words in the candidate texts, and form a keyword sub-graph of each candidate text;

[0045] For the relevant descriptions of Step S21 and Step S22, please refer to Step S11 and Step S12 in the above Figure 1 embodiment, which will not be elaborated here.

[0046] Step S23: Merge the keyword sub-graphs containing the same keywords to form a keyword cluster.

[0047] The keyword sub-graph is formed independently for each candidate text. In order to divide the related keywords into the same community during community division, the keyword sub-graphs containing the same keywords are merged to form a keyword cluster. For example, the keyword sub-graph of one candidate text includes nodes corresponding to three keywords A, B, and C, and the keyword sub-graph of another candidate text includes nodes corresponding to three keywords A, D, and E. Since there is the same node A, after merging, the keyword cluster includes nodes corresponding to five keywords A, B, C, D, and E. If there are no same keywords between other candidate texts and these two candidate texts, then other candidate texts and these two candidate texts do not belong to the same topic. In short, merging the keyword sub-graphs containing the same keywords makes the connection between the keywords close, while the difference is reflected between the keyword clusters.

[0048] Step S24: Use the community discovery method to divide the keyword cluster to form several communities.

[0049] The keyword clusters are divided using community detection methods to form several communities, such that each community includes several keywords, and the keywords in the same community are often associated with the same topic. The community detection method can be a modularity-based community detection method, or an overlapping community detection method based on hierarchical clustering, etc. Since there may be the same keywords in the keyword subgraphs corresponding to the candidate texts of different topics, if the modularity-based community detection method is used, the same keywords will be classified into one community while the other community does not have the same keyword. Therefore, the keyword clusters can be divided using the overlapping community detection method based on hierarchical clustering to form several communities, such that there may be the same keywords between different communities, that is, community overlap. The same keywords existing between different communities are often words that are more likely to appear in different candidate texts and have less impact on the topic. Therefore, this community division method is more suitable for the actual text scenario. For example, the keyword "Region B" may appear in many topics. Therefore, by placing the keyword "Region B" in different communities, it is avoided that two candidate texts of different topics are placed in the same community because of the keyword "Region B". Therefore, when performing community detection, the overlapping community detection method can be adopted, and the same keyword can appear in different communities, which is more suitable for the actual scenario and can improve the recall rate of the text.

[0050] Step S25: Calculate the relevance between the candidate text and the community, and add the candidate text corresponding to the maximum value of the relevance to the corresponding community respectively.

[0051] Calculate the relevance between the candidate text and each community, find the most relevant community, and add the candidate text to the corresponding community. The method for calculating the relevance between the candidate text and the community can be various measurement methods. In an open embodiment, the coincidence degree of the keyword set of the candidate text and the keyword set of each community is statistically calculated, and the candidate text is assigned to the community with the highest coincidence degree. To further improve the accuracy of topic-based clustering, first discard the candidate texts that do not exceed the preset coincidence degree, and then assign the remaining candidate texts to the community with the highest coincidence degree. Among them, the coincidence degree can be obtained by calculating the similarity coefficient.

[0052] Therefore, to divide the keywords into several communities based on the keyword subgraph and cluster the candidate texts into the communities according to the keywords of each candidate text respectively, first fuse the keyword subgraphs containing the same keywords to form keyword clusters; then use the community detection method to divide the keyword clusters to form several communities; finally, calculate the relevance between the candidate text and the community, and add the candidate text corresponding to the maximum value of the relevance to the corresponding community respectively. Through the above method, the candidate texts of the same topic are clustered into the same community, while the candidate texts of different topics belong to different communities respectively, realizing topic-based text clustering.

[0053] After completing topic-based text clustering and clustering candidate texts into different communities, in each community, based on the keyword subgraph, candidate texts describing the same event are clustered into the same event node. Specifically, refer to the following steps S26 and S27 to implement event clustering and cluster candidate texts onto event nodes.

[0054] In an actual scenario, the candidate text can be news text. When the user input requirement is to obtain candidate texts related to the latest progress of a certain event of Company A, for the candidate texts related to the first text of "the latest progress of a certain event of Company A", the recall rate is relatively high. However, at the same time, other candidate texts that are not related to the latest progress of a certain event of Company A will also be introduced. For example, according to Company A, the second text of "Company A obtains the first license in Region B" is screened out. As Figure 3 shown, Figure 3 This is a schematic diagram of community division in the event clustering method of this application. For example, the keyword subgraph of the first text of "the latest progress of a certain event of Company A" includes nodes corresponding to the keywords "Company A", "a certain", "event", and "latest progress", while the keyword subgraph of the second text of "Company A obtains the first license in Region B" includes nodes corresponding to the keywords "Company A", "obtains", "Region B", "first", and "license". Then, the node "Company A" is fused to form a keyword cluster. Although the first text and the second text have the same keyword "Company A", through the overlapping community discovery method based on hierarchical clustering, the keyword "Company A" will be assigned to two communities. Since different communities represent different topics, it is thus avoided to cluster the two texts of "Company A obtains the first license in Region B" and "the latest progress of a certain event of Company A" into the same community. It can be understood that since there are only two candidate texts in this actual scenario, and "Company A" appears in both candidate texts but has a relatively small impact on topic judgment. Therefore, through the overlapping community discovery method based on hierarchical clustering, the keyword "Company A" will be assigned to two communities. The keyword subgraphs of the first text and the second text and the keywords in the community remain unchanged. The actual scenario may be different due to the addition of other candidate texts and other reasons. Here, it only shows the processing of keywords in the overlapping community discovery method.

[0055] Step S26: Input the keyword subgraphs into the text modeling model respectively to obtain the semantic vector representation of each candidate text.

[0056] Since text similarity can be obtained by calculating the distance between semantic vector representations. For example, the text similarity between semantic vector representations with a closer distance may be higher. Therefore, the keyword subgraphs can be respectively input into the text modeling model to obtain the semantic vector representations of each candidate text, which is convenient for subsequent judgment on whether different candidate texts are clustered into the same event node. The text modeling model can be any model structure that can obtain semantic vector representations, and no specific limitation is made here. Additionally, when using the keyword subgraphs of candidate texts to model the candidate texts, the semantic encoding is not restricted by the text length.

[0057] Considering that although the candidate texts describing the same event may have slightly different expressed contents, their overall writing style and main idea are highly similar, while the similarity of candidate texts describing different events is slightly weaker. Therefore, by modeling the candidate texts, the semantic vector representations of each candidate text can be obtained, and the text similarity can be obtained by calculating the distance between the vectors. If the similarity meets certain conditions, the two candidate texts belong to the same event. However, the lengths of the candidate texts vary. Traditional deep learning models, such as long short-term memory network models, BERT models, etc., are difficult to model candidate texts with longer lengths. Therefore, in an open embodiment, the text modeling model is a graph neural network model, which can model candidate texts with longer lengths. The graph neural network model can include a graph autoencoder. When the keyword subgraphs are respectively input into the text modeling model to obtain the semantic vector representations of each candidate text, the graph autoencoder is used to encode the keyword subgraphs to obtain the semantic vector representations. The graph autoencoder is an end-to-end encoding-decoding structure. First, the encoding X of the keyword subgraph is obtained through a graph convolutional network. After the input feature X passes through the encoder, a new feature representation Z is obtained. Z is restored to X^ through the decoder, and the loss function is calculated. The loss function is optimized by gradient descent to make X and X^ as close as possible. Finally, Z is used as the new semantic representation vector of the input X^.

[0058] Step S27: Cluster several candidate texts whose semantic vector representations meet the text similarity conditions into the same event node.

[0059] After obtaining the semantic vector representations of each candidate text, several candidate texts whose semantic vector representations meet the text similarity conditions are clustered into the same event node to achieve event clustering.

[0060] In an open embodiment, when clustering a number of candidate texts whose semantic vector representations meet the text similarity condition into the same event node, calculate the similarity of the semantic vector representations, cluster the candidate texts with similarity greater than a preset threshold into the same event node, or use the spectral clustering method to cluster the candidate texts represented by the semantic vector representations. The specific steps of the calculation method of the similarity of the semantic vector representations and the spectral clustering method can refer to the prior art and will not be elaborated here.

[0061] Event clustering requires clustering the candidate texts that describe the same event among the candidate texts of each community into the same event node, that is, all candidate texts within the same event node tell the same event, and the contents of these candidate texts are slightly different. Since the number of event nodes is uncertain during event clustering and the number of candidate texts within each event node is different, traditional unsupervised clustering methods such as the k-means clustering algorithm and the hierarchical clustering algorithm are no longer applicable, and the number of candidate texts under each community is limited. Therefore, use a text modeling model to model the keyword subgraphs of the candidate texts, and adopt event clustering dominated by supervised learning, which can take into account the running time consumption of the algorithm while ensuring accuracy. In addition, it is proposed to model the keyword subgraphs of the candidate texts, represent each candidate text as a keyword subgraph, and combine the text modeling model to obtain a better text semantic representation for downstream tasks, which has better accuracy and efficiency compared with traditional supervised modeling methods.

[0062] Through the above method, first obtain the candidate texts, then extract the keywords of the candidate texts based on the structural features and semantic features of the words in the candidate texts to form the keyword subgraphs of each candidate text, and finally fuse the keyword subgraphs containing the same keywords to form keyword clusters, thus realizing keyword extraction and the construction of keyword clusters; use the community discovery method to divide the keyword clusters to form several communities, and then calculate the relevance between the candidate texts and the communities, and add the candidate texts corresponding to the maximum value of the relevance to the corresponding communities respectively, thus realizing text clustering based on topic granularity; input the keyword subgraphs into the text modeling model respectively to obtain the semantic vector representations of each candidate text, and cluster the candidate texts whose semantic vector representations meet the text similarity condition into the same event node, thus realizing text clustering based on event granularity.

[0063] Since then, in order to realize the text clustering of candidate texts, text clustering based on topic granularity can be realized, so that candidate texts of the same topic are clustered together, and under the same topic, text clustering based on event granularity can also be realized, so that candidate texts describing the same event are clustered together to form event nodes.

[0064] As an important branch in the field of information extraction, the event context extraction technology aims to automatically extract and summarize hot events from a vast amount of text, and track and display the context of events as they develop over time in a structured manner. It enables readers to have a clear understanding of the event development without referring to a large number of relevant texts and accurately grasp the key information during the event development process. After clustering candidate texts into events, it is expected to summarize the development context of the events, reveal the development process of the events and the different stages experienced during the event development, saving the time cost of collecting relevant events. In addition, in this way, it can assist relevant personnel in completing event supervision and analysis, tracking hot topics, etc. Therefore, this application also provides an event context construction method. After obtaining event nodes using the event clustering method of any of the above embodiments, the event nodes are structurally displayed to construct several story trees. For details, please refer to Figure 4 , Figure 4 which is a schematic flowchart of an embodiment of the event context construction method of this application. Specifically, it may include the following steps:

[0065] Step S41: Obtain event nodes.

[0066] After obtaining event nodes using the event clustering method of any of the above embodiments, each event node includes several candidate texts and their keywords.

[0067] Step S42: Determine the event time and event summary of the event node to form a brief text of the event node.

[0068] Since an event node may be composed of multiple candidate texts, and the candidate texts include a lot of content, for the convenience of reading, a brief text of the event node can be determined before forming a story tree using the event node. For step S42, other publicly disclosed embodiments can also be selected or discarded according to the situation. The method for determining the event time and event summary of the event node can be any method in the field of text processing. For example, in a publicly disclosed embodiment, when determining the event time of the event node, the event time of the candidate text with the earliest time in the event node is used as the event time of the event node, or after removing the event times of the candidate text with the earliest time and the candidate text with the latest time in the event node, the average value of the event times of the remaining candidate texts is used as the event time of the event node; when determining the event summary of the event node, the event summary of the candidate text with the earliest time in the event node is used as the event summary of the event node, or all candidate texts of the event node are input into a summary extraction model, and the output of the summary extraction model is used as the event summary of the event node. The summary extraction model can extract a generative summary to improve the convenience of determining the event summary.

[0069] Step S43: Take the event nodes as the event nodes to be assigned in sequence, and perform story clustering on the event nodes to be assigned, so as to cluster the event nodes under each community into different story trees.

[0070] For multiple event nodes under each community, the story tree under the community can be obtained by an online update method. Take the event nodes as the event nodes to be assigned in sequence. The first event node to be assigned is used as the node of the first story tree. Then continue to take the event nodes as the event nodes to be assigned, and judge whether the event node to be assigned matches the story tree. If it matches, cluster the event node to be assigned into the story tree; if it does not match, judge whether the event node to be assigned matches other story trees until the event node to be assigned is clustered into a story tree, or it is determined that there is no story tree that matches the event node to be assigned, then the event node to be assigned can be used as the node of a new story tree.

[0071] In an open embodiment, when taking the event nodes as the event nodes to be assigned in sequence and performing story clustering on the event nodes to be assigned so as to cluster the event nodes under each community into different story trees, first take the event nodes as the event nodes to be assigned in sequence based on the chronological order of the event times of the event nodes, and then obtain whether the compatibility between the keyword set of the event node to be assigned and the keyword set of the story tree exceeds a preset compatibility value. If so, judge whether the number of identical words included in the title of at least one candidate text in the event node to be assigned and the title of at least one candidate text in the story tree exceeds a preset number. At this time, if so and the story tree is unique, assign the event node to be assigned to the story tree. If so but the story tree is not unique, cluster the event node to be assigned into the story tree with the maximum compatibility and the maximum number of identical words. If the compatibility does not exceed the preset compatibility value and / or the number of identical words included in the title of at least one candidate text in the event node to be assigned and the title of at least one candidate text in the story tree does not exceed the preset number, then use the event node to be assigned as the node of a new story tree. The above preset number and preset compatibility value can be custom-set according to actual needs and will not be specifically limited here. By comparing the compatibility between the keyword set of the event node to be assigned and the keyword set of the story tree with the preset compatibility value, and comparing the number of identical words with the preset number, the story classification of the event node to be assigned is performed, and the story tree under the community can be obtained by an online update method.

[0072] Each candidate text has its own unique keyword set, and each story tree also has its corresponding keyword set. The compatibility in performing story clustering on the event nodes to be assigned can be to calculate the similarity coefficient between the keyword set of the candidate text and the keyword set of the story tree. The similarity coefficient is:

[0073]

[0074] Among them, the keyword set of the candidate text is C ε , and the keyword set of the story tree is Therefore, find the intersection and union of the keyword sets of the candidate text and the story tree, and use the quotient of the intersection and the union as the compatibility.

[0075] Step S44: In each story tree, sequentially use the event node as the event node to be assigned. Use the event node to be assigned as the root node of the story tree, or fuse, expand, or insert the event node to be assigned into the story tree.

[0076] Cluster the event nodes into different story trees, and then arrange the positions of the corresponding event nodes under the story tree. In each story tree, sequentially use the event node as the event node to be assigned. If it is the first event node to be assigned, use the event node to be assigned as the first node of the story tree. When continuing to use the event node as the event node to be assigned, fuse, expand, or insert the event node to be assigned into the story tree.

[0077] In an exemplary embodiment, in each story tree, when sequentially using the event node as the event node to be assigned and fusing, expanding, or inserting the event node to be assigned into the story tree, sequentially use the event nodes as the event nodes to be assigned based on the chronological order of the event times of the event nodes; respectively splice the candidate text of the event node to be assigned and the candidate text of the event node in the story tree to form a to-be-assigned splicing node and a candidate splicing node; use a similarity discrimination model to determine whether the to-be-assigned splicing node and the candidate splicing node meet the fusion condition; if so, add the candidate text of the event node to be assigned and its keywords to the event node corresponding to the candidate splicing node; if not, calculate the connection strength between the event node to be assigned and the event nodes in the story tree, and use the event node with the maximum connection strength as the parent node of the event node to be assigned to connect the event node to be assigned after the parent node. The similarity discrimination model can be implemented by a text modeling model in event clustering, or by other models for determining text similarity, or by an Internet application service engine, etc.

[0078] If the event node to be assigned does not meet the fusion condition, the position of the event node to be assigned is determined by calculating the connection strength. When calculating the connection strength between the event node to be assigned and the event nodes in the story tree, calculate the event compatibility, event consistency, and time penalty term between the event node to be assigned and the event nodes in the story tree respectively, and then use the product of the event compatibility, event consistency, and time penalty term as the connection strength. Specifically, the connection strength calculation formula is as follows:

[0079]

[0080] Among them, the event node to be assigned is The event nodes in the story tree are Event compatibility is The event consistency is The time penalty term is

[0081] Among them, event compatibility is The calculation formula is:

[0082]

[0083] in, Indicates the event node to be assigned The result of splicing all candidate texts in , that is, the splicing node to be assigned, d cj Represents an event node in the story tree The result of splicing all candidate texts in is the candidate splicing node. The vector representation of the splicing node to be assigned, TF(d cj ) represents the vector representation of the candidate splicing node, The modulus value of the vector representing the node to be assigned, The modulus value of the vector representation representing the candidate splice node.

[0084] Therefore, when calculating the event compatibility between the event node to be assigned and the event nodes in the story tree, respectively obtain the vector representation of the splicing node to be assigned and the candidate splicing node, and the modulus value of the vector representation; calculate the first product of the vector representation of the splicing node to be assigned and the candidate splicing node, and the second product of the modulus value of the vector representation; and take the ratio of the first product to the second product as the event compatibility.

[0085] Event Node The story line is defined as the story tree, from the root node S to the event node The path is recorded as Similarly, expand the event node Event nodes to be assigned on Its storyline path is represented as For a story line L, there are several event nodes express, It can also be represented as the root node S, and the event consistency is The calculation formula is:

[0086]

[0087] When calculating the event consistency between the event node to be assigned and the event node in the story tree, determine the story line between the root node of the story tree and the event node to be assigned, that is, According to the chronological order of the story line, obtain the sum of the compatibilities of adjacent event nodes among all event nodes from the root node to the event node to be assigned; determine the reciprocal of the number of event nodes from the root node to the event node to be assigned, that is Take the product of the reciprocal and the sum as the event consistency, that is

[0088] For each event node to be assigned The earliest event time can be taken as the event time of the event node to be assigned, and the event is represented in the form of a time stamp record, denoted as The calculation method of the time penalty term between two event nodes is as follows:

[0089]

[0090] Wherein, is an event node of the story line, is the event node to be assigned.

[0091] When calculating the time penalty term between the event node to be assigned and the event node in the story tree, determine the first time stamp of the event node to be assigned and the second time stamp of the event node in the story tree; judge whether the difference between the second time stamp and the first time stamp is less than zero; if so, take the difference as the exponent of the exponential function with base E, and take the obtained value as the time penalty term; if not, the time penalty term is zero.

[0092] Therefore, calculate the event compatibility, event consistency and time penalty term between the event node to be assigned and the event node in the story tree respectively, and then take the product of the event compatibility, event consistency and time penalty term as the connection strength. Select the event node with the largest connection strength as the parent node of the event node to be assigned to complete the update of the event context.

[0093] Step S45: Display the story tree and the brief text of its event nodes.

[0094] After fusing, expanding or inserting all event nodes into the story tree, several story trees are formed. The story tree can be displayed, and at the same time, the brief text of each event node in the story tree is displayed, so as to structurally display the context of event development.

[0095] Since candidate texts may be continuously added, in order to update the story tree, new event nodes can be obtained by reusing the event clustering method of any of the above embodiments, and the event time and event summary of the event nodes can be determined again, and the brief text of the event nodes and their subsequent steps can be formed to update the story tree, so as to realize the real-time tracking of events.

[0096] In an open embodiment, in order to better present the story tree, the definitions or concepts of the constituent elements in the story tree can be predefined and represented by characters. After obtaining the candidate text, the candidate text can be disassembled by elements or processed. The elements are as follows: The candidate text d includes a title, a release time, and a body. The keyword k is a specific concept or behavior extracted from the candidate text that can highly summarize the central content of the candidate text. The event node e refers to a specific event involving a specific time and place and relevant parties (such as people, organizations, institutions, etc.), which is represented by a quadruple <T e ,A e ,K e ,Docs set >, where T e represents the event time, A e represents the event summary, K e represents the keyword set of the event node, and Docs set = {d1, d2,... d n} represents the candidate text set of the event node. The story tree branch Branch consists of a seed event node and the related event nodes. Each Branch = <E, L, K branch ,T branch > reflects a sub - development stage of the event, where E = {e1, e2,..., e |E|} represents the event node set of the branch, L i,j = <e i ,e j > represents a directed edge from the event node e i to the event node e j , representing that the events have a chronological order; K branch represents the keyword set of the branch event; T branch represents the branch time, which can take the earliest time when the events in the branch occur. The story tree S is a structured display form of the event context, reflecting the evolution process of the current event, and contains quadruple information <T s ,A s ,K s ,Event set >, where Event set represents the event node set included in the story, T s represents the event segment when the story occurs, A s represents the brief text of the story, and K s represents the keyword set of the story tree. Each story tree S = {branch1, branch2,..., branch n} It is formed by connecting multiple branches in the chronological order of the events at the branch root nodes. Therefore, the task background of the event context is clearly described, and clear definitions and names are given to each part in the context structure of the story tree.

[0097] Through the above method, the event nodes can be displayed in a structured manner. Compared with simply connecting all event nodes in a line according to the event time in the event context of the event line structure, the story context of the story line structure in this application contains more sub-stage information of event development. The event context of the story line structure is clear, which can better improve the user experience. Moreover, in the existing method for representing event context based on the timeline, there are many hyperparameters that need to be manually preset in the state descriptions and evaluation functions in the reinforcement learning stage, and the model effect is significantly affected by the hyperparameter settings. However, the structure of the story tree in this application can more clearly reflect the evolution process of each event stage in the event context. Compared with reinforcement learning, it requires fewer hyperparameters to be manually designed and is less sensitive to hyperparameters. It can be understood that the steps in any embodiment of the above event context construction method can also be directly applied after clustering the candidate texts describing the same event into the same event node based on the keyword subgraph in any embodiment of the above event clustering method. That is, in some publicly disclosed embodiments of the event clustering method, after clustering the candidate texts describing the same event into the same event node based on the keyword subgraph, the event nodes can be structuredly displayed, and several story trees can be constructed, so as to construct the event context after event clustering.

[0098] Please refer to Figure 5 , Figure 5 which is a schematic framework diagram of an embodiment of the event clustering device 60 of this application.

[0099] The event clustering device 60 includes: a candidate text acquisition module 61, a formation module 62, a first clustering module 63, and a second clustering module 64. The candidate text acquisition module 61 is used to acquire candidate texts; the formation module 62 is used to extract the keywords of each candidate text based on the structural features and semantic features of the words in the candidate texts to form a keyword subgraph for each candidate text; the first clustering module 63 is used to divide the keywords into several communities based on the keyword subgraph and cluster the candidate texts into the communities according to the keywords of each candidate text; the second clustering module 64 is used to cluster the candidate texts describing the same event into the same event node in each community based on the keyword subgraph.

[0100] In some publicly disclosed embodiments, the second clustering module 64 is used to cluster the candidate texts describing the same event into the same event node based on the keyword subgraph, and is further used to input the keyword subgraphs into the text modeling model respectively to obtain the semantic vector representations of each candidate text; and cluster several candidate texts whose semantic vector representations meet the text similarity conditions into the same event node.

[0101] Therefore, by using a text modeling model to model the keyword subgraph of the candidate text and adopting event clustering dominated by supervised learning, the running time consumption of the algorithm can be taken into account while ensuring the accuracy. In addition, it is proposed to use the keyword subgraph of the candidate text for modeling, represent each candidate text as a keyword subgraph, and combine it with the text modeling model to obtain a better text semantic representation for downstream tasks, which has better accuracy and efficiency compared with the traditional supervised modeling method.

[0102] In some disclosed embodiments, the text modeling model is a graph neural network model. The graph neural network model includes a graph autoencoder. When the second clustering module 64 inputs the keyword subgraphs into the text modeling model respectively to obtain the semantic vector representation of each candidate text, it is also used to encode the keyword subgraphs by using the graph autoencoder to obtain the semantic vector representation. When the second clustering module 64 clusters several candidate texts whose semantic vector representations meet the text similarity condition into the same event node, it is also used to calculate the similarity of the semantic vector representations, cluster the candidate texts with a similarity greater than the preset threshold into the same event node, or use the spectral clustering method to cluster the candidate texts represented by the semantic vector representations.

[0103] Therefore, the text modeling model is a graph neural network model, which can model candidate texts with longer word counts, that is, use the keyword subgraph of the candidate text to model the candidate text, and the semantic encoding of the graph autoencoder is not limited by the text length.

[0104] In some disclosed embodiments, when the forming module 62 is used to extract the keywords of the candidate text based on the structural features and semantic features of the words in the candidate text and form the keyword subgraph of each candidate text, it is also used to extract the keywords of the candidate text by using a keyword extraction method based on statistical characteristics and use a named entity recognition model to extract the keywords of the candidate text; use the keywords as the nodes of the keyword subgraph, connect the keywords that meet the co-occurrence condition with edges, and discard the keywords that do not meet the co-occurrence condition to form the keyword subgraph.

[0105] Therefore, for keyword extraction, a method combining supervised and unsupervised methods is designed by using a keyword extraction method based on statistical characteristics and using a named entity recognition model to extract the keywords of the candidate text. The structural features and semantic features of the keywords are utilized. And compared with the existing semi-supervised extraction method, the method of combining supervised and unsupervised methods is used to extract the keywords, and the extracted keywords have better interpretability and improve the recall rate of keyword extraction.

[0106] In some disclosed embodiments, the co-occurrence condition is that the number of candidate texts containing two keywords simultaneously exceeds a preset number of texts, and the quotient of the number of candidate texts containing two keywords simultaneously and the sum of the number of candidate texts containing each of the two keywords respectively exceeds a preset conditional probability.

[0107] Therefore, by judging the above co-occurrence condition, the keywords that meet the co-occurrence condition are retained, and the keywords that do not meet the co-occurrence condition are discarded, improving the quality of the keywords and avoiding the influence of these keywords that do not represent candidate texts on subsequent clustering.

[0108] In some disclosed embodiments, when the first clustering module 63 is used to divide keywords into several communities based on the keyword subgraph and cluster candidate texts into communities according to the keywords of each candidate text respectively, it is also used to fuse the keyword subgraphs containing the same keywords to form keyword clusters; use the community discovery method to divide the keyword clusters to form several communities; calculate the relevance between the candidate texts and the communities, and add the candidate texts corresponding to the maximum value of the relevance to the corresponding communities respectively.

[0109] Therefore, fusing the keyword subgraphs containing the same keywords makes the connection between keywords close, while the differences are reflected between keyword clusters; using the community discovery method to divide the keyword clusters to form several communities, so that each community includes several keywords, and the keywords in the same community are often associated with the same topic; calculating the relevance between the candidate texts and the communities, and adding the candidate texts corresponding to the maximum value of the relevance to the corresponding communities respectively, clustering the candidate texts of the same topic into the same community, and the candidate texts of different topics belong to different communities respectively, realizing text clustering based on topics.

[0110] In some disclosed embodiments, the community discovery method includes an overlapping community discovery method based on hierarchical clustering.

[0111] Therefore, using the overlapping community discovery method based on hierarchical clustering to divide the keyword clusters to form several communities, so that there may be nodes corresponding to the same keywords between different communities. Since the same keywords existing between different communities are often words that are relatively easy to appear in different candidate texts and have little influence on the topic, the overlapping community discovery method based on hierarchical clustering is more in line with the actual text scenario and improves the recall rate of relevant candidate texts.

[0112] In some disclosed embodiments, when the candidate text acquisition module 61 is used to acquire candidate texts, it is also used to receive user input requirements; use a named entity recognition model to perform semantic parsing on the user input requirements to obtain index keywords; use the index keywords to screen out texts containing at least one index keyword in the text collection as candidate texts.

[0113] Therefore, screening candidate texts based on user input requirements, supporting user search, and implementing event clustering or subsequent event context construction based on candidate texts that match the user input requirements are what users truly care about, which enhances the user experience.

[0114] Please refer to Figure 6 , Figure 6 which is a schematic framework diagram of an embodiment of the event context construction method 70 of the present application.

[0115] The event context construction method 70 includes: an event node acquisition module 71 and a construction module 72. The event node acquisition module 71 is used to acquire event nodes; the construction module 72 is used to structurally display the event nodes and construct several story trees.

[0116] In some disclosed embodiments, when the construction module 72 is used to structurally display the event nodes and construct several story trees, it is further used to determine the event time and event summary of the event nodes to form a brief text of the event nodes; sequentially use the event nodes as the to-be-allocated event nodes, and perform story clustering on the to-be-allocated event nodes to cluster the event nodes under each community into different story trees; in each story tree, sequentially use the event nodes as the to-be-allocated event nodes, use the to-be-allocated event nodes as the root nodes of the story tree, or fuse, expand, or insert the to-be-allocated event nodes into the story tree; display the story tree and the brief text of its event nodes, or re-acquire new event nodes obtained by using the event clustering method embodiment of any of the above, and re-execute determining the event time and event summary of the event nodes to form the brief text of the event nodes and its subsequent steps to update the story tree.

[0117] Therefore, compared with simply connecting all event nodes into a line according to the event time in the event context of the event line structure, the story context of the story line structure of the present application contains more sub-stage information of event development. The event context of the story line structure is clear, which can better enhance the user experience. Moreover, when training a decision-making model with the existing event context representation method based on the timeline, there are many hyperparameters that need to be manually preset in the state description and evaluation function in the reinforcement learning stage, and the model effect is significantly affected by the hyperparameter settings. However, the structure of the story tree of the present application can more clearly reflect the evolution process of each event stage in the event context, and compared with reinforcement learning, it requires fewer hyperparameters to be manually involved in the design and is less sensitive to hyperparameters.

[0118] In some disclosed embodiments, the construction module 72 is configured to sequentially use the event nodes as the event nodes to be assigned, and when clustering the event nodes to be assigned for story clustering so as to cluster the event nodes under each community into different story trees, it is further configured to sequentially use the event nodes as the event nodes to be assigned based on the chronological order of the event times of the event nodes; obtain whether the compatibility between the keyword set of the event nodes to be assigned and the keyword set of the story tree exceeds a preset compatibility value; if so, determine whether the number of identical words contained in the title of at least one candidate text in the event nodes to be assigned and the title of at least one candidate text in the story tree exceeds a preset number; if so and the story tree is unique, assign the event nodes to be assigned to the story tree, and if so but the story tree is not unique, cluster the event nodes to be assigned into the story tree with the maximum compatibility and the maximum number of identical words.

[0119] Therefore, by comparing the compatibility between the keyword set of the event nodes to be assigned and the keyword set of the story tree with the preset compatibility value, and comparing the number of identical words with the preset number, story classification is performed on the event nodes to be assigned, and the story tree under the community can be obtained in an online update manner.

[0120] In some disclosed embodiments, the construction module 72 is configured to sequentially use the event nodes as the event nodes to be assigned in each story tree, and when fusing, expanding or inserting the event nodes to be assigned into the story tree, it is further configured to sequentially use the event nodes as the event nodes to be assigned based on the chronological order of the event times of the event nodes; respectively splice the candidate text of the event nodes to be assigned and the candidate text of the event nodes in the story tree to form a to-be-assigned splicing node and a candidate splicing node; use a similarity discrimination model to determine whether the to-be-assigned splicing node and the candidate splicing node meet the fusion condition; if so, add the candidate text of the event nodes to be assigned and its keywords to the event nodes corresponding to the candidate splicing node; if not, calculate the connection strength between the event nodes to be assigned and the event nodes in the story tree, and use the event node with the maximum connection strength as the parent node of the event nodes to be assigned to connect the event nodes to be assigned after the parent node.

[0121] Therefore, the similarity discrimination model can be used to determine whether the to-be-assigned splicing node and the candidate splicing node meet the fusion condition to achieve the fusion of the event nodes; and the position of the event nodes in the story tree can be determined by calculating the connection strength between the event nodes to be assigned and the event nodes in the story tree.

[0122] In some disclosed embodiments, when the construction module 72 is used to calculate the connection strength between the event node to be allocated and the event nodes in the story tree, it is also used to calculate the event compatibility, event consistency, and time penalty term between the event node to be allocated and the event nodes in the story tree respectively; and use the product of the event compatibility, event consistency, and time penalty term as the connection strength. Among them, calculating the event compatibility between the event node to be allocated and the event nodes in the story tree includes: respectively obtaining the vector representations of the splicing node to be allocated and the candidate splicing node, and the modulus values of the vector representations; calculating the first product of the vector representations of the splicing node to be allocated and the candidate splicing node, and the second product of the modulus values of the vector representations; and using the ratio of the first product to the second product as the event compatibility. Calculating the event consistency between the event node to be allocated and the event nodes in the story tree includes: determining the story line of the root node of the story tree and the event node to be allocated; according to the chronological order of the story line, obtaining the sum of the compatibilities of adjacent event nodes among all the event nodes from the root node to the event node to be allocated; determining the reciprocal of the number of event nodes from the root node to the event node to be allocated, and using the product of the reciprocal and the sum as the event consistency. Calculating the time penalty term between the event node to be allocated and the event nodes in the story tree includes: determining the first timestamp of the event node to be allocated and the second timestamp of the event node in the story tree; judging whether the difference between the second timestamp and the first timestamp is less than zero; if so, using the difference as the exponent of the exponential function with base E, and using the obtained value as the time penalty term; if not, the time penalty term is zero.

[0123] Therefore, by calculating the event compatibility, event consistency, and time penalty term between the event node to be allocated and the event nodes in the story tree respectively; and using the product of the event compatibility, event consistency, and time penalty term as the connection strength, it is convenient to determine the position of the event node in the story tree.

[0124] Please refer to Figure 7 , Figure 7 which is a schematic framework diagram of an embodiment of the electronic device 80 of the present application. The electronic device 80 includes a memory 81 and a processor 82 that are coupled to each other. Program instructions are stored in the memory 81, and the processor 82 is used to execute the program instructions to implement the steps in any of the above-mentioned event clustering method embodiments, or to implement the steps in any of the above-mentioned event context construction method embodiments. Specifically, the electronic device 80 may include, but is not limited to: desktop computers, laptop computers, servers, mobile phones, tablet computers, etc., which are not limited herein.

[0125] Specifically, the processor 82 is used to control itself and the memory 81 to implement the steps in any of the above audio optimization method embodiments. The processor 82 may also be referred to as a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip with signal processing capabilities. The processor 82 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 82 may be implemented jointly by integrated circuit chips.

[0126] In the above solution, after obtaining the candidate text, keywords of the candidate text are extracted respectively based on the structural features and semantic features of the words in the candidate text to form a keyword subgraph for each candidate text. The keywords are divided into several communities based on the keyword subgraph, and the candidate texts are clustered into the communities respectively according to the keywords of each candidate text. In each community, the candidate texts describing the same event are clustered into the same event node based on the keyword subgraph. Based on this, the keyword subgraph takes into account both structural features and semantic features, the interpretability of the keywords is better, and the connection between the keywords of the same topic and the same event node is closer. Therefore, community clustering and event node clustering are performed based on the keyword subgraph, which can improve the accuracy of event clustering.

[0127] Please refer to Figure 8 , Figure 8 is a schematic framework diagram of an embodiment of the computer-readable storage medium 90 of the present application. The computer-readable storage medium 90 stores program instructions 91 that can be run by a processor. The program instructions 91 are used to implement the steps in any of the above event clustering method embodiments, or to implement the steps in any of the above event context construction method embodiments.

[0128] In the above solution, after obtaining the candidate text, keywords of the candidate text are extracted based on the structural features and semantic features of the words in the candidate text to form a keyword subgraph for each candidate text. The keywords are divided into several communities based on the keyword subgraph, and the candidate texts are clustered into the communities according to the keywords of each candidate text. In each community, the candidate texts describing the same event are clustered into the same event node based on the keyword subgraph. Based on this, the keyword subgraph takes into account both structural features and semantic features, the interpretability of the keywords is better, and the connection between the keywords of the same topic and the same event node is closer. Therefore, community clustering and event node clustering based on the keyword subgraph can improve the accuracy of event clustering.

[0129] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0130] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0131] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of the device or unit may be in electrical, mechanical or other forms.

[0132] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0133] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0134] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

Claims

1. A method for constructing an event context, characterized in that, The method includes: Obtaining candidate texts; Respectively extracting keywords of the candidate texts based on the structural features and semantic features of the words in the candidate texts to form keyword subgraphs of each candidate text; Dividing the keywords into several communities based on the keyword subgraphs to obtain several event nodes; Structurally displaying the event nodes to construct several story trees, and including: Determining the event time and event summary of the event node to form a brief text of the event node; Successively taking the event nodes as to-be-allocated event nodes, and performing story clustering on the to-be-allocated event nodes to cluster the event nodes under each community into different story trees; In each story tree, successively taking the event nodes as the to-be-allocated event nodes based on the chronological order of the event times of the event nodes; Respectively splicing the candidate texts of the to-be-allocated event nodes and the candidate texts of the event nodes in the story tree to form a to-be-allocated splicing node and a candidate splicing node; Using a similarity discrimination model to determine whether the to-be-allocated splicing node and the candidate splicing node meet the fusion condition; If so, adding the candidate text and its keywords of the to-be-allocated event node to the event node corresponding to the candidate splicing node; If not, calculating the connection strength between the to-be-allocated event node and the event node in the story tree, and taking the event node with the maximum connection strength as the parent node of the to-be-allocated event node to connect the to-be-allocated event node after the parent node; Displaying the story tree and the brief text of its event nodes, or re-obtaining new event nodes and re-executing the steps of determining the event time and event summary of the event node to form the brief text of the event node and its subsequent steps to update the story tree.

2. The method according to claim 1, wherein The step of successively taking the event nodes as to-be-allocated event nodes and performing story clustering on the to-be-allocated event nodes to cluster the event nodes under each community into different story trees includes: Successively taking the event nodes as the to-be-allocated event nodes based on the chronological order of the event times of the event nodes; Obtaining whether the compatibility between the keyword set of the to-be-allocated event node and the keyword set of the story tree exceeds a preset compatibility value; If so, determining whether the number of identical words included in the title of at least one candidate text in the to-be-allocated event node and the title of at least one candidate text in the story tree exceeds a preset number; If so and the story tree is unique, allocating the to-be-allocated event node to the story tree, and if so but the story tree is not unique, clustering the to-be-allocated event node into the story tree with the maximum compatibility and the maximum number of identical words.

3. The method according to claim 1, wherein The step of calculating the connection strength between the to-be-allocated event node and the event node in the story tree includes: Respectively calculating the event compatibility, event consistency and time penalty term between the to-be-allocated event node and the event node in the story tree; Take the product of the event compatibility, event consistency, and time penalty term as the connection strength; Among them, calculating the event compatibility between the event node to be assigned and the event node in the story tree includes: Respectively obtain the vector representations of the splicing node to be assigned and the candidate splicing node, and the modulus values of the vector representations; Calculate the first product of the vector representations of the splicing node to be assigned and the candidate splicing node, and the second product of the modulus values of the vector representations; Take the ratio of the first product to the second product as the event compatibility; Calculating the event consistency between the event node to be assigned and the event node in the story tree includes: Determine the story line of the root node of the story tree and the event node to be assigned; According to the chronological order of the story line, obtain the sum of the compatibilities of adjacent two event nodes among all the event nodes from the root node to the event node to be assigned; Determine the reciprocal of the number of event nodes from the root node to the event node to be assigned, and take the product of the reciprocal and the sum as the event consistency; Calculating the time penalty term between the event node to be assigned and the event node in the story tree includes: Determine the first timestamp of the event node to be assigned and the second timestamp of the event node in the story tree; Judge whether the difference between the second timestamp and the first timestamp is less than zero; If so, take the difference as the exponent of the exponential function with base E, and take the obtained value as the time penalty term; if not, the time penalty term is zero.

4. The method according to claim 1, characterized in that The determining the event time of the event node includes: taking the event time of the candidate text with the earliest time in the event node as the event time of the event node, or after removing the event times of the candidate text with the earliest and latest times in the event node, taking the average of the event times of the remaining candidate texts as the event time of the event node; The determining the event summary of the event node includes: taking the event summary of the candidate text with the earliest time in the event node as the event summary of the event node, or inputting all the candidate texts of the event node into a summary extraction model, and taking the output of the summary extraction model as the event summary of the event node.

5. The method according to claim 1, wherein The obtaining a plurality of event nodes includes: Cluster the candidate texts into the communities according to the keywords of each candidate text respectively; In each community, input the keyword subgraphs into a text modeling model respectively, obtain the semantic vector representations of each candidate text, and cluster several candidate texts whose semantic vector representations meet the text similarity condition into the same event node.

6. The method according to claim 5, characterized in that, The text modeling model is a graph neural network model, and the graph neural network model includes a graph autoencoder. The inputting the keyword subgraphs into the text modeling model respectively to obtain the semantic vector representations of each candidate text includes: Use the graph autoencoder to encode the keyword subgraph to obtain the semantic vector representation; Clustering the several candidate texts whose semantic vector representations meet the text similarity condition into the same event node includes: Calculating the similarity of the semantic vector representations, clustering the candidate texts with the similarity greater than a preset threshold into the same event node, or clustering the candidate texts represented by the semantic vector representations by using a spectral clustering method.

7. The method according to claim 5, characterized in that, Extracting the keywords of each candidate text based on the structural features and semantic features of the words in the candidate text to form a keyword subgraph for each candidate text, includes: Extracting the keywords of the candidate text by a keyword extraction method based on statistical characteristics and using a named entity recognition model to extract the keywords of the candidate text; Taking the keywords as the nodes of the keyword subgraph, connecting the keywords that meet the co-occurrence condition with edges, and discarding the keywords that do not meet the co-occurrence condition to form the keyword subgraph.

8. The method according to claim 7, wherein The co-occurrence condition is that the number of candidate texts containing both keywords exceeds a preset number of texts, and the quotient of the number of candidate texts containing both keywords and the sum of the number of candidate texts respectively including the two keywords exceeds a preset conditional probability.

9. The method according to claim 5, characterized in that, Dividing the keywords into several communities based on the keyword subgraph, and clustering the candidate texts into the communities according to the keywords of each candidate text respectively, includes: Fusing the keyword subgraphs containing the same keywords to form a keyword cluster; Using a community discovery method to divide the keyword cluster to form the several communities; Calculating the relevance between the candidate text and the community, and adding the candidate text corresponding to the maximum value of the relevance to the corresponding community respectively.

10. The method according to claim 9, wherein The community discovery method includes an overlapping community discovery method based on hierarchical clustering.

11. The method according to claim 5, wherein Obtaining the candidate text includes: Receiving the user input requirement; Performing semantic parsing on the user input requirement by using a named entity recognition model to obtain index keywords; Using the index keywords to screen out the texts containing at least one of the index keywords in the text set as the candidate text.

12. An event context construction device, characterized in that An event node acquisition module, configured to acquire candidate texts, extract the keywords of each candidate text based on the structural features and semantic features of the words in the candidate text to form a keyword subgraph for each candidate text, and divide the keywords into several communities based on the keyword subgraph to obtain several event nodes; A construction module, configured to structurally display the event nodes, construct several story trees, and include determining the event time and event summary of the event node to form a brief text of the event node; Sequentially taking the event nodes as the to-be-allocated event nodes, and performing story clustering on the to-be-allocated event nodes to cluster the event nodes under each community into different story trees; In each story tree, sequentially taking the event nodes as the to-be-allocated event nodes based on the chronological order of the event times of the event nodes; Respectively splice the candidate text of the event node to be allocated and the candidate text of the event node in the story tree to form a to-be-allocated splicing node and a candidate splicing node; Use a similarity discrimination model to determine whether the to-be-allocated splicing node and the candidate splicing node meet the fusion conditions; If so, add the candidate text of the to-be-allocated event node and its keywords to the event node corresponding to the candidate splicing node; If not, calculate the connection strength between the to-be-allocated event node and the event node in the story tree, and use the event node with the maximum connection strength as the parent node of the to-be-allocated event node to connect the to-be-allocated event node after the parent node; Display the brief text of the story tree and its event nodes, or re-obtain new event nodes, and re-execute the determination of the event time and event summary of the event nodes to form the brief text of the event nodes and subsequent steps to update the story tree.

13. An electronic device, characterized in that, Comprising a mutually coupled memory and a processor, the processor is configured to execute program instructions stored in the memory to implement the event context construction method according to any one of claims 1 to 11.

14. A computer-readable storage medium storing program instructions executable by a processor, characterized in that, When the program instructions are executed by the processor, the event context construction method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Document clustering method and platform, server and computer-readable medium

    CN109522410A

  • File processing method and device based on graph neural network and storage medium

    CN112214993A

  • Text topic clustering method and device, equipment and storage medium

    CN112329460A

  • Financial news stream emergency detection method based on hierarchical clustering

    CN113449108A