An event context generation method and system incorporating deep semantic relation classification
By integrating the method of deep semantic relationship classification, using thematic model and dependent syntax analysis, combined with unsupervised clustering and two-layer clustering algorithms, the accuracy and coherence problems of event detection and vein generation are solved, and more refined event information extraction and readable event vein generation are achieved.
Patent Information
- Application Number
- CN202111530106.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The existing event detection methods are insufficient in accuracy and deep semantic feature extraction, and the event context generation method lacks consideration for the deep evolution relationship between events, resulting in inaccurate event detection and incoherent context display.
The method of integrating deep semantic relationship classification is adopted to obtain deep semantic features through topic models and dependent syntax analysis, and the event context is generated by combining unsupervised clustering and two-layer clustering algorithms.
It improves the accuracy of event detection and the consistency of context generation, reduces the pressure of manual labeling, and can more accurately describe the event development process.
Smart Images

Figure CN114265932B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an event context generation method and system incorporating deep semantic relationship classification, and belongs to the technical field of language processing. Background Art
[0002] Social networks have been widely used to publish news and report events. The real-time nature of information and the ability to spread quickly in social networks make it an important medium for obtaining information, and the short-text expression can also effectively convey key information. These characteristics of social networks have subverted the dominance of traditional media in information dissemination, making it a valuable source of data for monitoring events and their evolution. However, the rapid accumulation of text in social networks and the colloquial expression style pose great challenges to monitoring events and the evolution between events. Extracting events with the same theme and their evolution from social network text can greatly help us understand a certain event panoramically. For example: we expect to obtain information about all events (i.e., events) of the PyeongChang Winter Olympics and the progress of these events (i.e., event evolution). This requires us to first detect events, then cluster these events to obtain events with the same theme (i.e., stories), and finally present them in a user-friendly way (story context). In addition, deep learning and machine learning technologies have developed rapidly in recent years, but there are still some problems in the task of event context generation: 1) Events are represented by a set of texts and have a specific theme. How to extract a strongly relevant set of texts corresponding to the event from the set of texts is a key issue; 2) In the process of generating the event context structure, how to construct the event context from a global perspective and improve the coherence and integrity of the context structure is also an urgent problem to be solved.
[0003] The event context generation method can be divided into two parts: 1) event detection, 2) context generation. Event detection is to divide news descriptions of the same event into a group in a large collection of news data. Generally, the same event refers to the same time, place, entity, and accompanying results involved in multiple news descriptions; context generation is to track and reveal how events develop over time in a structured way. The event context shows the development process of a theme, that is, a set of a main event and its subsequent developing events.
[0004] Existing event detection methods mainly include two categories: document-based detection methods and keyword-based detection methods. Document-based detection methods are mainly based on news content features, and generally measure the relationship between events based on similarity. For example, Wu et al. calculated the cosine similarity using the document feature vectors extracted by TF-IDF, and divided events according to the similarity. Zhou et al. proposed a hybrid model based on term frequency-inverse event frequency (TF×IEF) and time distance cost factor, modeled events as vectors using TF×IEF, and then measured the similarity of event content according to the cosine similarity to complete event detection. In addition, Ozdikis et al. performed online processing on data within a time window. When calculating the similarity between tweets in the new time window and existing active clusters, they calculated the co-occurrence vector between words using the context within the current time window and one time window before and after, and then used this vector to calculate the similarity between words and multiply it by the TF-IDF value to generate the final vector representation, which is an extension of the TF-IDF vector representation and can well solve the concept drift problem that occurs over time. Finally, they analyzed the intensity evolution process of the event using the pattern of the frequency of a specific word changing over time within related events.
[0005] Keyword-based event detection methods mainly consider that when an event occurs, the frequency of certain feature words will increase sharply, and identify and discover events by analyzing these feature words. For example, Yang constructed a keyword co-occurrence graph based on the co-occurrence features of keywords, and selected a community detection algorithm to divide the keyword co-occurrence graph, and used the extracted topic feature words to achieve the division of topic events. Additionally, news is represented based on keywords, and clustering algorithms are used for topic or event detection. Common clustering methods include density-based clustering, partition-based clustering, hierarchical clustering, and incremental clustering.
[0006] Among the currently known context generation methods, the representation forms of event contexts mainly have the following three structures: time axis structure, plane structure, and graph structure. Among these three structures, the time axis structure directly connects events through the time evolution order of events, and the structure is relatively simple. This method directly generates an event context based on the time sequence of the obtained events; the plane structure diverges from a main event. This method mainly determines a core event, and all other events are considered to be the subsequent development of this event; the graph structure analyzes the association between events in different story branches and is relatively complex. In this method, based on the obtained events, a directed graph or an undirected graph is constructed, and the minimum spanning tree or the maximum spanning tree is used as the final event context structure. Summary of the Invention
[0007] The disadvantages of the prior art are as follows. (1) There are many deficiencies in the existing event detection methods: (a) In the existing keyword-based event detection methods, the effectiveness of keywords largely determines the accuracy of event detection. However, most of the current keyword methods are obtained by using methods such as textrank or TF-IDF. These methods mostly obtain keywords that tend to be some entity words, etc., and cannot fully reflect the meaning of events; (b) In clustering techniques, event detection methods based on TF-IDF vectors and word2vec word vectors are all aimed at the shallow semantic features of texts. Words are independent of each other and cannot reflect sequence information. In the process of solving the similarity of word vectors, it is difficult to distinguish synonymous problems, and fine and accurate event information cannot be obtained, resulting in inaccurate event detection and unable to accurately describe the development process of events; (c) In addition, in a large amount of news data, supervised methods are mostly used, which causes a great deal of manual pressure in this case, and the quality of such developing events cannot be guaranteed either. (2) The existing event context generation methods lack consideration of the deep evolution relationship between events. They simply determine the context branches of the current node according to the time sequence or according to the maximum similarity between the current node and all previous nodes, and cannot cope with the situation of topic deviation that is extremely deviated from the original event in the subsequent development of events, thus making it difficult to accurately show the evolution relationship.
[0008] The purpose of the present invention is to overcome the technical deficiencies existing in the prior art and propose an event context generation method and system incorporating deep semantic relationship classification to solve the following technical requirements: (1) In the stage of theme event division, based on the theme model to complete the theme event division, the neural theme model can effectively obtain the deep semantic features of texts, and at the same time adopt an unsupervised form, reducing the annotation pressure without reducing the accuracy; (2) In the event detection stage, choose dependency syntactic analysis to obtain keywords, and based on the deep semantic relationship, the core content described in the news can be more accurately described; (3) In the context generation stage, determine the branches according to the changes of keywords to generate the context, fully considering the development relationship of events.
[0009] The present invention specifically adopts the following technical solutions: An event context generation method incorporating deep semantic relationship classification, comprising the following steps:
[0010] A data preprocessing step, specifically including: performing word segmentation on the news data set D = [d 1 , d 2 , … d |D| , and generating a word document sequence v = [v 1 , v 2 , … v D after merging;
[0011] The topic clustering step specifically includes: training a topic model and using the trained topic model to complete topic clustering. For the news data set D = [d 1 , d 2 , … d |D| , after passing through the topic model, the probability p i of each news data for each topic is obtained. Finally, according to the probability p i , the news data set D is divided into multiple categories to obtain the topic clustering result T = {T 1 , T 2 , … T |T|}, where T i is a set of news data;
[0012] The event clustering step specifically includes: obtaining the keywords of the news data set D, and for each news t in the topic clustering result i , using the bert model to vectorize each news data, that is, after splicing all the keywords, input them into the bert model, and the final news text vector representation is the average of the vectors of all tokens; among them, w i is the i-th keyword of the news data,
[0013] The context generation step specifically includes: for all events obtained under each topic , determining the branches to obtain the branch set B = {branch 1 , branch 2 , … branch |B|} corresponding to each topic, where branch i is the event set corresponding to the i-th branch; connect the events in each branch in chronological order and connect the branches in chronological order, that is, connect them in the chronological order of the earliest events in the branches, and finally obtain the event context.
[0014] As a preferred embodiment, the training of the topic model specifically includes:
[0015] For the word document sequence v = [v 1 , v 2 , … v D , where D is the number of words included in this word document sequence, and v i ∈ {1, …, V} represents the position of the i-th word in the word document sequence in the word table, and V is the size of the corpus word table;
[0016] For the topic model, each vocabulary v i of the word document sequence has two hidden states containing context information, namely the forward hidden state and the backward hidden state The forward hidden state and the backward hidden state are obtained from the context information v i of v <i = [v 1 , …, v i-1 and v >i = [v i+1 , …, v D , and by introducing pre-trained word vectors as prior knowledge, that is includes the complete context information of v i ;
[0017]
[0018]
[0019] where g(.) is a non-linear activation function, and are bias vectors, H is the hidden layer size, that is, the number of topics, W is a parameter matrix, E is a pre-trained word vector matrix, γ is a weight coefficient, and represent the v j columns in matrices W and E respectively. Matrix W is a learnable parameter matrix, which represents the topic-word distribution of the topic model. Each row of W l,: encodes the topic information of the l-th latent topic, and each column is the vector representation of word v i ;
[0020] Secondly, the topic model decomposes the joint distribution p(v) of all words in the word-document sequence into the product of the conditional distributions of each word v i , that is and models the word-document sequence accordingly, where the forward and backward autoregressive conditions p(v i ) of each word are calculated by a neural network from the forward hidden state and the backward hidden state respectively:
[0021]
[0022]
[0023] where W ∈ {1, …, V}, are the backward and forward biases respectively;
[0024] Finally, the parameters are optimized by maximizing the log-likelihood function logp(v) to obtain the topic model.
[0025] As a preferred embodiment, the keywords for obtaining the news data set D include: obtaining keywords based on dependency syntactic analysis technology, extracting the subject-predicate relationship, verb-object relationship, indirect object relationship, and attributive-middle relationship in the news data set, and using these as the keywords for the news data set D for subsequent event clustering.
[0026] As a preferred embodiment, the event clustering step specifically includes:
[0027] Step 1) Use the first document as a seed to establish a topic;
[0028] Step 2) Calculate the similarity between the next document X and the cluster center news of all existing topics. Adopt the cosine distance measurement method to find the existing topic with the greatest similarity to document X; if the similarity value is greater than the threshold θ, add document X to the topic with the greatest similarity, and jump to step 4);
[0029] Step 3) If the similarity value is less than the threshold θ, then document X does not belong to any existing topic, and a new topic category needs to be created, and at the same time, the current text is attributed to the newly created topic category;
[0030] Step 4) The clustering ends, waiting for the next document to enter; after singlePass processing, each topic obtains multiple event sets where e i =<d, w> is the time set, d is all the news in the time set e i and w is the keyword set corresponding to the news.
[0031] As a preferred embodiment, the branch determination includes: for all events obtained under each topic First, obtain the high-frequency keywords of each event. For the high-frequency words of each event, compare the Jaccard similarity coefficients between the high-frequency words of each event, and select the top ten with the highest frequency as keywords for comparison. If the Jaccard similarity coefficient is less than the threshold δ, it is determined that the two do not belong to the same branch, otherwise it is determined that the two belong to the same branch.
[0032] The present invention also proposes an event context generation system incorporating deep semantic relationship classification, including:
[0033] A data preprocessing module, which specifically executes: performing word segmentation on the news data set D = [d 1 , d 2 , … d |D| , and generating a word document sequence v = [v 1 , v 2 , … v D after merging;
[0034] The theme clustering module specifically performs: training a theme model and using the trained theme model to complete theme clustering. For the news data set D = [d 1 , d 2 , … d |D| , after passing through the theme model, the probability p i of each news data for each theme is obtained. Finally, according to the probability p i , the news data set D is divided into multiple categories to obtain the theme clustering result T = {T 1 , T 2 , … T |T|}, where T i is a set of news data;
[0035] The event clustering module specifically performs: obtaining the keywords of the news data set D, and for each news t in the theme clustering result i , using the bert model to vectorize each news data, that is, after splicing all the keywords, input them into the bert model, and the final news text vector representation is the average of the vectors of all tokens; among them, w i is the i-th keyword of the news data,
[0036] The context generation module specifically performs: determining branches for all events obtained under each theme to obtain the branch set B = {branch 1 , branch 2 , … branch |B|} corresponding to each theme, where branch i is the set of events corresponding to the i-th branch; connect the events in each branch in chronological order, and also connect the branches in chronological order, that is, connect them in the chronological order of the earliest events in the branches, and finally obtain the event context.
[0037] As a preferred embodiment, the training of the theme model specifically includes:
[0038] For the word document sequence v = [v 1 , v 2 , … v D , where D is the number of words contained in the word document sequence, and v i ∈ {1, …, V} represents the position of the i-th word in the word document sequence in the word list, and V is the size of the corpus word list;
[0039] For the theme model, each vocabulary v iThere are two hidden states containing context information, namely the forward hidden state and the backward hidden state The forward hidden state and the backward hidden state are obtained from the context information v i of v <i = [v 1 , …, v i-1 and v >i = [v i+1 , …, v D and by introducing pre-trained word vectors as prior knowledge, that is including the complete context information of v i ;
[0040]
[0041]
[0042] where g(.) is a non-linear activation function, and are bias vectors, H is the hidden layer size, i.e., the number of topics, W is a parameter matrix, E is a pre-trained word vector matrix, γ is a weight coefficient, and respectively represent the v j columns in matrices W and E. Matrix W is a learnable parameter matrix, which represents the topic-word distribution of the topic model. Each row W l,: encodes the topic information of the l-th latent topic, and each column is the vector representation of the word v i ;
[0043] Secondly, the topic model decomposes the joint distribution p(v) of all words in the word-document sequence into the product of the conditional distributions of each word v i , that is and models the word-document sequence accordingly, where the forward and backward autoregressive conditions p(v i ) of each word are calculated by the neural network from the forward hidden state and the backward hidden state respectively:
[0044]
[0045]
[0046] where W ∈ {1, …, V}, are the backward and forward biases respectively;
[0047] Finally, the parameters are optimized by maximizing the log-likelihood function logp(v) to obtain the topic model.
[0048] As a preferred embodiment, the keywords for obtaining the news data set D include: obtaining keywords based on dependency syntax analysis technology, extracting the subject-predicate relationship, verb-object relationship, indirect object relationship, and attributive-middle relationship in the news data set, and using these as the keywords of the news data set D for subsequent event clustering.
[0049] As a preferred embodiment, the event clustering module specifically performs:
[0050] Step 1) Use the first document as a seed to establish a topic.
[0051] Step 2) Calculate the similarity between the next document X and the cluster center news of all existing topics. Adopt the cosine distance measurement method to find the existing topic with the greatest similarity to document X. If the similarity value is greater than the threshold θ, add document X to the topic with the greatest similarity, and jump to step 4).
[0052] Step 3) If the similarity value is less than the threshold θ, then document X does not belong to any existing topic, and a new topic category needs to be created, and at the same time, the current text is assigned to the newly created topic category.
[0053] Step 4) The clustering ends, waiting for the next document to enter; after singlePass processing, each topic obtains multiple event sets where e i =<d, w> is the time set, d is all the news in the time set e i and w is the keyword set corresponding to the news.
[0054] As a preferred embodiment, the branch determination includes: for all events obtained under each topic First, obtain the high-frequency keywords of each event. For the high-frequency words of each event, compare the Jaccard similarity coefficients between the high-frequency words of each event, and select the top ten with the highest frequency as keywords for comparison. If the Jaccard similarity coefficient is less than the threshold δ, it is determined that the two do not belong to the same branch, otherwise it is determined that the two belong to the same branch.
[0055] While applying deep learning methods, in the process of event detection, this invention makes full use of deep semantic features, and at the same time selects a two-layer clustering algorithm to make event detection more accurate. In addition, by comprehensively considering the evolution relationship of events, the coherence and readability of the event context are improved. Compared with the prior art, the advantages of this case are as follows: Advantage 1, a topic model is selected for topic detection. This topic model can effectively obtain the deep semantic features of the text, fully consider the context information, and at the same time selects an unsupervised model, which can accurately detect the topic while reducing manual annotation; Advantage 2, based on dependency syntactic analysis to determine keywords, solving the defect that most traditional keywords tend to be nouns or entity words. At the same time, in the clustering process, the cluster center is set as the latest document, fully considering the characteristics of event development and improving the accuracy of event detection; Advantage 3, in the process of event context, it is considered that event keywords will change with the change of the event focus, and branches are determined based on keyword changes, improving the readability of the event context. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is the topological schematic diagram of an event context generation method integrating deep semantic relationship classification of the present invention;
[0057] Figure 2 is the structural schematic diagram of a preferred embodiment of the topic model of the present invention;
[0058] Figure 3 is the schematic diagram of the form of the event context of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.
[0060] Embodiment 1: As Figure 1 shown, the present invention proposes an event context generation method integrating deep semantic relationship classification, including the following steps:
[0061] Data preprocessing step, specifically including: performing word segmentation on the news data set D = [d 1 , d 2 , … d |D| , and generating a word document sequence v = [v 1 , v 2 , … v D after merging; generally, jieba word segmentation is selected. In addition, in order to better understand the full text, after word segmentation, considering the co-occurrence degree between words, those with a co-occurrence degree greater than 80% are merged;
[0062] The topic clustering step specifically includes: training a topic model and using the trained topic model to complete topic clustering. For the news data set D = [d 1 , d 2 , … d |D| , after passing through the topic model, the probability p i of each news data for each topic is obtained. Finally, according to the probability p i , the news data set D is divided into multiple categories to obtain the topic clustering result T = {T 1 , T 2 , … T |T|}, where T i is a set of news data;
[0063] The event clustering step specifically includes: obtaining the keywords of the news data set D and, for each news t in the topic clustering result i , using the bert model to vectorize each news data, that is, after splicing all the keywords, inputting them into the bert model, and the final news text vector representation is the average of the vectors of all tokens; where w i is the i-th keyword of the news data,
[0064] The context generation step specifically includes: determining branches for all events obtained under each topic to obtain the branch set B = {branch , branch 1 , … branch 2 , … branch |B|}, where branch i is the set of events corresponding to the i-th branch; connecting the events in each branch in chronological order and connecting the branches in chronological order, that is, connecting them in the chronological order of the earliest events in the branches, and finally obtaining the event context.
[0065] As a preferred embodiment, the training of the topic model specifically includes:
[0066] This topic model is an unsupervised generative topic model. The structure of the topic model is as Figure 2 shown. This model extracts its latent features from the document and regenerates the text accordingly, with the log-likelihood function of the generated text as the final optimization objective.
[0067] For the word-document sequence v = [v 1 , v 2 , … v D , where D is the number of words included in this word-document sequence, v i∈{1,…,V} represents the position of the i-th word in the word-document sequence in the vocabulary, and V is the size of the vocabulary of the corpus;
[0068] For the topic model, each word v in the word-document sequence i has two hidden states containing context information, namely the forward hidden state and the backward hidden state The forward hidden state and the backward hidden state are obtained from the context information v i of v <i = [v 1 , …, v i-1 and v >i = [v i+1 , …, v D , as well as by introducing pre-trained word vectors as prior knowledge, that is contains the complete context information of v i ;
[0069]
[0070]
[0071] where g(.) is a non-linear activation function, and are bias vectors, H is the hidden layer size, that is, the number of topics, W is a parameter matrix, E is a pre-trained word vector matrix, γ is a weight coefficient, and respectively represent the v j th columns in the matrices W and E. The matrix W is a learnable parameter matrix, which represents the topic-word distribution of the topic model. Each row W l,: encodes the topic information of the l-th latent topic, and each column is the vector representation of the word v i ;
[0072] Secondly, the topic model decomposes the joint distribution p(v) of all words in the word-document sequence into the product of the conditional distributions of each word v i , that is and models the word-document sequence accordingly, where the forward and backward autoregressive conditions p(v i ) of each word are calculated by the forward hidden state and the backward hidden state through a neural network:
[0073]
[0074]
[0075] where \(W\in\{1,\ldots,V\}\), are the backward and forward biases respectively;
[0076] Finally, the parameters are optimized by maximizing the log-likelihood function \(\log p(v)\) to obtain the topic model.
[0077] As a preferred embodiment, the keywords for obtaining the news data set \(D\) include: Since traditional keywords tend to extract more nouns or entity words, but for a piece of news, it is impossible to accurately identify such fine-grained divisions of events based on these words alone. Based on the dependency parsing technology, keywords are obtained, and the subject-predicate relationship, verb-object relationship, indirect object relationship, and attributive-middle relationship in the news data set are extracted as the keywords of the news data set \(D\) for subsequent event clustering.
[0078] As a preferred embodiment, for all news text expressions, the singlePass one-pass text clustering algorithm is finally selected to implement event clustering. The cluster center is set as the latest document. It is found that this is more consistent with the development of events and can more accurately achieve event division compared with the latest news. The specific steps of the event clustering include:
[0079] Step 1) Use the first document as a seed to establish a topic;
[0080] Step 2) Calculate the similarity between the next document \(X\) and the cluster center news of all existing topics. Using the cosine distance metric method, find the existing topic with the greatest similarity to document \(X\); if the similarity value is greater than the threshold \(\theta\), add document \(X\) to the topic with the greatest similarity, and jump to Step 4);
[0081] Step 3) If the similarity value is less than the threshold \(\theta\), then document \(X\) does not belong to any existing topic, and a new topic category needs to be created, and at the same time, the current text is assigned to the newly created topic category;
[0082] Step 4) The clustering ends, waiting for the next document to enter; after being processed by singlePass, each topic obtains multiple event sets where \(e\) i \(=\langle d, w\rangle\) is the time set, \(d\) is all the news in the time set \(e\) i and \(w\) is the set of keywords corresponding to the news.
[0083] As a preferred embodiment, the branch determination includes: considering that there is a drift phenomenon during the event tracking process, the center of gravity of the event will change, and the event keywords will also change accordingly. For example, in the case of the rights protection event of Xi'an Mercedes-Benz, "finance" and "service fee" frequently appeared in the news on April 14, 2019, but never appeared in the previous event news. For all events obtained under each theme First, obtain the high-frequency keywords of each event. For the high-frequency words of each event, compare the Jaccard similarity coefficients between the high-frequency words of each event, and select the top ten with the highest frequency as keywords for comparison. If the Jaccard similarity coefficient is less than the threshold δ, it is determined that the two do not belong to the same branch, otherwise it is determined that the two belong to the same branch.
[0084] The present invention also proposes an event context generation system incorporating deep semantic relationship classification, including:
[0085] A data preprocessing module, which specifically executes: performing word segmentation on the news data set D = [d 1 , d 2 , … d |D| , and generating a word document sequence v = [v 1 , v 2 , … v D after merging;
[0086] A topic clustering module, which specifically executes: training a topic model, and using the trained topic model to complete topic clustering. For the news data set D = [d 1 , d 2 , … d |D| , after passing through the topic model, obtain the probability p i of each news data for each topic. Finally, according to the probability p i , divide the news data set D into multiple categories to obtain a topic clustering result T = {T 1 , T 2 , … T |T|}, where T i is a set of news data;
[0087] An event clustering module, which specifically executes: obtaining the keywords of the news data set D, and for each news t in each topic clustering result i , use the bert model to vectorize each news data, that is, splice all the keywords and input them into the bert model, and the final news text vector representation is the average of the vectors of all tokens; where w i is the i-th keyword of the news data,
[0088] The vein generation module specifically performs: for all events obtained under each theme perform branch determination to obtain a branch set B = {branch 1 , branch 2 , … branch |B|} corresponding to each theme, where branch i is the event set corresponding to the i-th branch; connect the events in each branch in chronological order, and also connect the branches in chronological order, that is, connect them in the chronological order of the earliest events in the branches, and finally obtain the event vein.
[0089] As a preferred embodiment, the training of the theme model specifically includes:
[0090] For the word document sequence v = [v 1 , v 2 , … v D , where D is the number of words contained in this word document sequence, and v i ∈{1, …, V} represents the position of the i-th word in the word document sequence in the word list, and V is the size of the corpus word list;
[0091] For the theme model, each vocabulary v i of the word document sequence has two hidden states containing context information, namely the forward hidden state and the backward hidden state The forward hidden state and the backward hidden state are obtained from the context information v i of v <i = [v 1 , …, v i-1 and v >i = [v i+1 , …, v D and by introducing pre-trained word vectors as prior knowledge, that is contains the complete context information of v i ;
[0092]
[0093]
[0094] where g(.) is a non-linear activation function, and are bias vectors, H is the size of the hidden layer, that is, the number of themes, W is the parameter matrix, E is the pre-trained word vector matrix, γ is the weight coefficient, and respectively represent v in matrices W and E jColumn, the matrix W is a learnable parameter matrix, which represents the topic-word distribution of the topic model. Each row of W l,: encodes the topic information of the l-th latent topic. Each column is the vector representation of the word v i ;
[0095] Secondly, the topic model decomposes the joint distribution p(v) of all words in the word-document sequence into the product of the conditional distributions of each word v i , that is and models the word-document sequence accordingly, where the forward and backward autoregressive conditions p(v i ) of each word are calculated by the neural network from the forward hidden state and the backward hidden state respectively:
[0096]
[0097]
[0098] where W ∈ {1, …, V}, are the backward and forward biases respectively;
[0099] Finally, the parameters are optimized by maximizing the log-likelihood function logp(v) to obtain the topic model.
[0100] As a preferred embodiment, the keywords for obtaining the news data set D include: obtaining keywords based on the dependency syntax analysis technology, extracting the subject-predicate relationship, verb-object relationship, indirect object relationship, and attributive-center relationship in the news data set, and using them as the keywords for the news data set D for subsequent event clustering.
[0101] As a preferred embodiment, the event clustering module specifically executes:
[0102] Step 1) Use the first document as a seed to establish a topic;
[0103] Step 2) Calculate the similarity between the next document X and the cluster center news of all existing topics. Using the cosine distance metric method, find the existing topic with the maximum similarity to the document X; if the similarity value is greater than the threshold θ, add the document X to the topic with the maximum similarity, and jump to Step 4);
[0104] Step 3) If the similarity value is less than the threshold θ, the document X does not belong to any existing topic, and a new topic category needs to be created, and at the same time, the current text is assigned to the newly created topic category;
[0105] Step 4) The clustering ends, waiting for the next document to enter; after singlePass processing, each topic obtains multiple event sets where e i =<d, w> is a time set, d is all the news in the time set e i and w is the set of keywords corresponding to the news.
[0106] As a preferred embodiment, the branch determination includes: for all events obtained under each theme firstly, obtain the high-frequency keywords of each event. For the high-frequency words of each event, compare the Jaccard similarity coefficients between the high-frequency words of each event, and select the top ten with the highest frequency as keywords for comparison. If the Jaccard similarity coefficient is less than the threshold δ, it is determined that the two do not belong to the same branch, otherwise it is determined that the two belong to the same branch.
[0107] It should be noted that the present invention is based on news data, and completes the generation process of the event context by combining the text clustering method based on the topic model and the event clustering method based on deep semantics, and constructs an accurate event context. Compared with the prior art, the key points to be protected in this case are as follows: Key point 1, in the process of topic clustering, an unsupervised topic model is selected, which can effectively obtain the deep semantic features of the text and fully consider the context information; Key point 2, in the process of event clustering, the dependency syntactic analysis is used to extract the keyword representatives of the events, the bert model is selected for vectorization, and in the clustering process, the characteristics of event development are fully considered, and the cluster center is set as the latest document, which greatly improves the accuracy of event detection; Key point 3, in the process of context generation, in addition to considering the time characteristics, the evolution relationship of the events is also considered, and the branches are determined based on the changes of high-frequency keywords, and finally the event context is formed.
[0108] Meaning of terms: Token is a string generated by the server as a token for the client to make requests. After the first login, the server generates a Token and returns this Token to the client. After that, the client only needs to bring this Token to request data, without having to bring the username and password again.
[0109] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0110] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. An event context generation method integrating deep semantic relationship classification, characterized in that, it includes the following steps: Data preprocessing steps, specifically including: performing word segmentation on the news data set D = [d 1 , d 2 , … d |D| , and generating a word document sequence v = [v 1 , v 2 , … v D after merging; The theme clustering steps specifically include: training a theme model and using the trained theme model to complete theme clustering. For the news data set D = [d 1 , d 2 , … d |D| , after passing through the theme model, the probability p i of each news data for each theme is obtained. Finally, according to the probability p i , the news data set D is divided into multiple categories to obtain the theme clustering result T = {T 1 , T 2 , … T |T|}, where T i is a set of news data; Event clustering steps, specifically including: obtaining the keywords of the news data set D, and for each news t in the topic clustering result in the i , using the bert model to vectorize each news data, that is, after splicing all the keywords, input them into the bert model, and the final news text vector representation is the average of the vectors of all tokens; where w i is the i-th keyword of the news data The vein generation step specifically includes: for all events obtained under each theme branch determination is performed to obtain a branch set B = {branch 1 , branch 2 , … branch |B|} corresponding to each theme, where branch i is the event set corresponding to the i-th branch; the events in each branch are connected in chronological order, and the branches are also connected in chronological order, that is, connected in the chronological order of the earliest events in the branches, and finally an event vein is obtained.
2. The event context generation method integrating deep semantic relationship classification according to claim 1, characterized in that, the training of the topic model specifically includes: For the word document sequence v = [v 1 , v 2 , … v D , where D is the number of words included in the word document sequence, v i ∈ {1, …, V} represents the position of the i-th word in the word document sequence in the vocabulary, and V is the size of the vocabulary of the corpus; For a topic model, each word v in the word-document sequence i has two hidden states containing context information, namely the forward hidden state and the backward hidden state The forward hidden state and the backward hidden state are obtained from the context information v i of v <i = [v 1 , …, v i-1 and v >i = [v i+1 , …, v D as well as by introducing pre-trained word vectors as prior knowledge, that is contains the complete context information of v i ; where g(.) is a non-linear activation function, and is a bias vector, H is the hidden layer size, i.e., the number of topics, W is a parameter matrix, E is a pre-trained word vector matrix, γ is a weight coefficient, and represent the v j -th columns in matrices W and E respectively. Matrix W is a learnable parameter matrix, which represents the topic-word distribution of the topic model. Each row of W l,: encodes the topic information of the l-th latent topic, and each column is the vector representation of word v i . Secondly, the topic model decomposes the joint distribution p(v) of all words in the word-document sequence into the product of the conditional distributions of each word v i , that is and models the word-document sequence accordingly, where the forward and backward autoregressive conditions p(v i ) of each word are calculated by the neural network from the forward hidden state and the backward hidden state respectively: where W ∈ {1, …, V}, are backward and forward biases, respectively; Finally, the parameters are optimized by maximizing the log-likelihood function logp(v) to obtain the topic model.
3. The event context generation method integrating deep semantic relationship classification according to claim 1, characterized in that, the keywords for obtaining the news data set D include: obtaining keywords based on dependency syntax analysis technology, extracting the subject-predicate relationship, verb-object relationship, indirect object relationship, and attributive-middle relationship in the news data set, and using these as the keywords of the news data set D for subsequent event clustering.
4. The event context generation method integrating deep semantic relationship classification according to claim 1, characterized in that, the event clustering step specifically includes: Step 1) Use the first document as a seed to establish a topic; Step 2) Calculate the similarity between the next document X and the centroid news of all existing topics. Adopt the cosine distance measurement method to find the existing topic with the greatest similarity to document X; if the similarity value is greater than the threshold θ, add document X to the topic with the greatest similarity, and jump to step 4); Step 3) If the similarity value is less than the threshold θ, then document X does not belong to any existing topic, and a new topic category needs to be created, and the current text is attributed to the newly created topic category; Step 4) Clustering ends, waiting for the next document to enter; after singlePass processing, multiple event sets are obtained for each topic where e i = <d, w> is the time set, d is all the news in the time set e i and w is the keyword set corresponding to the news.
5. The event context generation method integrating deep semantic relationship classification according to claim 1, characterized in that, The branch determination includes: for all events obtained under each topic First, obtain the high-frequency keywords of each event. For the high-frequency words of each event, compare the Jaccard similarity coefficients between the high-frequency words of each event, and select the top ten with the highest frequency as keywords for comparison. If the Jaccard similarity coefficient is less than the threshold δ, it is determined that the two do not belong to the same branch; otherwise, it is determined that the two belong to the same branch.
6. An event context generation system integrating deep semantic relationship classification, characterized in that, it includes: Data preprocessing module, specifically performing: segmenting the news data set D = [d 1 , d 2 , … d |D| , and generating a word document sequence v = [v 1 , v 2 , … v D after merging; The theme clustering module specifically performs: training a theme model, and using the trained theme model to complete theme clustering. For the news data set D = [d 1 , d 2 , … d |D| , after passing through the theme model, the probability p i of each news data for each theme is obtained. Finally, according to the probability p i , the news data set D is divided into multiple categories to obtain the theme clustering result T = {T 1 , T 2 , … T |T|}, where T i is a set of news data; The event clustering module specifically performs: obtaining the keywords of the news data set D, and for each news t in the topic clustering result i , using the bert model to vectorize each news data, that is, concatenating all the keywords and inputting them into the bert model, and the final news text vector representation is the average of the vectors of all tokens; where w i is the i-th keyword of the news data The vein generation module specifically performs: for all events obtained under each topic branch determination is carried out to obtain a branch set B = {branch 1 , branch 2 , … branch |B|} corresponding to each topic, where branch i is the event set corresponding to the i-th branch; the events in each branch are connected in chronological order, and the branches are also connected in chronological order, that is, connected in the chronological order of the earliest events in the branches, and finally an event vein is obtained.
7. The event context generation system integrating deep semantic relationship classification according to claim 6, characterized in that, the training of the topic model specifically includes: For the word document sequence v = [v 1 , v 2 , … v D , where D is the number of words contained in the word document sequence, v i ∈ {1, …, V} represents the position of the i-th word in the word document sequence in the word list, and V is the size of the corpus word list; For a topic model, each word v in the word-document sequence i has two hidden states containing context information, namely the forward hidden state and the backward hidden state The forward hidden state and the backward hidden state are obtained from the context information v i of v <i = [v 1 , …, v i-1 and v >i = [v i+1 , …, v D , as well as by introducing pre-trained word vectors as prior knowledge, that is contains the complete context information of v i ; where g(.) is a non - linear activation function, and is the bias vector, H is the hidden layer size, i.e., the number of topics, W is the parameter matrix, E is the pre - trained word vector matrix, γ is the weight coefficient, and represent the v - th j columns in matrices W and E respectively. Matrix W is a learnable parameter matrix, which represents the topic - word distribution of the topic model. Each row of W l,: encodes the topic information of the l - th latent topic, and each column is the vector representation of the word v i ; Secondly, the topic model decomposes the joint distribution p(v) of all words in the word-document sequence into the product of the conditional distributions of each word v i , that is and models the word-document sequence accordingly, where the forward and backward autoregressive conditions p(v i ) of each word are calculated by the forward hidden state and the backward hidden state through a neural network: where W ∈ {1, …, V}, are backward and forward biases, respectively; Finally, the parameters are optimized by maximizing the log-likelihood function logp(v) to obtain the topic model.
8. The event context generation system integrating deep semantic relationship classification according to claim 6, characterized in that, the keywords for obtaining the news data set D include: obtaining keywords based on dependency syntax analysis technology, extracting the subject-predicate relationship, verb-object relationship, indirect object relationship, and attributive-middle relationship in the news data set, and using these as the keywords of the news data set D for subsequent event clustering.
9. The event context generation system integrating deep semantic relationship classification according to claim 6, characterized in that, the event clustering module specifically executes: Step 1) Use the first document as a seed to establish a topic; Step 2) Calculate the similarity between the next document X and the centroid news of all existing topics. Adopt the cosine distance measurement method to find the existing topic with the greatest similarity to document X; if the similarity value is greater than the threshold θ, add document X to the topic with the greatest similarity, and jump to step 4); Step 3) If the similarity value is less than the threshold θ, then document X does not belong to any existing topic, and a new topic category needs to be created. At the same time, the current text is assigned to the newly created topic category; Step 4) The clustering ends, and wait for the next document to enter; After singlePass processing, multiple event sets are obtained for each topic where e i = <d, w> is the time set, d is all the news in the time set e i and w is the keyword set corresponding to the news 10. An event context generation system incorporating deep semantic relationship classification according to claim 6, wherein, The branch determination includes: for all events obtained under each topic First, obtain the high-frequency keywords of each event. For the high-frequency words of each event, compare the Jaccard similarity coefficients between the high-frequency words of each event, and select the top ten with the highest frequency as keywords for comparison. If the Jaccard similarity coefficient is less than the threshold δ, it is determined that the two do not belong to the same branch; otherwise, it is determined that the two belong to the same branch.
Citation Information
Patent Citations
News event evolution analysis method based on time sequence distribution information and topic model
CN103984681A
News text-oriented event line extraction method based on deep clustering model
CN111125520A