News story context generation method considering time-space association

Through optimal transmission and space-time information in news text, the time-space distance between news events is calculated, and a story context that takes into account time-space correlation is constructed, which solves the problem of inaccurate characterization of space-time evolution of news events generated by existing methods, and achieves a more accurate expression of space-time evolution of news events.

CN120579520APending Publication Date: 2025-09-02Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510700266.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing methods are difficult to accurately characterize the spatial and temporal evolution process of news events, and ignore the spatial and temporal attributes of news events, resulting in inaccurate story context.

Method used

Through optimal transmission and space-time information in news text, the time-space distance between news events is calculated, and a story context that takes into account time-space correlation is constructed. The named entity recognition and entity linking model are used to extract time and location information, and the maximum spanning tree algorithm is used to generate a story context with the largest time-space correlation.

Benefits of technology

More accurately expressing the development and evolution process of news events in the space-time dimension, providing more accurate event evolution detection and simulation tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579520A_ABST
    Figure CN120579520A_ABST
Patent Text Reader

Abstract

The invention discloses a time-space association-considered news story context generation method, which comprises the following steps of: firstly, designing a two-stage unsupervised story discovery method, namely, embedding a preliminarily aggregated news article according to document-level semantics of a news stream; the semantic related news is distributed to the same news story more finely through keyword distribution in the candidate story; then, respectively analyzing the time expression and the place name entity extracted from the news article into standard format time and position coordinates by utilizing regular matching and Wikii data, so as to mine space-time information in the news article; and finally, proposing a space-time distance calculation method based on optimal transmission, introducing a distance attenuation function to model an attenuation rule of space-time association, and constructing a story context considering the space-time association by using a maximum spanning tree. According to the method provided by the invention, the development and evolution process of the news event can be more accurately expressed in the space-time dimension.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of news data processing, and in particular to a method for generating news story context taking into account temporal and spatial associations. Background Art

[0002] News reports help people understand hot topics happening in the real world. Internet news, characterized by rapid updates and a wide range of sources, provides people with rich but fragmented information. Understanding the full story often requires exploring the story's evolution from its inception to its demise, drawing on news materials from diverse perspectives and genres. Leveraging the spatiotemporal information within news articles to construct a story context that considers both temporal and spatial connections can be used for event situational awareness and to study the spatiotemporal evolution of news stories. This can provide timely and effective information support to decision-makers in areas such as public opinion management, regional conflict analysis, and natural disaster emergency response.

[0003] Early research viewed story mining as a form of topic detection and tracking, using topic detection algorithms to extract event timelines from web news corpora. Santos et al. used Ripley's K function and the PageRank algorithm to filter important entities from data such as Twitter and GDELT, and then populated the entities and relationships into a time matrix to construct storylines. Azeemi et al. used time windows to partition news streams, then used the Birch algorithm to perform two clustering operations based on attributes such as title, topic, and location list to form event chains. Finally, similar event chains from different time windows were merged to achieve dynamic updating of event chains. Because a single timeline cannot easily depict the asynchronous evolution of news events, some studies have adopted hierarchical or tree generation techniques to construct more complex tree-structured storylines. Liu et al. (Liu B, Han FX, Niu D, et al. Story forest: extracting events and telling stories from breaking news [J]. ACM Transactions on Knowledge Discovery from Data, 2020, 14 (3): 1-28. DOI: 10.1145 / 3377939) proposed a tree-like storyline generation method called Story Forest. Through two community partitions, relevant news documents are assigned to events, and the story tree is dynamically generated and updated based on the semantic connection strength between events. Liu Dong used an improved hot word algorithm to construct a hierarchical tree-like storyline with a trunk and branches. Mao Tonglin used semantic similarity to construct a text graph, obtained events through the Louvain community partitioning algorithm and DBSCAN clustering, and constructed a weighted storyline graph based on timestamps and community similarity. Finally, the maximum spanning tree Prim algorithm was used to construct the storyline. These methods mainly consider word-level semantic associations and construct storylines based on the co-occurrence relationship or similarity of keywords in news texts, which are easily affected by noise words. In addition, although these methods take time factors into consideration, they ignore the changes in the spatial location attributes of events in news texts, making it difficult to effectively track the overall development of news events from the spatiotemporal dimensions.

[0004] To explore the spatiotemporal evolution of news events, some studies have attempted to model the spatiotemporal processes of news events. Yuan et al. proposed a two-layer structure to represent the storyline of a typhoon disaster. The first layer describes the main storyline that evolves over time and geographic location, while the second layer describes the local storylines at different locations. This storyline is dynamically updated using a dynamic Steiner tree. Jia Mengshu modeled a marine disaster ontology, dividing the disaster into three stages: prevention, response, and recovery. Based on the ontology model, she extracted disaster keywords from Weibo news and analyzed the spatiotemporal evolution of the disaster event in chronological order. Ye Peng modeled five types of object states in typhoon disaster events and, based on rules, extracted feature words from Weibo to determine the object state types. Finally, the temporal sequence was used to analyze the spatial position movement and attribute feature changes between states. These spatiotemporal process modeling methods often focus on analyzing event evolution in specific domains and lack cross-domain generalization capabilities. Furthermore, these disaster event ontology modeling methods rely too heavily on domain knowledge and predefined rules, requiring the participation of domain experts, making them difficult to generalize to other application scenarios.

[0005] With its powerful semantic representation ability, deep neural networks can effectively learn the semantic information of news texts and capture the semantic associations between events. Story mining based on neural networks has become a research hotspot in recent years. Wang Chongwei used the event graph convolution network (GCN) to discover the implicit relationships between events and generate story contexts based on the relationship strength and time sequence. Wang et al. further improved the event graph convolution model and combined GCN with PageRank to improve the neighborhood aggregation of the event graph. Peng et al. used a heterogeneous information network (HIN) to model event relationships, constructed positive and negative sample pairs from the original sample sampling, input the sample pairs into the GCN to update the event features, and clustered related events according to the meta-paths defined by it. Liu Xudong used the pre-trained T5 model and small sample transfer learning to generate news event summaries, and constructed story contexts based on the cosine similarity and time sequence of the summary embeddings. Yoon et al. (Yoon S, Meng Y, Lee D, et al. SCStory: self-supervised and continual online story discovery [C] / / Proceedings of the ACM Web Conference 2023, 2023: 1853-1864. DOI: 10.1145 / 3543507.3583507) proposed a self-supervised story discovery method, SCStory. This method uses self-supervised multi-head attention pooling to train sentence embeddings as news article embeddings. Finally, the articles are assigned to related stories based on the cosine similarity of the article embeddings. These methods use implicit semantic embeddings to mine semantic associations between news articles. However, deep learning-based methods may compress or lose spatiotemporal information when converting news text into vector representations, making it difficult to directly encode the spatiotemporal properties of news events. This method ignores the spatial proximity of news events and fails to represent the development and evolution of news stories in the spatiotemporal dimension. Summary of the Invention

[0006] To address the inability of prior work to mine spatiotemporal connections between news events with multiple spatiotemporal attributes from news articles, and to accurately express the developmental context of news stories from a spatiotemporal perspective, this paper proposes a method for generating news story context that takes both spatiotemporal and temporal relationships into account. By using optimal transport (OT) and spatiotemporal information in news texts, this method calculates the spatiotemporal distances between news events and constructs a story context that takes both spatiotemporal and temporal relationships into account. This method addresses the inaccurate depiction of the spatiotemporal evolution of news events by story contexts generated using existing methods. Experiments and case studies conducted on a publicly available story context generation dataset explore the differences in the ability of different methods to express spatiotemporal evolution in story context generation.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] A method for generating news story context taking into account temporal and spatial correlations, comprising:

[0009] We perform document-level embedding on news articles, assigning related news from different time periods to the same candidate story based on embedding vector similarity. We also use agglomerative hierarchical clustering to discover stories within the candidate stories based on the distribution vectors of news article keywords.

[0010] For each news article, we extract time expressions and place name entities through named entity recognition and rule matching, and use the entity linking model GENRE and Wikidata to parse the place name entities into location coordinates including latitude and longitude;

[0011] Using the acquired time and location coordinates, the spatiotemporal correlation between news articles is calculated through the optimal transmission distance and the corresponding attenuation function, and the maximum spanning tree algorithm is used to construct the story thread with the maximum spatiotemporal correlation.

[0012] Furthermore, assigning related news in different time slices to the same candidate story based on embedding vector similarity includes:

[0013] For the w i The news documents contained in the time slice are calculated, the cosine similarity between the news text embeddings is calculated and compared with the similarity threshold θ, and the news documents with similarity greater than θ are assigned to the same news story. The average pooling of these news text embeddings is used as the story embedding; then, the w-th time slice is calculated. i A time piece of new discovery story and the former w i-1 The cosine similarity between the embeddings of the discovered stories in the w-th time slice is calculated and compared with the similarity threshold θ. The newly discovered stories with a similarity greater than the similarity threshold θ are merged into the discovered stories, and the merged story embedding is updated by weighted pooling according to the number of news articles of the newly discovered stories and the discovered stories. i The same steps are used to process the wth time slicei+1 time slices until all time slices are processed and a set of candidate stories is obtained.

[0014] Furthermore, the news article keyword distribution vector is constructed as follows:

[0015]

[0016] Where: v j is the value of the jth dimension of the news article keyword distribution vector v, k is a keyword among all keywords K, k s is the TF-IDF value corresponding to k, L j Represents the jth keyword in the vocabulary L.

[0017] Furthermore, the method of discovering stories in candidate stories using agglomerative hierarchical clustering based on news article keyword distribution vectors includes:

[0018] First, a fixed step size is set and the distance threshold is gradually increased, and the number of clusters under different distance thresholds is recorded. Then, the optimal distance threshold is determined by maximizing the second-order derivative of the cluster number. Finally, based on the optimal distance threshold, the keyword distribution vectors of the news articles in the candidate stories are clustered in an agglomerative hierarchical clustering manner. By clustering all candidate stories, a set of news stories is finally obtained, completing the story discovery task.

[0019] Furthermore, the time expression in the news article is extracted as follows:

[0020] Perform named entity recognition on news articles, and use named entities of type "DATE" and "TIME" as time expressions;

[0021] Use the publishing timestamp of each news article as the base time and parse the time string by matching predefined rules with regular expressions;

[0022] Preserve the time points and time ranges in the parsing results;

[0023] The reciprocal of the time interval days is used as the weight score of the standard time, and the Softmax function is used to convert the weight score into the probability score of the standard time in the news article.

[0024] Furthermore, place name entity extraction and parsing are performed in the following manner:

[0025] Based on the word segmentation results of the article, the words with the part of speech of place names are counted. Based on the statistical results and regular matching with the place name dictionary, place name entities are extracted from the news article and given confidence scores. Using GENRE, the corresponding Wikidata entity ID is obtained based on the input place name entity. The coordinate attributes of the entity ID are queried using the online Wikidata query service interface. Finally, the Softmax function is used to convert the place name entity confidence score into the probability score of the place name coordinate in the news article.

[0026] Furthermore, the optimal transmission distance includes a time-optimal transmission distance and a space-optimal transmission distance;

[0027] The time cost function is constructed as follows:

[0028]

[0029] Where: TempCost(t i ,t j ) represents the time from standard time t i to t j The time cost function, and The standard time t i and t j The start time, and t i and t j The end time, It is t i and t j The absolute value of the difference in days between the start times, It is t i and t j The absolute value of the difference in days between the end times;

[0030] Based on the constructed time cost function, linear programming is used to calculate the time-optimal transmission distance;

[0031] The Haversine distance between the location coordinate points is used as the spatial cost function. Based on the constructed spatial cost function, linear programming is used to calculate the optimal spatial transmission distance.

[0032] Furthermore, the attenuation function includes a time distance attenuation function and a space distance attenuation function;

[0033] The time distance decay function is defined as:

[0034]

[0035] Where: Tdist is the time-optimal transmission distance, is the time distance attenuation coefficient;

[0036] The spatial distance decay function is defined as:

[0037]

[0038] Where: Sdist is the optimal spatial transmission distance, and δ is the spatial distance attenuation coefficient.

[0039] Furthermore, the spatiotemporal correlation between news events is calculated as follows:

[0040] The time distance decay function value and the space distance decay function value are respectively used as the time correlation and space correlation between news events;

[0041] For each news story s, calculate the pair of news events e according to the optimal transmission distance and attenuation function. i and e j The time relationship between TempRel(e i ,e j ) and spatial association SpatRel(e i ,e j ), and get the time correlation matrix A of story s t and the spatial incidence matrix A s , the weighted sum of the two is α·A t +β·A s As the spatiotemporal association weight matrix A of story s.

[0042] Furthermore, when constructing a story thread with maximum spatiotemporal correlation using the maximum spanning tree algorithm, each element a in A ij As an event i and e j The edge weights between them are compared by comparing the news events e i and e j The maximum probability start time is used to set the edge direction of the spanning tree to complete the story line generation based on spatiotemporal association.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] The present invention proposes a news story context generation method that takes into account both temporal and spatial correlations. By optimally transmitting the temporal and spatial information in the news text, the temporal and spatial distances between news events are calculated, and a story context that takes into account both temporal and spatial correlations is constructed. This solves the problem that story contexts generated by existing methods inaccurately depict the temporal and spatial evolution of news events.

[0045] Through relevant experiments, compared with the baseline methods StoryForest and SCStory, the method proposed in this invention performs better in terms of relevance, accuracy, and correlation. The method proposed in this invention can more accurately express the development and evolution process of news events in the temporal and spatial dimensions, providing a new tool for event evolution detection and simulation. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 Schematic diagram of the flow of a method for generating news story context that takes into account temporal and spatial associations according to an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments:

[0048] like Figure 1 As shown, the present invention proposes a method for generating news story context that takes into account temporal and spatial correlation, comprising:

[0049] We perform document-level embedding on news articles, assigning related news from different time periods to the same candidate story based on embedding vector similarity. We also use agglomerative hierarchical clustering to discover stories within the candidate stories based on the distribution vectors of news article keywords.

[0050] For each news article, we extract time expressions and place name entities through named entity recognition and rule matching, and use the entity linking model GENRE and Wikidata to parse the place name entities into location coordinates including latitude and longitude;

[0051] Using the acquired time and location coordinates, the spatiotemporal correlation between news events is calculated through the optimal transmission distance and the corresponding attenuation function, and the maximum spanning tree algorithm is used to construct the story context with the maximum spatiotemporal correlation.

[0052] The present invention calculates the spatiotemporal distances between news events through optimal transmission and the spatiotemporal information in news texts, and constructs a story context that takes into account both spatiotemporal and temporal correlations, thereby solving the problem that story contexts generated by existing methods inaccurately depict the spatiotemporal evolution of news events.

[0053] The method specifically includes:

[0054] 1 Problem Definition

[0055] In story mining research, in order to facilitate the analysis of the evolution and trends of news events, the original news corpus obtained from the media is usually processed into a document stream in a certain order.

[0056] Definition 1 News document stream D = {d1, d2, ..., d |D|}, represents a set of news documents arranged by release time, where d iRepresents a news article. Usually, it can be sliced ​​into certain time slices W = {w1,w2,...,w |W|} to divide the news document stream D, where w i Represents a time interval of fixed length.

[0057] Definition 2: A news article contains three features, denoted as d = (K, T, P), where K represents a keyword set, T represents a time set, and P is a location coordinate set.

[0058] The present invention considers that a news article is a report of a real-world news event, that is, a news article d reports a news event e.

[0059] Definition 3: A news story s is a news event set E = {e1, e2, ..., e |E|}, where news event e i It is the basic unit of a news story. Each news event has spatiotemporal attributes, which can reflect the development process of the news story in the spatiotemporal dimension. The news document stream D contains several news stories S under a specific topic = {s1, s2, ..., s |S|}.

[0060] Definition 4 Candidate story set C = {c1, c2, ..., c |C|}, is the result of preliminary aggregation of news stories contained in the news document stream D, where c i It represents a candidate story, which contains a set of possibly related news articles. A candidate story c can be further divided into multiple news stories.

[0061] Definition 5: Story context st is a tree structure consisting of all news events and the relationships between events in a news story s, which expresses the development of related events. Specifically, story context st = (E, R), where E = {e1, e2, ..., e |E|} is the news event set of story context st, R is the news event relationship set, and relationship r ij =(e i ,e j ) is an ordered pair, representing news event e i and e j Specifically, the relationship r ij Represents news event node e i to e j A directed edge of the news event e i The time is earlier than e j .

[0062] The goal of storyline generation is to mine semantically related news events from an unstructured news document stream D, and generate a structured storyline st that expresses the evolution of news stories based on the associations between events.

[0063] 2 Methodological framework

[0064] The present invention divides the original news corpus into fixed-length time slices, uses news article semantic embedding and keyword distribution to perform unsupervised story discovery, assigns discrete news articles in the news document stream to different news stories, and then calculates the spatiotemporal correlation between news articles based on the optimal spatiotemporal transmission distance and attenuation function to generate a story context that expresses the spatiotemporal evolution process. The basic process of the present invention is as follows: Figure 1 As shown, it mainly includes 3 steps:

[0065] (1) News story discovery. News articles are embedded at the document level. Based on the similarity of the embedding vectors, related news from different time slices are assigned to the same candidate story. In the candidate stories, the keyword vectors of the news articles are distributed, and story discovery is completed using agglomerative hierarchical clustering.

[0066] (2) Event spatiotemporal information extraction. For each news article, named entity recognition (NER) and rule matching are used to extract time expressions and place name entities, and the place name entities are parsed into latitude and longitude coordinates using the entity linking model GENRE and Wikidata data.

[0067] (3) Storyline generation. Using the acquired time and location coordinates, the spatiotemporal correlation between news events is calculated through the optimal transmission distance and the corresponding attenuation function, and the maximum spanning tree algorithm is used to construct a storyline with the maximum spatiotemporal correlation.

[0068] 2.1 News Story Discovery

[0069] This paper mines news stories from news document streams using a story discovery algorithm. Considering that news articles are updated rapidly in real applications, learning methods that rely on labeled data are not suitable for dynamic news streams. Therefore, this paper proposes a two-stage unsupervised story discovery method to discover a set of news stories S = {s1, s2, ..., s |S| In order to process a large amount of news flow within a period of time, the news flow is divided into time slices W = {w1, w2, ..., w |W|}, using a sliding window approach to process news documents in different time slices one by one.

[0070] First, starting from the overall semantics of the article, semantically related news documents are initially assigned to the same candidate story. The pre-trained language model trained on a large-scale dataset has excellent text embedding capabilities, so the pre-trained text embedding model M3E is used to encode each news article. The M3E model is based on the text embedding model Sentence-BERT and can convert text into dense vectors. For the wth i The news documents contained in the time slice are calculated, the cosine similarity between the news text embeddings is calculated and compared with the similarity threshold θ, and the news documents with similarity greater than θ are assigned to the same news story. The average pooling of these news text embeddings is used as the story embedding. Then, the w-th time slice is calculated. i A time piece of new discovery story and the former w i-1 The cosine similarity between the embeddings of the discovered stories in the time slice is calculated and compared with the similarity threshold θ. The newly discovered stories with a similarity greater than the similarity threshold θ are merged into the discovered stories, and the merged story embedding is updated using a weighted pooling method based on the number of news articles of the newly discovered stories and the discovered stories. The window then slides to the wth time slice. i+1 time slice, execute the wth i The same steps are performed for each time slice until all time slices are processed, and the candidate story set C = {c1, c2, ..., c |C|}.

[0071] In order to mine news stories from a more fine-grained perspective, this paper uses the keywords of news articles to further divide each candidate story. For candidate story c, the keywords of the news article and the corresponding keyword scores are obtained by calculating the Term Frequency-Inverse Document Frequency (TF-IDF). The keywords of all news articles in candidate story c constitute the vocabulary L of c. For news article d in c, according to all the keywords K and all the keyword scores K of the article, the vocabulary L is obtained. s and the vocabulary L of the candidate story c, construct the keyword distribution vector v of the news article d, whose length is the number of keywords in the vocabulary L, and the value v of the jth dimension j Defined as:

[0072]

[0073] Where: k is a keyword in K; k s is the keyword score corresponding to k, that is, the TF-IDF value of the keyword, L j Represents the jth keyword in the vocabulary L.

[0074] The present invention uses agglomerative hierarchical clustering to cluster the keyword distribution vectors of all news articles in candidate stories c. Agglomerative hierarchical clustering is a hierarchical clustering algorithm using a bottom-up clustering strategy. At each step, the algorithm merges the two closest clusters until a stopping condition is met. It can specify a distance threshold as the stopping condition, without presetting the number of clusters. Because candidate stories under different themes in the candidate story set C may have different numbers of news articles and keyword lists of different sizes, applying the same distance threshold to each candidate story cannot guarantee the quality of the clustering results. To automatically obtain the distance threshold for each candidate story, the present invention uses the elbow method to identify the inflection point of the cluster number, thereby obtaining the distance threshold for clustering each candidate story. Specifically, a fixed step size is first set, and the distance threshold is gradually increased. The number of clusters at different distance thresholds is then recorded. Subsequently, the optimal distance threshold is determined by maximizing the second-order derivative of the cluster number. Finally, based on the optimal distance threshold, agglomerative hierarchical clustering is performed on the keyword distribution vectors of the news articles in the candidate stories. By clustering all candidate stories, a set of news stories S = {s1, s2, ..., s |S| Through this process, semantically related news stories are extracted from a large number of news document streams D, providing a reliable foundation for subsequent story context generation.

[0075] 2.2 Event spatiotemporal information extraction

[0076] In order to calculate the spatiotemporal correlation between news events, the original news text must be preprocessed first, and time information and location information must be extracted from each news article. Through natural language processing technology and other reference data (such as Wikidata data), the extracted time expressions and place name entities are parsed into standard format time and place name coordinates respectively.

[0077] 2.2.1 Time Expression Extraction and Parsing

[0078] The present invention performs named entity recognition on news articles, and takes named entities of type "DATE" and "TIME" as time expressions. In order to parse the time expression described in natural language into the standard format of "yyyy-MM-dd HH:mm:ss" that can be used for calculation, the release timestamp of each news article is used as the reference time, and the time string is parsed by regular matching predefined rules. The parsing results include 4 categories, namely time point (time_point), time range (time_span), time period (time_delta) and time period (time_period). Most of the results are time points and time ranges. Table 1 takes "2024-11-1108:30:15" as the reference time, and lists 2 examples for each category to illustrate the results of each category. Among them, the two types of results, time point and time range, include the start standard time and the end standard time, which represent a time interval from start to end, so the present invention only retains the results of time point and time range. In addition, considering the semantic ambiguity of long time intervals, the present invention uses the reciprocal of the time interval in days as the weight score of the standard time, and uses the Softmax function to convert the weight score into the probability score of the standard time in the news article.

[0079] Table 1 Examples of four categories of time expression parsing

[0080]

[0081] 2.2.2 Place Name Entity Extraction and Parsing

[0082] Based on the word segmentation results of the article, the present invention collects statistics on words with the part-of-speech of place names. Based on the statistical results and regular matching with a place name dictionary, place name entities are extracted from the news article and given confidence scores. To parse the place name entities into longitude and latitude coordinates, the autoregressive entity linking model GENRE is used to obtain the corresponding Wikidata entity ID based on the input place name entity. The coordinate attributes of the entity ID are queried using the online Wikidata query service interface using the SPARQL query language. Finally, the softmax function is used to convert the place name entity confidence score into a probability score of the place name coordinate in the news article.

[0083] 2.3 Storyline Generation

[0084] The goal of news story context generation based on optimal spatiotemporal transmission is to calculate the paired news events e in the news event set E of s for each news story s. i and e j The time relationship between TempRel(e i ,e j ) and spatial association SpatRel(e i ,ej ), taking news events as nodes and the sum of temporal and spatial associations as the edge weights between nodes, a tree-structured story context st is constructed to maximize the temporal and spatial associations between news events in the story context. The formula is:

[0085]

[0086] Where: ɑ and β are weights to balance temporal and spatial correlations.

[0087] 2.3.1 Optimal Transmission Distance in Time and Space

[0088] The present invention studies the spatiotemporal correlation between document-level news events. Taking into account that a news article may contain multiple different temporal and spatial attributes, the optimal transmission method is introduced to calculate the differences in spatiotemporal distribution between events, thereby quantifying the spatiotemporal correlation between events. Optimal transmission can be solved by linear programming (LP). The goal of the optimal transmission problem is to find an optimal "transmission scheme" for two given probability distributions so that the "transportation cost" from the original distribution to the target distribution is minimized. The transportation cost is usually defined based on a cost function, which represents the cost required to move an object from its original position to its target position. In order to calculate the optimal transmission distance in time and space, it is necessary to introduce a time and space cost function. Taking into account that the retained time parsing results include the start time and the end time, the present invention defines the time cost function as:

[0089]

[0090] Where: and The standard time t i and t j The start time, and t i and t j The end time, It is t i and t j The absolute value of the difference in days between the start times, It is t i and t j The absolute value of the difference in days between the end times.

[0091] Haversine distance is a distance measurement method based on spherical geometry, which can effectively calculate the distance between two coordinate points p on the earth's spherical surface. i and p j Therefore, the present invention uses Haversine distance as the spatial cost function, and the formula is defined as:

[0092]

[0093] Where: x i and y i p i Longitude and latitude; x j and y j p j The longitude and latitude are both expressed in radians; r is the radius of the Earth, usually 6371 km.

[0094] For news events i and e j , calculate e i and e j The optimal transmission distance Tdist ij And the optimal transmission distance Sdist ij As shown in Algorithm 1.

[0095]

[0096] 2.3.2 Storyline Generation Based on Spatiotemporal Correlations

[0097] The optimal temporal and spatial transmission distances respectively measure the differences in the distribution of events in the temporal and spatial dimensions. However, directly combining these two optimal transmission distances can lead to irrational spatiotemporal correlations. For example, disaster events often span hundreds of kilometers in space, but may span only a few days. Large spatial distances can mask these temporal differences, making it impossible to accurately measure the true spatiotemporal correlations between events.

[0098] To address the above issues and model the attenuation law of spatiotemporal correlation, the present invention introduces a distance decay function, mapping the original optimal transmission distances in time and space to a value range of (0, 1], so that the distances in the time and space dimensions have consistent dimensions and comparability. Based on the spatiotemporal attributes of the original news events, the present invention introduces an exponential time distance decay function and a power-law space distance decay function, respectively, and uses the distance decay function values ​​as the temporal correlation TempRel and spatial correlation SpatRel between news events, respectively.

[0099] The time distance decay function is defined as:

[0100]

[0101] Where: Tdist is the time-optimal transmission distance, parameter is the time distance attenuation coefficient.

[0102] The spatial distance decay function is defined as:

[0103]

[0104] Where: Sdist is the optimal spatial transmission distance, and parameter δ is the spatial distance attenuation coefficient.

[0105] For each news story s, calculate the pair of news events e according to the optimal transmission distance and attenuation function. i and e j The time relationship between TempRel(e i ,e j ) and spatial association SpatRel(e i ,e j ), and get the time correlation matrix A of story s t and the spatial incidence matrix A s , the weighted sum of the two is α·A t +β·A s As the spatiotemporal association weight matrix A of story s. Each element a in A ij As an event i and e j The edge weights between them are used to generate a tree-shaped story thread st with the maximum spatiotemporal correlation using the maximum spanning tree algorithm. Finally, in order to indicate the order in which news events occur, the present invention compares the news events e i and e j The maximum probability start time is used to set the edge direction of the spanning tree to complete the story line generation based on spatiotemporal association.

[0106] 3 Experimental design and result analysis

[0107] 3.1 Experimental design

[0108] 3.1.1 Experimental data

[0109] This paper conducts experiments on the publicly available Chinese news datasets ChineseNewsEvents and ChineseNewsSameStory. ChineseNewsEvents, used for experimental testing, contains 11,748 Chinese news articles collected from internet platforms such as Tencent, Sina, WeChat, and Sohu. Each article is manually labeled with a true story tag by Tencent's product managers, resulting in a total of 3,728 news stories. ChineseNewsSameStory, used to adjust the cosine similarity threshold θ for the story discovery task, contains 33,503 news pairs collected from internet platforms. Each pair has a label used to determine whether the news pair belongs to the same news story.

[0110] 3.1.2 Experimental Setup

[0111] The present invention uses the Pytorch framework to implement the algorithm, and is configured on a computer platform with an AMD EPYC 7352 24-core processor and an NVIDIA RTX A6000 graphics processing unit. In the story discovery task, the present invention refers to the settings of Yoon et al. and sets the time slice size to 1d. After multiple experiments on the dataset ChineseNewsSameStory, the similarity threshold θ is set to 0.85. In the candidate story clustering, the step size of the elbow method to increase the distance threshold is set to 0.05. Time expression extraction uses the named entity recognition of the spaCy toolkit, and the semantic parsing tool of JioNLP is used to parse it into standard time. The place name recognition tool of JioNLP extracts place name entities. In the context generation task, the linear programming method of the POT toolkit is used to calculate the optimal spatiotemporal transmission distance, and the time distance decay coefficient The spatial distance attenuation coefficient δ is set to -100, the temporal association and spatial association weights α and β are both set to 1.

[0112] The method of the present invention is compared with two story discovery baseline methods. Among them, Story Forest adopts the default parameter settings. It realizes story context generation through two community divisions and one story assignment. The statistical results of the story collection are used for quantitative evaluation of story discovery. SCStory adopts self-supervised attention pooling sentence embedding and realizes story discovery through one story assignment. In order to avoid the performance differences caused by different pre-trained text embedding models, the present invention uses the M3E model for sentence embedding of SCStory, and the similarity threshold is set to 0.85. The rest are all based on the default settings of SCStory.

[0113] To verify the effectiveness of the proposed method, quantitative evaluation is used in the story discovery task. The experiment selects Adjusted Mutual Information (AMI), Adjusted Rand Index (ARI) and BCubed F1-Score (B 3 -F1) to evaluate the performance of story discovery. Combining these three evaluation indicators can comprehensively evaluate the quality of story discovery from multiple perspectives, making up for the shortcomings of a single indicator.

[0114] AMI is an adjustment of mutual information. It measures the amount of information shared between the story discovery prediction result U and the true story label V. It is suitable for the overall evaluation of story discovery results. A higher AMI indicates that the predicted clustering result can better capture the true story grouping of news data. The calculation formula is:

[0115]

[0116] Where: MI(U,V) is the mutual information between the story discovery results U and V, E(MI(U,V)) is the expected mutual information between U and V, H(U) and H(V) are the entropies of U and V respectively. The calculation formulas of MI(U,V) and H(U) are:

[0117]

[0118] Where: P(i) is the ratio of the number of news documents containing the i-th story in U to the total number of news documents, P(j) is the ratio of the number of news documents containing the j-th story in V to the total number of news documents. P(i,j) is the ratio of the number of news documents contained in the intersection of the i-th story in U and the j-th story in V to the total number of news documents. |U| and |V| are the number of stories in the story discovery results U and V, respectively.

[0119] ARI is an adjustment to the Rand Index (RI), focusing on the matching of paired news events. A higher ARI value indicates that the predicted story and the actual story have a high degree of consistency in the matching relationship between paired news events. The formula for calculating ARI and RI is:

[0120]

[0121] Where: E(RI) is the expected Rand index, TP is the number of news document logs where the prediction and the true value are the same story, TN is the number of news document logs where the prediction and the true value are different stories, FP is the number of news document logs where the prediction is the same story but the true value is different stories, and FN is the number of news document logs where the prediction is different stories but the true value is the same story.

[0122] B 3 -F1 is the harmonic mean of BCubed Precision (P) and BCubed Recall (R), which comprehensively considers the precision and recall of the story discovery prediction results. 3 -F1 indicates that for each news document, the predicted clustering result can ensure that most of the news documents in the story it is assigned to belong to the same story, and also ensure that most of the news documents belonging to the same story as the news document are assigned to the same story, thus achieving both high accuracy and completeness. The calculation formula is:

[0123]

[0124]

[0125] Where: U(i) is the predicted story containing news document i; V(i) is the true story containing news document i; N is the total number of news documents in the news article corpus.

[0126] For the context generation task, manual evaluation and qualitative visual analysis were used for comparison. Since SCStory lacks a context generation step, the maximum spanning tree was used to generate the story context based on the cosine similarity of keyword distribution vectors.

[0127] 3.2 Experimental results and analysis

[0128] 3.2.1 Story Discovery Results and Analysis

[0129] Table 2 lists the quantitative evaluation results of story discovery of the proposed method and the baseline method on the ChineseNewsEvents dataset, where the bold values ​​represent the optimal values ​​and the underlined values ​​represent the suboptimal values. The experimental results show that the proposed method outperforms the other baseline methods by 0.147 and 0.103 in AMI and ARI respectively, and 3 -F1 is similar to the performance of SCStory. The main reason is that the present invention uses agglomerative hierarchical clustering for the keyword distribution vectors of candidate stories. Compared with the local iterative optimization based on the spherical distribution assumption of SphericalKmeans, agglomerative hierarchical clustering constructs a hierarchical structure by calculating the distance relationship of all sample pairs from the bottom up, which can more accurately capture the global relationship between news articles. Therefore, it scores higher on global evaluation indicators such as AMI and ARI. However, agglomerative hierarchical clustering is more sensitive to small clusters or boundary areas, and is not as good as SphericalKmeans used by SCStory in handling local similarity relationships. Therefore, it is not as effective in evaluating B for local relationships. 3 -F1 performance failed to surpass SCStory. Furthermore, using TF-IDF to extract keywords and construct keyword distribution vectors ignores the semantic information between words and cannot handle synonyms or polysemy, which may affect story discovery performance.

[0130] Table 2 Comparison of story discovery performance

[0131]

[0132]

[0133] 3.2.2 Context Generation Results and Analysis

[0134] In order to verify the effectiveness of the method of the present invention in the task of generating story threads, the present invention visualizes the story threads generated by different methods, and qualitatively analyzes the expression effect of each method on the development and evolution of news stories. The method of the present invention and the baseline method were verified on the theme story of "Typhoon Aere in 2016". The method of the present invention has fewer omissions and incorrect mergers of events, and the edge relationships of the story threads are the most coherent in the spatiotemporal dimension, showing the development of Typhoon Aere in the eastern waters of Hainan on the 7th, affecting Chaozhou and eastern Guangdong to the northeast on the 8th, and turning back to the southwest on the 9th to affect the Pearl River Delta. The edge relationships constructed for SCStory using keyword similarity and news timestamps also show the development of Typhoon Aere, but compared with the method of the present invention, the edge relationships are looser in the spatial dimension, which illustrates the advantage of the present invention in using optimal transmission to calculate spatiotemporal associations. Furthermore, SCStory's story discovery tended to generate more stories related to Typhoon Aere. It not only correctly assigned events like "entering the South China Sea on the 6th" and "heavy rain in Fujian on the 9th," which were missed by our method, but also mistakenly merged other events, such as the disasters in Nanchang and Wenzhou caused by Typhoon Megi and the damage to Florida caused by Hurricane Matthew, into the news story about Typhoon Aere. Story Forest, on the other hand, assigned the news event about Typhoon Aere to different stories, failing to generate a complete storyline.

[0135] In summary, the present invention starts from document-level news events and proposes a news story context generation method that takes into account the temporal and spatial associations. The method first assigns semantically related news articles to the same news story based on the overall semantics of the article and the distribution of keywords. In order to make full use of the multiple temporal and spatial attributes of news events, the temporal and spatial associations between events are mined, the temporal and spatial information of each news article is extracted and parsed, the temporal and spatial associations between paired news events are calculated by the optimal transmission method and the distance decay function, and the maximum spanning tree is used to generate a story context with the maximum temporal and spatial association weight. The present invention conducts quantitative experiments and manual evaluation on a public Chinese news dataset. The results show that in story discovery, the method of the present invention improves by 0.147 and 0.103 in clustering indicators AMI and ARI compared with the baseline method, respectively, and in B 3 -F1 achieved comparable performance. In terms of context generation, the story context constructed by the method of the present invention performed better in terms of relevance, accuracy and correlation. The visualization results show that the story context generation results of the present invention more accurately express the real evolution process of the news story than the baseline method. Taking Typhoon "Aere" in 2016 as an example, an ablation experiment analysis was conducted on the story context generation, and it was found that the story context based on the optimal space-time transmission distance showed a more coherent development process in the space-time dimension. This study provides an effective space-time calculation method for news story context generation, which can provide technical support for the analysis of the space-time evolution process of news stories based on maps.

[0136] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for generating news story context that takes into account temporal and spatial correlation, characterized in that: include: We perform document-level embedding on news articles, assign related news from different time periods to the same candidate story based on embedding vector similarity, and use agglomerative hierarchical clustering to discover stories based on the distribution vectors of news article keywords within the candidate stories. For each news article, we extract time expressions and place name entities through named entity recognition and rule matching, and use the entity linking model GENRE and Wikidata to parse the place name entities into location coordinates including latitude and longitude; Using the acquired time and location coordinates, the spatiotemporal correlation between news events is calculated through the optimal transmission distance and the corresponding attenuation function, and the maximum spanning tree algorithm is used to construct the story context with the maximum spatiotemporal correlation.

2. The method for generating news story context taking into account the temporal and spatial associations according to claim 1, characterized in that: The method of assigning related news of different time slices to the same candidate story based on embedding vector similarity includes: For the w i The news documents contained in the time slice are calculated, the cosine similarity between the news text embeddings is calculated and compared with the similarity threshold θ, and the news documents with similarity greater than θ are assigned to the same news story. The average pooling of these news text embeddings is used as the story embedding; then, the w-th time slice is calculated. i A time piece of new discovery story and the former w i-1 The cosine similarity between the embeddings of the discovered stories in the w-th time slice is calculated and compared with the similarity threshold θ. The newly discovered stories with a similarity greater than the similarity threshold θ are merged into the discovered stories, and the merged story embedding is updated by weighted pooling according to the number of news articles of the newly discovered stories and the discovered stories. i The same steps are used to process the wth time slice i+1 time slices until all time slices are processed and a set of candidate stories is obtained.

3. The method for generating news story context taking into account the temporal and spatial associations according to claim 1, characterized in that: The news article keyword distribution vector is constructed as follows: Where: v j is the value of the jth dimension of the news article keyword distribution vector v, k is a keyword among all keywords K, k s is the TF-IDF value corresponding to k, L j Represents the jth keyword in the vocabulary L.

4. The method for generating news story context taking into account the temporal and spatial associations according to claim 1, characterized in that: The method of discovering stories using agglomerative hierarchical clustering based on news article keyword distribution vectors in candidate stories includes: First, a fixed step size is set and the distance threshold is gradually increased, and the number of clusters under different distance thresholds is recorded. Then, the optimal distance threshold is determined by maximizing the second-order derivative of the cluster number. Finally, based on the optimal distance threshold, the keyword distribution vectors of the news articles in the candidate stories are clustered in an agglomerative hierarchical clustering manner. By clustering all candidate stories, a set of news stories is finally obtained, completing the story discovery task.

5. The method for generating news story context taking into account the temporal and spatial associations according to claim 1, characterized in that: Extract the time expression in the news article as follows: Perform named entity recognition on news articles, treating named entities of type "DATE" and "TIME" as time expressions; Use the publishing timestamp of each news article as the base time and parse the time string by matching predefined rules with regular expressions; Preserve the time points and time ranges in the parsing results; The reciprocal of the time interval days is used as the weight score of the standard time, and the Softmax function is used to convert the weight score into the probability score of the standard time in the news article.

6. The method for generating news story context taking into account the temporal and spatial associations according to claim 1, characterized in that: Place name entity extraction and parsing are performed in the following manner: Based on the word segmentation results of the article, the words with the part of speech of place names are counted. Based on the statistical results and regular matching with the place name dictionary, place name entities are extracted from the news article and given confidence scores. Using GENRE, the corresponding Wikidata entity ID is obtained based on the input place name entity. The coordinate attributes of the entity ID are queried using the online Wikidata query service interface. Finally, the Softmax function is used to convert the place name entity confidence score into the probability score of the place name coordinate in the news article.

7. The method for generating news story context taking into account the temporal and spatial associations according to claim 1, characterized in that: The optimal transmission distance includes a time-optimal transmission distance and a space-optimal transmission distance; The time cost function is constructed as follows: Where: TempCost(t i ,t j ) represents the time from standard time t i to t j The time cost function, and The standard time t i and t j The start time, and t i and t j The end time, It is t i and t j The absolute value of the difference in days between the start times, It is t i and t j The absolute value of the difference in days between the end times; Based on the constructed time cost function, linear programming is used to calculate the time-optimal transmission distance; The Haversine distance between the location coordinate points is used as the spatial cost function. Based on the constructed spatial cost function, linear programming is used to calculate the optimal spatial transmission distance.

8. The method for generating news story context taking into account the temporal and spatial associations according to claim 7, characterized in that: The attenuation function includes a time distance attenuation function and a space distance attenuation function; The time distance decay function is defined as: Where: Tdist is the time-optimal transmission distance, is the time distance attenuation coefficient; The spatial distance decay function is defined as: Where: Sdist is the optimal spatial transmission distance, and δ is the spatial distance attenuation coefficient.

9. The method for generating news story context taking into account the temporal and spatial associations according to claim 8, characterized in that: The spatiotemporal correlation between news events is calculated as follows: The time distance decay function value and the space distance decay function value are respectively used as the time correlation and space correlation between news events; For each news story s, calculate the pair of news events e according to the optimal transmission distance and attenuation function. i and e j The time relationship between TempRel(e i ,e j ) and spatial association SpatRel(e i ,e j ), and get the time correlation matrix A of story s t and the spatial incidence matrix A s , the weighted sum of the two is α·A t +β·A s As the spatiotemporal association weight matrix A of story s.

10. The method for generating news story context taking into account the temporal and spatial associations according to claim 1, characterized in that: When constructing a story thread with maximum spatiotemporal association using the maximum spanning tree algorithm, each element a in A ij As an event i and e j The edge weights between them are compared by comparing the news events e i and e j The maximum probability start time is used to set the edge direction of the spanning tree to complete the story line generation based on spatiotemporal association.