A Topic Selection Method Based on Knowledge Graph-Based Heat Value Propagation Path Analysis

By constructing a knowledge graph to calculate composite heat values ​​and conducting causal inference analysis, the problems of blind spots in heat transmission paths and lack of causal insight in existing technologies have been solved, enabling in-depth tracing and prediction of news topics and generating more insightful topic suggestions.

CN122365411APending Publication Date: 2026-07-10DANGKANG DATA INTELLIGENCE TECHNOLOGY (GUANGDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DANGKANG DATA INTELLIGENCE TECHNOLOGY (GUANGDONG) CO LTD
Filing Date
2026-06-10
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies cannot effectively reveal the dynamic transmission relationship of popularity between entities, resulting in a lack of causal insight into topic selection, blind spots in transmission paths, and distorted trend judgments, making it difficult to identify deep connections and generate high-quality topics from massive news events.

Method used

We construct and maintain a domain knowledge graph, calculate composite heat values ​​through real-time data collection and entity linking, identify heat value propagation paths by combining causal inference analysis, and generate news topic suggestions using visualization and large language models.

Benefits of technology

It achieves accurate identification and quantification of heat conduction paths, discovers hidden cross-domain correlations, filters short-term noise, provides in-depth source tracing and prediction basis, and generates more insightful news topics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122365411A_ABST
    Figure CN122365411A_ABST
Patent Text Reader

Abstract

This invention discloses a method for topic selection based on the propagation path analysis of heatmap values ​​using knowledge graphs, relating to the fields of information technology and news media technology. The method includes: constructing and maintaining a domain knowledge graph containing entities and relationships; calculating a composite heatmap value for each entity, integrating real-time heatmap values ​​and periodic heatmap coefficients; starting with entities with high composite heatmap values, performing causal inference analysis based on the time series of heatmap values ​​of these entities and their associated entities, identifying and quantifying the heatmap value propagation path and its confidence level; mapping the composite heatmap value to node colors and the propagation path to edge styles for integrated visualization rendering; and generating optimized news topic suggestions through a rule engine and a large language model in collaboration with the propagation path characteristics. This invention solves the problems of lack of causal insight, blind spots in transmission paths, and distorted trend judgment in existing technologies, achieving a closed loop from hotspot tracking and causal tracing to intelligent topic selection, thus improving the depth and efficiency of news production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of information technology and news media technology, and in particular to a topic selection method based on the heat value propagation path analysis of knowledge graphs. Background Technology

[0002] In the information explosion era, quickly identifying trending topics from massive amounts of news events, understanding their underlying connections, and generating high-quality storylines are core needs for news media and content creators. Existing related technical solutions mainly fall into two categories: One type is a static knowledge graph popularity system based on entity co-occurrence frequency. It extracts entities and constructs a graph using natural language processing technology, and calculates popularity based on statistical indicators such as mention frequency, inverse document frequency, and source authority. However, its core drawback is that it is based solely on quantitative statistics and cannot reveal the dynamic transmission relationship of popularity between entities. The other type is a propagation analysis system based on social graphs. It constructs a propagation tree with users as nodes and social relationships as edges, and analyzes the propagation structure characteristics through graph traversal algorithms. However, it is limited to the user behavior dimension, detached from the semantic context of entities, and cannot explain the reasons for the popularity association of semantically related entities.

[0003] The aforementioned existing technologies have three major flaws: 1. Lack of causal insight, making it impossible to determine the causal relationship of entity popularity, resulting in topic selection remaining on the surface of phenomena; 2. Blind spots in transmission paths, making it impossible to discover cross-domain related topics, easily falling into homogeneous competition; 3. Distorted trend judgment, subject to short-term noise interference, making it difficult to identify long-term valuable topics.

[0004] In summary, existing technologies have shortcomings in deeply integrating dynamic propagation path analysis with semantically rich knowledge graphs. This results in editors being unable to understand the causal path of trend transmission, making it easy for topic selection to remain superficial and hindering in-depth tracing and forward-looking prediction. Summary of the Invention

[0005] The purpose of this invention is to solve the problems mentioned in the background art by proposing a topic selection method based on the heat value propagation path analysis of knowledge graphs.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A topic selection method based on knowledge graph-based heat value propagation path analysis, characterized by the following steps: S1. Construct and maintain a domain knowledge graph, which contains several entities and relationships between entities; continuously collect multi-source data, associate the data with the corresponding entities in the knowledge graph through entity linking technology, and update the entity attributes; S2. For each entity in the knowledge graph, calculate its composite heat value; the composite heat value is generated by fusing real-time heat values ​​and periodic heat coefficients; the real-time heat value is calculated based on the entity's dissemination volume, social interaction, search attention, and content feature indicators within a preset time window; the periodic heat coefficient is calculated based on the entity's trend persistence, content evolution degree, influence quality, and historical correlation indicators. S3. Starting with the current entity with high composite thermal value, obtain its associated entities in the knowledge graph, and obtain the thermal value time series of the current entity and each associated entity in the historical time period; perform causal inference analysis based on the thermal value time series, identify and confirm the thermal value propagation path from the source entity to the target entity, and calculate the confidence of each path. S4. Map the composite thermal value as a visual attribute to render the entity nodes in the knowledge graph, and use the propagation path and its confidence level identified in step S3 as a visual attribute to render the relationship edges between entities. S5. Based on the type, direction and confidence level of the propagation path identified in step S3, and combined with the type and heat state of the entities involved in the path, a basic topic selection framework is generated through a predefined topic selection rule engine. The basic topic selection framework is then semantically deepened and optimized based on a large language model, and the final news topic suggestions are output.

[0007] As a further aspect of the present invention: the calculation of the historical correlation degree in step S2 includes the following sub-steps: S21. Calculate the direct co-occurrence correlation degree: based on the co-occurrence frequency of the current entity and historical hot events in the recent time window, the salience weight of the historical events, and the time decay factor; S22. Calculate semantic similarity correlation: Calculate based on the similarity of the text embedding vectors of the current entity and historical hot events, as well as the degree of reactivation of historical events; S23. Calculate the correlation degree of periodic recurrence: Calculate based on the intensity of the periodic pattern of the current entity's historical thermal value time series, and the matching degree between the current time and the historical periodic points; S24. The direct co-occurrence correlation, semantic similarity correlation, and periodic recurrence correlation are weighted and summed to obtain the historical correlation.

[0008] As a further aspect of the present invention: the causal inference analysis in step S3 is specifically a Granger causality test, including: For the time series of thermodynamic values ​​of the current entity Eh and any related entity Ei, test the null hypothesis that "past values ​​of Ei do not contain information that helps predict future values ​​of Eh". If the test results reject the null hypothesis at the preset significance level, it is determined that there is a thermal value propagation path from Ei to Eh; The confidence level of the path is derived from the p-value of the test.

[0009] As a further aspect of the present invention: the causal inference analysis in step S3 employs dynamic time warping and peak lead-lag analysis, including: Calculate the dynamic time-warped distance between the current entity and the time series of thermal values ​​of related entities; Detect the peak points of the time series of both parties and calculate the peak time difference; If the dynamic time warping distance is less than a preset threshold, and the peak time difference indicates that the peak value of the associated entity is ahead of the peak value of the current entity, then it is determined that there is a propagation path from the associated entity to the current entity.

[0010] As a further aspect of the present invention: in step S4, the visualized attributes include: The color of the physical node is mapped in a preset color spectrum according to the level of composite thermal value; The style of the relationship edges between entities, including edge width, color, and / or dynamic particle flow effects, is set according to the confidence level and / or direction of the propagation path.

[0011] As a further aspect of the present invention: In step S5, the predefined topic selection rule engine matches different topic selection types according to the pattern of the propagation path, wherein the topic selection types include at least: Source tracing type topic selection: triggered when there is a propagation path from source entity Ei to high-popularity entity Eh with high confidence; Exploratory topic selection: triggered when there is a propagation path from high-heat entity Eh to potential beneficiary entity Ej and the heat value of Ej is on the rise; Comparative impact-based topic selection: triggered when there is a propagation path from a high-profile entity Eh to its competitor or related entity Ek.

[0012] As a further aspect of the present invention: in step S5, the semantic deepening and optimization based on the large language model includes the following inputs: the basic topic selection framework, the structured information of the dissemination path, the attribute information of relevant entities, and industry background knowledge; the large language model generates more insightful and expressive news topic titles and angles based on the inputs.

[0013] Compared with existing technologies, the advantages of this invention are as follows: it overcomes the problem of lack of causal insight in existing technologies, and achieves accurate identification and quantification of heat conduction paths through time-series causal testing, providing in-depth tracing and prediction basis for topic selection; it solves the problem of blind spots in conduction paths, and discovers hidden cross-domain correlations through the fusion analysis of knowledge graph topology and heat time series, thus expanding the perspective of topic selection; and it avoids distortion of trend judgment, and effectively filters short-term noise and highlights long-term valuable topics through the composite calculation of real-time and periodic heat indicators. Attached Figure Description

[0014] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] A topic selection method based on the propagation path analysis of heat values ​​using knowledge graphs is proposed. Its core lies in combining dynamic heat transmission causal analysis with semantically rich static knowledge graphs, and transforming data insights into actionable news creation strategies through an intelligent topic generation engine.

[0017] The technical solution of the present invention will be described in detail below with reference to an example of a "breakthrough in solid-state battery technology".

[0018] First, the system needs to construct and maintain a domain knowledge graph, which pre-contains entities related to the news domain (such as "solid-state battery," "lithium mine," "new energy vehicle," and "traditional lithium battery") and semantic relationships between entities (such as "raw material supply," "application field," and "technology substitution"). The system continuously collects data from multiple sources, including news websites, social media platforms, and search engines, and aggregates it through a distributed message queue (such as Apache Kafka). It then consumes this data using a stream processing engine (such as Apache Flink) and calls a pre-trained natural language processing model (such as a BERT-based entity recognition and disambiguation model) for real-time entity linking. For example, different expressions such as "lithium battery" and "Li-ionbattery" appearing in the text are uniformly linked to the "traditional lithium battery" entity node in the knowledge graph, and the relevant attributes and associated data streams of this node are updated in real time, laying the foundation for subsequent computations.

[0019] While updating the knowledge graph, the system calculates a comprehensive composite heat value for each entity. This value is generated by fusing real-time heat values ​​with periodic heat coefficients, aiming to simultaneously reflect the immediate intensity of heat and long-term value potential. The calculation of the real-time heat value is based on multi-dimensional indicators within a short time window (e.g., the past hour): including dissemination volume (e.g., the product of the number of news documents and media authority weights), social interaction (e.g., the weighted sum of mentions and reposts, taking into account the publisher's influence), search attention index (normalized), and content characteristics (e.g., sentiment polarity and suddenness indicators). These indicators are weighted, summed, and normalized according to preset weights to obtain the real-time heat value (H_r).

[0020] Real-time thermal value (H_r) calculation: Dissemination volume V: The number of non-duplicate news documents mentioning the entity within the window, multiplied by the average authority weight of the publishing media (e.g., central media = 1.0, portals = 0.8).

[0021] Social Interaction S: The weighted sum of mentions, reposts, likes, and comments on social media within the window, multiplied by the logarithm of the publisher's number of followers for influence weighting.

[0022] Search Attention E: The normalized value of the entity's search index in search engines and news clients within the window.

[0023] Content Feature C: The sentiment score of the content derived from a sentiment analysis model (such as TextCNN), and whether it contains sudden keywords.

[0024] Example of calculation formula: H_r=(α1*V+α2*S+α3*E+α4*C)*100, where α1-α4 are configurable weight coefficients.

[0025] The periodic heat coefficient (H_p) is calculated offline or through low-frequency batch processing to assess the quality and persistence of heat. Its calculation integrates trend persistence (the slope of regression analysis based on historical H_r sequences), content evolution (the number of associated new entities), influence quality (the proportion of authoritative sources), and historical relevance. Historical relevance is a comprehensive indicator, obtained by weighted summation of direct co-occurrence relevance (considering recent co-occurrence frequency, historical event saliency, and time decay), semantic similarity relevance (based on cosine similarity of text embedding vectors and the degree of reactivation of historical events), and periodic recurrence relevance (based on the periodic analysis results of the entity's own heat sequence).

[0026] Calculation of periodic thermal coefficient (H_p): Trend persistence T: Linear regression analysis was performed on the H_r time series over the past 30 days, and the normalized value of the trend slope was taken. The stronger the positive trend, the higher the score.

[0027] Content Evolution Degree D: This quantifies the number of new entities that establish a direct connection with the core entity in the knowledge graph for the first time within the past week, reflecting the topic's ability to generate derivative content.

[0028] Influence Quality Q: The proportion of mentions from authoritative sources (preset whitelist) to the total number of mentions.

[0029] Historical correlation R: This is a comprehensive indicator, and its calculation methods include: Direct co-occurrence correlation: Calculates the co-occurrence frequency of the current entity and each historical hot event in recent documents, weighted by the peak popularity and time decay factor of the historical events.

[0030] Semantic similarity correlation: Use a pre-trained language model (such as BERT) to obtain the semantic vectors of the current entity and historical events, calculate the cosine similarity, and combine it with the degree to which the historical event has been recently mentioned.

[0031] Periodic recurrence correlation: Fourier analysis is performed on the historical popularity sequence of an entity to detect periodic patterns and assess the degree of matching between the current time point and historical periodic points.

[0032] Finally, the correlation of the above three dimensions is weighted and summed to obtain the historical correlation R.

[0033] Example of calculation formula: H_p=β1*T+β2*D+β3*Q+β4*R, where β1-β4 are weighting coefficients.

[0034] Ultimately, the composite thermal value (H) is generated by combining H_r and H_p through multiplication and other methods, such as: H = H_r * H_p. This results in a short-term speculative entity (high H_r, low H_p) being downweighted, while a steadily rising entity with depth (medium to high H_r, high H_p) is highlighted.

[0035] When the combined thermal value of an entity (such as a "solid-state battery") reaches a high level, the system initiates a propagation path analysis process to explore the source of its heat and the direction of its influence. This process starts with the currently high-heat entity and queries a graph database (such as Neo4j) for related entities within its multi-hop range as candidate sources or targets. Subsequently, it extracts high-granularity thermal value time series of the current entity and each candidate entity over a historical time period (such as the past 24 hours) from a time-series database (such as InfluxDB). Based on these time series, the system uses a causal inference algorithm to verify the propagation relationship.

[0036] In a preferred embodiment, a Granger causality test is employed: for the sequence of the current entity "solid-state battery" (Eh) and the candidate entity "lithium mine" (Ei), the null hypothesis that "past values ​​of Ei do not contain information helpful in predicting future values ​​of Eh" is tested; if the p-value is less than the significance level (e.g., 0.05), the null hypothesis is rejected, and a propagation path from "lithium mine" to "solid-state battery" is determined, with the path confidence set to (1-p-value). In an alternative embodiment, a dynamic time warping algorithm can be used to calculate sequence similarity, combined with the time lead-lag relationship of peak points for judgment. The analysis result is a set of propagation paths including the source entity, target entity, propagation direction, and confidence level.

[0037] To visually represent the analysis results, the system employs integrated visualization rendering. Entity nodes are colored based on their composite thermal value (H) using a color interpolation function, mapping them to a continuous spectrum from cool colors (such as blue) to warm colors (such as red). For example, the "solid-state battery" node is rendered as deep red. Simultaneously, for propagation paths confirmed by causal analysis, special rendering is applied to the relationship edges of the knowledge graph: arrowed curves clearly indicate the direction of conduction; the visual width of the curve is proportional to the path confidence; and particle animations flowing along the arrow direction can be added to enhance the dynamic visual metaphor of "thermal conduction." This allows editors to easily identify hotspots, trace their origins, and predict their flow.

[0038] Finally, the system automatically generates news topic suggestions based on the identified propagation paths. This process employs a two-layer architecture that combines a rule engine and a Large Language Model (LLM). First, the rule engine matches predefined topic type templates based on the structured features of the propagation path. For example, when a high-confidence "lithium mine - solid-state battery" path is detected, a "source tracing" topic rule is triggered, generating the basic framework "Behind the [Solid-State Battery] Craze, the [Lithium Mine] Field is Quietly Changing"; when a "Solid-State Battery - Traditional Lithium Battery" path is detected and the latter is a competitor, a "comparative impact" rule is triggered, generating the basic framework "[Solid-State Battery] Emerges as a Dark Horse, [Traditional Lithium Battery] Faces Impact".

[0039] Then, the system combines this basic framework, relevant path details (entities, relationships, confidence levels, trends), and industry background knowledge into detailed prompts and inputs them into a large language model (such as GPT-4). Leveraging its deep semantic understanding capabilities, LLM further refines the basic framework semantically, enriches its perspectives, and optimizes its expression, ultimately outputting more insightful and widely disseminated topics, such as optimizing the aforementioned source-tracing framework into "Behind the Hot Concept of Solid-State Batteries, the Battle for Upstream Lithium Mineral Resources Quietly Escalates."

[0040] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A topic selection method based on knowledge graph-based heat value propagation path analysis, characterized in that, Includes the following steps: S1. Construct and maintain a domain knowledge graph, which contains several entities and relationships between entities; continuously collect multi-source data, associate the data with the corresponding entities in the knowledge graph through entity linking technology, and update the entity attributes; S2. For each entity in the knowledge graph, calculate its composite thermal value; the composite thermal value is generated by fusing real-time thermal values ​​and periodic thermal coefficients. The real-time heat value is calculated based on a weighted average of the entity's propagation volume, social interaction, search attention, and content feature indicators within a preset time window. The periodic thermal coefficient is calculated based on the entity's trend persistence, content evolution, influence quality, and historical correlation indicators. S3. Starting with the current entity with high composite thermal value, obtain its associated entities in the knowledge graph, and obtain the thermal value time series of the current entity and each associated entity in the historical time period. Causal inference analysis is performed based on the thermal value time series to identify and confirm the thermal value propagation path from the source entity to the target entity, and the confidence level of each path is calculated. S4. Map the composite thermal value as a visual attribute to render the entity nodes in the knowledge graph, and use the propagation path and its confidence level identified in step S3 as a visual attribute to render the relationship edges between entities. S5. Based on the type, direction and confidence level of the propagation path identified in step S3, and combined with the type and heat state of the entities involved in the path, a basic topic selection framework is generated through a predefined topic selection rule engine. The basic topic selection framework is then semantically deepened and optimized based on a large language model, and the final news topic suggestions are output.

2. The method for selecting topics based on the propagation path analysis of heat values ​​using knowledge graphs according to claim 1, characterized in that, The calculation of historical correlation degree in step S2 includes the following sub-steps: S21. Calculate the direct co-occurrence correlation degree: based on the co-occurrence frequency of the current entity and historical hot events in the recent time window, the salience weight of the historical events, and the time decay factor; S22. Calculate semantic similarity correlation: Calculate based on the similarity of the text embedding vectors of the current entity and historical hot events, as well as the degree of reactivation of historical events; S23. Calculate the correlation degree of periodic recurrence: Calculate based on the intensity of the periodic pattern of the current entity's historical thermal value time series, and the matching degree between the current time and the historical periodic points; S24. The direct co-occurrence correlation, semantic similarity correlation, and periodic recurrence correlation are weighted and summed to obtain the historical correlation.

3. The method for selecting topics based on the propagation path analysis of heat values ​​using knowledge graphs according to claim 1, characterized in that, The causal inference analysis described in step S3 specifically refers to the Granger causality test, which includes: For the time series of thermodynamic values ​​of the current entity Eh and any associated entity Ei, test the null hypothesis that "past values ​​of Ei do not contain information that helps predict future values ​​of Eh". If the test results reject the null hypothesis at the preset significance level, it is determined that there is a thermal value propagation path from Ei to Eh; The confidence level of the path is derived from the p-value of the test.

4. The method for selecting topics based on the propagation path analysis of heat values ​​using knowledge graphs according to claim 1, characterized in that, The causal inference analysis described in step S3 employs dynamic time warping and peak lead-lag analysis, including: Calculate the dynamic time-warped distance between the current entity and the time series of thermal values ​​of related entities; Detect the peak points of the time series of both parties and calculate the peak time difference; If the dynamic time warping distance is less than a preset threshold, and the peak time difference indicates that the peak value of the associated entity is ahead of the peak value of the current entity, then it is determined that there is a propagation path from the associated entity to the current entity.

5. The method for selecting topics based on the propagation path analysis of heat values ​​using knowledge graphs according to claim 1, characterized in that, In step S4, the visualized attributes include: The color of the physical node is mapped in a preset color spectrum according to the level of composite thermal value; The style of the relationship edges between entities, including edge width, color, and / or dynamic particle flow effects, is set according to the confidence level and / or direction of the propagation path.

6. The method for selecting topics based on the propagation path analysis of heat values ​​using knowledge graphs according to claim 1, characterized in that, In step S5, the predefined topic selection rule engine matches different topic types based on the pattern of the propagation path, and the topic types include at least: Source tracing type topic selection: triggered when there is a propagation path from source entity Ei to high-popularity entity Eh with high confidence; Exploratory topic selection: triggered when there is a propagation path from high-heat entity Eh to potential beneficiary entity Ej and the heat value of Ej is on the rise; Comparative impact-based topic selection: triggered when there is a propagation path from a high-profile entity Eh to its competitor or related entity Ek.

7. The method for selecting topics based on the propagation path analysis of heat values ​​using knowledge graphs according to claim 1, characterized in that, In step S5, the semantic deepening and optimization based on the large language model includes the following inputs: the basic topic selection framework, the structured information of the propagation path, the attribute information of relevant entities, and industry background knowledge. The large language model generates more insightful and expressive news headlines and angles based on the input.