News background relevance intelligent analysis method and system based on knowledge graph

By constructing an intelligent analysis method for news background correlation based on knowledge graphs, the problem of lack of causal logic in the news recommendation system is solved, deep logical correlation matching and personalized recommendation of news events are achieved, and the correlation and interpretability of news information are improved.

CN120234424AInactive Publication Date: 2025-07-01NANJING FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510293664.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing news recommendation system lacks causal logic and it is difficult to automatically construct the causal relationship between news events, resulting in the recommended news that cannot help users establish a complete awareness of the event background.

Method used

By constructing an intelligent analysis method for news background relevance based on knowledge graphs, direct causal relationships, indirect causal relationships and counterfactual causal relationships are established, the causal reasoning model is used to analyze the causal relationships of news events, and personalized background news recommendations are made based on user portraits.

Benefits of technology

It realizes the deep logical correlation matching of news events, can personalize recommendations according to the user's needs level, improves the relevance and interpretability of news information, and meets users' needs for high-quality news reading experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234424A_ABST
    Figure CN120234424A_ABST
Patent Text Reader

Abstract

The invention discloses a news background relevance intelligent analysis method and system based on a knowledge graph, and the method comprises the steps: S1, collecting a news text from a plurality of news sources, and generating a standardized news text; s2, obtaining structured news information; s3, constructing a news knowledge graph under the unified naming system based on the structured news information; s4, constructing a news event causal reasoning model by utilizing the news knowledge graph, and generating a causal relationship set; s5, collecting news reading history, interest preference and behavior data of the user to construct a personalized user portrait, and analyzing a demand level of the user for news background information based on the user portrait; and S6, in combination with the causal relationship set and the personalized user portrait, a recommendation algorithm is adopted to sort the background news information with causal reasoning support according to user demand levels, and then the sorted background news information is pushed for use. By constructing the direct causal relationship, the indirect causal relationship and the anti-fact causal relationship, deep logical association matching of news events is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of news technology, and in particular, to an intelligent analysis method and system for news background relevance based on a knowledge graph. Background Art

[0002] With the accelerating development of Internet news dissemination, the news reading mode has gradually evolved from traditional linear browsing to personalized recommendation and intelligent analysis. The acquisition of news information is no longer limited to a single source, and users usually obtain news content through multiple channels such as news websites, social media, and government announcements.

[0003] Currently, mainstream news retrieval and recommendation systems mainly rely on keyword matching, text similarity calculation based on statistical models, or content recommendation based on deep learning. However, traditional methods have the following problems in the intelligent relevance analysis of news background information:

[0004] First of all, existing news recommendation methods lack logical relevance. Traditional news recommendation systems mainly push content based on keyword matching of news titles or texts. Although they can provide similar news to a certain extent, these news often only have lexical similarity and cannot reflect the causal relationship of events. For example, when a user reads a news article about a company's bankruptcy, a common recommendation system may recommend articles with the keyword "bankruptcy" without considering whether the company's bankruptcy is related to market trends, management decisions, or policy changes. Due to the lack of causal reasoning, the recommended news often cannot help users establish a complete understanding of the event background.

[0005] Secondly, existing technologies have limitations in the personalized recommendation of news background information. The background information provided by most news platforms is usually standardized, such as linking to encyclopedia entries or historical event timelines. However, different users have different needs for background information. For example, ordinary readers may need concise background information, while financial analysts or policy researchers may need more detailed data support and in-depth interpretation. Existing technologies are difficult to dynamically adjust the presentation of background information according to users' reading habits and interests, resulting in lack of pertinence and depth in information recommendation.

[0006] In addition, existing methods are difficult to automatically construct the causal relationship between news events. Although knowledge graph technology has been applied to news data analysis to a certain extent, most current methods mainly rely on static entity relationship mining and lack the ability of causal reasoning for the evolution of news events over time. Existing methods are difficult to capture the dynamic changes of news events and deduce the causal chain of events at different time points. For example, the release of a policy may affect multiple industries, but without an effective causal reasoning mechanism, the recommendation system is difficult to automatically mine these implicit relationships, thus affecting the relevance and interpretability of news information.

[0007] In summary, there are obvious deficiencies in the existing technology in the intelligent analysis of news background information, including the lack of causal logic in news recommendations, limited ability to recommend personalized background information, and the difficulty in automatically constructing the causal relationship of news events, which makes it difficult to meet the user's demand for a high-quality news reading experience. Summary of the Invention

[0008] An object of the present invention is to propose an intelligent analysis method and system for news background relevance based on a knowledge graph. The present invention realizes deep logical association matching of news events by constructing direct causal relationships, indirect causal relationships, and counterfactual causal relationships.

[0009] An intelligent analysis method for news background relevance based on a knowledge graph according to an embodiment of the present invention includes the following steps:

[0010] S1. Collect news texts from multiple news sources and generate standardized news texts;

[0011] S2. Perform word segmentation, named entity recognition, relationship extraction, and event extraction on the standardized news texts to obtain structured news information;

[0012] S3. Construct a news knowledge graph under a unified naming system based on the structured news information;

[0013] S4. Use the news knowledge graph to construct a causal reasoning model for news events, and perform causal relevance analysis on the news events corresponding to the target news to generate a set of causal relationships;

[0014] S5. Collect the news reading history, interest preferences, and behavior data of users to construct a personalized user profile, and analyze the user's demand level for news background information based on the user profile;

[0015] S6. Combine the set of causal relationships with the personalized user profile, and use a recommendation algorithm to sort and push the background news information supported by causal reasoning to users according to the user's demand level.

[0016] Optionally, the S1 includes the following steps:

[0017] S11. Collect news text data from news websites, social media, government announcements, and industry reports. The collected news text data includes news titles, news texts, release times, news sources, and additional multimedia information;

[0018] S12. Perform format conversion on the collected news text data to make the news text data from different sources consistent in character encoding, punctuation format, line break processing, and special character encoding, and uniformly convert it to a preset standard encoding format;

[0019] S13. Perform duplicate removal processing on the news text data based on the text hashing matching algorithm, calculate the hash value of each news text data for detecting redundant data. If the hash values of two news text data are the same, they are determined as duplicate news and merged or deleted.

[0020] S14. Clear the noise information in the news text data, where the noise information includes advertisement text, hyperlinks, unstructured tags, redundant punctuation, and garbled characters.

[0021] S15. For the collected news text data, perform text tokenization, lemmatization, stop word removal, and synonym replacement, and optimize the text representation using a language standardization algorithm based on a statistical model to generate standardized news text.

[0022] Optionally, the S2 includes the following steps:

[0023] S21. Perform tokenization processing on the standardized news text, parse the word boundaries in the news text, and optimize the tokenization result through the maximum probability path method to maximize the occurrence probability of the tokenization sequence W of the news text:

[0024]

[0025] where T represents the news text, w i represents the i-th tokenization unit, P(W|T) represents the probability of the text T under the tokenization sequence W, and n is the number of news texts.

[0026] S22. Perform named entity recognition on the tokenized news text, extract entities including people, organizations, place names, times, policies, and industry terms, and assign category labels to each entity to define a named entity set.

[0027] S23. Extract the relationships between named entities, identify the association relationships R between entities in the news text, where the association relationships include subordination relationships, cooperation relationships, competition relationships, and causal relationships.

[0028] S24. Extract events from the news text, where the events include policy announcements, market changes, corporate mergers and acquisitions, and emergencies.

[0029] S25. Based on the tokenization processing, named entity recognition, relationship extraction, and event extraction results in steps S21 - S24, construct a structured news information data set and parse the news text into standardized knowledge units.

[0030] Optionally, the S3 includes the following steps:

[0031] S31. Based on the standardized knowledge units and structured news information, for each news event node E t :

[0032] E t =(e a ,e b ,r c ,T d ,L e );

[0033] Among them, e a and e b respectively represent the subject and object of the event, r c represents the event type, T d represents the event occurrence time, and L e represents the event occurrence location;

[0034] Design a dynamic node embedding formula to achieve the joint representation of multi-dimensional features of news events:

[0035] v Et =σ(W e ·f(e a ,e b )+W r ·g(r c )+W t ·h(T d )+W l ·j(L e )+b);

[0036] Among them, represents the embedding vector of the news event E t , used to capture the comprehensive semantic features of the event, f(e a ,e b ) is a feature extraction function based on the event subject information, which jointly maps e a and e b to the vector space, g(r c ) is an event type feature mapping function, h(T d ) is a time information encoding function, which converts the event occurrence time T d into a vector through timestamp conversion and periodic feature extraction, j(L e ) is a location information embedding function, W e ,W r ,W t ,W l are the weight matrices of the above features respectively, b is the bias vector, and σ is the activation function;

[0037] S32. For any two news events and Design an evolution propagation scoring formula to capture the temporal and semantic transmission relationships between news events:

[0038]

[0039] Among them, represents the evolution propagation score from news event to , and respectively represent the occurrence times of news events and , is the time difference between the two, λ is the time decay constant, which regulates the impact of the time difference on the evolution score, is the cosine similarity between event embedding vectors, used to measure the semantic similarity of events, and η is the similarity influence index, used to adjust the weight of semantic similarity in the score;

[0040] S33. Construct an enhanced matching function for the matching degree between each news event E t and external data D k :

[0041]

[0042] Among them, M ext (E t , D k ) represents the matching degree between news event E t and external data D k , and the value range is (0, 1), is the embedding vector of external data D k , generated by a dedicated text encoder, is the cosine similarity between the two vectors, reflecting the text semantic similarity, and are respectively the logical relationship vectors of news event E t and external data D k , generated by the event trigger word and relationship extraction module, κ is the logical relationship weight factor, and μ is the scale adjustment parameter of the matching degree;

[0043] S34. Based on the event evolution score, design a causal inference probability formula:

[0044]

[0045] Among them, represents the probability that news event becomes the causal precursor of news event , α is the evolution score emphasis index, controlling P evolInfluence in causal reasoning represents the set of all subsequent events that have an evolutionary association with the event , β is the attenuation coefficient of semantic and logical differences represents the news event and The comprehensive semantic and logical relationship difference is defined as:

[0046]

[0047] Among them represents the Euclidean distance between event embedding vectors, measuring the semantic difference between events, and λ r is the weight factor of the logical relationship difference

[0048] Optionally, the S4 includes the following steps:

[0049] S41. Based on the news knowledge graph, use the event embedding vector to calculate the semantic similarity between the news event and :

[0050]

[0051] Among them, M represents the number of dimensions of the semantic feature space and respectively represent the representations of the news event in multiple dimensions of the semantic space, and ω m is the weight factor of different semantic dimensions m. cos(·,·) calculates the cosine similarity of vectors, indicating the degree of semantic proximity of news events. The larger the value, the more similar the event content is

[0052] S42. Based on the causal reasoning model, calculate the probability P that the news event has a direct causal relationship with dir ;

[0053] S43. Based on direct causal reasoning, calculate the probability P evol that the news event indirectly causes the news event to occur through other news events ind :

[0054]

[0055] Among them represents that the event indirectly causes the event The probability of occurrence, P represents the number of propagation paths from the news event to , and Q p represents the number of events on path p. Generally, the shorter the path, the stronger the causal relevance. P dir (E q-1 , E q ) represents the direct causal probability between adjacent events in the path;

[0056] S44. Calculate the counterfactual causal relationship C sem based on the event similarity R and semantic change cf :

[0057]

[0058] Among them, represents the impact amount of the news event on the probability of occurrence, represents the probability of occurrence of the event when is forced to occur, represents the probability of occurrence of the event in the natural state;

[0059] S45. Generate a causal relationship set R causal based on the causal reasoning results of S42 - S44 and update the news knowledge graph:

[0060]

[0061] Among them, R represents the category of causal relationships, and θ1, θ2, and θ3 are the first threshold, the second threshold, and the third threshold respectively. represents the probability that the news event directly causes the news event .

[0062] Optionally, the S6 includes the following steps:

[0063] S61. Construct a personalized user profile based on the user's news reading history, interest preferences, and behavior data to generate a user profile vector u U ;

[0064] S62. Generate an embedding vector v of the candidate news for the candidate background news information set using the same embedding method as the news event I ;

[0065] S63. Calculate the causal support score F for the candidate background news I based on the causal relationship set R causal causal(I):

[0066]

[0067] wherein, R I represents a set of causal relationships associated with candidate background news I, δ is a direct causal weight factor with a value range in [0, 1], used to balance the contributions of direct and indirect causal relationships;

[0068] S64. Based on the matching degree between the candidate background news embedding vector and the personalized user profile vector and the causal support score F causal (I), use a recommendation scoring function to calculate the information recommendation score S of each candidate news rec (I, u U ):

[0069] S rec (I, u U ) = λ1 · cos(v I , u U ) + (1 - λ1) · F causal (I);

[0070] wherein, S rec (I, u U ) represents the recommendation score of candidate background news I for user U, cos(v I , u U ) is the cosine similarity between the candidate news and the user profile vector, measuring the tightness of the match between the news content and the user's interests, and λ1 is the balance coefficient of user needs and causal support in the recommendation score, with a value range in [0, 1];

[0071] S65. Sort each candidate news in the candidate background news information set according to the recommendation score S rec (I, u U ), and push the sorted news information to the user according to the user's need level.

[0072] A news background relevance intelligent analysis system based on a knowledge graph, used to execute a news background relevance intelligent analysis method based on a knowledge graph, including the following modules:

[0073] A news data collection module, used to obtain news text data from multiple news sources, the news sources including news websites, social media, government announcements, and industry reports. The news data collection module formats, de-duplicates, clears noise, and performs language standardization processing on the collected news data, and generates standardized news text data;

[0074] A news information parsing module for performing text parsing on standardized news text data, including word segmentation, named entity recognition, relationship extraction, and event extraction. The news information parsing module extracts news events, news subjects, time, and location information from news texts based on natural language processing technology and generates structured news information;

[0075] A news knowledge graph construction module for constructing a news knowledge graph based on the structured news information. The news knowledge graph contains news events, news subjects, time, location, and their mutual relationships. The news knowledge graph construction module dynamically models news events in an event-driven manner and optimizes the relevance of news entities by combining a time decay mechanism;

[0076] A causal reasoning analysis module for performing causal reasoning analysis on news events in the news knowledge graph. The causal reasoning analysis module calculates the causal relevance of news events based on direct causal relationships, indirect causal relationships, and counterfactual causal relationships and generates a set of causal relationships. The causal reasoning analysis module uses a causal reasoning model to calculate the direct causal probability, indirect causal propagation path, and counterfactual intervention impact of events and stores the reasoning results in the news knowledge graph;

[0077] A personalized user profile construction module for collecting a user's news reading history, interest preferences, and behavior data and constructing a user profile based on the data;

[0078] A background news recommendation module for calculating the information recommendation scores of candidate background news in combination with the causal reasoning analysis results and the personalized user profile and sorting the background news according to the recommendation scores. The background news recommendation module determines the background news that best meets the user's needs by calculating the semantic similarity between the background news and the user profile and the causal support score of the news event and pushes it hierarchically.

[0079] The beneficial effects of the present invention are:

[0080] (1) Based on the news knowledge graph and causal reasoning method, the present invention realizes deep logical association matching of news events by constructing direct causal relationships, indirect causal relationships, and counterfactual causal relationships, calculates the correlation degree between news events using a causal reasoning model, and analyzes the upstream and downstream logic of events by combining causal propagation path analysis. It can analyze the historical evolution path of the policy, the affected industry dynamics, and key market changes based on the news knowledge graph, thereby intelligently supplementing background news with causal logic.

[0081] (2) The present invention combines user portrait modeling and personalized recommendation algorithms with the user's reading history, interest preferences, and behavior data to construct a personalized user portrait vector, and ranks personalized background news based on the causal inference support score. It can combine the user's preferences for different types of background information to achieve multi-level information recommendation, and ensure that the recommendation results can meet the needs of different users by introducing a user semantic preference matching item into the recommendation scoring function.

[0082] (3) The present invention introduces a time decay factor and dynamic event propagation path analysis in the news background recommendation process. It optimizes the timeliness of news background matching by calculating the time correlation and propagation impact of events, and ensures that the recommended content preferentially matches the latest and most influential events by calculating the time decay factor for candidate background news, while avoiding the interference of outdated information. Description of the Drawings

[0083] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0084] Figure 1 is a flowchart of an intelligent analysis method and system for news background relevance based on a knowledge graph proposed by the present invention. Detailed Embodiments

[0085] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.

[0086] Refer to Figure 1 , an intelligent analysis method for news background relevance based on a knowledge graph, includes the following steps:

[0087] S1. Collect news texts from multiple news sources and generate standardized news texts;

[0088] S2. Perform word segmentation, named entity recognition, relation extraction, and event extraction on the standardized news texts to obtain structured news information;

[0089] S3. Construct a news knowledge graph under a unified naming system based on the structured news information;

[0090] S4. Use the news knowledge graph to construct a news event causal inference model, and perform causal relevance analysis on the news events corresponding to the target news to generate a causal relationship set;

[0091] S5. Collect the user's news reading history, interest preferences, and behavior data to construct a personalized user portrait, and analyze the user's demand level for news background information based on the user portrait;

[0092] S6. Combine the causal relationship set with the personalized user profile, and use a recommendation algorithm to sort the background news information supported by causal reasoning according to the user's need hierarchy and then push it to the user.

[0093] In this embodiment, S1 includes the following steps:

[0094] S11. Collect news text data from news websites, social media, government announcements, and industry reports. The collected news text data includes news titles, news texts, release times, news sources, and additional multimedia information;

[0095] S12. Perform format conversion on the collected news text data to make the news text data from different sources consistent in character encoding, punctuation format, line break character processing, and special character encoding, and uniformly convert it to a preset standard encoding format;

[0096] S13. Perform duplicate removal processing on the news text data based on the text hash matching algorithm. Calculate the hash value of each news text data for detecting redundant data. If the hash values of two news text data are the same, they are determined to be duplicate news, and merge or deletion processing is performed;

[0097] S14. Clear the noise information in the news text data. The noise information includes advertisement texts, hyperlinks, unstructured tags, redundant punctuation, and garbled characters;

[0098] S15. For the collected news text data, perform text tokenization, lemmatization, stop word removal, and synonym replacement, and optimize the text representation using a language standardization algorithm based on a statistical model to generate standardized news text.

[0099] In this embodiment, S2 includes the following steps:

[0100] S21. Perform tokenization processing on the standardized news text, parse the word boundaries in the news text, and optimize the tokenization result through the maximum probability path method to maximize the occurrence probability of the tokenization sequence W of the news text:

[0101]

[0102] where T represents the news text, w i represents the i-th tokenization unit, P(W|T) represents the probability of the text T under the tokenization sequence W, and n is the number of news texts;

[0103] S22. Perform named entity recognition on the tokenized news text, extract entities including people, organizations, place names, times, policies, and industry terms, assign category labels to each entity, and define a named entity set;

[0104] S23. Extract the relationships between named entities, identify the association relationships R between entities in the news text, where the association relationships include subordinate relationships, cooperation relationships, competition relationships, and causal relationships;

[0105] S24. Extract events from the news text, where the events include policy announcements, market changes, corporate mergers and acquisitions, and emergencies;

[0106] S25. Based on the word segmentation processing, named entity recognition, relationship extraction, and event extraction results in steps S21 - S24, construct a structured news information dataset, and parse the news text into standardized knowledge units.

[0107] In this embodiment, S3 includes the following steps:

[0108] S31. Based on the standardized knowledge units and structured news information, for each news event node E t :

[0109] E t =(e a , e b , r c , T d , L e );

[0110] Among them, e a and e b respectively represent the subject and object of the event, r c represents the event type, T d represents the event occurrence time, and L e represents the event occurrence location;

[0111] Design a dynamic node embedding formula to achieve the joint representation of multi - dimensional features of news events:

[0112]

[0113] Among them, represents the embedding vector of the news event E t , which is used to capture the comprehensive semantic features of the event. f(e a , e b ) is a feature extraction function based on the event subject information, which jointly maps e a and e b to the vector space. g(r c ) is an event type feature mapping function, h(T d ) is a time information encoding function, which converts the event occurrence time T d into a vector through timestamp conversion and periodic feature extraction. j(L e ) is a location information embedding function, We , W r , W t , W l are the weight matrices of the above features respectively, b is the bias vector, and σ is the activation function;

[0114] S32. For any two news events and design an evolutionary propagation scoring formula to capture the temporal and semantic transmission relationships between news events:

[0115]

[0116] where represents the evolutionary propagation score from news event to , and represent the occurrence times of news events and respectively, is the time difference between them, λ is the time decay constant, which regulates the influence of the time difference on the evolutionary score, is the cosine similarity between event embedding vectors, used to measure the similarity of event semantics, and η is the similarity influence index, used to adjust the weight of semantic similarity in the score;

[0117] S33. Construct an enhanced matching function for the matching degree between each news event E t and external data D k :

[0118]

[0119] where M ext (E t , D k ) represents the matching degree between news event E t and external data D k , and its value range is (0, 1), is the embedding vector of external data D k , generated by a dedicated text encoder, is the cosine similarity between the two vectors, reflecting the text semantic similarity, and are the logical relationship vectors of news event E t and external data D k respectively, generated by the event trigger word and relationship extraction module, κ is the logical relationship weight factor, and μ is the scale adjustment parameter of the matching degree;

[0120] S34. Based on the event evolution score, design a causal inference probability formula:

[0121]

[0122] Among them, represents a news event becomes a news event The probability of the causal precursor, α is the evolution score emphasis index, controlling P evol 's influence in causal reasoning, C(E ti ) represents the set of all subsequent events that have an evolutionary association with event , β is the attenuation coefficient of semantic and logical differences, represents the news event and The comprehensive semantic and logical relationship difference, defined as:

[0123]

[0124] Among them, represents the Euclidean distance between event embedding vectors, measuring the semantic difference between events, λ r is the weight factor of logical relationship difference.

[0125] In this embodiment, S4 includes the following steps:

[0126] S41. Based on the news knowledge graph, use the event embedding vector Calculate the semantic similarity between the news event and

[0127]

[0128] Among them, M represents the number of dimensions of the semantic feature space, and respectively represent the representations of the news event in multiple dimensions of the semantic space, ω m is the weight factor for different semantic dimensions m, cos(·,·) calculates the cosine similarity of vectors, indicating the degree of semantic proximity of the news event, and the larger the value, the more similar the event content;

[0129] S42. Based on the causal reasoning model, calculate the probability P that the news event has a direct causal relationship with dir ;

[0130] S43. Based on direct causal reasoning, based on the event propagation path P evol Calculate the news event indirectly causes the news event Probability of occurrence \(P\) ind :

[0131]

[0132] wherein, represents the event indirectly causes event to occur through a number of intermediate events, \(P\) represents the number of propagation paths from news event to , \(Q\) p represents the number of events on path \(p\). The shorter the path, the stronger the causal relevance usually is. \(P\) dir (\(E\) q-1 , \(E\) q ) represents the direct causal probability between adjacent events in the path;

[0133] S44. Calculate the counterfactual causal relationship \(C\) sem based on the event similarity \(R\) and semantic change cf :

[0134]

[0135] wherein, represents the influence amount of news event on the occurrence probability, represents the occurrence probability of event when is forcibly set to occur, represents the occurrence probability of event in the natural state;

[0136] S45. Generate a causal relationship set \(R\) causal based on the causal reasoning results of S42 - S44 and update the news knowledge graph:

[0137]

[0138] wherein, \(R\) represents the category of causal relationships, and \(\theta_1\), \(\theta_2\) and \(\theta_3\) are the first threshold, the second threshold and the third threshold respectively, represents the probability that news event directly causes news event .

[0139] In this embodiment, S6 includes the following steps:

[0140] S61. Construct a personalized user profile according to the user's news reading history, interest preferences and behavior data, and generate a user profile vector \(u\) U ;

[0141] S62. Generate the embedding vector v of the candidate news by using the same embedding method as the news event for the candidate background news information set I ;

[0142] S63. Calculate the causal support score F causal (I) for the candidate background news I based on the causal relationship set R causal (I):

[0143]

[0144] where R I represents the set of causal relationships associated with the candidate background news I, and δ is the direct causal weight factor, with a value range in [0,1], used to balance the contributions of direct and indirect causal relationships;

[0145] S64. Calculate the information recommendation score S causal (I) for each candidate news by using a recommendation scoring function based on the matching degree between the candidate background news embedding vector and the personalized user profile vector and the causal support score F rec (I,u U ):

[0146] S rec (I,u U ) = λ1·cos(v I ,u U )+(1 - λ1)·F causal (I);

[0147] where S rec (I,u U ) represents the recommendation score of the candidate background news I for the user U, cos(v I ,u U ) is the cosine similarity between the candidate news and the user profile vector, measuring the tightness of the match between the news content and the user's interests, and λ1 is the balance coefficient of the user's needs and causal support in the recommendation scoring, with a value range in [0,1];

[0148] S65. Sort the candidate news in the candidate background news information set according to the recommendation score S rec (I,u U ), and push the sorted news information to the user according to the user's demand level.

[0149] A news background relevance intelligent analysis system based on a knowledge graph, used to execute a news background relevance intelligent analysis method based on a knowledge graph, includes the following modules:

[0150] A news data collection module is used to obtain news text data from multiple news sources, including news websites, social media, government announcements, and industry reports. The news data collection module formats, deduplicates, clears noise, and standardizes the collected news data, and generates standardized news text data;

[0151] A news information parsing module is used to perform text parsing on the standardized news text data, including word segmentation, named entity recognition, relation extraction, and event extraction. Based on natural language processing technology, the news information parsing module extracts news events, news subjects, time, and location information from the news text, and generates structured news information;

[0152] A news knowledge graph construction module is used to construct a news knowledge graph based on the structured news information. The news knowledge graph contains news events, news subjects, time, location, and their interrelationships. The news knowledge graph construction module dynamically models news events in an event-driven manner and optimizes the relevance of news entities by combining a time decay mechanism;

[0153] A causal reasoning analysis module is used to perform causal reasoning analysis on the news events in the news knowledge graph. The causal reasoning analysis module calculates the causal relevance of news events based on direct causal relationships, indirect causal relationships, and counterfactual causal relationships, and generates a set of causal relationships. The causal reasoning analysis module uses a causal reasoning model to calculate the direct causal probability, indirect causal propagation path, and counterfactual intervention impact of events, and stores the reasoning results in the news knowledge graph;

[0154] A personalized user profile construction module is used to collect the news reading history, interest preferences, and behavior data of users, and construct user profiles based on the data;

[0155] A background news recommendation module is used to combine the results of causal reasoning analysis and the personalized user profile, calculate the information recommendation scores of candidate background news, and rank the background news according to the recommendation scores. The background news recommendation module determines the background news that best meets the user's needs by calculating the semantic similarity between the background news and the user profile, as well as the causal support score of the news event, and pushes it hierarchically.

[0156] Example 1:

[0157] On June 1, 2023, a news portal published a news about the new policy on enterprise new energy subsidies, which quickly attracted wide attention from society. Preliminary monitoring data shows that within 6 hours after the news was released, the reading volume of the news reached 1.2 million times, among which users in the "new energy vehicle consumers" category accounted for 46.2%, users in the "financial market investors" category accounted for 38.7%, and users in the "policy researchers" category accounted for 15.1%.

[0158] Thirty minutes after the news release, the system detected a significant increase in the volume of relevant discussions about the news on social media. The main keywords included "New Energy Subsidy Adjustment", "Car Purchase Incentives", "Stock Market Reaction", and "Policy Impact". At the same time, multiple media outlets also successively published relevant reports, covering topics such as new energy vehicle sales forecasts, investment opportunities in the charging pile industry, and enterprises' long-term plans for the new energy industry.

[0159] The system processed relevant news from 12:00 to 18:00 on June 1, 2023, extracted key information, and constructed a news knowledge graph. The core events extracted included:

[0160] E1 (Policy Release): Adjustment of new energy subsidy policy (released by enterprises);

[0161] E2 (Market Reaction): Forecast of increased new energy vehicle sales (predicted by market analysts);

[0162] E3 (Industry Impact): Increased demand for charging piles (fluctuations in the stock prices of related enterprises);

[0163] E4 (Stock Market Reaction): Fluctuations in the stocks of the new energy industry chain (analyzed by the stock market);

[0164] The system automatically calculated the causal paths and obtained the following causal reasoning chain:

[0165] E1→E2 (Confidence level 0.92): After the release of the new energy subsidy policy, the market expects sales to increase;

[0166] E2→E3 (Confidence level 0.87): The increase in new energy vehicle sales drives up the demand for charging piles;

[0167] E3→E4 (Confidence level 0.78): The increase in the demand for charging piles affects the stock prices of related enterprises;

[0168] Within 6 hours after the news release, the system detected differences in the reading behaviors of different user groups:

[0169] Ordinary consumers (Target keywords: "How much is the new energy subsidy?", "Car purchase incentives?")

[0170] Reading behavior: Average stay time of 3 minutes and 12 seconds, 80% of users are interested in news related to policy interpretations;

[0171] Recommended news: "How Does the New Energy Subsidy Policy Affect Car Purchase Costs?", "2023 New Energy Vehicle Purchase Guide".

[0172] Investors (Target keywords: "Fluctuations in the new energy sector?", "How does the policy affect the stock market?")

[0173] Reading behavior: average stay time 5 minutes and 48 seconds, 60% of users click on market analysis news

[0174] Recommended news: "New energy subsidies drive car sales growth, which companies benefit?", "Analysis of investment opportunities in the new energy sector".

[0175] Policy researcher (target keywords: "new energy industry planning?" "enterprise subsidy trends?")

[0176] Reading behavior: average stay time is 7 minutes and 21 seconds, 90% of users read policy interpretation and historical review articles

[0177] Recommended news: "Review of New Energy Subsidy Policies over the Years", "The Economic Logic Behind New Energy Industry Policies".

[0178] In order to verify the effectiveness of the method of the present invention, the laboratory set up two groups of control experiments:

[0179] Experimental group: using the news background correlation analysis method based on causal reasoning of the present invention

[0180] Control group: using the traditional recommendation method based on TF-IDF keyword matching

[0181] During the experiment, the system tracked and analyzed the click behavior of 10,000 users and obtained the following results:

[0182] Evaluation Index Traditional Method Method of the Present Invention Improvement Rate Click-Through Rate (CTR) of Recommended News 12.3% 31.2% +18.9% Accuracy of Recommended News Matching 46.7% 72.9% +26.2% Average Reading Duration (seconds) 21.5 48.5 +26.9 seconds Coverage of User News Interests 41.3% 78.6% +37.3%

[0183] During the news dissemination process, the system continuously monitors user behavior data and dynamically adjusts causal reasoning parameters. Within 24 hours after the policy is released, ordinary users are more concerned about the impact of the subsidy policy on car purchases, while after 48 hours, market investors are more concerned about the market performance of the new energy industry chain. Therefore, the system automatically adjusts the recommendation model to make it easier for investors to obtain relevant market analysis news, while ordinary consumers can still get priority access to car purchase-related content.

[0184] This example verifies the practical application effect of the present invention in the intelligent analysis of news background information. The present invention can analyze the upstream and downstream logical relationships of news events based on causal reasoning, and provide personalized and dynamic news background recommendations in combination with user portraits. Experimental results show that the present invention is significantly superior to traditional methods in terms of news matching accuracy, user reading experience, and recommendation system effectiveness, providing a new technical path for intelligent recommendations in the news industry.

[0185] Based on a news knowledge graph and a causal reasoning method, the present invention realizes deep logical association matching of news events by constructing direct causal relationships, indirect causal relationships, and counterfactual causal relationships. It uses a causal reasoning model to calculate the association degree between news events, and combines causal propagation path analysis to identify the upstream and downstream logic of events. It can analyze the historical evolution path of the policy, the affected industry dynamics, and the key market changes based on the news knowledge graph, so as to intelligently supplement background news with causal logic.

[0186] The present invention combines user portrait modeling and personalized recommendation algorithms with the user's reading history, interest preferences, and behavior data to construct a personalized user portrait vector, and ranks personalized background news based on the causal reasoning support score. It can combine the user's preferences for different types of background information to achieve multi-level information recommendation, and ensure that the recommendation results can meet the needs of different users by introducing a user semantic preference matching item in the recommendation scoring function.

[0187] The present invention introduces a time decay factor and dynamic event propagation path analysis in the news background recommendation process, optimizes the timeliness of news background matching by calculating the time correlation and propagation influence of events, and ensures that the recommended content preferentially matches the latest and most influential events by calculating the time decay factor for candidate background news, while avoiding the interference of outdated information.

[0188] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A news background relevance intelligent analysis method based on knowledge graph, characterized in that: The steps include: S1. Collect news texts from multiple news sources and generate standardized news texts; S2. Perform word segmentation, named entity recognition, relationship extraction and event extraction on the standardized news text to obtain structured news information; S3. Build a news knowledge graph under a unified naming system based on structured news information; S4. Use the news knowledge graph to build a news event causal reasoning model, and perform causal correlation analysis on the news events corresponding to the target news to generate a causal relationship set; S5. Collect the user's news reading history, interest preferences and behavior data to build a personalized user portrait, and analyze the user's demand level for news background information based on the user portrait; S6. Combining the causal relationship set with personalized user portraits, a recommendation algorithm is used to sort the background news information supported by causal reasoning according to the user demand level and then push it to the user.

2. According to claim 1, a news background relevance intelligent analysis method based on knowledge graph is characterized in that: The S1 comprises the following steps: S11. Collect news text data from news websites, social media, government announcements and industry reports. The collected news text data includes news title, news text, release time, news source and additional multimedia information; S12. Convert the collected news text data to a format that makes the news text data from different sources consistent in character encoding, punctuation format, line break processing and special character encoding, and convert them uniformly into a preset standard encoding format; S13. De-duplicate the news text data based on the text hash matching algorithm, calculate the hash value of each news text data, and use it to detect redundant data. If the hash values ​​of two news text data are the same, they are determined to be duplicate news and merged or deleted; S14. Clearing noise information in the news text data, wherein the noise information includes advertising text, hyperlinks, unstructured tags, redundant punctuation and garbled characters; S15. For the collected news text data, perform text segmentation, word form normalization, stop word removal and synonym replacement, use the language standardization algorithm based on the statistical model to optimize the text representation, and generate standardized news text.

3. According to claim 1, a news background relevance intelligent analysis method based on knowledge graph is characterized in that: The S2 comprises the following steps: S21. Perform word segmentation processing on the standardized news text, parse the word boundaries in the news text, and optimize the word segmentation results through the maximum probability path method to maximize the probability of occurrence of the word segmentation sequence W of the news text: Where T represents the news text, w i represents the i-th word segmentation unit, P(W|T) represents the probability of text T under the word segmentation sequence W, and n is the number of news texts; S22. Perform named entity recognition on the news text after word segmentation, extract entities including people, institutions, place names, time, policies and industry terms, assign category labels to each entity, and define a named entity set; S23. extracting the relationship between named entities, identifying the association relationship R between entities in the news text, wherein the association relationship includes a subordinate relationship, a cooperative relationship, a competitive relationship, and a causal relationship; S24. extracting events from news texts, wherein the events include policy releases, market changes, corporate mergers and acquisitions, and emergencies; S25. Based on the word segmentation, named entity recognition, relationship extraction and event extraction results of steps S21-S24, a structured news information dataset is constructed to parse the news text into standardized knowledge units.

4. According to claim 1, a news background relevance intelligent analysis method based on knowledge graph is characterized in that: The S3 comprises the following steps: S31. Based on standardized knowledge units and structured news information, for each news event node E t : E t =(e a ,e b ,r c ,T d ,L e ); Among them, e a and e b Respectively represent the subject and object of the event, r c Indicates the event type, T d Indicates the time when the event occurred, L e Indicates the location where the event occurred; Design a dynamic node embedding formula to achieve the joint representation of multi-dimensional features of news events: in, Indicates news event E t The embedding vector is used to capture the comprehensive semantic features of the event, f(e a ,e b ) is a feature extraction function based on event subject information, e a With e b Jointly mapped to vector space, g(r c ) is the event type feature mapping function, h(T d ) is the time information encoding function, which converts the event occurrence time T into d Converted into a vector, j(L e ) is the location information embedding function, W e ,W r ,W t ,W l are the weight matrices of the above features, b is the bias vector, and σ is the activation function; S32. For any two news events and Design an evolutionary propagation scoring formula to capture the temporal and semantic transmission relationship between news events: in, Indicates news events and The evolutionary propagation score of and Represents news events and The time of occurrence, is the time difference between the two, λ is the time decay constant, and the influence of the time difference on the evolution score is regulated. is the cosine similarity between event embedding vectors, which is used to measure the semantic similarity of events. η is the similarity influence index, which is used to adjust the weight of semantic similarity in scoring. S33. For each news event E t With external data D k The matching degree between them is used to construct an enhanced matching function: Among them, M ext (E t ,D k ) indicates news event E t With external data D k The matching degree of is in the range of (0,1). For external data D k The embedding vector of is generated by a dedicated text encoder, is the cosine similarity between two vectors, reflecting the semantic similarity of texts. and News Event E t and external data D k The logical relationship vector is generated by the event trigger word and relationship extraction module, κ is the logical relationship weight factor, and μ is the scale adjustment parameter of the matching degree; S34. Based on the event evolution scoring, design the causal reasoning probability formula: in, Indicates news events Become a news event The probability of the causal predecessor, α is the evolution score emphasis index, and controls P evol The influence of causal reasoning, Represents all events There is a set of subsequent events with evolutionary associations, β is the attenuation coefficient of semantic and logical differences, Indicates news events and The comprehensive semantic and logical relationship difference is defined as: in, represents the Euclidean distance between event embedding vectors, measuring the semantic difference between events, and λ r is the weight factor of the logical relationship difference.

5. According to claim 1, a news background relevance intelligent analysis method based on knowledge graph is characterized in that: The S4 comprises the following steps: S41. Using event embedding vectors based on news knowledge graph Counting News Events and The semantic similarity between Among them, M represents the dimension of the semantic feature space, and Respectively represent the representation of news events in multiple dimensions of semantic space, ω m is the weight factor of different semantic dimensions m, cos(·,·) calculates the cosine similarity of the vectors, indicating the semantic proximity of news events. The larger the value, the more similar the event contents are; S42. Calculating news events based on causal inference models right The probability of direct causal relationship P dir ; S43. Based on direct causal reasoning, based on the event propagation path P evol Counting News Events Indirectly leading to news events through other news events The probability of occurrence P ind : in, Indicates an event Indirectly leading to an event through several intermediate events The probability of occurrence, P represents the probability of occurrence from the news event arrive The number of propagation paths, Q p represents the number of events on path p. The shorter the path, the stronger the causal relationship. dir (E q-1 ,E q ) represents the direct causal probability between adjacent events in the path; S44. Based on event similarity R sem and semantic changes Calculating counterfactual causality C cf : in, Indicates news events right The impact of the probability of occurrence, Indicates that the mandatory setting In case of occurrence, the event The probability of occurrence, Indicates an event Probability of occurrence in natural state; S45. Generate causal relationship set R based on the causal reasoning results of S42-S44 causal And update the news knowledge graph: Where R represents the category of causal relationship, θ1, θ2 and θ3 are the first threshold, the second threshold and the third threshold respectively. Indicates news events Directly lead to news events probability.

6. The news background relevance intelligent analysis method based on knowledge graph according to claim 1 is characterized in that: The S6 comprises the following steps: S61. Build a personalized user portrait based on the user's news reading history, interest preferences and behavior data, and generate a user portrait vector u U ; S62. Generate an embedding vector v for candidate news using the same embedding method as the news event for the candidate background news information set I ; S63. Based on causal relationship set R causal Calculate the causal support score F for the candidate background news I causal (I) Among them, R I represents the set of causal relationships associated with the candidate background news I, δ is the direct causal weight factor, which ranges from [0,1] and is used to balance the contribution of direct and indirect causal relationships; S64. Matching degree and causal support score F based on candidate background news embedding vector and personalized user portrait vector causal (I) The recommendation score function is used to calculate the information recommendation score S of each candidate news. rec (I,u U ): S rec (I,u U )=λ1·cos(v I ,u U )+(1-λ1)·F causal (I); Among them, S rec (I,u U ) represents the recommendation score of candidate background news I to user U, cos(v I ,u U ) is the cosine similarity between the candidate news and the user portrait vector, which measures the closeness of the match between the news content and the user's interest. λ1 is the balance coefficient between user demand and causal support in the recommendation score, and its value range is [0,1]. S65. According to the recommended score S rec (I,u U ) Sort the candidate news in the candidate background news information set, and push the sorted news information to the user according to the user demand level.

7. A news background relevance intelligent analysis system based on knowledge graph, used to execute a news background relevance intelligent analysis method based on knowledge graph according to any one of claims 1 to 6, characterized in that: Includes the following modules: A news data collection module is used to obtain news text data from multiple news sources, including news websites, social media, government announcements and industry reports. The news data collection module formats, removes duplicates, removes noise and performs language standardization on the collected news data, and generates standardized news text data; The news information parsing module is used to perform text parsing on standardized news text data, including word segmentation, named entity recognition, relationship extraction and event extraction. The news information parsing module extracts news events, news subjects, time and location information from news texts based on natural language processing technology, and generates structured news information; A news knowledge graph construction module is used to construct a news knowledge graph based on the structured news information. The news knowledge graph includes news events, news subjects, time, location and their mutual relationships. The news knowledge graph construction module uses an event-driven approach to dynamically model news events and optimizes the relevance of news entities in combination with a time decay mechanism. The causal reasoning analysis module is used to perform causal reasoning analysis on news events in the news knowledge graph. The causal reasoning analysis module calculates the causal correlation of news events based on direct causal relationships, indirect causal relationships, and counterfactual causal relationships, and generates a causal relationship set. The causal reasoning analysis module uses a causal reasoning model to calculate the direct causal probability, indirect causal propagation path, and counterfactual intervention impact of an event, and stores the reasoning results in the news knowledge graph; A personalized user portrait building module is used to collect the user's news reading history, interest preferences and behavior data, and build a user portrait based on the data; The background news recommendation module is used to combine the results of causal reasoning analysis and personalized user portraits to calculate the information recommendation scores of candidate background news and sort the background news according to the recommendation scores. The background news recommendation module determines the background news that best meets user needs by calculating the semantic similarity between background news and user portraits, as well as the causal support score of news events, and pushes them in layers.