Knowledge graph-combined context-aware text abstract generation method and system
By combining the knowledge graph to analyze quantitative statements in text and comparing them with authoritative data, a text summary with confidence labels is generated, which solves the problem of insufficient credibility assessment in the prior art, and realizes high-reliability text summary generation and dynamic optimization of knowledge graphs.
Patent Information
- Application Number
- CN202510415661.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing text summary technology lacks the credibility assessment ability of quantifiable statements and cannot verify its consistency with authoritative data. The conflict handling between multiple data sources and user feedback optimization may be insufficient, resulting in factual deviations in the summary results.
By combining the knowledge graph, the quantitative statements in the text are automatically parsed as the main body-predicate-value triplets, the pre-constructed knowledge graph is used to find authoritative data sources for comparison and verification, the deviation between the declared numerical value and the reference value is calculated to generate a credibility label, and a user feedback mechanism is introduced to dynamically adjust the data source weight or correct the knowledge base content.
It realizes the credibility marking of text summary, quickly identify exaggerated or error information, adapts to authoritative needs in different scenarios, and continuously optimizes the accuracy and reliability of the knowledge graph through user feedback.
Smart Images

Figure CN120336519A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of language processing, and specifically to a context-aware text summarization generation method and system combined with a knowledge graph. Background Art
[0002] With the explosive growth of Internet information, the automatic text summarization technology has become an important research direction in the field of information processing. Traditional text summarization methods mainly rely on statistical features or deep learning models, focusing on extracting key sentences of the text or generating concise semantic summaries. However, these methods usually lack the ability to evaluate the credibility of specific statements (such as numerical statements) in the text, resulting in possible factual biases or misleading information in the summary results.
[0003] In fields such as finance, healthcare, and technology, the text often contains a large number of quantifiable statements (such as "the company's annual revenue increased by 20%" or "the drug efficacy rate is 85%"), and their accuracy directly affects the quality of decision-making. Although existing technologies can extract such information, they cannot verify its consistency with authoritative data, nor can they dynamically adjust the credibility annotation of the summary. In addition, problems such as conflict handling between multiple data sources and closed-loop optimization of user feedback have not been effectively solved. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention provides a context-aware text summarization generation method and system combined with a knowledge graph.
[0005] To achieve the above purpose, the technical solution of the present invention is as follows:
[0006] In the first aspect, the present invention discloses a context-aware text summarization generation method and system combined with a knowledge graph. The context-aware text summarization generation method combined with a knowledge graph includes the following steps:
[0007] Obtain the original statement data D1 in the text to be summarized. The original statement data D1 includes at least one quantifiable statement item, and the quantifiable statement item includes a statement subject, a predicate, and a statement value;
[0008] According to the statement subject and the predicate, retrieve the matching knowledge graph data D2 in the pre-constructed domain knowledge graph. The knowledge graph data D2 includes the authoritative fact value corresponding to the quantifiable statement item;
[0009] And when the authoritative fact value in the knowledge graph data D2 comes from multiple authoritative data sources, perform a weighted average on the confidence levels of the authoritative data sources according to a preset weight, and the preset weight is positively correlated with the authoritative level of the data source;
[0010] Calculate the deviation degree between the statement value and the authoritative fact value, and generate a credibility label D3;
[0011] Perform hierarchical annotation processing on the original statement data according to the preset threshold interval to which the confidence value of the credibility label D3 belongs, and generate a text summary with credibility annotation;
[0012] Receive the historical feedback data D4 of the user on the credibility label D3. When the difference between the historical feedback data D4 and the knowledge graph data D2 exceeds the preset error threshold, dynamically adjust the weight value of the authoritative fact value or correct the authoritative fact value;
[0013] The retrieved and matched knowledge graph data D2 includes:
[0014] Locate the associated authoritative data source according to the entity node of the statement subject in the knowledge graph;
[0015] Convert the predicate into a knowledge graph attribute field through a predicate mapping table, and the predicate mapping table stores the correspondence between natural language predicates and knowledge graph attributes.
[0016] In a second aspect, the present invention discloses a context-aware text summary generation system combined with a knowledge graph, which implements the above-mentioned context-aware text summary generation method combined with a knowledge graph, including:
[0017] An original statement data parsing module, configured to extract the original statement data D1 from the input text to be summarized, identify and filter non-quantitative description statements through a semantic parser, and convert numerical statements into a triple form of <statement subject, predicate, statement value>;
[0018] A knowledge graph retrieval and matching module, configured to retrieve and match the knowledge graph data D2 in a pre-constructed domain knowledge graph according to the statement subject and predicate, locate the associated authoritative data source through the entity node, convert the natural language predicate into a knowledge graph attribute field by using the predicate mapping table, and return the corresponding authoritative fact value; if there are multiple data sources for the same attribute, perform weighted averaging on the confidence levels of the authoritative data sources according to preset weights, and the weight value is positively correlated with the authoritative level of the data source;
[0019] A deviation calculation and credibility label generation module, configured to calculate the deviation between the original statement value and the authoritative fact value, generate the credibility label D3, and store the calculation result as a structured label, including the numerical source, time stamp, and data source reference link;
[0020] A hierarchical annotation and summary generation module, configured to perform hierarchical annotation on the original statement data D1 according to the preset threshold interval of the confidence value; dynamically determine the threshold interval, and generate a text summary with an interactive label, and the label can trigger the display of the data source details and the update time;
[0021] User feedback dynamic optimization module, which is used to receive the historical feedback data D4 of users on the credibility label D3, including the corrected statement values or adjusted data source ratings. The module updates the data source weights through a formula; if the feedback difference exceeds the preset error threshold, it triggers the dynamic correction or weight adjustment of the authoritative fact values, and the correction results are synchronized to the knowledge graph in real time;
[0022] Authoritative data conflict detection module, which is used to detect the conflicting authoritative fact values in the knowledge graph data D2, preferentially adopts the value with the latest timestamp, and attaches a conflict warning mark to the credibility label D3; if a conflict is detected, it automatically retrieves the similar user correction records in the historical feedback data D4, generates auxiliary decision-making suggestions and updates them to the knowledge graph;
[0023] User interaction behavior monitoring module, which is used to continuously monitor the interaction behavior data D5 of users on the credibility label D3; when the number of interactions for the same statement item exceeds the preset activity threshold, it improves the display priority of this statement item in the abstract and optimizes the dynamic sorting logic of the abstract content;
[0024] Calibration instruction processing and review module, which is used to receive the calibration instruction D6 submitted by users; count the veto ratio, if it exceeds the preset dispute threshold, freeze the current abstract and trigger the manual review process. The review results are updated in reverse to the authoritative rating of the knowledge graph to ensure the dynamic iteration of the credibility of the data source;
[0025] Knowledge graph dynamic maintenance module, which is used to update the authoritative fact values and data source weights of the knowledge graph data D2 in real time; the module integrates the user feedback data D4, conflict detection results and manual review conclusions, updates the entity-attribute-value triples through an automated script, and synchronously adjusts the binding relationship of the authoritative data source reference links to ensure the timeliness and accuracy of the knowledge graph;
[0026] Abstract output and interaction interface module, which is used to integrate all processing results, generate the final text abstract with credibility annotation, and output it through the API or the front-end interface; support user interaction operations, and record the user behavior data to the log file for subsequent optimization analysis.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] 1. By automatically comparing the authoritative data sources in the knowledge graph, it can quickly identify exaggerated, incorrect or outdated information, provide an abstract with credibility annotation, and assist users in judging the reliability of key statements;
[0029] 2. Grade and annotate the abstract content according to the confidence threshold interval, and introduce user feedback to dynamically optimize the knowledge graph. Through the user feedback mechanism, the continuous iteration of the knowledge graph and the dynamic adjustment of the data source weights are realized to adapt to the authoritative demand differences in different scenarios. Brief Description of the Drawings
[0030] The disclosure of the present invention will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them:
[0031] Figure 1 is a flowchart of the method steps of the present invention;
[0032] Figure 2 is a schematic diagram of the method operation of the present invention;
[0033] Figure 3 is a process diagram of the knowledge graph data retrieved and matched by the present invention;
[0034] Figure 4 is a schematic diagram of the module operation of the present invention;
[0035] Figure 5 is a classification diagram of the confidence values of the credibility labels of the present invention;
[0036] Figure 6 is a schematic diagram of the verification of the numerical conflict of authoritative facts of the present invention;
[0037] Figure 7 is a feedback regulation diagram of the interaction behavior data of the present invention;
[0038] Figure 8 is a feedback regulation diagram of the calibration instruction of the present invention. Detailed Embodiments
[0039] It is easy to understand that according to the technical solution of the present invention, without changing the essence of the present invention, those of ordinary skill in the art can propose various interchangeable structural ways and implementation ways. Therefore, the following detailed embodiments and the accompanying drawings are only exemplary descriptions of the technical solution of the present invention, and should not be regarded as all of the present invention or as a limitation or restriction of the technical solution of the present invention.
[0040] Application Overview
[0041] As described above, as a structured knowledge representation method, the knowledge graph has been used to enhance semantic understanding, but its application in text summarization is still limited to entity linking or relationship reasoning, and has not been fully used for claim verification and credibility grading.
[0042] The existing solutions have the following deficiencies:
[0043] (1) Lack of a structured parsing and verification mechanism for quantitative claims;
[0044] (2) Do not consider the dynamic weighted fusion of multi-source authoritative data;
[0045] (3) The credibility assessment is disconnected from user feedback, making it difficult to achieve continuous optimization.
[0046] Therefore, there is an urgent need for a context-aware text summarization method that combines a knowledge graph. By means of structured verification of quantified statements, dynamic fusion of multi-source data, and a user feedback-driven optimization mechanism, it generates high-quality summaries with credibility annotations to meet the requirements of high-precision information processing.
[0047] In view of the above defects in the prior art, the basic concept of this application stems from a profound insight into the three core problems existing in the existing text summarization technology, namely numerical blind spots, static knowledge limitations, and feedback breaks. The credibility of the summary is improved by comparing structured data. The technical solution first automatically parses the quantified statements in the text into "subject - predicate - value" triples, and then looks up authoritative data sources through a pre-constructed knowledge graph for comparison and verification.
[0048] The reference value is automatically weighted and calculated according to the authority level of different data sources, and the deviation degree between the declared value and the reference value is calculated to generate a credibility score. According to the score, the system will adopt three processing methods: directly annotating high credibility, displaying medium credibility side by side, and requiring manual confirmation for low credibility. User feedback will be continuously collected. When it is found that the difference between the user's corrected value and the system's reference value is large, the data source weight will be automatically adjusted or the knowledge graph will be updated. This significantly improves the reliability and practicality of the summary and effectively solves the core pain point of insufficient verification of numerical statements in traditional summary technology.
[0049] After introducing the basic concept of the present invention, the embodiments of the present invention will be specifically introduced below with reference to the accompanying drawings.
[0050] Embodiment 1
[0051] As Figure 1 、 Figure 2 shown, the context-aware text summarization generation method that combines a knowledge graph includes the following steps:
[0052] Step 1: Obtain the original statement data D1 in the text to be summarized. The original statement data D1 includes at least one quantifiable statement item, and the quantifiable statement item includes a statement subject, a predicate, and a statement value.
[0053] The essence of the original statement data D1 is a set of assertions extracted from the text that can be quantitatively verified, and it is the input basis for subsequent credibility analysis.
[0054] The quantifiable statement item excludes subjective descriptions (such as "significantly improved" and "good effect"), and only retains the statements containing explicit numerical values.
[0055] The statement subject, predicate, and statement value are the three essential elements required for each quantifiable statement item;
[0056] Among them, the statement subject is the object of the assertion, for example: Company A, Drug X.
[0057] The predicate is an action or attribute that describes a relationship, for example: revenue growth rate, cure rate.
[0058] The statement value is a quantified value, which is a specific numerical value, such as 20%, 80%.
[0059] The steps to obtain the original statement data in the text to be summarized are as follows:
[0060] Extract the numerical statements in the text to be summarized parsed by the semantic parser;
[0061] Filter out the non-quantified description statements from the numerical statements and convert them into the triple form;
[0062] The triple form is 〈statement subject, predicate, statement value〉.
[0063] Example:
[0064] Input text: "Company B announced that its net profit in the first half of 2023 reached 520 million yuan, a year-on-year increase of 18%, and its market share increased to 22%."
[0065] The triple form obtained after extraction is:
[0066] <Company B, net profit, 520 million yuan>;
[0067] <Company B, year-on-year growth rate of net profit, 18%>;
[0068] <Company B, market share, 22%>.
[0069] Step 2: According to the statement subject and predicate, retrieve the matching knowledge graph data D2 in the pre-constructed domain knowledge graph. The knowledge graph data D2 includes the authoritative fact values corresponding to the quantifiable statement items.
[0070] As Figure 3 shown, the process of retrieving the matching knowledge graph data D2 is as follows:
[0071] Locate the associated authoritative data source according to the entity node of the statement subject in the knowledge graph;
[0072] Convert the predicate into a knowledge graph attribute field through the predicate mapping table, and the predicate mapping table stores the corresponding relationship between the natural language predicate and the knowledge graph attribute.
[0073] Specifically:
[0074] Precisely map the declarative subject in the text (such as "Enterprise A") to the specific entity node in the knowledge graph, and distinguish homonymous entities through context (such as "Apple Inc." and "fruit apple"). For example, combined with auxiliary information such as industry type and related events, if there is no exactly matching entity in the knowledge graph, the system will attempt to associate with near synonyms or upper category nodes (such as "XX Technology (Shenzhen) Co., Ltd." → "XX Technology Group"). Finally, obtain the unique entity identifier in the knowledge graph.
[0075] The predefined predicate mapping table is equivalent to a "translation dictionary", which converts daily expressions (such as "revenue growth") into the standard attribute fields of the knowledge graph (such as revenue_growth_rate).
[0076] After determining the entity node and attribute field, obtain numerical values from at least one of the following two types of data sources:
[0077] Main data source: High-authority sources (such as government financial reports, academic journals), with the highest priority;
[0078] Auxiliary data source: Supplementary sources (such as industry reports, corporate self-statements), and the credibility weights need to be marked.
[0079] When the data from different sources is inconsistent, by default, adopt the numerical value from the source with the highest weight, but retain the data from other sources for future reference.
[0080] Automatically filter the data with the latest timestamp and mark the information that has exceeded the validity period (such as "This data has not been updated for 12 months").
[0081] Example: Verification of new energy vehicle battery performance claims.
[0082] Input claim: <Brand X, driving range, 800 km>.
[0083] Retrieval process:
[0084] Entity location: Confirm that "Brand X" corresponds to the CarBrand: X node in the knowledge graph;
[0085] Predicate conversion: Map "driving range" to the attribute field max_range;
[0086] Data acquisition:
[0087] Authoritative data: 750 km (source: Ministry of Industry and Information Technology catalog, weight 0.95);
[0088] Enterprise data: 820 km (source: brand official website, weight 0.6).
[0089] System output: Using the data of 750 km from the Ministry of Industry and Information Technology, generate confidence labels and prompt: "The enterprise claims 800 km, reference value 750 km (source: the 12th batch of catalogs of the Ministry of Industry and Information Technology in 2023)".
[0090] The pre - constructed domain knowledge graph includes:
[0091] Extract entity - attribute - value triples from the structured database as the core nodes;
[0092] Bind at least one authoritative data source reference link to each attribute field, and the reference link points to the original data file.
[0093] Extract data from structured databases (such as MySQL tables, CSV files) or semi - structured data (such as JSON - formatted API responses), ensuring that the original data has clear field definitions.
[0094] Map database fields to knowledge graph triples:
[0095] Entity: An object with a unique identifier;
[0096] Attribute: A field with a standardized name;
[0097] Value: Retain the original numerical value and unit.
[0098] Each attribute value must be associated with at least one traceable original data file, and the link types include:
[0099] Direct link: PDF / Excel files on government open data platforms (such as the annual reports of the National Bureau of Statistics);
[0100] API endpoint: Real - time data sources called through APIs (such as the RESTful interface of the SEC Edgar database);
[0101] Versioned storage: Files with version numbers in the enterprise intranet document management system.
[0102] When the authoritative fact values in the knowledge graph data D2 come from multiple authoritative data sources, weighted average the confidence levels of the authoritative data sources according to the preset weights, and the preset weights are positively correlated with the authoritative levels of the data sources;
[0103] The authority of the data source is predefined by levels manually and linearly mapped to weight values.
[0104] An example of assigning authority level weights is shown in the following table:
[0105] Authority Level Data Source Example Preset Weight Level 5 National Bureau of Statistics, SEC Financial Reports 1.0 Level 4 Industry White Papers, Nature Papers 0.8 Level 3 Reports from Well - known Consulting Agencies 0.6 Level 2 Data from Local Regulatory Departments 0.4 Level 1 Self - disclosed Documents of Enterprises 0.2
[0106] High-authority data sources (such as government agencies) usually undergo strict review and have a low error probability, so they are given higher weights.
[0107] Multi-source data conflict handling:
[0108] Scene classification examples are shown in the following table:
[0109] Conflict Type Processing Method The numerical difference is < 5% Direct Weighted Average Numerical Difference 5% - 15% Weighted Average + Mark "Needs Review" Numerical Difference > 15% Trigger Manual Review, Pause Automatic Processing
[0110] Example: Statistical analysis of the debt ratio of listed companies;
[0111] Input data:
[0112] Central bank credit investigation system (weight 1.0): Debt ratio = 62%;
[0113] Enterprise annual report (weight 0.6): Debt ratio = 58%;
[0114] Securities firm research report (weight 0.8): Debt ratio = 65%.
[0115] Weighted calculation:
[0116]
[0117] Output result:
[0118] 62% (weighted result, mainly based on: central bank data 62%, securities firm data 65%).
[0119] Step 3: Calculate the deviation degree between the claimed value and the authoritative fact value, and generate a credibility label D3;
[0120] The formula for calculating the deviation degree is:
[0121] By calculating the proportion of the absolute difference between the claimed value and the authoritative value to the claimed value, the deviation degree is quantified; the characteristics of the formula are shown in the following table:
[0122]
[0123] The formula for calculating the confidence value of the credibility label D3 is: Confidence value = (1 - deviation degree) × 100%.
[0124] Reverse map the deviation degree to the confidence level to form an intuitive credibility evaluation:
[0125] Deviation degree 0% → Confidence level 100%;
[0126] Deviation degree 50% → Confidence level 50%;
[0127] Deviation degree 100% → Confidence level 0%.
[0128] Example: Verification of ESG statements of listed companies;
[0129] Input data:
[0130] Statement value: "Carbon emissions reduced by 40%";
[0131] Authoritative value: 32% of the monitoring data of the Environmental Protection Bureau;
[0132] Calculation process:
[0133] Deviation = |40 - 32| / 40 = 20%;
[0134] Confidence level = (1 - 0.2) × 100% = 80%;
[0135] Output result:
[0136] Annotation: "The enterprise claims a 40% emission reduction, referring to 32% of the environmental protection data (confidence level 80%)."
[0137] As Figure 5 shown, for the confidence level value of the calculated credibility label D3, there are three sets of thresholds to classify it, specifically:
[0138] When the confidence level value ≥ the first preset threshold, retain the original statement data D1 in the abstract and append the first credibility annotation;
[0139] When the confidence level value < the first preset threshold and ≥ the second preset threshold, display both the original statement data D1 and the knowledge graph data D2 in the abstract, and append the second credibility annotation to the original statement data D1;
[0140] When the confidence level value < the second preset threshold, send verification request data to the client to obtain the user's selection result, and append the third credibility annotation to the original statement data D1;
[0141] A third preset threshold is set between the first preset threshold and the second preset threshold, and when the confidence level value is between the second preset threshold and the third preset threshold, send verification request data Q1 to the client, and dynamically adjust the content of the finally displayed abstract according to the selection result submitted by the user through the verification request data Q1.
[0142] The meanings of the three sets of thresholds are shown in the following table:
[0143] Threshold Level Technology Name Suggested Value Functional Positioning First Preset Threshold High - credibility Threshold (T1) 90% Statement is Basically Consistent with Authoritative Data Second Preset Threshold Medium - credibility Threshold (T2) 70% Statement has Acceptable Deviations Third Preset Threshold Low - credibility Threshold (T3) 50% Statement May be Seriously Inaccurate
[0144] The thresholds can be adjusted according to industry needs. For example, in the financial field, they can be set as T1 = 95% and T2 = 80%.
[0145] The first credibility annotation, the second credibility annotation, and the third credibility annotation are all interactive tags, and the interactive tags are used to trigger the display of the source and update timestamp of the corresponding knowledge graph data D2 in response to user operations.
[0146] Example: Summary of pharmaceutical clinical data;
[0147] Input statement: "The effective rate of drug Y is 92% (clinical trial data)".
[0148] System verification:
[0149] Knowledge graph data: 87% (source: FDA database, updated in 2024-01)
[0150] Confidence calculation: (1 - |92 - 87| / 92) × 100% = 94.6%;
[0151] Processing flow:
[0152] 94.6% > T1 (90%) → Retain the statement and annotate:
[0153] "Effective rate 92% ('First credibility annotation' verified, 5.4% deviation from FDA data)"
[0154] Click on 'First credibility annotation' to display:
[0155] Data comparison: Company statement 92% vs FDA official 87%;
[0156] Deviation range: within the industry allowable error (<6%);
[0157] Final verification: 2025-02-20.
[0158] Step 4. According to the preset threshold interval to which the confidence value of the credibility label D3 belongs, perform hierarchical annotation processing on the original statement data to generate a text summary with credibility annotation.
[0159] Example: Disclosure of clinical data by pharmaceutical companies;
[0160] Input statement:
[0161] "The effective rate of drug X is 92% (based on phase III clinical trials)";
[0162] System processing:
[0163] Knowledge graph matching value: 87% (source: FDA database);
[0164] Confidence calculation: 82% (deviation 5.4%) > T1 (80%) → Retain the statement and annotate;
[0165] Processing result:
[0166] The effective rate of drug X is 92% (verified by FDA data, with the deviation within the allowable range).
[0167] Step 5: Receive the historical feedback data D4 of the user on the credibility label D3. When the difference between the historical feedback data D4 and the knowledge graph data D2 exceeds the preset error threshold, dynamically adjust the weight value of the authoritative fact value or correct the authoritative fact value;
[0168] Receiving the historical feedback data D4 of the user on the credibility label D3 includes:
[0169] Obtain the statement value or data source rating corrected by the user;
[0170] The formula for updating the data source weight is:
[0171] The feedback data reception and classification results are shown in the following table:
[0172] Feedback Data Type Data Structure Example Processing Method Numerical Correction - type Feedback {Original Statement ID: "A123", User - corrected Value: 18%} Trigger Recalculation of Weights Data Source Rating - type Feedback {Data Source ID: "SEC", User Rating: 4.2 / 5} Weighted Adjustment of Original Weights
[0173] The industry differentiation threshold is set according to the industry, and the example values are shown in the following table:
[0174]
[0175]
[0176] Example calculation of the weight update formula:
[0177] Original weight: 0.8; Knowledge graph value: 15%; User correction value: 18%;
[0178] The new weight is: The formula analysis is shown in the following table:
[0179] Scenario Magnitude of Weight Change Design Intention User Correction = Knowledge Graph Value Maintain Original Weight Confirm Data Source Accuracy User Correction > Knowledge Graph Value Linear Decrease Penalty Apply Stronger Constraints on Overestimated Data User Correction < Knowledge Graph Value Same - magnitude Penalty (Taking Absolute Value) Fairly Handle Underestimation / Overestimation
[0180] Example: Dispute over ROE data of listed companies;
[0181] 1. Initial state:
[0182] Knowledge graph value: 12% (source: Wind, weight 0.7), enterprise statement: 15%;
[0183] System confidence: T2 > 60% > T3 (triggering the 'Second credibility annotation' annotation);
[0184] 2. User feedback:
[0185] The accountant submits an audit report: 13.5%;
[0186] Difference calculation: |13.5 - 12| / 13.5 = 11.1% > financial threshold of 5%;
[0187] Weight update: New weight = 0.7 × (1 - 0.111) = 0.622;
[0188] 3. System actions:
[0189] Reduce the weight of Wind data to 0.62;
[0190] Annotation update: "Audit confirmed 13.5% (original authoritative data 12%, weight adjusted)". As Figure 6 shown, this solution also includes:
[0191] Detect whether there are conflicting authoritative fact values in the knowledge graph data D2;
[0192] The detected conflict situation of authoritative fact values is shown in the following table:
[0193]
[0194]
[0195] If there is a conflict, preferentially adopt the most recent authoritative fact value and attach a conflict warning mark to the credibility label D3;
[0196] And when a conflict is detected, retrieve the user correction records of similar declarations from the historical feedback data D4 as an auxiliary decision-making basis.
[0197] The conflict resolution priority strategy is shown in the following table:
[0198] Priority Decision - making Factor Execution Action 1 Data Timeliness Force Adoption of Data with the Latest Timestamp 2 User's Historical Correction Records Refer to Correction Values Adopted by More than 70% of Users 3 Data Source Weight Select the Data Source with the Highest Weight from Multiple Concurrent Data
[0199] Example: Dispute over the incidence rate of drug side effects;
[0200] Conflict detection:
[0201] Data from the Drug Administration in February 2024: Incidence rate 2.1%;
[0202] Data from a medical journal in January 2024: Incidence rate 4.3%;
[0203] Difference rate: 104% > threshold of 5%;
[0204] Resolution process:
[0205] Preferentially adopt the updated data of 2.1% from the Drug Administration;
[0206] Retrieval found:
[0207] User correction records in the past 3 months: 2.5% (adoption rate 72%);
[0208] Median of the data of the last 6 papers: 2.8%;
[0209] Final output:
[0210] Incidence rate 2.1% (major conflict warning, user correction suggestions 2.5%, literature reference range 2.3% - 3.1%).
[0211] As Figure 7 shown, it also includes:
[0212] Continuously monitor the interaction behavior data D5 of users with respect to the credibility label D3;
[0213] The types of interaction behavior data D5 include the number of label clicks, hover duration, and feedback submission frequency, and the data collection can be realized through front - end buried point statistics, mouse movement trajectory tracking, and form submission log analysis respectively.
[0214] When the number of interactions for the same statement item exceeds the preset active threshold, increase the display priority of this statement item in the abstract.
[0215] Automatically decay the historical interaction count by 20% every 30 days to prevent old data from occupying the priority for a long time.
[0216] Use a formula to calculate the new priority, specifically:
[0217] New priority = base weight + type coefficient × log 10 (1 + number of interactions);
[0218] Among them, the base weight is default set to 0.5 to ensure the basic sorting of un - interacted data; the type coefficient is adjusted according to the data industry type to control the promotion amplitude of different content types.
[0219] Example: New energy battery safety statement;
[0220] 1. Initial state:
[0221] Statement: "The passing rate of battery acupuncture test is 99%";
[0222] Base priority: 0.5 (ranked 8th);
[0223] 2. User interaction:
[0224] Interaction data accumulated within two weeks:
[0225] Click to view details: 43 times;
[0226] Hover for more than 5 seconds: 28 times;
[0227] User feedback: 12 items;
[0228] 3. System response:
[0229] Calculate new priority: 0.5+log 10 (1+83)×0.2≈1.18;
[0230] The ranking was raised to No. 2, and a highlighted background was added;
[0231] The update frequency of this node in the knowledge graph is increased to hourly synchronization.
[0232] like Figure 8 As shown, it also includes:
[0233] Obtaining a calibration instruction D6 submitted by a user, wherein the calibration instruction D6 is a rejection mark for the confidence value or an additional local authoritative data file;
[0234] The instruction types are rejection mark and local authoritative file. The rejection mark includes the user's questioning of the confidence value generated by the system, and must be accompanied by a preset rejection reason code. The local authoritative file contains file hash value verification and metadata extraction. The user uploads supplementary evidence with legal effect (such as quality inspection reports and audit reports with official seals), and the system automatically extracts key data fields.
[0235] When the rejection ratio exceeds the preset dispute threshold, the summary is frozen and the manual review process is triggered;
[0236] When the manual review process is triggered, a structured review work order is generated, which includes:
[0237] Summary of dispute focus: Extract the top 3 most frequently rejected reasons;
[0238] Data comparison view: Display declared value, knowledge graph value, and user-submitted value side by side;
[0239] Urgency score: calculated based on (dispute ratio × field risk factor).
[0240] The preset dispute threshold is dynamically set according to different industries and time periods; the following table is an example of dynamic threshold setting in actual applications:
[0241] Data Domain Base Threshold Adjustment Factor Actual Effective Threshold Financial Data 15% +5% (Annual Report Season) 20% Medical Data 20% -3% (Involving FDA Approval) 17% Consumer Product Evaluation 25% No Adjustment 25%
[0242] The review results are reversely updated to the authoritative rating of the knowledge graph data D2.
[0243] After review and confirmation, the data source rating is updated according to the formula:
[0244] New rating = Original rating × (1 - Number of confirmed disputes × 0.1) + Review correction value (±0.2).
[0245] Positive confirmation: When a dispute is valid, the credit score of the declarant is deducted;
[0246] False alarm handling: When a dispute is invalid, the user's authority level is adjusted accordingly.
[0247] Example 2: As Figure 4 shown, a context-aware text summarization generation system combined with a knowledge graph is disclosed, implementing the above-mentioned context-aware text summarization generation method based on the combination of a knowledge graph, including:
[0248] Original claim data parsing module, used to extract the original claim data D1 from the input text to be summarized, identify and filter non-quantitative description statements through a semantic parser, and convert numerical statements into a triple form of <declarant, predicate, claim value>;
[0249] Knowledge graph retrieval and matching module, used to retrieve and match the knowledge graph data D2 in the pre-constructed domain knowledge graph according to the declarant and predicate, locate the associated authoritative data source through the entity node, convert the natural language predicate into a knowledge graph attribute field using the predicate mapping table, and return the corresponding authoritative fact value; if there are multiple data sources for the same attribute, the confidence of the authoritative data sources is weighted and averaged according to the preset weight, and the weight value is positively correlated with the authority level of the data source;
[0250] Deviation calculation and credibility label generation module, used to calculate the deviation between the original claim value and the authoritative fact value, generate the credibility label D3, and store the calculation result as a structured label, including the value source, timestamp, and data source reference link;
[0251] Grading annotation and summary generation module, used to perform grading annotation on the original claim data D1 according to the preset threshold interval of the confidence value; dynamically judge the threshold interval, and generate a text summary with interactive labels, and the labels can trigger the display of the data source details and the update time;
[0252] User feedback dynamic optimization module, used to receive the historical feedback data D4 of the user on the credibility label D3, including correcting the claim value or adjusting the data source rating, and the module updates the data source weight through a formula; if the feedback difference exceeds the preset error threshold, trigger the dynamic correction or weight adjustment of the authoritative fact value, and the correction result is synchronized to the knowledge graph in real time;
[0253] The authoritative data conflict detection module is used to detect the conflicting authoritative fact values in the knowledge graph data D2, preferentially adopt the values with the latest timestamp, and attach a conflict warning mark to the credibility label D3; if a conflict is detected, it automatically retrieves the similar user correction records in the historical feedback data D4, generates auxiliary decision-making suggestions and updates them to the knowledge graph;
[0254] The user interaction behavior monitoring module is used to continuously monitor the interaction behavior data D5 of the user with the credibility label D3; when the number of interactions for the same statement item exceeds the preset active threshold, it increases the display priority of this statement item in the abstract and optimizes the dynamic sorting logic of the abstract content;
[0255] The calibration instruction processing and review module is used to receive the calibration instruction D6 submitted by the user; count the rejection ratio, and if it exceeds the preset dispute threshold, freeze the current abstract and trigger the manual review process. The review results are updated in reverse to the authoritative rating of the knowledge graph to ensure the dynamic iteration of the credibility of the data source;
[0256] The knowledge graph dynamic maintenance module is used to update the authoritative fact values and data source weights of the knowledge graph data D2 in real time; the module integrates the user feedback data D4, conflict detection results and manual review conclusions, updates the entity-attribute-value triples through an automated script, and synchronously adjusts the binding relationship of the authoritative data source reference links to ensure the timeliness and accuracy of the knowledge graph;
[0257] The abstract output and interaction interface module is used to integrate all processing results, generate the final text abstract with credibility annotation, and output it through the API or the front-end interface; support user interaction operations, and at the same time record the user behavior data to the log file for subsequent optimization analysis;
[0258] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0259] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0260] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0261] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0262] The technical scope of the present invention is not limited to the content described above. Those skilled in the art can make various deformations and modifications to the above embodiments without departing from the technical idea of the present invention, and these deformations and modifications should all fall within the protection scope of the present invention.
Claims
1. A context-aware text summarization generation method combined with a knowledge graph, characterized in that: It includes the following steps: Obtain the original statement data D1 in the text to be summarized. The original statement data D1 includes at least one quantifiable statement item, and the quantifiable statement item includes a statement subject, a predicate, and a statement value; Retrieve the matching knowledge graph data D2 in the pre-constructed domain knowledge graph according to the statement subject and the predicate. The knowledge graph data D2 includes the authoritative fact value corresponding to the quantifiable statement item; And when the authoritative fact values in the knowledge graph data D2 come from multiple authoritative data sources, the confidence levels of the authoritative data sources are weighted and averaged according to preset weights, and the preset weights are positively correlated with the authoritative levels of the data sources; Calculate the deviation degree between the statement value and the authoritative fact value, and generate a credibility label D3; Perform a hierarchical annotation process on the original statement data according to the preset threshold interval to which the confidence value of the credibility label D3 belongs, and generate a text summary with credibility annotations; Receive the historical feedback data D4 of the user on the credibility label D3. When the difference between the historical feedback data D4 and the knowledge graph data D2 exceeds the preset error threshold, dynamically adjust the weight value of the authoritative fact value or correct the authoritative fact value; The retrieved matching knowledge graph data D2 includes: Locate the associated authoritative data source according to the entity node of the statement subject in the knowledge graph; Convert the predicate into a knowledge graph attribute field through a predicate mapping table, and the predicate mapping table stores the corresponding relationship between natural language predicates and knowledge graph attributes.
2. The context-aware text summarization generation method combining a knowledge graph according to claim 1, characterized in that: The obtaining of the original statement data in the text to be summarized includes: Extract the numerical statements in the text to be summarized parsed by a semantic parser; Filter the non-quantitative description statements from the numerical statements and convert them into a triple form; The triple form is <statement subject, predicate, statement value>; The pre-constructed domain knowledge graph includes: Extract entity-attribute-value triples from a structured database as core nodes; Bind at least one authoritative data source reference link to each attribute field, and the reference link points to the original data file.
3. The context-aware text summarization generation method combined with a knowledge graph according to claim 1, characterized in that: The calculation formula for the deviation degree is as follows: The calculation formula for the confidence value of the credibility label D3 is: Confidence value = (1 - deviation degree) × 100%.
4. The method for generating context-aware text summaries by combining knowledge graphs according to claim 1, characterized in that: When the confidence value ≥ the first preset threshold, retain the original statement data D1 in the summary and attach a first credibility annotation; When the confidence value < the first preset threshold and ≥ the second preset threshold, display both the original statement data D1 and the knowledge graph data D2 in the summary, and attach a second credibility annotation to the original statement data D1; When the confidence value < the second preset threshold, send verification request data to the user terminal to obtain the user's selection result, and attach a third credibility annotation to the original statement data D1; A third preset threshold is set between the first preset threshold and the second preset threshold. When the confidence value is between the second preset threshold and the third preset threshold, send verification request data Q1 to the user terminal, and dynamically adjust the content of the finally displayed summary according to the selection result submitted by the user through the verification request data Q1.
5. The context-aware text summarization generation method combining a knowledge graph according to claim 4, characterized in that: The first credibility annotation, the second credibility annotation, and the third credibility annotation are all interactive tags, and the interactive tags are used to respond to user operations and trigger the display of the source and update timestamp of the corresponding knowledge graph data D2.
6. The context-aware text summarization generation method combining a knowledge graph according to claim 1, characterized in that: The historical feedback data D4 received from the user for the credibility label D3 includes: Obtaining the declared value or data source rating corrected by the user; The formula for updating the data source weight is as follows:
7. The context-aware text summarization generation method combined with a knowledge graph according to claim 1, characterized in that: It also includes: Detecting whether there are conflicting authoritative fact values in the knowledge graph data D2; If there is a conflict, preferentially adopt the most recent authoritative fact value and attach a conflict warning mark to the credibility label D3; And when a conflict is detected, retrieve the user correction records of similar declarations from the historical feedback data D4 as an auxiliary decision-making basis.
8. The context-aware text summarization generation method combining a knowledge graph according to claim 1, characterized in that: It also includes: Continuously monitoring the interactive behavior data D5 of the user for the credibility label D3; When the number of interactions for the same declaration item exceeds the preset activity threshold, increase the display priority of this declaration item in the abstract.
9. The context-aware text summarization generation method combining a knowledge graph according to claim 1, characterized in that: It also includes: Obtaining the calibration instruction D6 submitted by the user, where the calibration instruction D6 is a veto mark for the confidence value or an attached local authoritative data file; When the veto ratio exceeds the preset dispute threshold, freeze the abstract and trigger an artificial review process; Update the authority rating of the knowledge graph data D2 in reverse according to the review result.
10. A context-aware text summarization generation system combined with a knowledge graph, characterized in that: Implementing the context-aware text abstract generation method based on the combined knowledge graph as shown in any one of claims 1 to 9, including: An original declaration data parsing module, which is used to extract the original declaration data D1 from the input text to be abstracted, identify and filter non-quantified description statements through a semantic parser, and convert numerical statements into a triple form of <declaration subject, predicate, declared value>; A knowledge graph retrieval and matching module, which is used to retrieve and match the knowledge graph data D2 in a pre-constructed domain knowledge graph according to the declaration subject and predicate, locate the associated authoritative data source through the entity node, convert the natural language predicate into a knowledge graph attribute field by using a predicate mapping table, and return the corresponding authoritative fact value; if there are multiple data sources for the same attribute, perform a weighted average of the confidence levels of the authoritative data sources according to the preset weights, and the weight value is positively correlated with the authority level of the data source; A deviation calculation and credibility label generation module, which is used to calculate the deviation between the original declared value and the authoritative fact value, generate the credibility label D3, and store the calculation result as a structured label, including the value source, timestamp, and data source reference link; A hierarchical annotation and abstract generation module, which is used to perform hierarchical annotation on the original declaration data D1 according to the preset threshold interval of the confidence value; dynamically judge the threshold interval, and generate a text abstract with interactive tags, and the tags can trigger the display of the data source details and the update time; A user feedback dynamic optimization module, which is used to receive the historical feedback data D4 from the user for the credibility label D3, including correcting the declared value or adjusting the data source rating, and the module updates the data source weight through a formula; if the feedback difference exceeds the preset error threshold, trigger the dynamic correction or weight adjustment of the authoritative fact value, and the correction result is synchronized to the knowledge graph in real time; The authoritative data conflict detection module is used to detect the conflicting authoritative fact values in the knowledge graph data D2, preferentially adopt the value with the latest timestamp, and attach a conflict warning mark to the credibility label D3; if a conflict is detected, it automatically retrieves the similar user correction records in the historical feedback data D4, generates auxiliary decision-making suggestions and updates them to the knowledge graph; The user interaction behavior monitoring module is used to continuously monitor the interaction behavior data D5 of the user with the credibility label D3; when the interaction times of the same statement item exceed the preset active threshold, the display priority of this statement item in the summary is increased, and the dynamic sorting logic of the summary content is optimized; The calibration instruction processing and review module is used to receive the calibration instruction D6 submitted by the user; count the veto ratio, if it exceeds the preset dispute threshold, freeze the current summary and trigger the manual review process. The review result is updated in reverse to the authoritative rating of the knowledge graph to ensure the dynamic iteration of the credibility of the data source; The knowledge graph dynamic maintenance module is used to update the authoritative fact values and data source weights of the knowledge graph data D2 in real time; the module integrates the user feedback data D4, conflict detection results and manual review conclusions, updates the entity-attribute-value triples through an automated script, and synchronously adjusts the binding relationship of the authoritative data source reference link to ensure the timeliness and accuracy of the knowledge graph; The summary output and interaction interface module is used to integrate all processing results, generate the final text summary with credibility annotations, and output it through the API or the front-end interface; support user interaction operations, and at the same time record the user behavior data to the log file for subsequent optimization analysis.
Citation Information
Cited By
Information generation method and device, electronic equipment and storage medium
CN121880413A
Intelligent generation method and system for building decision, and medium
CN121996768A