Document argumentation method and device for dynamic deduction and automatic argumentation review
By constructing a directed probabilistic logic graph and a human-computer collaborative interaction mechanism, the problems of uncertainty in the argumentation process and insufficient cross-document correlation capabilities in existing technologies have been solved, realizing efficient and dynamic document argumentation analysis and improving review accuracy and decision support capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies cannot effectively quantify the uncertainties in the verification process, lack cross-document correlation capabilities, and have a passive human-machine collaboration mechanism, making it difficult to improve the accuracy of review.
A directed probabilistic logic graph is constructed. By identifying the logical relationships between content nodes and quantifying the strength of arguments, logical inference is optimized by combining human-computer collaborative interaction mechanisms, and interpretable AI technology is used to dynamically adjust the logic ontology library.
It enables dynamic representation of argument relationships and cross-document association, improving the efficiency and depth of analysis, providing forward-looking decision-making basis, and solving the static and passive problems of traditional models.
Smart Images

Figure CN121766296A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of document processing, and in particular relates to a document verification method and apparatus for dynamic deduction and automated verification review. Background Technology
[0002] In key areas such as financial risk control decision-making, judicial evidence chain review, and scientific hypothesis verification, documents, as the core carriers of information transmission and logical argumentation, directly determine the quality of decisions based on the completeness of their content and the rigor of their logic. With the acceleration of digitalization, the scale of documents is growing exponentially, and traditional manual review methods can no longer meet the dual demands of efficiency and accuracy in modern applications. For example, statistics from Fujitsu show that document review in the financial and insurance industry during the architecture design phase takes an average of 5.5 to 6.4 hours per KSLOC (thousand lines of code corresponding to document volume), and manual review, due to fatigue and experience limitations, easily leads to the omission of more than 35% of logical gaps or conflicts. Therefore, developing AI-based automated document review and argumentation technology has become a key direction for industry breakthroughs.
[0003] Currently, various technical solutions have emerged in the field of intelligent document processing, but all have significant limitations. One type of solution is based on pre-trained language models (LLM), such as using BERT or GPT series models for basic semantic verification. The Fujitsu team applied GPT-3.5-turbo and GPT-4 to perform consistency checks on software design documents, but its capabilities are limited to basic review within a single document. It is powerless to handle advanced needs such as complex logical connections across documents and quantification of argument strength, and its performance drops significantly when processing very long documents. Another type of solution uses knowledge graphs for structured modeling, constructing static relationship networks by extracting arguments and evidence. Although related research has achieved high accuracy in component identification, its models are inherently static and deterministic, unable to handle implicit connections across documents, nor can they quantify uncertainties such as "fluctuations in the credibility of evidence" or "gradual changes in the strength of argument support" that are common in the real world.
[0004] Furthermore, while vertical review systems targeting specific fields, such as the Jinxiandai Intelligent Document Processing Platform which integrates industry standard libraries or the Fadada Intelligent Contract Review Platform which has built-in contract templates, perform excellently in specific scenarios, their core relies on manually constructed rule bases. When facing rapidly evolving and complex businesses such as financial derivatives risk control, the rule base updates lag behind, and the system lacks autonomous logical reasoning capabilities. It cannot automatically trace the evolutionary relationships across multiple financial statements, such as "revenue data - profit forecast - risk warning," and still requires deep expert intervention.
[0005] In summary, existing technologies suffer from three major bottlenecks. First, the logical representation is static; neither LLM text analysis nor knowledge graph node association can effectively quantify the uncertainties in the argumentation process. Second, cross-document association capabilities are lacking; existing solutions are mostly limited to single-document environments and fail to establish a cross-document inference mechanism that supports "viewpoint inheritance, evidence supplementation, and refutation evolution," which is precisely the model relied upon by over 80% of key arguments. Third, the human-machine collaboration mechanism is passive; the system only passively accepts manual corrections after output, failing to construct an evolutionary closed loop of "model actively discovering uncertainties - expert-directed feedback - model iterative optimization," making it difficult to continuously improve review accuracy. Therefore, developing a new document argumentation model that can achieve dynamic inference, cross-document association, quantitative argumentation, and support human-machine collaborative evolution has significant theoretical and practical implications for overcoming current technological limitations. Summary of the Invention
[0006] In view of this, this application provides a document argumentation method and apparatus for dynamic deduction and automated argumentation review, which aims to achieve dynamic representation of document logic, cross-document association, and improve argumentation accuracy through proactive human-machine collaboration mechanism.
[0007] Firstly, this application provides a document verification method for dynamic deduction and automated verification review, including: Parse one or more target documents, identify the content elements of the target documents, and establish several content nodes based on the content elements; A logical argument ontology is established, and directed logical relationships between several content nodes are identified based on the logical argument ontology. The argument strength of each directed logical relationship is quantified, and then several content nodes are connected to obtain a directed probabilistic logic graph. Verify the uncertainty of logical inference in directed probabilistic logic graphs, and optimize the directed logical relationships in directed probabilistic logic graphs by combining human-computer collaborative interaction mechanisms; Actively detect key logical patterns in directed probabilistic logic graphs and generate analytical hypotheses.
[0008] Optionally, the steps of identifying directed logical relationships between several content nodes, quantifying the argument strength of each directed logical relationship, and then connecting several content nodes to obtain a directed probabilistic logic graph include: Define a set of logical relations, which includes several logical relations; From a number of content nodes, select any pair of content nodes as target content nodes; use each logical relationship in the set of logical relationships as the logical relationship between target content nodes, and calculate the argument strength of each logical relationship between target content nodes based on the semantic information, visual information and layout information of target content nodes. The logical relationship with the strongest argument is taken as the target logical relationship between target content nodes, and the target logical relationship is taken as the type of directed edge between target content nodes; Repeat the above steps until the logical relationships and argument strengths between all content nodes are determined. Then, connect several content nodes with directed edges to obtain a directed probabilistic logic graph.
[0009] Optionally, it also includes: Each time the human-computer collaborative interaction mechanism is used to optimize the directed logical relationship of the directed probabilistic logic graph, incremental learning samples are generated based on user feedback. Then, based on the incremental learning samples formed by multiple human-computer collaborative interaction mechanisms, an incremental learning sample set is formed. When the number of samples in the incremental learning sample set is greater than a preset value, the logical argument ontology is adjusted using the incremental learning sample set.
[0010] Optionally, it also includes: Unsupervised learning algorithms are used to cluster relation instances that the model failed to identify with high confidence in order to discover new potential relation type clusters; Interpretable AI technology is used to extract and present the core semantic or structural features of each relation type cluster to the user in order to explain the reasons for its clustering. In response to the user's interactive definition of the new relation type, the new relation type is dynamically added to the argument logic ontology library.
[0011] Optionally, the key logical patterns include logical conflicts between core arguments, missing links in the argumentation of key conclusions, or inconsistencies in cross-document chains of evidence.
[0012] Optionally, if a user's query is received, a simulation inference algorithm is executed on the directed probability logic graph to perform counterfactual analysis or scenario simulation and output quantitative inference results; and the directed probability logic graph is adjusted according to the inference results to make the directed probability logic graph more relevant to the user.
[0013] Optionally, when processing multiple target documents, a directed probabilistic logic graph is built for each target document, and after performing directed logic optimization on each directed probabilistic logic graph, several directed probabilistic logic graphs are merged into one.
[0014] Secondly, this application provides a document verification device for dynamic deduction and automated verification review, comprising: The recognition module is used to parse one or more target documents, identify the content elements of the target documents, and establish several content nodes based on the content elements. A module is established to build a logical argument ontology library. Based on the logical argument ontology library, directed logical relationships between several content nodes are identified, and the argument strength of each directed logical relationship is quantified. Then, several content nodes are connected to obtain a directed probabilistic logic graph. The optimization module is used to verify the uncertainty of logical inference in the directed probabilistic logic graph. Combined with the human-computer collaborative interaction mechanism, it optimizes the directed logical relationships in the directed probabilistic logic graph. The detection module is used to actively detect key logic patterns in the directed probabilistic logic graph and generate analytical hypotheses.
[0015] Thirdly, this application provides an electronic device, including the document verification device described above for dynamic deduction and automated verification review.
[0016] Fourthly, this application provides a computer-readable storage medium storing at least one piece of program code, which is executed by a processor to implement the document verification method for dynamic deduction and automated verification review as described in any of the preceding claims.
[0017] The beneficial effects of the technical solution provided in this application include: (1) This application replaces the traditional static logic structure by constructing a directed probabilistic logic graph. This not only quantifies the uncertainty of the argument relationship, but more importantly, it provides a mathematical basis for performing simulation deduction. Through counterfactual analysis and "what-if" scenario simulation functions, the system can quantitatively predict forward-looking questions such as "what will happen if specific conditions change", providing direct and powerful decision-making basis for complex decision-making scenarios such as financial risk control and judicial analysis.
[0018] (2) This application enables the system to autonomously discover deep logical problems that are not easily noticed by humans through an active logical conflict and argument missing loop detection mechanism, and to generate structured analytical hypotheses. This new model of "model active discovery - human collaborative verification" fundamentally changes the nature of human-computer interaction, transforming AI from a tool waiting for instructions into an intelligent partner that can inspire thinking and guide analysis, greatly improving the efficiency and depth of analysis work.
[0019] (3) In one embodiment, by deeply applying interpretable AI (XAI) technology to the self-evolution process of the argument logic ontology library, this application makes the discovery and definition process of new knowledge completely transparent to the user. The user can understand "why" the model discovers new logical relationships, thereby establishing trust and making efficient and accurate interactive definitions. This solves the trust deficit problem caused by the "black box learning" of traditional models and builds a sustainable and trustworthy human-machine collaborative knowledge evolution system.
[0020] (4) In one embodiment, by supporting the association and integration from micro-content (such as sentences and data series) to macro-knowledge networks (through meta-graph synthesis), the present invention can construct a panoramic view of domain knowledge. This provides users with a comprehensive and structured data foundation for cross-document knowledge tracing, viewpoint evolution analysis, and domain trend judgment. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1 A flowchart of a document verification method for dynamic deduction and automated verification review provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a directed probabilistic logic graph provided in an embodiment of this application; Figure 3 A flowchart illustrating an active learning and human-computer collaborative interaction mechanism provided in an embodiment of this application; Figure 4 A flowchart illustrating the self-evolution mechanism of a logical ontology library provided in an embodiment of this application; Figure 5 A structural block diagram of a document verification device for dynamic deduction and automated verification review provided in an embodiment of this application; Figure 6 This is a structural block diagram of an electronic device provided in an embodiment of this application.
[0023] The attached figures are labeled as follows: 11: Identification module; 12: Establishment module; 13: Optimization module; 14: Detection module; 21: Processor; 22: Memory. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0025] Figure 1This is a flowchart illustrating a document verification method for dynamic deduction and automated verification review, provided as an embodiment of this application. See also... Figure 1 ,include: S101. Parse one or more target documents, identify the content elements of the target documents, and establish several content nodes based on the content elements.
[0026] In some examples, the purpose of step S101 is to perform multimodal parsing on the target document and then structure the content elements into content nodes containing content data and metadata.
[0027] In some examples, the steps for performing multimodal adaptive parsing of a document include: By using a deep learning-based document object detection model (e.g., a model based on YOLOv5 or LayoutLM architecture), the document is processed page by page to accurately segment and identify logically independent content elements.
[0028] In some examples, content elements include text paragraphs, headings at various levels, tables (such as profit and loss statements), charts (such as revenue growth trend graphs), and footnotes.
[0029] In some examples, taking a listed company's annual financial report in PDF format as an example, the process of creating content nodes is as follows: First, based on the target document's metadata (such as filename) and visual features (such as homepage layout and company logo), its document type is identified as "formal financial statement document." Then, a deep learning-based document object detection model (e.g., a model based on YOLOv5 or LayoutLM architecture) is invoked to process the target document page by page to accurately segment and identify logically independent content elements, including text paragraphs, headings at various levels, tables (such as profit and loss statements), charts (such as revenue growth trend graphs), and footnotes. Each identified content element is encapsulated as a standardized content node data object.
[0030] In some examples, the content node contains a unique in-document ID, type (TEXT, TABLE, FIGURE, etc.), content data (e.g., text in a text block, JSON-formatted data in a table, image data in a chart, and OCR text), and metadata (e.g., page numbers and bounding box coordinates in the original document). (including the chapter / section it belongs to).
[0031] S102. Establish a logical argument ontology library, identify directed logical relationships between several content nodes based on the logical argument ontology library, quantify the argument strength for each directed logical relationship, and then connect several content nodes to obtain a directed probabilistic logic graph.
[0032] In some examples, the argument logic ontology can be pre-defined by domain experts using a set of basic logical relation types (such as support, refutation, exemplification, inference, etc.), or generated by cold-starting from a large-scale labeled argument corpus through unsupervised or semi-supervised learning.
[0033] In some examples, step S102 includes: S1021, Define a set of logical relations, which includes several logical relations.
[0034] In some examples, logical relation types include support, refutation, exemplification, inference, etc.
[0035] The types of logical relationships described above are merely examples provided in this application. Those skilled in the art should understand that the set of logical relationships defined in this application should include any existing logical relationship or select a small number of core logical relationships from existing logical relationships based on their importance, thereby forming the set of logical relationships in this application.
[0036] The set of logical relationships defined in this application is used in subsequent steps to specify the logical relationships between content nodes. By analyzing the logical relationships between content nodes, it is possible to identify whether there are logical contradictions between the content expressed by a large number of content nodes.
[0037] S1022, select any pair of content nodes from a number of content nodes as target content nodes; take each logical relationship in the set of logical relationships as the logical relationship between target content nodes, and calculate the argument strength of each logical relationship between target content nodes based on the semantic information, visual information and layout information of target content nodes.
[0038] In some examples, the purpose of step S1022 is to select the optimal logical relationship for each pair of target content nodes. The selection principle is to determine the optimal logical relationship for each pair of content nodes by calculating the argument strength of all logical relationships.
[0039] In some examples, for any pair of content nodes (i.e., the selected node) and nodes (A pair of target content nodes) are formed, and a pre-trained multimodal machine learning model is used to infer the logical relationship.
[0040] The model first performs feature encoding on the two nodes. This encoding integrates semantic information of nodes (text semantics obtained through language models such as BERT), visual information (graphic image features extracted through CNN), and layout information (relative position, distance, and alignment between nodes).
[0041] Subsequently, the encoded feature pairs The input is fed into the model's classification layer, and the output is a set of predefined logical relation types. (For example, =Data support, =Explanation, =Contradictions, etc., conditional probability distributions on the conditional probability distribution.
[0042] This process can be represented by the following formula:
[0043] in Indicates the model parameters Defined classification function, express A certain type of logical relationship, for example It can be For example, It can be .
[0044] It should be noted that the value of the conditional probability distribution corresponds to the strength of the argument for this type of relationship.
[0045] S1023 takes the logical relationship with the strongest argument as the target logical relationship between target content nodes, and takes the target logical relationship as the type of directed edge between target content nodes.
[0046] The system selects the logical relation type with the highest probability. The type of the edge is used as the probability value of the argument strength. ,Right now:
[0047] Finally, at the node and Establish a directed edge between them, the edge consisting of a tuple Provide a complete description.
[0048] S1024, Repeat the above steps until the logical relationships and argument strengths between all content nodes are determined. Then, connect several content nodes with directed edges to obtain a directed probabilistic logic graph.
[0049] Reference Figure 2 This is a schematic diagram of a directed probabilistic logic graph generated according to an embodiment of the present invention. Nodes 1, 2, and 3 represent content nodes of different types, while edge 1 represents the type and probabilistic argument strength between them. The probabilistic logical relationship.
[0050] It should be noted that the directed probabilistic logic graph constructed in this invention is the technical foundation for realizing the dynamic deduction function in subsequent steps. This directed probabilistic logic graph can be formally regarded as a Bayesian network, where the directed edges between content nodes represent the direction of influence or dependence between arguments, and the probability value of the edge (i.e., the strength of the argument) quantifies the strength of this influence.
[0051] This modeling approach transforms the graph from a mere description of the static logic of a document into a computable simulation model. Based on this model, by intervening in the belief states of specific nodes, their impact on downstream nodes can be simulated, making it possible to perform counterfactual reasoning and 'what-if' analysis.
[0052] In some examples, the argument logic ontology provided in this application is a dynamically evolving argument logic ontology. More specifically, dynamic evolution means that the argument logic ontology automatically adjusts its logical relationships during the execution of the method. For example, it identifies and discovers new logical relationships, or it dynamically adjusts the identification model used by the argument logic ontology (which will be explained in detail in step S103).
[0053] See Figure 3 In some examples, one self-evolutionary mechanism of the argument logic ontology for identifying and discovering new logical relations is as follows: The first step is to periodically collect model confidence scores (i.e., the strength of the argument). ) Instances of relationships below the threshold.
[0054] The second step involves using an unsupervised clustering algorithm (such as k-means) to analyze the feature vectors of these instances. Analysis revealed potential new clusters of relationships.
[0055] The third step involves using an interpretable AI (XAI) technique, such as SHAP (SHapley Additive exPlanations), for each new cluster to calculate the contribution of each input feature dimension to the decision that "the instance belongs to this cluster." The feature with the highest contribution (such as "node A is directly above node B" or "node A contains the word 'therefore'") is taken as the core explanation for that cluster and presented to domain experts.
[0056] In the fourth step, after understanding the AI's "clustering rationale", the experts interactively named and defined the new relationship types, and finally dynamically added them to the ontology library.
[0057] In some examples, the established directed probabilistic logic graph can be further subdivided to optionally provide a more refined analytical perspective. The specific process is as follows: A parent node pair with an established logical relationship is divided into finer-grained child nodes based on its internal structure (such as sentences in a paragraph or data series in a chart), and the structured relationships between these child nodes are provided to the model inference and insight module to support micro-drill-down analysis of macro-logical relationships.
[0058] In some examples, when processing multiple target documents, a directed probabilistic logic graph is built for each target document. After performing directed logic optimization on each directed probabilistic logic graph, several directed probabilistic logic graphs are merged into one. The specific steps are as follows: The first step is to generate independent probabilistic logic graphs for each target document, and then identify and associate nodes in different graphs that point to the same real-world concept or entity through entity linking or concept alignment techniques. The second step is to infer cross-document meta-relationships, such as "follow-up research" and "evolution of viewpoints," between aligned nodes based on inter-document metadata.
[0059] The third step is to link the independent multiple directed probabilistic logic graphs through meta-relations and combine them into a semantic meta-graph that can display the knowledge context and the overall picture of the domain.
[0060] In subsequent steps S103 and S104, if multiple directed probabilistic logic graphs are merged into a semantic meta-graph, then steps S103 and S104 are performed on the semantic meta-graph.
[0061] S103 verifies the uncertainty of logical inference in the directed probabilistic logic graph, and optimizes the directed logical relationship of the directed probabilistic logic graph by combining human-computer collaborative interaction mechanism.
[0062] In some of these examples, the core objective of this step is not only to generate a static snapshot of the map, but also to establish a dynamic, robust, verifiable, and self-improving argumentation model.
[0063] In some examples, this step can be broken down into three sub-steps: graph integration, uncertainty-based active learning, and iterative optimization of the model. S1031, all generated content nodes (as vertices of the graph) ) and all probabilistic logical relation edges generated (as directed edges of the graph) This is integrated to form a complete, directed, weighted graph structure. .
[0064] To facilitate subsequent calculations and queries, the graph structure can be serialized into a standard graph data format, such as GraphML or JSON-LD, or a set of Cypher or GQL (Graph Query Language) creation statements can be directly generated and imported into graph databases such as Neo4j.
[0065] S1032, to proactively eliminate ambiguity in the model's relation inference, the system incorporates an active learning mechanism. The core of this mechanism lies in quantifying the uncertainty of the model's inference.
[0066] A preferred quantization method is to calculate the conditional probability distribution of the relation type output in step S1022. Information entropy :
[0067] in, It is the total number of predefined logical relation types. The model inferred as the first The probability of each logical relation type Represents the predefined first Types of logical relationships. Information entropy. The larger the value, the closer the probability distribution is to uniformity, and the higher the uncertainty of the model.
[0068] when Exceeding a preset dynamic threshold When this occurs, the system marks the inference as a highly ambiguous inference.
[0069] Subsequently, the system will proactively push an interactive query task to the user interface. This task clearly presents the two ambiguous content nodes and lists the ones with the highest probability. One (e.g.) or The system displays candidate relationship types and their confidence levels, requesting users to make final confirmations, selections, or corrections.
[0070] S1033, after receiving feedback from the user through the interactive interface, the system performs a dual optimization operation.
[0071] The first optimization operation is to optimize the graph structure: for the currently instantiated directed probabilistic logic graph... The system immediately confirms the correct relationship as verified by the user. Update the corresponding edge and its proof strength. Adjust to a high confidence value (e.g., 1.0) to ensure the accuracy of the current analysis.
[0072] See Figure 4Specifically, it illustrates the process of generating proactive query tasks (human-computer interaction) based on confidence level (uncertainty), thereby generating incremental learning samples, and adjusting the directed probabilistic logistic graph based on user feedback.
[0073] The second optimization operation is to optimize the argument logic ontology (which is also the second self-evolution mechanism of the argument logic ontology): each time the human-computer collaborative interaction mechanism is used to optimize the directed logical relationship of the directed probabilistic logic graph, incremental learning samples are generated based on user feedback, and then an incremental learning sample set is formed based on the incremental learning samples formed by multiple human-computer collaborative interaction mechanisms; when the number of samples in the incremental learning sample set is greater than a preset value, the logic argument ontology is adjusted using the incremental learning sample set.
[0074] More specifically, each piece of user feedback, such as Each of these is considered a high-quality manually labeled dataset and is stored in a dedicated incremental training sample library.
[0075] The system will periodically (e.g., at the end of each workday) or when the sample size reaches a certain scale, use this sample library to test the multimodal machine learning model in step S1022. Perform incremental training or fine-tuning. This further improves the multimodal machine learning model. The accuracy, thus when the argument logic ontology library is used through a multimodal machine learning model It will be more accurate when determining the logical relationships between content nodes.
[0076] This process adjusts the model parameters using optimization algorithms such as gradient descent. The aim is to minimize the model's prediction loss on these high-quality samples. Through this closed-loop mechanism, the model's inference ability is continuously improved, and the possibility of ambiguity when dealing with similar scenarios in the future will be significantly reduced, thus achieving dynamic iterative optimization of the graph and the model.
[0077] It should be further explained that, in addition to the aforementioned uncertainty-based active learning, the human-computer collaborative interaction mechanism of this application is also deeply coupled with the automated argument structure review mechanism in step S105. When the system generates analytical hypotheses (such as detecting logical conflicts) in S104 and presents them to the user, the user's confirmation, rejection, or explanation of the hypothesis is also regarded as a high-level interactive feedback and is used to optimize the system's built-in argument logic ontology library or further train the model, thereby realizing the collaborative evolution of the argument logic ontology library between the model and the user.
[0078] In some examples, logical deduction and counterfactual analysis can also be performed on the directed probabilistic logic graph. The specific steps are as follows: if a user's query is received, a simulation deduction algorithm is executed on the directed probabilistic logic graph to perform counterfactual analysis or scenario simulation and output quantitative deduction results; and the directed probabilistic logic graph is adjusted according to the deduction results to make the directed probabilistic logic graph more relevant to the user.
[0079] More specifically, the constructed directed probabilistic logic graph can be viewed as a Bayesian network or a structured logic model (SCM). When a user's "what-if" query is received, for example, "assuming 'R&D investment' (corresponding node...),"... The value of 'Company's Future Growth Potential' (corresponding node) doubled, and the value of 'Company's Future Growth Potential' doubled. How will the confidence level of ) change? The system formalizes this query as an intervention.
[0080] This intervention can be represented using Judea Pearl's do() operator. The objective of the system computation is the posterior probability. ,in This represents the new state after intervention. The computation process involves nodes on a graph model. Force a value assignment to sever the influence of all its parent nodes on it. Then, based on the network structure and the conditional probability table of each edge (derived from the argument strength), propagate the probability update brought about by this intervention forward until the target node is calculated. The new probability distribution.
[0081] In this way, each user query will adjust the directed probability logic graph, making the directed probability logic graph more closely match the user's needs.
[0082] S104 actively detects key logic patterns in the directed probabilistic logic graph and generates analytical hypotheses.
[0083] In some examples, the key logical patterns include logical conflicts between core arguments, missing links in the argumentation of key conclusions, or inconsistencies in cross-document chains of evidence.
[0084] Reference Figure 4 It predefines a series of key logical patterns, such as conflict patterns: Make The relationship type is "support", and The relationship type is "contradictory"; missing cycle pattern: The core argument is presented, but the strength of the arguments for all its input edges is... All are below the preset threshold .
[0085] The system continuously scans the map to detect these patterns. Once a detection is successful, the system generates an analytical hypothesis described in natural language.
[0086] This hypothesis is presented as a structured text output, such as: "For example: A [support-contradiction] conflict was detected between the arguments of node [financial report data] and node [analyst report] regarding node [company profit expectations]. Cross-validation is recommended." This is then proactively pushed to the user to complete in-depth analysis.
[0087] Figure 5 This is a structural block diagram of a document verification device for dynamic deduction and automated verification review, provided as an embodiment of this application. See also... Figure 5 ,include: The recognition module 11 is used to parse one or more target documents, identify the content elements of the target documents, and establish several content nodes based on the content elements. Module 12 is used to establish a logical argument ontology library, identify directed logical relationships between several content nodes based on the logical argument ontology library, quantify the argument strength for each directed logical relationship, and then connect several content nodes to obtain a directed probabilistic logic graph. Optimization module 13 is used to verify the uncertainty of logical inference in the directed probabilistic logic graph and optimize the directed logical relationship of the directed probabilistic logic graph by combining the human-computer collaborative interaction mechanism. The detection module 14 is used to actively detect key logic patterns in the directed probabilistic logic graph and generate analytical hypotheses.
[0088] It should be noted that the document verification device provided in this application for dynamic deduction and automated verification review is used to perform, for example... Figure 1 The method described herein, therefore, can be selectively applied by those skilled in the art. Figure 1 The module included in the adaptive partitioning device of the method is not limited in this application.
[0089] Figure 6 This is a structural block diagram of an electronic device provided according to an embodiment of this application. See also... Figure 6 Electronic devices may include: Figure 5The document verification device described above is used for dynamic deduction and automated verification review. Typically, the electronic device includes a processor 21 and a memory 22. The processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is used to process data in the wake-up state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. The memory 22 may include one or more computer-readable storage media, which may be non-transitory. The memory 22 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage medium in memory 22 is used to store at least one instruction, which is executed by processor 21 to implement the document argumentation method for dynamic deduction and automated argumentation review provided by an electronic device in the method embodiments of this application.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A document verification method for dynamic deduction and automated verification review, characterized in that, include: Parse one or more target documents, identify the content elements of the target documents, and establish several content nodes based on the content elements; A logical argument ontology is established, and directed logical relationships between several content nodes are identified based on the logical argument ontology. The argument strength of each directed logical relationship is quantified, and then several content nodes are connected to obtain a directed probabilistic logic graph. Verify the uncertainty of logical inference in directed probabilistic logic graphs, and optimize the directed logical relationships in directed probabilistic logic graphs by combining human-computer collaborative interaction mechanisms; Actively detect key logical patterns in directed probabilistic logic graphs and generate analytical hypotheses.
2. The document verification method for dynamic deduction and automated verification review according to claim 1, characterized in that, The steps of identifying directed logical relationships between several content nodes, quantifying the argument strength of each directed logical relationship, and then connecting several content nodes to obtain a directed probabilistic logic graph include: Define a set of logical relations, which includes several logical relations; From a number of content nodes, select any pair of content nodes as target content nodes; use each logical relationship in the set of logical relationships as the logical relationship between target content nodes, and calculate the argument strength of each logical relationship between target content nodes based on the semantic information, visual information and layout information of target content nodes. The logical relationship with the strongest argument is taken as the target logical relationship between target content nodes, and the target logical relationship is taken as the type of directed edge between target content nodes; Repeat the above steps until the logical relationships and argument strengths between all content nodes are determined. Then, connect several content nodes with directed edges to obtain a directed probabilistic logic graph.
3. The document verification method for dynamic deduction and automated verification review according to claim 1, characterized in that, Also includes: Each time the human-computer collaborative interaction mechanism is used to optimize the directed logical relationship of the directed probabilistic logic graph, incremental learning samples are generated based on user feedback. Then, based on the incremental learning samples formed by multiple human-computer collaborative interaction mechanisms, an incremental learning sample set is formed. When the number of samples in the incremental learning sample set is greater than a preset value, the logical argument ontology is adjusted using the incremental learning sample set.
4. The document verification method for dynamic deduction and automated verification review according to claim 1, characterized in that, Also includes: Unsupervised learning algorithms are used to cluster relation instances that the model failed to identify with high confidence in order to discover new potential relation type clusters; Interpretable AI technology is used to extract and present the core semantic or structural features of each relation type cluster to the user in order to explain the reasons for its clustering. In response to the user's interactive definition of the new relation type, the new relation type is dynamically added to the argument logic ontology library.
5. The document verification method for dynamic deduction and automated verification review according to claim 1, characterized in that, The key logical patterns include logical conflicts between core arguments, missing links in the argumentation of key conclusions, or inconsistencies in cross-document evidence chains.
6. The document verification method for dynamic deduction and automated verification review according to claim 1, characterized in that, If a user's query is received, a simulation inference algorithm is executed on the directed probability logic graph to perform counterfactual analysis or scenario simulation and output quantitative inference results; and the directed probability logic graph is adjusted according to the inference results to make the directed probability logic graph more relevant to the user.
7. The document verification method for dynamic deduction and automated verification review according to claim 1, characterized in that, When processing multiple target documents, a directed probabilistic logic graph is built for each target document. After performing directed logic optimization on each directed probabilistic logic graph, several directed probabilistic logic graphs are merged into one.
8. A document verification device for dynamic deduction and automated verification review, characterized in that, include: The recognition module is used to parse one or more target documents, identify the content elements of the target documents, and establish several content nodes based on the content elements. A module is established to build a logical argument ontology library. Based on the logical argument ontology library, directed logical relationships between several content nodes are identified, and the argument strength of each directed logical relationship is quantified. Then, several content nodes are connected to obtain a directed probabilistic logic graph. The optimization module is used to verify the uncertainty of logical inference in the directed probabilistic logic graph. Combined with the human-computer collaborative interaction mechanism, it optimizes the directed logical relationships in the directed probabilistic logic graph. The detection module is used to actively detect key logic patterns in the directed probabilistic logic graph and generate analytical hypotheses.
9. An electronic device, characterized in that, Includes the document verification device for dynamic deduction and automated verification review as described in claim 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is executed by a processor to implement the document verification method for dynamic deduction and automated verification review as described in any one of claims 1 to 7.