A provenance tracing method and system

By constructing a Bayesian network and combining it with a causal database and expert argumentation, the problem of determining causal relationships for multi-object, non-specific events was solved. This enabled accurate determination of causal relationships and identification of key causes, thereby improving the scientific rigor and impartiality of the determination system.

CN114091675BActive Publication Date: 2026-04-14HEFEI UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2019-01-15
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies cannot effectively determine the causal relationship of at least two objects experiencing at least one non-specific event, especially when the event type is unclear. They lack accuracy and comprehensiveness, and existing methods mainly rely on qualitative judgments, which cannot handle causal relationships that do not meet the four judgment rules.

Method used

A causal relationship determination system based on Bayesian networks is constructed. By building modules to analyze object factors and establishing undirected graph structure constraints, and combining causal databases and expert demonstrations for correction, a Bayesian network is constructed using a heuristic search algorithm, and causal indicators are calculated using Pearl's principle to output a causal relationship spectrum.

Benefits of technology

It enables accurate causal relationship determination for multiple non-specific events. As the number of events increases, the accuracy of the determination results continuously improves, providing more scientific causal analysis and reasoning, and supporting fair adjudication in medical disputes and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114091675B_ABST
    Figure CN114091675B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of traceability inference method and system, the system at least includes: construction module, for constructing bayesian network;Causal library, for establishing historical original literature library, characterized in that, construction module is configured to in the case where causal library carries out deep learning to the result of expert argumentation, the relationship between factor pair L m And L n Number is corrected based on causal library according to factor pair L m And L n Pair.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention is a divisional application of application number 201910034540.4, filed on January 15, 2019, of the invention type, entitled "A Method for Determining Causality Based on Bayesian Networks". Technical Field

[0002] This invention belongs to the field of cause tracing technology and relates to a method and system for abductive reasoning. Background Technology

[0003] Causal tracing involves reasoning step by step to uncover the key and root causes hidden within an unexpected event or effect, revealing its complex causal relationships.

[0004] For example, Chinese Patent Publication No. CN109063253A discloses a Bayesian network-based reliability modeling method for aviation power systems. This method includes: establishing input-output relationship description tables for each component of the aviation power system; constructing corresponding Bayesian network nodes for each row of data in N sub-tables; determining the parent node of each node in the constructed Bayesian network; constructing a parent node with the same name for component C in step 2, and determining the conditional probability distribution of each node based on the determined state of each node in the constructed Bayesian network; determining target nodes for power supply busbars and calculating the power supply reliability of the aviation power system; and determining corresponding target nodes for each power supply busbar and calculating the power supply reliability of the aviation power system. This invention improves the efficiency of aviation power system reliability calculation. This method only involves a single event occurring in a single object.

[0005] For example, a cause tracing method disclosed in Chinese Patent Publication No. CN105468703A includes the following steps: initializing a causal relationship knowledge base, which includes anomalies of a class of objects and the causes of the anomalies, as well as the causal relationship between the anomalies and their causes; selecting anomalies with known current states from the list of anomalies; forming a new causal relationship knowledge base based on the causal relationships in the causal relationship knowledge base and recording the traced causes; and outputting the traced causes as result information.

[0006] In existing technologies, cause tracing is applicable when a single entity experiences an abnormal or unexpected effect, but it cannot be performed when at least two objects experience at least one specific event. Furthermore, such cause tracing methods only involve qualitative determinations. Therefore, a system or method is urgently needed to address how to determine the cause of a specific event when at least two objects have experienced at least one specific event.

[0007] The prior art, patent document CN107563596A, proposes a causal relationship determination system based on Bayesian networks. It determines the causal relationship between variables through conditional independence testing. The specific implementation method is as follows: Select a variable Xi from either the top or bottom layer. Simultaneously, in the intermediate layer, select another node variable Xj connected to the node of variable Xi via an undirected edge EAij. Test the conditional independence between variables Xi and Xj. If there exists another variable Xk such that, given variable Xk, variables Xi and Xj are conditionally independent, then delete the undirected edge EAij between variables Xi and Xj; otherwise... Retain the undirected edge EAij; repeat this process until all variables in the top and bottom layers have passed the conditional independence test; for those undirected edges that are still retained and connected to variables in the top or bottom layers, their direction is either from the top layer variable to the intermediate layer variable or from the intermediate layer variable to the bottom layer variable; select those nodes where directed edges have been established and perform conditional independence tests, determine the direction of the undirected edge between the two nodes according to the causal discovery rule, and repeatedly apply the causal discovery rule until all existing undirected edges have passed the conditional independence test and are marked with a definite or possible direction; for edges between variables whose direction still cannot be determined after the above steps, leave them unchanged.

[0008] The aforementioned existing technologies only determine the direction of undirected edges between two nodes using four main causal discovery rules. For edges between variables whose direction cannot be determined after the judgment steps, no further judgment operation is performed. The judgment capability of the aforementioned existing technologies is limited, and they cannot judge causal relationships that do not meet the four judgment rules. In contrast, this invention can determine the causal relationship through expert demonstration when the causal relationship retrieval of the target object fails, and update the causal relationship to the causal database. After the causal database is updated, the causal relationship of the target object is retrieved and corrected again. This achieves more accurate causal relationship analysis, and as the number of target objects increases and the causal database is continuously updated, the retrieval results become more accurate, and the number of retrieval failures decreases. Summary of the Invention

[0009] To address the shortcomings of existing technologies, this invention provides a causal relationship determination system based on Bayesian networks, comprising: a construction module for constructing a Bayesian network; a determination module for generating and outputting a causal relationship spectrum based on the requested Bayesian network; and a causal library for establishing a historical original literature database. When at least two objects experience at least one non-specific event, the construction module analyzes the self-factors of the objects causing the non-specific event to construct a factor set and constructs an undirected graph structure constraint based on the factor set. Furthermore, the construction module can modify the undirected graph structure constraint based on the causal library to establish the Bayesian network. The determination module calculates causal indicators of the self-factors causing the non-specific event based on the Bayesian network and outputs a causal relationship spectrum of the non-specific event based on the causal indicators, thereby determining the key causes of the non-specific event.

[0010] According to a preferred embodiment, the construction module defines at least one domain based on the non-specific event, and retrieves at least one exclusive factor related to the non-specific event in the domain based on the causal library; if the exclusive factor does not belong to the self-factor, the construction module prompts the third party to verify the evidence; if the third party determines that the exclusive factor objectively exists, the construction module adds the exclusive factor to the undirected graph structure constraint to further modify the undirected graph structure constraint.

[0011] According to a preferred embodiment, the construction module performs factor pair L based on the causal library. m and L n If the retrieval of the relationship between the factors fails, the construction module prompts the third party to check the relationship between the factors and L. m and L n The relationship between the factors is verified by experts, and the results of the expert verification are fed back to the causal database. The causal database then uses deep learning to refine itself based on the results of the expert verification. When the causal database uses deep learning to refine itself based on the results of the expert verification, the construction module adjusts the L factor according to the factors in the causal database. m and L n Numbering the factors to L m and L n The relationship between them needs to be corrected.

[0012] According to a preferred embodiment, the construction module establishes a dataset D based on the self-factors, and divides the self-factors into several factor sets L according to the dataset D, forming factor pairs L by pairing them up in a pairwise manner. m and L n And for the dataset D and the factor pair L m and Ln Numbering is performed; the construction module, based on the causal database, assigns factor pairs to factor pairs according to their numbering. m and L n The relationship between L is searched, and correction factors are applied to L based on the search results. m and L n The relationship between L and m →L n Relationship reliability score, L n →L m Relationship reliability score and L n ⊥L m The relational confidence values ​​are assigned to construct the undirected graph structure constraints.

[0013] According to a preferred embodiment, based on Bayes' theorem, the construction module constructs a Bayesian network evaluation function based on the dataset D and the factor pair set L. This evaluation function is used to iteratively generate the Bayesian network with the highest evaluation index from several candidate Bayesian networks, based on the undirected graph structure constraints and when the construction module uses a heuristic search algorithm. Specifically, the construction module first constructs an initial first candidate Bayesian network based on the undirected graph structure constraints and evaluates it using the Bayesian network evaluation function to obtain a first evaluation index. Subsequently, the construction module uses the heuristic search algorithm to locally modify the second candidate Bayesian network based on the undirected graph structure constraints and evaluates it again using the Bayesian network evaluation function to obtain a second evaluation index. The construction module can obtain at least two candidate Bayesian networks and their corresponding evaluation indices through iterative iteration using the heuristic search algorithm. When at least two candidate Bayesian networks and at least two evaluation indices are obtained, the construction module outputs the final Bayesian grid with the optimal evaluation index as the output causal relationship spectrum of non-specific events.

[0014] According to a preferred embodiment, the determination module calculates the effect of each factor on L based on the final Bayesian grid and Pearl's principle. m and L n The causal indicators between factors L are used to output the causal relationship spectrum, wherein for factor L m The relationship between the factors L and the undirected graph structure constraints is obtained through traversal. m The connected nodes form its node set; and the relationship between each node and factor L is calculated sequentially. m The correlation between nodes is analyzed, and the nodes with the highest correlation are selected. Independence assumptions are then made, and nodes with the highest correlation in the given dataset D are removed. m Independent nodes are used to improve the decision-making efficiency of the decision-making module; Factor L n With factor Lm The independence between them is measured by mutual information:

[0015]

[0016] When the mutual information exceeds the mutual information threshold, then factor L n With factor L m They are correlated but not very independent; if the mutual information does not exceed the threshold of mutual information, then factor L... n With factor L m They are not related and are independent.

[0017] According to a preferred embodiment, the causal database is established as follows: the causal database is constructed by classifying numerous relevant documents containing multiple historical attributes according to the technical field to form several document units, thereby constructing the original document database, and mining the reliability values ​​of the relationships between attributes through data patterns; wherein, the document layer in the causal database counts the frequency of words / phrases in each document, and obtains the joint occurrence probability of the words / phrases according to the independence assumption; the document layer calculates the correlation strength of the words / phrases; the document layer constructs the correlation reduced coordinates of the documents, and classifies the relevant documents according to the form of an iterative algorithm based on the correlation reduced coordinates and correlation strength of all the relevant documents to form several document units.

[0018] According to a preferred embodiment, when the data layer in the causal database obtains the document unit, the data layer obtains the dataset by pairing two historical attributes; the data layer extracts the relationship between two historical attributes for each relevant document using syntactic analysis of natural language processing to establish a relationship knowledge base for the two historical attributes, the relationship between the two historical attributes including positive, negative, and vertical relationships; furthermore, the data layer searches for documents containing the two histories within the document unit based on the relationship knowledge table and obtains the relationship reliability values ​​of the two historical attributes in a fusion manner to establish a relationship reliability value database for the two historical attributes, the relationship between the two historical attributes including positive, negative, and vertical relationship reliability values; thus, the data layer constructs the historical dataset based on the relationship knowledge base and relationship reliability value database established by pairing all histories.

[0019] According to a preferred embodiment, the present invention also discloses a causal relationship determination method based on Bayesian networks. The method includes: a construction module constructing a Bayesian network; a determination module generating and outputting a causal relationship spectrum based on the requested Bayesian network; a causal library establishing a historical original literature library; in the case that at least two objects experience at least one non-specific event, the construction module analyzes the self-factors of the objects causing the non-specific event to construct a factor set and constructs an undirected graph structure constraint based on the factor set; and the construction module can modify the undirected graph structure constraint based on the causal library to establish the Bayesian network; the determination module calculates the causal index of the self-factors causing the non-specific event based on the Bayesian network and outputs the causal relationship spectrum of the non-specific event based on the causal index, thereby determining the key cause of the non-specific event.

[0020] According to a preferred embodiment, the method further includes: the construction module defining at least one domain based on the non-specific event, and retrieving at least one exclusive factor related to the non-specific event in the domain based on the causal library; if the exclusive factor does not belong to the self-factor, the construction module prompts the third party to verify the evidence; if the third party determines that the exclusive factor objectively exists, the construction module adds the exclusive factor to the undirected graph structure constraint to further modify the undirected graph structure constraint.

[0021] The main advantage of this invention lies in constructing an undirected graph constraint for a Bayesian network of factors related to at least one non-specific event occurring between at least two objects. This undirected graph constraint is then modified based on a literature knowledge base and an expert knowledge base. A heuristic search algorithm is then used to construct the Bayesian network based on this constraint. The structure learning of this Bayesian network is primarily used to reveal qualitative relationships between variables, while also revealing quantitative relationships. However, constructing a Bayesian network solely from a data perspective presents numerous difficulties. Relying on non-specific events between objects and their historical events can increase the accuracy of the Bayesian network construction. Finally, causal analysis and reasoning are implemented based on the constructed Bayesian network. In summary, the key problems this invention aims to solve are: integrating object knowledge to construct a reasonable Bayesian network and causal analysis and reasoning; determining the causal relationship spectrum of non-specific events; and identifying key causes. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a preferred embodiment of the determination method provided by the present invention; and

[0023] Figure 2 This is a schematic diagram of a preferred module of the determination system provided by the present invention.

[0024] List of reference numerals

[0025] 1: Construction Module 2: Decision Module

[0026] 3: Causal Database Detailed Implementation

[0027] The following is in conjunction with the appendix Figure 1 and 2 Please provide a detailed explanation.

[0028] In the description of this invention, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," and "third" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0029] Example 1

[0030] This embodiment provides a causal relationship determination system based on Bayesian networks, aiming to output the causal relationship spectrum when at least two objects have experienced at least one non-specific event. The main advantage of this invention lies in constructing an undirected graph constraint for the Bayesian network between factors of at least two objects for at least one non-specific event, and modifying this undirected graph constraint based on a literature knowledge base and an expert knowledge base. Then, a heuristic search algorithm is used to construct the Bayesian network based on this undirected graph constraint. The structure learning of this Bayesian network is mainly used to reveal qualitative relationships between variables, while also revealing quantitative relationships. However, constructing a Bayesian network solely from a data perspective presents many difficulties. Relying on non-specific events between objects and their historical events can increase the accuracy of the Bayesian network construction. Finally, causal analysis and reasoning are realized based on the constructed Bayesian network. In summary, the key problems to be solved by this invention are: integrating object knowledge to construct a reasonable Bayesian network and causal analysis and reasoning, determining the causal relationship spectrum of non-specific events, and identifying key causes.

[0031] Specifically, the system includes a construction module 1, a decision module 2, and a causal database 3. Construction module 1 is used to construct a Bayesian network. Decision module 2 is used to generate and output a causal relationship spectrum based on the requested Bayesian network. Causal database 3 is used to establish a historical original literature database. For at least two objects in at least one non-specific event, construction module 1 analyzes the intrinsic factors of the objects causing the non-specific event to construct a factor set and builds an undirected graph structure constraint based on the factor set. Furthermore, construction module 1 can modify the undirected graph structure constraint based on causal database 3 to establish a Bayesian network. Decision module 2 calculates the causal indices of the non-specific events caused by the intrinsic factors based on the Bayesian network and outputs the causal relationship spectrum of the non-specific events based on the causal indices, thereby determining the key causes of the non-specific events. In this invention, non-specific events can be real events, such as a collision between two vehicles. Non-specific events can also be events anticipated by researchers, such as a spacecraft docking failure.

[0032] Preferably, the construction module 1 defines at least one domain based on non-specific events, and retrieves at least one other factor related to the non-specific events in the domain based on the causal database 3. If the other factor is not a self-factor, the construction module 1 prompts a third party to verify the evidence; if the third party confirms the objective existence of the other factor, the construction module 1 adds the other factor to the undirected graph structure constraint to further modify the undirected graph structure constraint.

[0033] Preferably, module 1 constructs factor pairs L based on the causal library 3. m and L n If the retrieval of the relationship between factors fails, Module 1 prompts a third party to check the relationship between factors and L. m and L n The relationships between factors are verified by experts, and the results of the expert verification are fed back to causal database 3. Causal database 3 then uses deep learning to refine its causal database based on the expert verification results. With causal database 3 using deep learning to refine its causal database based on the expert verification results, module 1 constructs modules based on causal database 3 to classify L according to factors. m and L n Numbering of factors to L m and L n The relationship between them needs to be corrected.

[0034] Preferably, module 1 establishes a dataset D based on its own factors, and divides its own factors into several factor sets L according to dataset D, forming factor pairs L by pairing them up in a pairwise manner. m and L n And for dataset D and factors on L m and L n Numbering is performed. Module 1, based on the causal library 3, assigns factor pairs to L based on their numbering. m and Ln The relationship between L is searched, and correction factors are applied to L based on the search results. m and L n The relationship between L and m →L n Relationship reliability score, L n →L m Relationship reliability score and L n ⊥L m The relational confidence values ​​are assigned to construct the undirected graph structure constraints.

[0035] Preferably, according to Bayes' theorem, construction module 1 constructs a Bayesian network evaluation function based on dataset D and factor pair set L. This evaluation function is used to iteratively generate the Bayesian network with the highest evaluation index from several candidate Bayesian networks, based on undirected graph structure constraints and with construction module 1 employing a heuristic search algorithm. Preferably, construction module 1 first constructs an initial first candidate Bayesian network based on undirected graph structure constraints and evaluates it using the Bayesian network evaluation function to obtain a first evaluation index. Subsequently, construction module 1 uses a heuristic search algorithm to locally modify the second candidate Bayesian network based on undirected graph structure constraints, and evaluates the second candidate Bayesian network again using the Bayesian network evaluation function to obtain a second evaluation index. Construction module 1 can obtain at least two candidate Bayesian networks and their corresponding evaluation indices through an iterative process using a heuristic search algorithm. Having obtained at least two candidate Bayesian networks and at least two evaluation indices, construction module 1 outputs the final Bayesian grid with the optimal evaluation index as the output causal relationship spectrum for non-specific events.

[0036] Preferably, the determination module 2 calculates the impact of each factor on L based on the final Bayesian grid and Pearl's principle. m and L n The causal indicators between factors are used to output a causal relationship spectrum. Specifically, for factor L... m By traversing through undirected graphs and using structural constraints, factors L are obtained. m Connected nodes form its node set. The relationship between each node and factor L is calculated sequentially. m The correlation between nodes is analyzed, and the nodes with the highest correlation are selected. Independence assumptions are then made, and nodes with the highest correlation in the given dataset D are removed. m Independent nodes are used to improve the decision-making efficiency of decision-making module 2. Factor L n With factor L m The independence between them is measured by mutual information:

[0037]

[0038] When mutual information exceeds the mutual information threshold, then factor Ln With factor L m They are correlated but not very independent; if the mutual information does not exceed the mutual information threshold, then factor L... n With factor L m They are not related and are independent.

[0039] Preferably, the causal database 3 is established in the following manner: the causal database 3 is constructed by classifying numerous relevant documents with multiple historical attributes according to the technical field to form several document units, thereby constructing an original document database, and mining the reliability values ​​of the relationships between attributes through data patterns; wherein, the document layer in the causal database 3 counts the frequency of words / phrases in each document, and obtains the joint occurrence probability of words / phrases according to the independence assumption; the document layer calculates the correlation strength of words / phrases; the document layer constructs the correlation reduction coordinates of documents, and classifies the relevant documents according to the form of an iterative algorithm based on the correlation reduction coordinates and correlation strength of all relevant documents to form several document units.

[0040] Preferably, when the data layer in the causal database 3 has obtained the document unit body, the data layer obtains the dataset by pairing two historical attributes; the data layer extracts the relationship between two historical attributes for each relevant document using syntactic analysis of natural language processing to establish a relationship knowledge base for the two historical attributes, the relationship between the two historical attributes includes positive relationship, negative relationship and vertical relationship; furthermore, the data layer searches for documents containing two histories within the document unit body based on the relationship knowledge table and obtains the relationship reliability value of the two historical attributes in a fusion manner to establish a relationship reliability value database for the two historical attributes, the relationship between the two historical attributes includes positive relationship reliability value, negative relationship reliability value and vertical relationship reliability value; thus, the data layer constructs the historical dataset based on the relationship knowledge base and relationship reliability value database established by pairing all histories.

[0041] Example 2

[0042] This embodiment discloses a causal relationship determination method based on Bayesian networks for determining liability in medical disputes. Where there is no conflict or contradiction, the whole and / or parts of preferred embodiments of other embodiments can be used as supplements to this embodiment. Preferably, this method can be implemented by the method of this invention and / or other alternative modules.

[0043] "Medical disturbances" refer to acts by individuals hired by patients involved in medical disputes to exert pressure on hospitals and profit from them by severely disrupting medical order, escalating the situation, and negatively impacting the hospital. The direct consequence of medical disturbances is a significant loss of medical personnel in my country, both directly and indirectly, with extremely serious and negative consequences, severely hindering the development of my country's medical profession. Defining the causes of medical disputes is difficult. Medical disputes refer to disputes occurring within legally qualified medical enterprises or institutions in sectors such as medical and health care, preventive healthcare, and cosmetic medicine. Currently, medical disputes in China are particularly difficult to handle. The root cause is that medical disputes are usually caused by medical errors and negligence. Medical errors are mistakes made by medical personnel during diagnosis and nursing care. Medical negligence refers to errors made by medical personnel in medical activities such as diagnosis and treatment. These errors often lead to patient dissatisfaction or harm, thus causing medical disputes. Besides medical disputes arising from medical malpractice and negligence, disputes can also arise sometimes even when the medical staff has made no mistakes or errors in their actions, but rather due to the patient's unilateral dissatisfaction. These disputes can stem from the patient's lack of basic medical knowledge, misunderstanding of proper medical treatment, the natural course of disease, unavoidable complications, and unexpected medical accidents; or they can arise from the patient's unreasonable accusations. These are sometimes referred to as medical malpractice disputes, namely, disputes between the provider and recipient of medical services regarding whether medical actions and their consequences constitute infringement and the liability for such infringement. Therefore, to provide medical professionals with a comfortable and healthy working environment and patients with a fair and just explanation, the causes of such medical disputes need to be presented in a transparent manner, ensuring transparency and fairness.

[0044] Therefore, this embodiment provides a Bayesian network-based causality determination method to assist judicial departments in resolving medical disputes. In judicial practice regarding medical disputes, the adoption of evidence or factors asserted by both parties is often based on a legal perspective, while the relationship between evidence or factors is often assessed qualitatively. This is one of the reasons for the entanglement between doctors and patients, and also one of the reasons that judicial personnel struggle to make rulings or judgments. In cases of entanglement between doctors and patients, both parties may question or even appeal the rulings or judgments, consuming excessive legal resources. In an era that advocates the rule of law values ​​of "fairness and justice," resolving conflicts between doctors and patients through "data-driven" methods, restoring a relaxed working environment for doctors to save lives, providing patients or their families with a convincing explanation, and providing judicial institutions with a scientific reference document—this is the important value of this invention.

[0045] Specifically, the method mainly includes:

[0046] S1: Module 1: Constructing a Bayesian network. Specifically, Module 1 constructs an undirected graph structure constraint based on factors claimed by the medical staff and the patient. Furthermore, in the case of third-party intervention, Module 1 modifies the undirected graph structure constraint based on the causal library 3 to establish the Bayesian network.

[0047] S2: Judgment Module 2: Outputs a causal relationship spectrum based on a Bayesian network. Judgment Module 2 calculates the causal indicators between each pair of factors based on a Bayesian network and outputs a causal relationship spectrum of medical disputes based on the causal indicators. This ensures that the causal relationship spectrum takes into account both the factors claimed by both the doctor and the patient and considers scientific validity, effectively preventing the escalation of disputes between the doctor and the patient and providing data support for third parties to make rulings or judgments.

[0048] Preferably, step S1 specifically includes the following steps:

[0049] S11: Establish a dataset D = (D1, D2, ..., D4) based on the factors claimed by the medical staff and the patient. i ) represents several sets of attributes. L = (L1, L2, ..., L...) n The specific factor pairing of a certain set of attributes. For example, the factors claimed by the medical staff include the patient's illness time, medical treatment time, disease severity, and disease type. The dataset can then be constructed using time attributes and disease attributes to create a dataset D = (time attribute, disease attribute). The specific factor pairing L represents the specific attributes corresponding to the illness time and medical treatment time of the time factor, and these attributes are numbered.

[0050] S12: Construct an undirected graph structure constraint based on the relationships between the factors claimed by the medical staff and the patient. The relationships between specific factor pairs are determined by retrieving the relationships between the factors claimed by the medical staff and the patient, and the factor pairs are numbered. Preferably, the relationships between factor pairs include positive relationships, negative relationships, and vertical relationships, i.e., attribute L. m Influence attribute L n , denoted as L m →L n Attribute L m and attribute L n The possible relationship is an inverse relationship, i.e., attribute L. n Influence attribute L m , denoted as L n →L m Attribute L m and attribute L n The relationship that appears may be a vertical relationship, i.e., attribute L m With attribute L n They do not affect each other. n ⊥L mFor example, if the patient believes that the medication's components worsened their cerebral palsy, then the medication's components have a positive relationship with worsening cerebral palsy, recorded as medication components → worsening cerebral palsy. Alternatively, one could consider the worsening of cerebral palsy to have a negative relationship with the medication's components, also recorded as medication components → worsening cerebral palsy.

[0051] S13: Modify the structural constraints of the undirected graph based on the causal database. Retrieve factor pairs by their IDs in the causal database, and then adjust the factor pairs L based on the retrieved values. m and L n The relationship was corrected, and L m →L n Relationship reliability score, L n →L m Relationship reliability score and L n ⊥L m The relationship reliability value is assigned accordingly. For example, if the patient believes that a drug component will worsen cerebral palsy, but a search of the causal database shows that the drug component will not worsen cerebral palsy, then the relationship "drug component → worsening cerebral palsy" is corrected to "drug component ⊥ worsening cerebral palsy". For example, if the drug component is L1 and the worsening of cerebral palsy is L2, then the relationship "drug component → worsening cerebral palsy" is numbered 12.

[0052] Preferably, the construction module 1 defines at least one request domain based on the factors claimed by the medical staff and the patient, and retrieves at least one other factor related to the factors claimed by the medical staff and the patient based on the request domain. If the at least one factor is not among the factors claimed by the medical staff and the patient, the construction module 1 prompts a third party to investigate the evidence. If the third party determines that the at least one factor objectively exists, the construction module 1 adds the at least one factor to the undirected graph structure constraint to further modify the undirected graph structure constraint. For example, in the event of neonatal death, the factors claimed by the medical staff and the patient include amniotic fluid embolism, hypoxia, and multiple pregnancies. The request domain defined by the construction module 1 based on these factors is childbirth. The construction module 1 then searches for other factors related to the incident in the corresponding childbirth domain in the causal database 3. For example, it retrieves "thin uterine wall," but "thin uterine wall" does not appear among the factors claimed by the medical staff and the patient. The construction module 1 prompts a third party to investigate whether the mother has a thin uterine wall. If this condition objectively exists, the construction module 1 adds the factor of "thin uterine wall" to further modify the undirected graph structure constraint. In step S13, there is another possibility that at least one factor that neither party has provided evidence for is an important factor affecting the result. In this case, the construction module 1 will prompt the third party to investigate the at least one factor. If the at least one factor has objectively occurred, the at least one factor needs to be added to the undirected graph structure constraint to modify it, so as to increase the reliability and scientific nature of the result, enhance the fairness of the third party's judgment or ruling, and demonstrate the rigor, fairness and responsibility of the third party.

[0053] Preferably, the causal database 3 may not contain a certain factor pair or the relationship between some factor pairs and its relationship reliability value. This is to ensure that the factors claimed by both parties are supported. That is, module 1 constructs the factor pair L based on the causal database 3. m and L n If the retrieval of the relationship between factors fails, Module 1 prompts a third party to check the relationship between factors and L. m and L n The relationship between the factors needs to be verified by experts. For example, the patient's family argues that the size of the newborn's head was not the cause of the newborn's death. However, the causal database 3 does not contain a relationship between the newborn's head size and the newborn's death. Therefore, module 1 will prompt a third party to conduct an expert verification. The results of the expert verification will be fed back to causal database 3, which will then use deep learning to refine its causal database. With causal database 3 using deep learning to refine its causal database, module 1 will then adjust the L factor classification based on causal database 3. m and L n Numbering of factors to L m and L n The relationship between them needs to be corrected.

[0054] Preferably, module 1 constructs a Bayesian network evaluation function based on the causal library 3, dataset D, and factor pair set L:

[0055] logP(G,D,K L )=logP(G)+logP(D|G)+logP(K L |G)

[0056] Preferably, module 1 constructs a request Bayesian network based on the Bayesian network evaluation function and undirected graph structure constraints. In the formula, G is the Bayesian grid, and its values ​​include values ​​in the order L = (L1, L2...L...). n A directed acyclic graph (DAG) is a graph of nodes formed by combining specific factors of a certain set of attributes. Here, P(G) is the prior distribution. Based on existing knowledge, the maximum value of logP(G) + logP(D|G) is equivalent to logP(G|D). logP(G|D) can be scored according to the Bayesian Information Criterion (BIC). In the formula,

[0057]

[0058] In this case, any edge in structure G is represented as L. m →L n ,but KL(L m →L nThis represents the relational reliability value. The summation in the formula sums the literature knowledge reliability of all directed relations corresponding to the directed edges in structure G. For a given dataset D, for any factor in D, the relational reliability is calculated for L. m and L n A Bayesian network is constructed based on the Bayesian network evaluation function and the undirected graph structure constraints. After determining the undirected graph structure constraints of the Bayesian network, a heuristic search algorithm, such as the K2 algorithm, can be executed to seek the network structure with the optimal evaluation function. The general steps are as follows: start the search from the initial model. At each step of the search, first, the current model is locally modified using search operators to obtain a series of candidate models. Then, the score of each candidate model is calculated, and the optimal candidate model is compared with the current model. If the score of the optimal candidate model is larger, it is adopted as the next current model, and the search continues; otherwise, the search stops, and the current model is returned. According to the Bayesian principle, the candidate model with the largest score is the Bayesian network. Preferably, the Bayesian network evaluation function is constructed based on the established Bayesian network and Bayesian rules. The Bayesian network evaluation function can be constructed based on classic heuristic structure learning algorithms, such as the K2 algorithm, the Max-Min Parents and Children algorithm, and Markov chain Monte Carlo search, etc.

[0059] Preferably, the determination module 2 calculates the causal indicators between each pair of factors based on request Bayesian networks and Pearl's principle, thereby outputting a causal relationship spectrum. The determination module 2 is based on mining causal indicators between attributes through data pattern analysis, enabling it to determine whether the causal indicators indicate whether the attributes constitute complications or comorbidities. When determining causal indicators, the determination module 2 calculates the causal indicators between attributes based on Pearl's principle and Bayesian network structure. When exploring whether event X is the cause of event Y, Pearl needs to intervene in X to implement event X, calculating E(Y|do(X)), that is, if the average change of event Y under the intervention of X is greater than the significance level, then X is considered the cause of Y.

[0060] In the decision module 2, based on mining causal indicators between attributes through data patterns, the sheer volume of literature results in a massive Bayesian grid. Therefore, a backdoor criterion is employed to calculate the causal indicators. The backdoor criterion states that the Bayesian grid G ​​is a directed acyclic graph (L...). m L n () is a pair of nodes in G, and the set of nodes Z is (L) m L n The backdoor of Z is such that none of the nodes in Z are descendants of Z, and Z blocks all pointers to L. m L connection m To L n Therefore, the causal relationship between factors on Lm and Ln can be inferred using the backdoor principle.

[0061] To simplify the undirected graph constraints without affecting the causal relationship between factor pairs, module 2 uses an independence test. For example, the independence test can be a chi-square independence test. In this invention, the independence test can also be performed as follows: for factor L... m The L-axis is obtained by compiling the structure of the undirected graph. m Connected nodes form its node set. The relationship between each node and factor L is calculated sequentially. m The correlation between nodes is analyzed, and the nodes with the highest correlation are selected and their independence is assumed. Nodes in the given request subset D are then removed. i Below and L m Independent nodes. In this invention, entropy is used to measure the factor set L. m Uncertainty. Given factor L m In the case of factor L n The uncertainty can be measured using conditional entropy as follows:

[0062]

[0063] Factor L n With L m The degree of correlation between them can be measured by mutual information:

[0064]

[0065] If the mutual information exceeds the mutual information threshold, then L is considered to be... n With L m It is correlated. If the mutual information does not exceed the mutual information threshold, then L is considered to be... n With L m It is not relevant.

[0066] Preferably, the causal database 3 is established as follows: Based on numerous relevant documents with multiple historical attributes acquired according to the technical field, the causal database 3 is classified into several document units to construct the original document database, thereby mining the reliability values ​​of relationships between historical attributes through data patterns. Specifically, the document layer in the causal database 3 counts the frequency of words / phrases in each document and obtains the joint occurrence probability of words / phrases according to the independence assumption. The document layer calculates the correlation strength of words / phrases; the document layer constructs the correlation reduction coordinates of documents, and based on the correlation reduction coordinates and correlation strength of all relevant documents, it classifies the relevant documents according to an iterative algorithm to form several document units.

[0067] Preferably, the document layer is based on numerous relevant documents containing various historical attributes. The document layer categorizes these relevant documents into several document units to construct the original document database. These relevant documents include medical records, research reports, conference proceedings, journal articles, books, academic papers, and patents. Given such a large volume of documents, they need to be classified using specific methods. Document classification aims to effectively observe the relationships between historical attributes and reduce the system load. For example, they can be classified according to digestive diseases, cardiovascular diseases, and neurological diseases. They can also be classified according to academic fields, such as rehabilitation medicine and psychology. However, given the large volume of documents, accurate and efficient classification directly affects the differentiation of complications and comorbidities. Preferably, document classification can employ Bayesian methods, SVM methods, and k-NN methods.

[0068] Preferably, the relevant literature classification is performed as follows: the frequency of words / phrases in each literature is counted at the literature level, and the joint occurrence probability of words / phrases is obtained according to the independence assumption. For example, the joint occurrence probability distribution of a specific literature can be calculated using the Naive Bayes method.

[0069] Preferably, the document layer calculates the correlation strength of words / phrases. Calculating the correlation strength reflects the relevance of words / phrases, which is suitable for document classification. Preferably, in classification, N is defined as the set of document samples, and V is the set of document types. i It is a subset of the i-th document type. W is the set of words / phrases. i It is a subset of the i-th word / phrase. In V i Contains S j There are n samples, where the reduced coordinates T of the p-th sample are... p It is an n-dimensional array:

[0070]

[0071] Where, k i The number of occurrences of the i-th word in (i = 1, 2, 3, ..., n) Normalization coefficient.

[0072] In V i The correlation vector is all V i The average of the reduced coordinates of the mid-sample associations reflects the strength of the associations between words / phrases in the literature.

[0073]

[0074] Preferably, the document layer obtains the reduced coordinates of the documents and, based on the classification function constructed from the reduced coordinates of all related documents, classifies the related documents according to an iterative algorithm to form several document units. Preferably, the reduced coordinates of any document are:

[0075]

[0076] In the formula, q i This represents the number of times the i-th word appears in the document. During classification, the document to be classified is compared with each category of documents, V. i The support points (b1, b2, ..., b) n The distance is denoted as:

[0077]

[0078] Based on the strength of the correlation, construct a document classification function:

[0079]

[0080] In the formula, γ i Related to the strength of the correlation.

[0081] Preferably, the iterative algorithm can employ a minimization iterative algorithm, a minimum optimization iterative algorithm, or an expectation-maximization iterative algorithm. Preferably, the classification function can be based on the sample size of relevant literature for deep learning, thereby enhancing the accuracy of the literature layer.

[0082] Preferably, when the data layer in the causal database 3 has obtained document units, the data layer obtains the historical dataset by pairing two historical attributes. The data layer extracts the relationship between two historical attributes for each relevant document using syntactic analysis in natural language processing to establish a relationship knowledge base for the two historical attributes. The relationship between the two historical attributes includes positive, negative, and vertical relationships. Furthermore, based on the relationship knowledge table, the data layer searches within the document unit for documents containing two histories and obtains the relationship reliability values ​​of the two historical attributes through fusion to establish a relationship reliability value database for the two historical attributes. The relationship between the two historical attributes includes positive, negative, and vertical relationship reliability values. Thus, the data layer constructs the historical dataset based on the relationship knowledge base and relationship reliability value database established by pairing all histories.

[0083] Preferably, the data layer can obtain key feature parameters based on document units and construct historical datasets based on these key feature parameters. This reduces the interference of numerous feature parameters formed by numerous related documents on the causal relationships between historical attributes and improves the utilization value of the original document database. Preferably, when the data layer obtains document units, it obtains historical datasets by pairing two historical attributes. The data layer extracts the relationship between two historical attributes for each related document using syntactic analysis in natural language processing to establish a relationship knowledge base for the two historical attributes. The relationship between the two historical attributes includes positive, negative, and vertical relationships. Furthermore, based on the relationship knowledge table, the data layer searches for documents containing two historical attributes within the document unit and obtains the relationship reliability values ​​of the two historical attributes in a fusion manner to establish a relationship reliability value database for the two historical attributes. The relationship between the two historical attributes includes positive, negative, and vertical relationship reliability values. Thus, the data layer constructs historical datasets based on the relationship knowledge base and relationship reliability value database established by pairing all historical attributes. For example, several historical attributes, such as historical attribute L1, historical attribute L2, historical attribute L3, and historical attribute L4, are obtained from relevant literature. Based on the relationships between these historical attributes, a relationship knowledge base can be constructed for historical attributes L1 and L3, and another for historical attributes L2 and L3, and so on. Then, within a unit document, a relationship reliability value database is constructed based on the aforementioned relationship knowledge base according to the content of different documents. Preferably, the sum of the positive relationship reliability value, the negative relationship reliability value, and the vertical relationship reliability value is normalized. That is, within a unit document, all documents are traversed and queried, and the positive relationship reliability value, the negative relationship reliability value, and the vertical relationship reliability value are weighted according to frequency. The data layer inputs the aforementioned relationship knowledge base and relationship reliability value database into the judgment module 2 for the next step.

[0084] Preferably, for journal articles, the reliability value of the positive relationship between L1 and L2 can also be defined as follows:

[0085]

[0086] Where C(Xi) represents the credibility of document Xi, calculated as: C(Xi) = (IFi+1) × (CIi+1), where Xi represents the i-th document, IFi is the standardized impact factor of the journal containing document Xi, and CIi is the standardized citation count. If no document has an L1-L2 relationship, then KL(L1→L2) = 0, KL(L2→L1) = 0, and KL(L1⊥L2) = 1. Other types of documents can be defined in the same way; for example, medical records can be defined based on the doctor's authority, and conference articles can be defined based on the conference's authority, and so on.

[0087] Example 3

[0088] This embodiment discloses a causal relationship determination system based on Bayesian networks. Without causing conflict or contradiction, the whole and / or parts of preferred embodiments of other embodiments can be used as supplements to this embodiment. Preferably, this method can be implemented by the method of this invention and / or other alternative modules.

[0089] This embodiment provides a causal relationship determination method based on Bayesian networks. The method includes: a construction module 1 constructing a Bayesian network; a determination module 2 generating and outputting a causal relationship spectrum based on the requested Bayesian network; and a causal database 3 establishing a historical original literature database. When at least two objects experience at least one non-specific event, the construction module 1 analyzes the self-factors of the objects causing the non-specific event to construct a factor set and constructs an undirected graph structure constraint based on the factor set; the construction module 1 can also modify the undirected graph structure constraint based on the causal database 3 to establish a Bayesian network; the determination module 2 calculates the causal index of the non-specific event caused by its own factors based on the Bayesian network and outputs the causal relationship spectrum of the non-specific event based on the causal index, thereby determining the key causes of the non-specific event.

[0090] Preferably, the construction module 1 defines at least one domain based on non-specific events, and retrieves at least one exclusive factor related to the non-specific events within the domain based on the causal database 3. If the exclusive factor is not an intrinsic factor, the construction module 1 prompts a third party to verify the evidence. If the third party confirms the objective existence of the exclusive factor, the construction module 1 adds the exclusive factor to the undirected graph structure constraint to further modify the undirected graph structure constraint.

[0091] Preferably, the construction module 1 used in this invention is a server with a search engine and computing capabilities. The judgment module 2 is a data server with computing capabilities. The causal database 3 is a server with a search engine, computing capabilities, and storage. The construction module 1, judgment module 2, and causal database 3 are interconnected via wired or wireless means such as optical fiber, data cable, Bluetooth, Wi-Fi, and / or 4G.

[0092] It should be noted that the specific embodiments described above are exemplary, and those skilled in the art can devise various solutions inspired by the disclosure of this invention. These solutions all fall within the scope of this invention and its protection. Those skilled in the art should understand that this specification and its accompanying drawings are illustrative and not intended to limit the scope of the claims. The scope of protection of this invention is defined by the claims and their equivalents.

Claims

1. An abductive reasoning system, comprising at least: Module (1) is used to build Bayesian networks; Causal database (3) is used to establish a historical original document database. The feature is that the construction module (1) is configured to perform deep learning on the results of expert argumentation based on the causal library (3), and then, based on the causal library (3), classify L according to factors. m and L n Numbering of factors to L m and L n The relationship between objects is modified. In the case that at least two objects experience at least one non-specific event, the construction module (1) analyzes the self-factors of the objects that caused the non-specific events to construct a factor set and constructs an undirected graph structure constraint based on the factor set; and the construction module (1) can modify the undirected graph structure constraint based on the causal library (3) to establish a Bayesian network. The construction module (1) constructs a dataset D based on the self-factors of the objects, and divides the self-factors into several factor sets L according to the dataset D as the unit and forms factor pairs L in a pairwise pairing manner. m and L n And for the dataset D and the factor pair L m and L n Number them.

2. The abductive reasoning system according to claim 1, characterized in that, The system also includes a determination module (2), which calculates the causal index of non-specific events caused by its own factors based on a Bayesian network and outputs the causal relationship spectrum of non-specific events based on the causal index, thereby determining the key causes of non-specific events.

3. The abductive reasoning system according to claim 2, characterized in that, Module (1) is built based on the causal library (3) according to the factor pair numbering pair factor pair L m and L n The relationship between L is searched, and correction factors are applied to L based on the search results. m and L n The relationship between L and m →L n Relationship reliability score, L n →L m Relationship reliability score and L n ⊥L m Assign a value to the relationship reliability score.

4. The abductive reasoning system according to claim 3, characterized in that, The data layer of the causal library (3) retrieves documents containing two histories within the document unit based on the relational knowledge table and obtains the relational reliability values ​​of the two historical attributes in a fusion manner to establish a relational reliability value library of the two historical attributes. Thus, the data layer constructs a historical dataset based on the relational knowledge library and relational reliability value library established by pairing all histories.

5. The abductive reasoning system according to claim 4, characterized in that, According to Bayes' theorem, the building module (1) constructs a Bayesian network evaluation function based on the dataset D and the factor pair set L. The evaluation function is used to iteratively generate the Bayesian network with the highest evaluation index from several candidate Bayesian networks based on the undirected graph structure constraints and with the heuristic search algorithm enabled in the building module (1).

6. The abductive reasoning system according to claim 5, characterized in that, The decision module (2) calculates the impact of each factor on L based on the final Bayesian grid and Pearl's principle. m and L n The causal indicators between them are used to output a causal relationship spectrum.

7. The abductive reasoning system according to claim 6, characterized in that, Factor L n With factor L m The independence between them is measured by mutual information: When mutual information exceeds the mutual information threshold, then factor L n With factor L m They are correlated but not very independent; if the mutual information does not exceed the mutual information threshold, then factor L... n With factor L m They are not related and are independent.

8. An abductive reasoning method, characterized in that, At least including: Constructing Bayesian networks; Establish a historical original document database using causal databases; With deep learning applied to the causal database based on the results of expert argumentation, L is categorized by factors according to the causal database. m and L n Numbering of factors to L m and L n The relationship between them needs to be corrected; When at least two objects experience at least one non-specific event, the intrinsic factors of the objects causing the non-specific events are analyzed to construct a factor set, and an undirected graph structure constraint is constructed based on the factor set. Furthermore, the undirected graph structure constraint is modified based on a causal library to establish a Bayesian network. A dataset D is constructed based on the intrinsic factors of the objects, and the intrinsic factors are divided into several factor sets L according to the dataset D, forming factor pairs L in a pairwise pairing manner. m and L n And for the dataset D and the factor pair L m and L n Number them.

Citation Information

Patent Citations

  • Reason tracing method

    CN105468703A

  • Reliability modeling method of aviation power system based on Bayesian network

    CN109063253A

  • Bayesian network and ontology combined reasoning method capable of self-perfecting network structure

    CN102360457A

  • Evaluation indicator equilibrium state analysis method based on Bayesian causal network

    CN107563596A