Knowledge spectrum analysis visualization processing system and method based on multi-source information fusion
Through the multi-source information fusion knowledge spectrum analysis visual processing system, the problem of difficult medical knowledge graph update is solved, efficient knowledge management and utilization is achieved, and data redundancy and waste of algorithm resources are reduced.
Patent Information
- Application Number
- CN202411397517.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-10-09
AI Technical Summary
The update of existing medical knowledge graphs is difficult to cope with the problems of in-depth medical research and rapid data growth, resulting in data redundancy and waste of algorithm resources, affecting the promotion and utilization of knowledge graphs.
The knowledge spectrum analysis visual processing system based on multi-source information fusion is adopted. The database module recognizes feature entities and divides the minimum storage space. The feature relationship processing module depicts undirected relationship edges and attaches relationship feature labels. The feature relationship analysis module calculates edge weights, and performs periodic updates and stability evaluations through the graph update analysis module to update the knowledge graph simultaneously.
Effectively manage and utilize medical knowledge, reduce data redundancy, improve knowledge graph update efficiency, save resource space, and promote the promotion and reuse of knowledge graphs.
Smart Images

Figure CN119358650B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical spectrum technology, and specifically to a knowledge spectrum analysis visualization processing system and method based on multi-source information fusion. Background Art
[0002] Knowledge graph is a technology that displays the relationships between entities in a graphical way. It provides a new perspective and solution for the medical field by integrating and analyzing large amounts of data. Knowledge graphs in the medical field usually contain entities such as drugs, diseases, symptoms, treatment methods, and depict the relationships between entities. This structured data display method not only helps to better understand complex medical information, but also plays an important role in drug development, clinical decision support, disease prevention, etc.
[0003] In the prior art, the update of knowledge graphs mainly includes the addition, deletion and modification of entities, the addition, deletion and modification of relationships, and the optimization of graph structures. However, with the deepening of medical research and the continuous growth of data, medical knowledge has shown explosive growth, which has put forward higher requirements for the management and utilization of medical knowledge. In view of this, in order to ensure the latest and readiness of data, it is necessary to update the data source at all times. The huge amount of data system will inevitably require more data resource space. At the same time, it is difficult to avoid the waste of algorithm resources caused by redundant data, which is not conducive to the promotion of knowledge graphs in the medical field. Summary of the invention
[0004] The purpose of the present invention is to provide a knowledge spectrum graph analysis visualization processing system and method based on multi-source information fusion to solve the problems raised in the above background technology.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] A knowledge spectrum graph analysis visualization processing system based on multi-source information fusion, the system includes: a database module, a feature relationship processing module, a feature relationship analysis module and a spectrum update analysis module;
[0007] The database module is used to identify characteristic entities in the knowledge graph, divide the minimum storage space based on the number of characteristic entities, and store all electronic medical records related to the characteristic entities;
[0008] The feature relationship processing module is used to depict undirected relationship edges in the knowledge graph to visually connect the feature entities; capture the same electronic medical records between the minimum storage space, and add relationship feature labels to the undirected relationship edges;
[0009] The feature relationship analysis module calculates the feature relevance and feature independence of the undirected relationship edge based on the relationship feature label; and calculates the edge weight of the undirected relationship edge based on the feature relevance and feature independence;
[0010] The graph update analysis module updates the edge weights by periodic updating, and evaluates the relationship stability of the undirected relationship edges under the influence of the periodic updating, determines whether to recalculate the edge weights of the undirected relationship edges in the next cycle, and synchronously updates the knowledge graph according to the judgment results.
[0011] Further, the database module includes a feature entity unit and an electronic medical record unit;
[0012] The characteristic entity unit, based on the multi-source information of pathological characteristics, integrates the characteristic entities of the knowledge spectrum, and uniformly numbers the characteristic entities. The characteristic entities include the name of the disease type, the name of the drug type, the name of the symptom type, and the name of the treatment method type. The i-th characteristic entity is denoted as S i ;
[0013] The electronic medical record unit obtains I minimum storage spaces by dividing the minimum storage space based on the number of feature entities, where I represents the number of feature entities, i∈[1,I], and divides the feature entity S i The corresponding minimum storage space is recorded as DS (S i );
[0014] The electronic medical records are uniformly numbered, and the ath electronic medical record is recorded as EM a , through text recognition technology, identify electronic medical records EM a If the feature entity is identified as S i , then the electronic medical record EM a Stored in the minimum storage space DS (S i )middle.
[0015] Furthermore, the feature relationship processing module includes a feature relationship generating unit and a feature label processing unit;
[0016] The feature relationship generating unit is used to establish an undirected relationship edge between the feature entities, and the undirected relationship edge is used to visually connect the feature entities. i and feature entity S j The undirected relationship edge between them is denoted as r ij :RS(S i -S j ), j is the characteristic entity S j The serial number of , j∈[1,I];
[0017] The feature label processing unit captures the number of identical electronic medical records in the minimum storage space based on the feature entity, and processes the undirected relationship edge r based on the capture behavior. ij Add relationship feature labels. The specific process is as follows:
[0018] Based on the undirected relationship edge r ij Connected feature entity S i and feature entity S j , retrieve the feature entity S i The corresponding minimum storage space DS(S i ) and feature entity S j The corresponding minimum storage space DS(Sj);
[0019] Establish an iterative correlation analysis model, the initialization input of which is the electronic medical record library L = {EM a |a∈[1, A]}, A represents the total number of electronic medical records, and the output of the association iteration analysis model is the associated medical record sample set RB(r ij ) and non-associated sample set NRB(r ij );
[0020] Let the input of the kth iteration be L k =L-RB(r ij )∪NRB(r ij ), and the input of the first iteration is L1=L; at the kth iteration, from the input L k Choose any electronic medical record EM a , if EM a ∈DS(S i ) and EM a ∈DS(S j ), then the electronic medical record EM a Recorded in the associated medical record sample set SB(r ij ), otherwise, the electronic medical record EM a Recorded in non-associated sample set NRB(r ij )middle;
[0021] When NUM(L k )=0, the iteration stops, NUM(L k ) represents the input L of the kth iteration k the number of electronic medical records included;
[0022] Then for the undirected relationship edge r ij Additional relational feature labels, denoted as Q[RB(r ij ), NRB(r ij )];
[0023] In particular, if i = j, then RB(r ij )=DS(S i )=DS(S j ).
[0024] Further, the characteristic relationship analysis module includes a first characteristic parameter analysis unit and a second characteristic parameter analysis unit;
[0025] The first feature parameter analysis unit is based on the relationship feature label Q[RB(r ij ), NRB(r ij )], calculate the undirected relationship edge r ij The feature correlation Calculate the undirected relationship edge r ij The feature independence
[0026] In the formula, NUM[RB(r ij )] represents the associated medical record sample set RB(r ij ), NUM[NRB(r ij )] represents the non-correlated sample set NRB(r ij ) contains the number of electronic medical records, and NUM[RB(r ij )]+NUM[NRB(r ij )] = A, NUM[DS(S i )∪DS(S j )] represents the minimum storage space DS(S i ) and the minimum storage space DS(S j )The number of electronic medical records included in the intersection result;
[0027] The second characteristic parameter analysis unit is based on the undirected relationship edge r ij The feature correlation and feature independence of the undirected relationship edge r ij The edge weight
[0028] Further, the graph update analysis module includes a feature evaluation unit and an update analysis unit;
[0029] The feature evaluation unit uses the scale T as a cycle to evaluate the undirected relationship edge r ij The edge weight EW(r ij ) is updated, and the edge weight EW(r ij ) is recorded as EW t (r ij );
[0030] Based on the update records of edge weights, analyze and evaluate the undirected relationship edges rij The relationship stability is as follows:
[0031]
[0032] In the formula, QW(r ij |t) represents an undirected relationship edge r ij The stability of the relationship in the tth period, p represents the expected value of the preset edge weight;
[0033] The update analysis unit analyzes and determines whether to perform the t+1th period update based on the relationship stability:
[0034]
[0035] Where G(t) represents the stability boundary value of the tth cycle, g represents the preset boundary reference value, F[] represents the counting function, and F[if: QW(r ij |t)≥g] means that if the relationship stability QW(r ij |t) is greater than or equal to the boundary reference value g, then let F[if:QW(r ij |t)≥g]=1, otherwise let F[if:QW(r ij |t)≥g]=0;
[0036] The preset lower limit value is set. If the stability limit value of the tth period is less than the lower limit value, it is determined to update the t+1th period. At the end of the t+1th period, the undirected relationship edge r is recalculated. ij The edge weight EW(r ij ), and update the compiled knowledge graph synchronously.
[0037] A knowledge spectrum analysis visualization processing method based on multi-source information fusion, the method comprises the following steps:
[0038] Step S1: Identify feature entities in the knowledge graph, divide the minimum storage space based on the number of feature entities, and store all electronic medical records related to the feature entities;
[0039] Step S2: depicting undirected relationship edges in the knowledge graph to visualize the connection of the feature entities; capturing the same electronic medical records in the minimum storage space, and attaching relationship feature labels to the undirected relationship edges;
[0040] Step S3: Based on the relationship feature labels, the feature relevance and feature independence of the undirected relationship edges are calculated respectively; based on the feature relevance and feature independence, the edge weight of the undirected relationship edge is calculated;
[0041] Step S4: Update the edge weights in a periodic updating manner, and evaluate the relationship stability of the wireless relationship edges under the influence of the periodic updating, determine whether to recalculate the edge weights of the undirected relationship edges in the next cycle, and synchronously update the knowledge graph based on the determination result.
[0042] Furthermore, the specific implementation process of step S1 includes:
[0043] Based on the multi-source information of pathological characteristics, the characteristic entities of the knowledge spectrum are integrated and compiled, and the characteristic entities are uniformly numbered. The characteristic entities include the name of the disease type, the name of the drug type, the name of the symptom type and the name of the treatment method type. The i-th characteristic entity is denoted as S i ;
[0044] The number of feature entities is used as the basis for the minimum storage space division, and a total of I minimum storage spaces are obtained, where I represents the number of feature entities, i∈[1,I]. The feature entity S i The corresponding minimum storage space is recorded as DS (S i );
[0045] The electronic medical records are uniformly numbered, and the ath electronic medical record is recorded as EM a , through text recognition technology, identify electronic medical records EM a If the feature entity is identified as S i , then the electronic medical record EM a Stored in the minimum storage space DS (S i )middle.
[0046] Furthermore, the specific implementation process of step S2 includes:
[0047] An undirected relationship edge is established between the feature entities, and the undirected relationship edge is used to visually connect the feature entities. i and feature entity S j The undirected relationship edge between them is denoted as r ij :RS(S i -S j ), j is the characteristic entity S j The serial number of , j∈[1,I];
[0048] Based on the feature entity, the number of identical electronic medical records is captured in the minimum storage space, and based on the capture behavior, the undirected relationship edge r ij Add relationship feature labels. The specific process is as follows:
[0049] Based on the undirected relationship edge r ij Connected feature entity S i and feature entity S j, retrieve the feature entity S i The corresponding minimum storage space DS(S i ) and feature entity S j The corresponding minimum storage space DS(S j );
[0050] Establish an iterative correlation analysis model, the initialization input of which is the electronic medical record library L = {EM a |a∈[1, A]}, A represents the total number of electronic medical records, and the output of the association iteration analysis model is the associated medical record sample set RB(r ij ) and non-associated sample set NRB(r ij );
[0051] Let the input of the kth iteration be L k =L-RB(r ij )∪NRB(r ij ), and the input of the first iteration is L1=L; at the kth iteration, from the input L k Choose any electronic medical record EM a , if EM a ∈DS(S i ) and EM a ∈DS(S j ), then the electronic medical record EM a Recorded in the associated medical record sample set SB(r ij ), otherwise, the electronic medical record EM a Recorded in the non-associated sample set NRB(rij);
[0052] When NUM(L k )=0, the iteration stops, NUM(L k ) represents the input L of the kth iteration k the number of electronic medical records included;
[0053] Then for the undirected relationship edge r ij Additional relational feature labels, denoted as Q[RB(r ij ), NRB(r ij )];
[0054] In particular, if i = j, then RB(r ij )=DS(S i )=DS(S j ).
[0055] Furthermore, the specific implementation process of step S3 includes:
[0056] Based on the relational feature label Q[RB(r ij ), NRB(rij )], calculate the undirected relationship edge r ij The feature correlation Calculate the undirected relationship edge r ij The feature independence
[0057] In the formula, NUM[RB(r ij )] represents the associated medical record sample set RB(r ij ), NUM[NRB(r ij )] represents the non-correlated sample set NRB(r ij ) contains the number of electronic medical records, and NUM[RB(r ij )]+NUM[NRB(r ij )] = A, NUM[DS(S i )∪DS(S j )] represents the minimum storage space DS(S i ) and the minimum storage space DS(S j )The number of electronic medical records included in the intersection result;
[0058] Based on the undirected relationship edge r ij The feature correlation and feature independence of the undirected relationship edge r ij The edge weight
[0059] According to the above method, at the end of a certain cycle, the number of collected electronic medical records is limited. In the limited electronic medical record library, relevant feature entities can be quickly found through iteration. The electronic medical record includes multiple feature entities. The similarities between feature entities make the electronic medical records correlated. The associated medical record sample set can quickly collect related electronic medical records, and the unassociated sample set can quickly collect unrelated electronic medical records. For two feature entities, there is either correlation or irrelevance of the electronic medical records; the feature correlation is the probability that the related electronic medical records appear in the data space formed by the two special diagnosis entities. The greater the probability of occurrence, the greater the correlation; the feature irrelevance is the probability that the unrelated electronic medical records generated by the two feature entities appear in the data space formed by all feature entities. The greater the probability of occurrence, the greater the irrelevance.
[0060] Furthermore, the specific implementation process of step S4 includes:
[0061] With scale T as the cycle period, for the undirected relationship edge r ij The edge weight EW(r ij ) is updated, and the edge weight EW(r ij ) is recorded as EWt (r ij );
[0062] Based on the update records of edge weights, analyze and evaluate the undirected relationship edges r ij The relationship stability is as follows:
[0063]
[0064] In the formula, QW(r ij |t) represents an undirected relationship edge r ij The stability of the relationship in the tth period, p represents the expected value of the preset edge weight;
[0065] Based on the relationship stability, analyze and determine whether to perform the t+1th cycle update:
[0066]
[0067] Where G(t) represents the stability boundary value of the tth cycle, g represents the preset boundary reference value, F[] represents the counting function, and F[if: QW(r ij |t)≥g] means that if the relationship stability QW(r ij |t) is greater than or equal to the boundary reference value g, then let F[if:QW(r ij |t)≥g]=1, otherwise let F[if:QW(r ij |t)≥g]=0;
[0068] The preset lower limit value is set. If the stability limit value of the tth period is less than the lower limit value, it is determined to update the t+1th period. At the end of the t+1th period, the undirected relationship edge r is recalculated. ij The edge weight EW(r ij ), and update the compiled knowledge graph synchronously;
[0069] According to the above method, with periodic updates, if the correlation and irrelevance are closer, the edge weight volatility between the two feature entities is smaller. The periodic law of volatility is used to analyze whether to update the graph in the next cycle, thereby achieving targeted topological updates, isolating the impact between the expansion of the data space and the knowledge graph update algorithm, improving the utilization of algorithm resources, and facilitating the optimization of redundant data in the early stage of data space accumulation, thereby improving the efficiency of knowledge graph updates in the medical field.
[0070] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: in the knowledge graph analysis visualization processing system and method based on multi-source information fusion provided by the present invention, feature entities are identified in the knowledge graph, and the minimum storage space is divided based on the number of feature entities to store all electronic medical records related to the feature entities; undirected relationship edges are depicted, and the same electronic medical records are captured between the minimum storage space to attach relationship feature labels to the undirected relationship edges; the feature relevance, feature independence and edge weight of the undirected relationship edges are calculated respectively; under the influence of periodic updates, the relationship stability of the undirected relationship edges is evaluated, and it is determined whether to recalculate the edge weights of the undirected relationship edges in the next cycle, and the knowledge graph is synchronously updated according to the judgment results; thereby saving data resource space while avoiding the waste of algorithm resources caused by redundant data, improving the efficiency of updating the knowledge graph in the medical field, and promoting the promotion and reuse of the knowledge graph in the medical field. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0072] Figure 1 It is a structural schematic diagram of the knowledge spectrum graph analysis visualization processing system based on multi-source information fusion of the present invention;
[0073] Figure 2 It is a schematic diagram of the steps of the knowledge spectrum graph analysis visualization processing method based on multi-source information fusion of the present invention. DETAILED DESCRIPTION
[0074] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0075] See also Figure 1 , in the first embodiment: a knowledge spectrum graph analysis visualization processing system based on multi-source information fusion is provided, the system comprising: a database module, a feature relationship processing module, a feature relationship analysis module and a spectrum update analysis module;
[0076] A database module is used to identify characteristic entities in the knowledge graph, divide the minimum storage space based on the number of characteristic entities, and store all electronic medical records related to the characteristic entities;
[0077] Preferably, the database module includes a feature entity unit and an electronic medical record unit;
[0078] The characteristic entity unit is based on the multi-source information of pathological characteristics, and integrates the characteristic entities of the knowledge spectrum. The characteristic entities are uniformly numbered. The characteristic entities include the name of the disease type, the name of the drug type, the name of the symptom type, and the name of the treatment method type. The i-th characteristic entity is recorded as S i ;
[0079] The electronic medical record unit is divided into I minimum storage spaces by taking the number of feature entities as the basis for the minimum storage space division. I represents the number of feature entities, i∈[1,I]. The feature entity S i The corresponding minimum storage space is recorded as DS (S i ) ; Unify the electronic medical records and record the ath electronic medical record as EM a , through text recognition technology, identify electronic medical records EM a If the feature entity is identified as S i , then the electronic medical record EM a Stored in the minimum storage space DS (S i )middle;
[0080] The feature relationship processing module is used to depict undirected relationship edges in the knowledge graph to visualize the connection of feature entities; capture the same electronic medical records between the minimum storage space and attach relationship feature labels to the undirected relationship edges;
[0081] Preferably, the feature relationship processing module includes a feature relationship generating unit and a feature label processing unit;
[0082] The feature relationship generation unit is used to establish undirected relationship edges between feature entities. The undirected relationship edges are used to visually connect feature entities. i and feature entity S j The undirected relationship edge between them is denoted as r ij :RS(S i -S j ), j is the characteristic entity S j The serial number of , j∈[1,I];
[0083] The feature label processing unit captures the number of identical electronic medical records in the minimum storage space based on the feature entity, and based on the capture behavior, it processes the undirected relationship edges r ij Add relationship feature labels. The specific process is as follows:
[0084] Based on the undirected relationship edge r ij Connected feature entity S i and feature entity S j , retrieve the feature entity S i The corresponding minimum storage space DS(S i) and feature entity S j The corresponding minimum storage space DS(Sj);
[0085] Establish an association iteration analysis model. The initialization input of the association iteration analysis model is the electronic medical record library L = {EM a |a∈[1, A]}, A represents the total number of electronic medical records, and the output of the association iteration analysis model is the associated medical record sample set RB(r ij ) and non-associated sample set NRB(r ij );
[0086] Let the input of the kth iteration be Lk=L-RB(rij)∪NRB(rij), and the input of the first iteration be L1=L; at the kth iteration, from the input L k Choose any electronic medical record EM a , if EM a ∈DS(S i ) and EM a ∈DS(S j ), then the electronic medical record EM a Recorded in the associated medical record sample set SB(r ij ), otherwise, the electronic medical record EM a Recorded in the non-associated sample set NRB(rij);
[0087] When NUM(L k )=0, the iteration stops, NUM(L k ) represents the input L of the kth iteration k the number of electronic medical records included;
[0088] Then for the undirected relationship edge r ij Additional relational feature labels, denoted as Q[RB(r ij ), NRB(r ij )];
[0089] In particular, if i = j, then RB(r ij )=DS(S i )=DS(S j );
[0090] The feature relationship analysis module calculates the feature relevance and feature irrelevance of the undirected relationship edge based on the relationship feature label; and calculates the edge weight of the undirected relationship edge based on the feature relevance and feature irrelevance;
[0091] Preferably, the characteristic relationship analysis module includes a first characteristic parameter analysis unit and a second characteristic parameter analysis unit;
[0092] The first feature parameter analysis unit, based on the relationship feature label Q[RB(rij ), NRB(r ij )], calculate the undirected relationship edge r ij The feature correlation Calculate the undirected relationship edge r ij The feature independence
[0093] In the formula, NUM[RB(r ij )] represents the associated medical record sample set RB(r ij ), NUM[NRB(r ij )] represents the non-correlated sample set NRB(r ij ) contains the number of electronic medical records, and NUM[RB(r ij )]+NUM[NRB(r ij )] = A, NUM[DS(S i )∪DS(S j )] represents the minimum storage space DS(S i ) and the minimum storage space DS(S j )The number of electronic medical records included in the intersection result;
[0094] The second characteristic parameter analysis unit is based on the undirected relationship edge r ij The feature correlation and feature independence of the undirected relationship edge r ij The edge weight
[0095] The graph update analysis module updates edge weights through periodic updates, and evaluates the stability of undirected edge relationships under the influence of periodic updates, determines whether to recalculate the edge weights of undirected edge relationships in the next period, and synchronously updates the knowledge graph based on the judgment results;
[0096] Preferably, the graph update analysis module includes a feature evaluation unit and an update analysis unit;
[0097] The feature evaluation unit uses a scale T as a cycle to evaluate the undirected relationship edge r ij The edge weight EW(r ij ) is updated, and the edge weight EW(r ij ) is recorded as EW t (r ij );
[0098] Based on the update records of edge weights, analyze and evaluate the undirected relationship edges r ij The relationship stability is as follows:
[0099]
[0100] In the formula, QW(r ij |t) represents an undirected relationship edge r ij The stability of the relationship in the tth period, p represents the expected value of the preset edge weight;
[0101] The update analysis unit analyzes and determines whether to perform the t+1th cycle update based on the relationship stability:
[0102]
[0103] Where G(t) represents the stability boundary value of the tth cycle, g represents the preset boundary reference value, F[] represents the counting function, and F[if: QW(r ij |t)≥g] means that if the relationship stability QW(r ij |t) is greater than or equal to the boundary reference value g, then let F[if:QW(r ij |t)≥g]=1, otherwise let F[if:QW(r ij |t)≥g]=0;
[0104] The preset lower limit value is set. If the stability limit value of the tth period is less than the lower limit value, it is determined to update the t+1th period. At the end of the t+1th period, the undirected relationship edge r is recalculated. ij The edge weight EW(r ij ), and update the compiled knowledge graph synchronously.
[0105] See also Figure 2 In the second embodiment, a knowledge spectrum analysis visualization processing method based on multi-source information fusion is provided, and the method includes the following steps:
[0106] Step S1: Identify feature entities in the knowledge graph, divide the minimum storage space based on the number of feature entities, and store all electronic medical records related to the feature entities;
[0107] For example, based on the multi-source information of pathological characteristics, the characteristic entities of the knowledge spectrum are integrated and compiled, and the characteristic entities are uniformly numbered. The characteristic entities include the name of the disease type, the name of the drug type, the name of the symptom type, and the name of the treatment method type. The i-th characteristic entity is recorded as S i ;
[0108] The number of feature entities is used as the basis for the minimum storage space division, and a total of I minimum storage spaces are obtained, where I represents the number of feature entities, i∈[1,I]. The feature entity S i The corresponding minimum storage space is recorded as DS (S i );
[0109] The electronic medical records are uniformly numbered, and the ath electronic medical record is recorded as EMa , through text recognition technology, identify electronic medical records EM a If the feature entity is identified as S i , then the electronic medical record EM a Stored in the minimum storage space DS (S i )middle.
[0110] Step S2: Draw undirected relationship edges in the knowledge graph to visualize the connected feature entities; capture the same electronic medical records in the minimum storage space, and attach relationship feature labels to the undirected relationship edges;
[0111] For example, undirected relationship edges are established between feature entities, and the undirected relationship edges are used to visually connect feature entities. i and feature entity S j The undirected relationship edge between them is denoted as r ij :RS(S i -S j ), j is the characteristic entity S j The serial number of , j∈[1,I];
[0112] Based on the feature entity, the number of identical electronic medical records is captured in the minimum storage space, and based on the capture behavior, the undirected relationship edge r ij Add relationship feature labels. The specific process is as follows:
[0113] Based on the undirected relationship edge r ij Connected feature entity S i and feature entity S j , retrieve the feature entity S i The corresponding minimum storage space DS(S i ) and feature entity S j The corresponding minimum storage space DS(S j );
[0114] Establish an association iteration analysis model. The initialization input of the association iteration analysis model is the electronic medical record library L = {EM a |a∈[1, A]}, A represents the total number of electronic medical records, and the output of the association iteration analysis model is the associated medical record sample set RB(r ij ) and non-associated sample set NRB(r ij );
[0115] Let the input of the kth iteration be L k =L-RB(r ij )∪NRB(r ij ), and the input of the first iteration is L1=L; at the kth iteration, from the input L kChoose any electronic medical record EM a , if EM a ∈DS(S i ) and EM a ∈DS(S j ), then the electronic medical record EM a Recorded in the associated medical record sample set SB(r ij ), otherwise, the electronic medical record EM a Recorded in the non-associated sample set NRB(rij);
[0116] When NUM(L k )=0, the iteration stops, NUM(L k ) represents the input L of the kth iteration k the number of electronic medical records included;
[0117] Then for the undirected relationship edge r ij Additional relational feature labels, denoted as Q[RB(r ij ), NRB(r ij )];
[0118] In particular, if i = j, then RB(r ij )=DS(S i )=DS(S j ).
[0119] Step S3: Based on the relationship feature labels, the feature relevance and feature independence of the undirected relationship edges are calculated respectively; based on the feature relevance and feature independence, the edge weight of the undirected relationship edge is calculated;
[0120] For example, based on the relationship feature label Q[RB(r ij ), NRB(r ij )], calculate the undirected relationship edge r ij The feature correlation Calculate the undirected relationship edge r ij The feature independence
[0121] In the formula, NUM[RB(r ij )] represents the associated medical record sample set RB(r ij ), NUM[NRB(r ij )] represents the non-correlated sample set NRB(r ij ) contains the number of electronic medical records, and NUM[RB(r ij )]+NUM[NRB(r ij )] = A, NUM[DS(S i )∪DS(S j)] represents the minimum storage space DS(S i ) and the minimum storage space DS(S j )The number of electronic medical records included in the intersection result;
[0122] Based on the undirected relationship edge r ij The feature correlation and feature independence of the undirected relationship edge r ij The edge weight
[0123] Step S4: Update edge weights in a periodic update manner, and evaluate the relationship stability of undirected relationship edges under the influence of periodic updates, determine whether to recalculate the edge weights of undirected relationship edges in the next period, and synchronously update the knowledge graph based on the determination results;
[0124] For example, with scale T as the cycle period, for the undirected relationship edge r ij The edge weight EW(r ij ) is updated, and the edge weight EW(rij) updated in the tth cycle is recorded as EW t (rij);
[0125] Based on the update records of edge weights, analyze and evaluate the undirected relationship edges r ij The relationship stability is as follows:
[0126]
[0127] In the formula, QW(r ij |t) represents an undirected relationship edge r ij The stability of the relationship in the tth period, p represents the expected value of the preset edge weight;
[0128] Based on the relationship stability, analyze and determine whether to perform the t+1th cycle update:
[0129]
[0130] Where G(t) represents the stability boundary value of the tth cycle, g represents the preset boundary reference value, F[] represents the counting function, and F[if: QW(r ij |t)≥g] means that if the relationship stability QW(r ij |t) is greater than or equal to the boundary reference value g, then let F[if:QW(r ij |t)≥g]=1, otherwise let F[if:QW(r ij |t)≥g]=0;
[0131] The preset lower limit value is set. If the stability limit value of the tth period is less than the lower limit value, it is determined to update the t+1th period. At the end of the t+1th period, the undirected relationship edge r is recalculated. ij The edge weight EW(r ij ), and update the compiled knowledge graph synchronously.
[0132] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0133] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A knowledge spectrum graph analysis visualization processing method based on multi-source information fusion, characterized in that: The method comprises the following steps: Step S1: Identify feature entities in the knowledge graph, divide the minimum storage space based on the number of feature entities, and store all electronic medical records related to the feature entities; Step S2: depicting undirected relationship edges in the knowledge graph to visualize the connection of the feature entities; capturing the same electronic medical records in the minimum storage space, and attaching relationship feature labels to the undirected relationship edges; Step S3: Based on the relationship feature labels, the feature relevance and feature independence of the undirected relationship edges are calculated respectively; based on the feature relevance and feature independence, the edge weight of the undirected relationship edge is calculated; Step S4: updating the edge weights in a periodic updating manner, and evaluating the relationship stability of the undirected relationship edges under the influence of the periodic updating, determining whether to recalculate the edge weights of the undirected relationship edges in the next period, and synchronously updating the knowledge graph according to the determination result; In step S1, the i-th feature entity is recorded as S i , i∈[1,I], I represents the number of feature entities; In step S2, set j to be the feature entity S j The serial number of , j∈[1,I]; The specific implementation process of step S3 includes: Based on the relational feature label Q[RB(r ij ), NRB(r ij )], calculate the undirected relationship edge r ij The feature correlation Calculate the undirected relationship edge r ij The feature independence In the formula, NUM[RB(r ij )] represents the associated medical record sample set RB(r ij ), NUM[NRB(r ij )] represents the non-correlated sample set NRB(r ij ) contains the number of electronic medical records, and NUM[RB(r ij )]+NUM[NRB(r ij )] = A, NUM[DS(S i )∪DS(S j )] represents the minimum storage space DS(S i ) and the minimum storage space DS(S j )The number of electronic medical records included in the intersection result; Based on the undirected relationship edge r ij The feature correlation and feature independence of the undirected relationship edge r ij The edge weight The specific implementation process of step S4 includes: With scale T as the cycle period, for the undirected relationship edge r ij The edge weight EW(r ij ) is updated, and the edge weight EW(r ij ) is recorded as EW t (r ij ); Based on the update records of edge weights, analyze and evaluate the undirected relationship edges r ij The relationship stability is as follows: In the formula, QW(r ij |t) represents an undirected relationship edge r ij The stability of the relationship in the tth period, p represents the expected value of the preset edge weight; Based on the relationship stability, analyze and determine whether to perform the t+1th cycle update: Where G(t) represents the stability boundary value of the tth cycle, g represents the preset boundary reference value, F[] represents the counting function, and F[if: QW(r ij |t)≥g] means that if the relationship stability QW(r ij |t) is greater than or equal to the boundary reference value g, then let F[if:QW(r ij |t)≥g]=1, otherwise let F[if:QW(r ij |t)≥g]=0; The preset lower limit value is set. If the stability limit value of the tth period is less than the lower limit value, it is determined to update the t+1th period. At the end of the t+1th period, the undirected relationship edge r is recalculated. ij The edge weight EW(r ij ), and update the compiled knowledge graph synchronously.
2. The method for visualizing knowledge spectrum analysis based on multi-source information fusion according to claim 1 is characterized in that: The specific implementation process of step S1 includes: Based on the multi-source information of pathological characteristics, the characteristic entities of the knowledge spectrum are integrated and compiled, and the characteristic entities are uniformly numbered. The characteristic entities include the name of the disease type, the name of the drug type, the name of the symptom type, and the name of the treatment method type; The number of feature entities is used as the basis for the minimum storage space division, and a total of I minimum storage spaces are obtained. The feature entities S i The corresponding minimum storage space is recorded as DS (S i ); The electronic medical records are uniformly numbered, and the ath electronic medical record is recorded as EM a , through text recognition technology, identify electronic medical records EM a If the feature entity is identified as S i , then the electronic medical record EM a Stored in the minimum storage space DS (S i )middle.
3. The knowledge spectrum analysis visualization processing method based on multi-source information fusion according to claim 2 is characterized in that: The specific implementation process of step S2 includes: An undirected relationship edge is established between the feature entities, and the undirected relationship edge is used to visually connect the feature entities. i and feature entity S j The undirected relationship edge between them is denoted as r ij :RS(S i -S j ); Based on the feature entity, the number of identical electronic medical records is captured in the minimum storage space, and based on the capture behavior, the undirected relationship edge r ij Add relationship feature labels. The specific process is as follows: Based on the undirected relationship edge r ij Connected feature entity S i and feature entity S j , retrieve the feature entity S i The corresponding minimum storage space DS(S i ) and feature entity S j The corresponding minimum storage space DS(S j ); Establish an iterative correlation analysis model, the initialization input of which is the electronic medical record library L = {EM a |a∈[1, A]}, A represents the total number of electronic medical records, and the output of the association iteration analysis model is the associated medical record sample set RB(r ij ) and non-associated sample set NRB(r ij ); Let the input of the kth iteration be L k =L-RB(r ij )∪NRB(r ij ), and the input of the first iteration is L1=L; at the kth iteration, from the input L k Choose any electronic medical record EM a , if EM a ∈DS(S i ) and EM a ∈DS(S j ), then the electronic medical record EM a Recorded in the associated medical record sample set RB(r ij ), otherwise, the electronic medical record EM a Recorded in non-associated sample set NRB(r ij )middle; When NUM(L k )=0, the iteration stops, NUM(L k ) represents the input L of the kth iteration k the number of electronic medical records included; Then for the undirected relationship edge r ij Additional relational feature labels, denoted as Q[RB(r ij ), NRB(r ij )]; In particular, if i = j, then RB(r ij )=DS(S i )=DS(S j ).
4. A knowledge graph analysis and visualization processing system based on multi-source information fusion, characterized by: The system includes: a database module, a feature relationship processing module, a feature relationship analysis module and a graph update analysis module; The database module is used to identify characteristic entities in the knowledge graph, divide the minimum storage space based on the number of characteristic entities, and store all electronic medical records related to the characteristic entities; The feature relationship processing module is used to depict undirected relationship edges in the knowledge graph to visually connect the feature entities; capture the same electronic medical records between the minimum storage space, and add relationship feature labels to the undirected relationship edges; The feature relationship analysis module calculates the feature relevance and feature independence of the undirected relationship edge based on the relationship feature label; and calculates the edge weight of the undirected relationship edge based on the feature relevance and feature independence; The graph update analysis module updates the edge weights by periodic updating, and evaluates the relationship stability of the undirected relationship edges under the influence of the periodic updating, determines whether to recalculate the edge weights of the undirected relationship edges in the next period, and synchronously updates the knowledge graph according to the determination result; The database module includes a feature entity unit and an electronic medical record unit; In the characteristic entity unit, the i-th characteristic entity is denoted as S i ; Set I in the electronic medical record unit to represent the number of feature entities, i∈[1,I]; The feature relationship processing module includes a feature relationship generating unit and a feature label processing unit; In the feature relationship generation unit, set j to be the feature entity S j The serial number of , j∈[1,I]; The characteristic relationship analysis module includes a first characteristic parameter analysis unit and a second characteristic parameter analysis unit; The first feature parameter analysis unit is based on the relationship feature label Q[RB(r ij ), NRB(r ij )], calculate the undirected relationship edge r ij The feature correlation Calculate the undirected relationship edge r ij The feature independence In the formula, NUM[RB(r ij )] represents the associated medical record sample set RB(r ij ), NUM[NRB(r ij )] represents the non-correlated sample set NRB(r ij ) contains the number of electronic medical records, and NUM[RB(r ij )]+NUM[NRB(r ij )] = A, NUM[DS(S i )∪DS(S j )] represents the minimum storage space DS(S i ) and the minimum storage space DS(S j )The number of electronic medical records included in the intersection result; The second characteristic parameter analysis unit is based on the undirected relationship edge r ij The feature correlation and feature independence of the undirected relationship edge r ij The edge weight The graph update analysis module includes a feature evaluation unit and an update analysis unit; The feature evaluation unit uses the scale T as a cycle to evaluate the undirected relationship edge r ij The edge weight EW(r ij ) is updated, and the edge weight EW(r ij ) is recorded as EW t (r ij ); Based on the update records of edge weights, analyze and evaluate the undirected relationship edges r ij The relationship stability is as follows: In the formula, QW(r ij |t) represents an undirected relationship edge r ij The stability of the relationship in the tth period, p represents the expected value of the preset edge weight; The update analysis unit analyzes and determines whether to perform the t+1th period update based on the relationship stability: Where G(t) represents the stability boundary value of the tth cycle, g represents the preset boundary reference value, F[] represents the counting function, and F[if: QW(r ij |t)≥g] means that if the relationship stability QW(r ij |t) is greater than or equal to the boundary reference value g, then let F[if:QW(r ij |t)≥g]=1, otherwise let F[if:QW(r ij |t)≥g]=0; The preset lower limit value is set. If the stability limit value of the tth period is less than the lower limit value, it is determined to update the t+1th period. At the end of the t+1th period, the undirected relationship edge r is recalculated. ij The edge weight EW(r ij ), and update the compiled knowledge graph synchronously.
5. The knowledge spectrum analysis visualization processing system based on multi-source information fusion according to claim 4 is characterized by: The characteristic entity unit, based on the multi-source information of pathological characteristics, integrates the characteristic entities of the knowledge spectrum and uniformly numbers the characteristic entities, wherein the characteristic entities include the name of the disease type, the name of the drug type, the name of the symptom type and the name of the treatment method type; The electronic medical record unit obtains a total of I minimum storage spaces by dividing the minimum storage space based on the number of characteristic entities. i The corresponding minimum storage space is recorded as DS (S i ) ; Unify the electronic medical records and record the ath electronic medical record as EM a , through text recognition technology, identify electronic medical records EM a If the feature entity is identified as S i , then the electronic medical record EM a Stored in the minimum storage space DS (S i )middle.
6. The knowledge spectrum analysis visualization processing system based on multi-source information fusion according to claim 5 is characterized by: The feature relationship generating unit is used to establish an undirected relationship edge between the feature entities, and the undirected relationship edge is used to visually connect the feature entities. i and feature entity S j The undirected relationship edge between them is denoted as r ij :RS(S i -S j ); The feature label processing unit captures the number of identical electronic medical records in the minimum storage space based on the feature entity, and processes the undirected relationship edge r based on the capture behavior. ij Add relationship feature labels. The specific process is as follows: Based on the undirected relationship edge r ij Connected feature entity S i and feature entity S j , retrieve the feature entity S i The corresponding minimum storage space DS(S i ) and feature entity S j The corresponding minimum storage space DS(S j ); Establish an iterative correlation analysis model, the initialization input of which is the electronic medical record library L = {EM a |a∈[1, A]}, A represents the total number of electronic medical records, and the output of the association iteration analysis model is the associated medical record sample set RB(r ij ) and non-associated sample set NRB(r ij ); Let the input of the kth iteration be L k =L-RB(r ij )∪NRB(r ij ), and the input of the first iteration is L1=L; at the kth iteration, from the input L k Choose any electronic medical record EM a , if EM a ∈DS(S i ) and EM a ∈DS(S j ), then the electronic medical record EM a Recorded in the associated medical record sample set RB(r ij ), otherwise, the electronic medical record EM a Recorded in non-associated sample set NRB(r ij )middle; When NUM(L k )=0, the iteration stops, NUM(L k ) represents the input L of the kth iteration k the number of electronic medical records included; Then for the undirected relationship edge r ij Additional relational feature labels, denoted as Q[RB(r ij ), NRB(r ij )]; In particular, if i = j, then RB(r ij )=DS(S i )=DS(S j ).
Citation Information
Patent Citations
Gradually-increased knowledge graph entity extraction method and system for electric power customer service questions and answers
CN112199488A
Similar medical record recommendation method and system
CN115841861A