A medical electronic medical record data access traceability management method

CN122842829APending Publication Date: 2026-09-29SOOCHOW UNIV AFFILIATED CHILDRENS HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611177160.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-05
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]现有访问控制技术在医疗数据实际运作中通常围绕账号身份、角色权限、机构归属、访问时间、数据类别等静态条件执行规则匹配,访问请求只要满足预设权限范围即可进入电子病历资源,访问理由同患者当前诊疗需求之间是否存在真实关联往往缺少结构化核验,医生拥有某类病历访问权限并不代表每次访问均具备合理医疗目的,例如非接诊医生凭借科室权限查询无诊疗关系患者的历史病历时,传统规则仍可能判定为合法访问,进而形成越权浏览、隐私窥探及内部滥用风险

Benefits of technology

[0034]与现有技术相比,本发明的优点和积极效果在于:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842829A_ABST
    Figure CN122842829A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of access control, in particular to a medical electronic medical record data access traceability management method, comprising the following steps: receiving an electronic medical record access request and an associated access intention text transmitted by a request party organization gateway node, constructing an intention subgraph, extracting the current visit chief complaint and the past medical history record text of the target patient pointed to by the electronic medical record access request, and generating a comprehensive medical history subgraph.In the present application, the control rule judgment path, the organization gateway node digital signature public key and the intermediate layer activation feature jointly generate a traceability log summary, which is archived in a traceability graph database, and the risk judgment basis, the participating organization identity, the key calculation state and the final judgment result can be associated to establish a verifiable association, so that when the log content is replaced, deleted, spliced, the summary is inconsistent, thereby enhancing the abnormal investigation capability of the cross-organization electronic medical record access process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of access control technology, and in particular to a method for tracing and controlling access to medical electronic medical record data. Background Technology

[0002] Access control technology involves technologies related to healthcare information security, medical data sharing, identity authentication, data traceability, and privacy protection.

[0003] Existing access control technologies in the actual operation of medical data typically rely on static conditions such as account identity, role permissions, institutional affiliation, access time, and data category to execute rule matching. Access requests that meet the preset permission range can access electronic medical record resources. However, the connection between the reason for access and the patient's current treatment needs often lacks structured verification. A doctor's access to a certain type of medical record does not guarantee a legitimate medical purpose for every access. For example, if a non-attending physician uses departmental permissions to access the historical medical records of patients with whom they have no medical relationship, traditional rules may still classify it as legitimate access, leading to risks of unauthorized browsing, privacy breaches, and internal abuse. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for tracing and controlling access to electronic medical record data.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for accessing and tracing medical electronic medical record data, comprising the following steps:

[0006] The system receives the electronic medical record access request and associated access intent text transmitted by the gateway node of the requesting institution, constructs an intent subgraph, extracts the current chief complaint and past medical history records of the target patient pointed to by the electronic medical record access request, and generates a comprehensive medical history subgraph.

[0007] Calculate the subgraph isomorphic matching degree of the intent subgraph in the comprehensive medical history subgraph, and map the subgraph isomorphic matching degree to an intent difference penalty factor;

[0008] Locate the requesting doctor node and the target patient node in the local knowledge graph, and generate local node embedding vectors and intermediate layer activation features;

[0009] Based on the knowledge anchors mapped from the standardized medical terminology set, cross-domain feature alignment is performed on the local node embedding vectors of the requesting doctor node and the target patient node to calculate an initial topological similarity; the initial topological similarity is converted into an initial topological difference, and the cross-domain topological association distance is generated by combining the initial topological difference with the intent difference penalty factor.

[0010] Extract the historical operation duration and historical access frequency parameters of the requesting doctor node, and concatenate them with the cross-domain topology association distance to generate a multi-dimensional access behavior feature vector;

[0011] The multidimensional access behavior feature vector is input into the risk decision tree model, which outputs a real-time risk quantification score for the electronic medical record access request.

[0012] Preferably, the method further includes:

[0013] The real-time risk quantification score is compared with the risk stratification limit in the multi-level access control rule base, and a control identifier corresponding to the electronic medical record access request is assigned and the corresponding access control operation is triggered.

[0014] Extract the control rule determination path, the digital signature public key of the agency gateway node participating in the calculation, and the activation feature of the intermediate layer to generate a traceability log digest, and archive the traceability log digest into the traceability graph database.

[0015] Preferably, the steps for obtaining the intention sub-map and the comprehensive medical history sub-map are as follows:

[0016] The system receives an electronic medical record access request and associated access intent text transmitted by the requesting institution's gateway node. It reads the request identifier, target patient identifier, access field range, and access timestamp from the electronic medical record access request. It binds the access intent text according to the target patient identifier, locating the start and end positions of characters corresponding to disease names, symptom names, examination items, treatment behaviors, drug names, and body parts sentence by sentence. It unifies the medical terminology names corresponding to synonyms, determines semantic relationships based on the word order association, modification direction, and treatment orientation between adjacent medical entities, sets medical entities as nodes, and sets semantic relationships as node connection edges to establish an intent subgraph. Simultaneously, it retrieves the chief complaint text and past medical history record text from the current consultation triage system according to the target patient identifier, locating the medical history entities in the current consultation chief complaint and past medical history record text sentence by sentence. It establishes node connection edges according to the diagnostic, medication, examination, and symptom relationships between medical history entities, obtaining an intent subgraph and a comprehensive medical history subgraph.

[0017] Preferably, the step of obtaining the intent difference penalty factor is as follows:

[0018] The medical term name, medical entity category, and attribute label of each node are read separately. The semantic relationship type, connection direction, and associated node number of each node connection edge are read separately. The nodes in the intent subgraph are mapped one by one to the candidate nodes in the comprehensive medical history subgraph that have the same medical term name or the same synonym unification result. The candidate node mappings that do not have the same medical entity category are deleted. The mapping is extended layer by layer to the adjacent nodes along the candidate node mapping. The semantic relationship type, connection direction, and node adjacency order in the extension path are verified. The mapping combination that holds both the node correspondence and the connection edge correspondence is retained. The number of matching nodes, the number of matching connection edges, and the continuous matching level in all mapping combinations are compared. The mapping combination with the most matching nodes and the largest number of matching connection edges is selected. The subgraph isomorphic matching degree is obtained by quantifying the selected mapping combination according to the proportion of the selected mapping combination to all nodes and all connection edges in the intent subgraph.

[0019] Read the preset matching degree segment boundaries, the penalty benchmark value corresponding to each matching degree segment, and the decreasing change range corresponding to each matching degree segment. Compare the subgraph isomorphic matching degree with the matching degree segment boundaries in descending order to determine the unique matching degree segment corresponding to the subgraph isomorphic matching degree. Call the penalty benchmark value corresponding to the matching degree segment. Subtract the subgraph isomorphic matching degree from the upper boundary of the matching degree segment to obtain the matching difference value. Amplify the matching difference value according to the decreasing change range corresponding to the matching degree segment. Superimpose the amplified matching difference value onto the penalty benchmark value. The superimposed result exceeding the preset penalty upper boundary is defined as the penalty upper boundary, and the superimposed result below the preset penalty lower boundary is defined as the penalty lower boundary. Complete the numerical mapping according to the correspondence that the superimposed result increases when the subgraph isomorphic matching degree decreases to obtain the intent difference penalty factor.

[0020] Preferably, the step of obtaining the initial topological dissimilarity is as follows:

[0021] The requesting doctor identifier, requesting institution code, target patient identifier, and target institution code in the electronic medical record access request are parsed. Personnel nodes are retrieved in the local knowledge graph of the requesting institution gateway node according to the requesting doctor identifier, and patient nodes are retrieved in the local knowledge graph of the target institution gateway node according to the target patient identifier. The occupational type marker of the personnel node and the identity type marker of the patient node are verified to determine the requesting doctor node and the target patient node. The first-order adjacent nodes, progressively expanded adjacent nodes, connection edge directions, connection edge categories, node degrees, and adjacent path lengths of the requesting doctor node and the target patient node are read respectively. The node's own attribute components, adjacent node attribute components, and connection edge category codes are aggregated layer by layer according to the corresponding levels. Dimension compression, component weighting, and nonlinear mapping are performed on the aggregation results of each layer. The local plaintext graph data is always kept inside the corresponding institution gateway node. Local node embedding vectors and intermediate layer activation features are generated in real time or cached according to a preset update cycle.

[0022] The standardized medical terminology numbers jointly stored by the requesting institution gateway node and the target institution gateway node are read. Standardized medical terminology numbers that can be mapped to adjacent nodes of the requesting doctor node and the target patient node are selected. A knowledge anchor correspondence sequence is established according to the standardized medical terminology numbers. The local node embedding vectors of the requesting doctor node and the target patient node are split into multiple random numerical shares. An encrypted communication tunnel is established between the requesting institution gateway node and the target institution gateway node. Only the random numerical shares that cannot be independently restored from the local node embedding vectors are exchanged. The cross-product of the random numerical shares is calculated in the ciphertext state using a homomorphic encryption algorithm. The vector component arrangement position is unified according to the knowledge anchor correspondence sequence. The scale of the corresponding vector component is corrected according to the number of associated nodes of each knowledge anchor. The vector components that have completed the position unification and scale correction are multiplied and accumulated one by one to obtain the initial topological similarity. Then, according to the preset mapping order that the difference value decreases as the initial topological similarity increases, interval positioning, boundary constraints, and reverse numerical transformation are performed on the initial topological similarity to obtain the initial topological difference.

[0023] Preferably, the step of obtaining the cross-domain topological association distance is as follows:

[0024] Read the lower and upper bounds of the initial topology difference degree and the lower and upper bounds of the intention difference penalty factor. Calculate the relative positions of the initial topology difference degree and the intention difference penalty factor within their respective value boundaries. Convert the two relative positions to the same numerical range. Read the preset topology difference weight and intention penalty weight. Verify the sum of the topology difference weight and intention penalty weight. If the sum of weights is not equal to a unit value, perform weight correction according to the proportion of each weight to the total weight. Multiply the converted initial topology difference degree by the corrected topology difference weight, and multiply the converted intention difference penalty factor by the corrected intention penalty weight. Summate the two products. Values ​​exceeding the upper boundary of the summation are defined as the upper boundary, and values ​​below the lower boundary are defined as the lower boundary, generating the cross-domain topology association distance.

[0025] Preferably, the steps for obtaining the multidimensional access behavior feature vector are as follows:

[0026] The historical operation records corresponding to the requesting doctor node are read from the electronic medical record access request. The start timestamp, end timestamp, access event identifier, and target patient identifier of each historical operation record are extracted. The duration of each historical operation is calculated in the order of end timestamp minus start timestamp. The number of historical operation records with different access event identifiers within a preset statistical period is counted to form a historical access frequency parameter. The historical operation duration, the historical minimum value, and the historical maximum value of the historical access frequency parameter are read respectively. The difference obtained by subtracting the historical minimum value from the historical operation duration and dividing by the historical maximum value minus the historical minimum value is calculated. The historical access frequency parameter is converted to a unified numerical range in the same order. The feature array is concatenated according to the fixed arrangement order of the historical operation duration normalization result, the historical access frequency parameter normalization result, and the cross-domain topological association distance to generate a multi-dimensional access behavior feature vector.

[0027] Preferably, the steps for obtaining the real-time risk quantification score are as follows:

[0028] The normalized results of historical operation duration, the normalized results of historical access frequency parameters, and the cross-domain topological association distance are read according to the fixed arrangement order in the multidimensional access behavior feature vector. The multidimensional access behavior feature vector is input into the pre-trained risk decision tree. The feature dimension number, quantization segmentation limit, left branch pointing identifier, and right branch pointing identifier bound to the root node are read. The corresponding feature dimension value is compared with the quantization segmentation limit. If it is less than the quantization segmentation limit, the left branch pointing node is entered. If it is greater than or equal to the quantization segmentation limit, the right branch pointing node is entered. The feature dimension number and quantization segmentation limit bound to the current node are read repeatedly until the leaf node corresponding to the branch pointing identifier is reached. The feature dimension value, quantization segmentation limit, branch pointing identifier, and leaf node number used in each quantization comparison are recorded to form a risk judgment path.

[0029] The risk weights are read according to the leaf node numbers recorded in the risk assessment path. It is verified whether the feature dimension values ​​in the risk assessment path have been quantified and compared. It is verified whether the adjacent branch pointing indicators are continuously connected according to the previous node pointing result. Invalid risk weights corresponding to missing feature dimension values, missing quantization segmentation limits, and interrupted branch pointing indicators are deleted. The remaining risk weights are accumulated item by item according to the recording order of the leaf node numbers. The risk weights corresponding to duplicate leaf node numbers are only retained once. The values ​​of the accumulated result below the lower boundary of the preset risk score are limited to the lower boundary of the preset risk score. The values ​​of the accumulated result above the upper boundary of the preset risk score are limited to the upper boundary of the preset risk score. The real-time risk quantification score for the current electronic medical record access request is output.

[0030] Preferably, the step of obtaining the source log summary is as follows:

[0031] The real-time risk quantification score is read level by level according to the preset multi-level access control rule base, and the risk stratification limit is arranged in intervals according to the numerical order of the risk stratification limit. The real-time risk quantification score is compared with each of the risk stratification limits in turn. When the real-time risk quantification score is in the release interval, a release control flag is written. When the real-time risk quantification score is in the verification interval, a second authentication activation control flag is written. When the real-time risk quantification score is in the blocking interval, a blocking control flag is written. The control flags along with the electronic medical record access request are sent to the corresponding institutional gateway node. The institutional gateway node opens the access session according to the release control flag, terminates the access link according to the blocking control flag, and suspends the access session and initiates identity credential verification according to the second authentication activation control flag, thus forming the access control operation result.

[0032] Based on the access control operation results, the feature dimension numbers, quantization segmentation limits, branch pointing identifiers, and leaf node numbers traversed sequentially from the root node to the hit leaf node during the risk determination process are read. A control rule determination path is formed according to the node traversal order. The digital signature public keys used by the requesting institution gateway node and the target institution gateway node in this calculation are read respectively. The digital signature public keys are arranged according to the institution coding order. The intermediate layer activation features generated by the requesting doctor node and the target patient node at each propagation level are read and arranged according to the node identifier, propagation level, and feature component position. The control rule determination path, digital signature public keys, and intermediate layer activation features are converted into character sequences of uniform length. The first and last parts are concatenated according to the fixed field order. The concatenated character sequences are compressed and mapped segment by segment. The compression result of the previous segment is merged with the next segment of the character sequence until all character sequences are processed, generating a traceability log summary with anti-tampering characteristics.

[0033] The system reads the request identifier, requesting doctor node identifier, target patient node identifier, requesting institution code, target institution code, and access timestamp corresponding to the electronic medical record access request. It also reads the control identifier, identity verification status, and access link execution status from the access control operation result. The request identifier is set as the unique index of the access event node. The requesting doctor node identifier and target patient node identifier are bound to the access event node respectively. The requesting institution code and target institution code are written to the institution association field. The access timestamp is written to the time sequence field. The control identifier, identity verification status, and access link execution status are combined into a judgment result. The source tracing log summary is written to the summary field. An association edge is established between the access event node and the requesting doctor node, target patient node, and institution gateway node. After verifying that the unique index does not have duplicates, the record is written to the storage record, forming an archived record in the source tracing graph database.

[0034] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0035] This invention transforms the access intent text associated with electronic medical record access requests into intent subgraphs, and organizes the target patient's current consultation complaint and past medical history records into a comprehensive medical history subgraph. This moves the access purpose beyond keyword verification at the textual level, transforming it into a structured semantic relationship comprised of medical entities, treatment relationships, and medical history associations. Furthermore, an intent difference penalty factor is generated based on the degree of isomorphic matching between the intent subgraph and the comprehensive medical history subgraph. This quantifies the deviation between the request content and the patient's actual treatment background, reducing the interference caused by vague access reasons, terminology substitution, and semantic packaging on access permission determination. Under the condition that the requesting institution's gateway node and the target institution's gateway node each retain the plaintext of their local knowledge graphs, local node embedding vectors and intermediate layer activation features are extracted around the requesting doctor node and the target patient node. Then, cross-domain feature alignment is completed by using knowledge anchors corresponding to standardized medical terminology sets. This allows differences in node naming, graph structure, and terminology encoding between different institutions to be mapped to a unified comparison space. Furthermore, the cross-domain topological association distance is formed through the joint calculation of the initial topological difference degree and the agreement graph difference penalty factor. This ensures that the degree of professional relevance, the rationality of the diagnosis and treatment relationship, and the access intention between the visiting subject and the target patient are considered. Figure 1 The degree of consensus is evaluated collaboratively; the cross-domain topological association distance is concatenated with parameters such as historical operation duration and historical access frequency to form a multi-dimensional access behavior feature vector, which can simultaneously reflect the rationality of the relationship and the abnormality of the behavior, avoiding the risk level being determined by a single identity permission or a single access frequency. The risk decision tree accumulates risk weights based on the branch comparison of each dimension of features, and can output a real-time risk quantification score for the current request. Based on the risk stratification limit, control identifiers corresponding to allow, block, and secondary authentication are assigned, so that the access control results have continuous hierarchical capabilities. The control rule judgment path, the digital signature public key of the institutional gateway node, and the activation features of the intermediate layer jointly generate a traceability log summary, which is archived to the traceability graph database. It can establish a verifiable association between the risk judgment basis, the identity of the participating institutions, the key calculation status, and the final judgment result. When the log content is replaced, deleted, or concatenated, inconsistencies in the summary will be formed, thereby enhancing the ability to trace anomalies in the cross-institutional electronic medical record access process. Attached Figure Description

[0036] Figure 1 A graph showing the relationship between subgraph isomorphic matching degree and intent difference penalty factor;

[0037] Figure 2 A graph showing the relationship between cross-domain topological association distance, normalized results of historical operation duration, and real-time risk quantification scores. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0039] Please see Figure 1-2 This invention provides a technical solution, a method for accessing and tracing medical electronic medical record data, comprising the following steps:

[0040] The system receives the electronic medical record access request and associated access intent text transmitted by the gateway node of the requesting institution, constructs an intent subgraph, extracts the current chief complaint and past medical history records of the target patient pointed to by the electronic medical record access request, and generates a comprehensive medical history subgraph; calculates the subgraph isomorphic matching degree of the intent subgraph in the comprehensive medical history subgraph, and maps the subgraph isomorphic matching degree to the intent difference penalty factor.

[0041] In the local knowledge graph, the requesting doctor node and the target patient node are located, and local node embedding vectors and intermediate layer activation features are generated. Based on the knowledge anchors mapped by the standardized medical terminology set, cross-domain feature alignment is performed on the local node embedding vectors of the requesting doctor node and the target patient node to calculate the initial topological similarity. The initial topological similarity is converted into the initial topological difference, and the cross-domain topological association distance is generated by combining the initial topological difference with the intent difference penalty factor.

[0042] Extract the historical operation duration and historical access frequency parameters of the requesting doctor node, and concatenate them with the cross-domain topology association distance to generate a multi-dimensional access behavior feature vector; input the multi-dimensional access behavior feature vector into the risk decision tree model, and output the real-time risk quantification score for electronic medical record access requests;

[0043] The real-time risk quantification score is compared with the risk stratification limit in the multi-level access control rule base, and the control identifier corresponding to the electronic medical record access request is assigned and the corresponding access control operation is triggered. The control rule judgment path, the digital signature public key of the institutional gateway node participating in the calculation, and the intermediate layer activation feature are extracted to generate the source tracing log summary, and the source tracing log summary is archived to the source tracing graph database.

[0044] The steps for obtaining the intent subplot and the comprehensive medical history subplot are as follows:

[0045] The system receives electronic medical record access requests and associated access intent text transmitted by the requesting institution's gateway node. It reads the request identifier, target patient identifier, access field range, and access timestamp from the electronic medical record access request. It binds the access intent text according to the target patient identifier, locating the start and end positions of characters corresponding to disease names, symptom names, examination items, treatment behaviors, drug names, and body parts sentence by sentence. It unifies the medical terminology names corresponding to synonyms, determines semantic relationships based on the word order association, modification direction, and treatment orientation between adjacent medical entities, sets medical entities as nodes, and sets semantic relationships as node connection edges to establish an intent subgraph. Simultaneously, it retrieves the chief complaint text and past medical history record text from the current consultation triage system according to the target patient identifier, locating the medical history entities in the current consultation chief complaint and past medical history record text sentence by sentence. It establishes node connection edges according to the diagnostic, medication, examination, and symptom relationships between medical history entities, obtaining an intent subgraph and a comprehensive medical history subgraph.

[0046] Specifically, in the received electronic medical record access request and associated access intent text, the access intent text is first bound to the target patient based on metadata such as request identifier, target patient identifier, access field range, and access timestamp. Then, a deep learning-based natural language processing flow is used to parse the access intent text. Specifically, a ClinicalBERT-CRF model pre-trained on a Chinese medical text corpus (e.g., CMeEE, CMeKG) is used for named entity recognition. This model consists of a 12-layer Transformer encoder and a conditional random field decoder layer. The input is the sentence-by-sentence segmented intent text, and the model outputs Begin, Inside, Outside, character by character. End and Single tags are used to locate the start and end positions of characters in seven categories of medical entities in the text, such as disease names, drug names, examination items, and body parts, including "hypertension," "cephalosporins," "CT scan," and "left lower limb." Using a medical terminology thesaurus, the identified entities (e.g., "hypertension" and "primary hypertension") are standardized into the standard term "hypertension." After entity standardization, for each pair of adjacent medical entities in the text, a relationship classification task is constructed. The two entities and their intervening context text are input into a classification model also based on ClinicalBERT. The model output layer is a Softmax classifier used to determine whether there are predefined semantic relationships between two entities, such as "complications", "treatment drugs", "examination items", "symptoms" and 12 other types. Finally, standardized medical entities are used as nodes in the graph, and the identified semantic relationships are used as directed edges connecting the nodes to build an intent subgraph. At the same time, the same named entity recognition and relationship extraction process is used to process the patient's current visit complaint text retrieved from the triage system and the previous medical history record text from the electronic medical record system to construct a comprehensive medical history subgraph containing more comprehensive historical information, thus obtaining the intent subgraph and the comprehensive medical history subgraph.

[0047] The steps to obtain the intent difference penalty factor are as follows:

[0048] The medical term name, medical entity category, and attribute label of each node are read separately. The semantic relationship type, connection direction, and associated node number of each node's connecting edge are read separately. The nodes in the intent subgraph are mapped one by one to the candidate nodes in the comprehensive medical history subgraph that have the same medical term name or the same synonym unification result. The candidate node mappings that do not match the medical entity category are deleted. The mapping is extended layer by layer to the adjacent nodes along the candidate node mapping. The semantic relationship type, connection direction, and node adjacency order in the extension path are verified. The mapping combination that holds both the node correspondence and the connecting edge correspondence is retained. The number of matching nodes, the number of matching connecting edges, and the continuous matching level in all mapping combinations are compared. The mapping combination with the most matching nodes and the largest number of matching connecting edges is selected. The subgraph isomorphic matching degree is obtained by quantifying the proportion of the selected mapping combination to all nodes and all connecting edges in the intent subgraph.

[0049] The system reads the preset matching degree segment boundaries, the penalty benchmark value corresponding to each matching degree segment, and the decreasing change range corresponding to each matching degree segment. It compares the subgraph isomorphic matching degree with the matching degree segment boundaries in descending order to determine the unique matching degree segment corresponding to the subgraph isomorphic matching degree. It calls the penalty benchmark value corresponding to the matching degree segment, subtracts the subgraph isomorphic matching degree from the upper boundary of the matching degree segment to obtain the matching difference value, amplifies the matching difference value according to the decreasing change range corresponding to the matching degree segment, and superimposes the amplified matching difference value onto the penalty benchmark value. The superimposed result exceeding the preset penalty upper boundary is defined as the penalty upper boundary, and the superimposed result below the preset penalty lower boundary is defined as the penalty lower boundary. The numerical mapping is completed according to the correspondence that the superimposed result increases when the subgraph isomorphic matching degree decreases, thus obtaining the intent difference penalty factor.

[0050] Specifically, based on the intention subgraph obtained in the previous step and comprehensive medical history sub-diagram The process begins by calculating the subgraph isomorphism matching degree. This involves finding a subgraph in a large graph that is similar in structure and properties to a smaller graph. Considering that exact subgraph isomorphism is an NP-complete problem, a heuristic search algorithm is used for approximate matching, first traversing the intended subgraph. Each node in In the comprehensive medical history sub-diagram Search for all results related to Nodes whose medical terminology names are exactly the same or whose synonyms have the same unification result are formed as nodes. candidate mapping set Then, for the candidate mapping set Perform pruning if candidate nodes Medical entity categories (such as "diseases", "drugs") and nodes If the entity categories are inconsistent, then from After initial node mapping is completed, a constrained breadth-first search (BFS) is performed, starting from each node in the intent subgraph, to find the maximum common subgraph from an established mapping pair. (in Departure for inspection All adjacent nodes and the edges connecting them and try in Find matching nodes among the adjacent nodes that satisfy the same entity category, edge type, and edge direction. Continue expanding until no more matches can be found, recording all combinations of nodes and edges that successfully match during this process. After comparing all possible mapping combinations, select the one containing the most matched nodes and edges as the optimal mapping combination. Then, the subgraph isomorphic matching degree is obtained by quantification according to the following formula. :

[0051] ;

[0052] in, and These are the best mapping combinations The number of matched nodes and edges in the data. and It represents the total number of nodes and edges in the intentional subgraph. This is the weight balancing factor for node matching and edge matching. It is a preset hyperparameter, which is set to 0.5 based on experience, indicating that nodes and structures are equally important.

[0053] Read the subgraph isomorphic matching degree calculated in the previous step It is then non-linearly mapped to an intent difference penalty factor based on a set of preset configuration parameters. These configuration parameters are jointly formulated based on risk analysis of historical access data and expert security strategies, including the matching degree segment boundaries, the penalty baseline value corresponding to each segment, and the decreasing change range. For example, the preset matching degree segment boundaries are... Matching degree The range of values Divided into four intervals: (High matching zone) (Middle Matching Zone) (Low matching zone) and (Very low matching zone), each zone corresponds to a penalty baseline value. For example, respectively At the same time, each interval also corresponds to a decreasing range of change. For example, respectively During calculation, the subgraph isomorphic matching degree is first calculated. Compare the segments with the segment boundaries from largest to smallest to determine the unique segment in which it belongs. For example, if Then it falls into The interval, the upper boundary of the interval. Penalty benchmark Decreasing change range Then, calculate the matching difference within the segment, amplify it using the change amplitude, and finally superimpose it onto the baseline value. The calculation formula is as follows:

[0054] ;

[0055] Substitute the values ​​from the example, This calculation method results in a larger increase in the penalty value as the matching degree is lower within the same segment. To ensure the numerical stability of the penalty factor, boundary constraints are applied to limit all calculation results to the preset upper and lower boundaries of the penalty. Within this range, any calculation result less than 0 is reset to 0, and any result greater than 1 is reset to 1, ultimately resulting in a range of... Intentional Difference Punishment Factor .

[0056] The steps for obtaining the initial topological dissimilarity are as follows:

[0057] The system parses the requesting doctor identifier, requesting institution code, target patient identifier, and target institution code from the electronic medical record access request. Based on the requesting doctor identifier, it retrieves personnel nodes from the local knowledge graph of the requesting institution's gateway node, and based on the target patient identifier, it retrieves patient nodes from the local knowledge graph of the target institution's gateway node. It verifies the occupational type marker of the personnel nodes and the identity type marker of the patient nodes to determine the requesting doctor node and the target patient node. It reads the first-order adjacent nodes, progressively expanding adjacent nodes, connection edge directions, connection edge categories, node degrees, and adjacent path lengths of the requesting doctor node and the target patient node, respectively. It aggregates the node's own attribute components, adjacent node attribute components, and connection edge category codes layer by layer according to the corresponding levels. It performs dimensionality compression, component weighting, and non-linear mapping on each aggregation result. The local plaintext graph data is always retained within the corresponding institution's gateway node. It combines a preset update cycle for caching or real-time generation of local node embedding vectors and intermediate layer activation features.

[0058] The system reads the standardized medical terminology numbers jointly stored by the requesting institution's gateway node and the target institution's gateway node, filters out standardized medical terminology numbers that can be mapped to adjacent nodes of the requesting doctor node and the target patient node respectively, establishes a knowledge anchor point correspondence sequence according to the standardized medical terminology numbers, and splits the local node embedding vectors of the requesting doctor node and the target patient node into multiple random numerical shares respectively. An encrypted communication tunnel is established between the requesting institution's gateway node and the target institution's gateway node, exchanging only the random numerical shares that cannot be independently restored from the local node embedding vectors. The cross-product of the random numerical shares is calculated in the ciphertext state using a homomorphic encryption algorithm. The vector component arrangement position is unified according to the knowledge anchor point correspondence sequence, and the scale of the corresponding vector component is corrected according to the number of associated nodes of each knowledge anchor point. The vector components that have completed position unification and scale correction are multiplied and accumulated one by one to obtain the initial topological similarity. Then, according to the preset mapping order that the difference value decreases as the initial topological similarity increases, interval positioning, boundary constraints, and reverse numerical transformation are performed on the initial topological similarity to obtain the initial topological difference.

[0059] Specifically, after parsing the requesting doctor identifier, requesting institution code, target patient identifier, and target institution code from the electronic medical record access request, node location is performed within the local knowledge graph deployed in each institution's gateway node. The local knowledge graph is a heterogeneous graph composed of nodes such as personnel, departments, diseases, operations, and equipment, and the relationships between them. The requesting doctor identifier retrieves the personnel node labeled "doctor" by occupation type and identifies it as the requesting doctor node. The target patient identifier retrieves the node labeled "patient" by identity type and identifies it as the target patient node. Then, local node embedding vectors containing their neighborhood structure information are generated for these two nodes. This process uses the GraphSAGE graph neural network model, which contains three graph convolutional layers. Each layer performs neighbor information aggregation and self-information update operations. Specifically, for a node... In its first Layer embedding vector The generation process is as follows:

[0060] ;

[0061] ;

[0062] in, Represents a node The set of first-order adjacent nodes, It is the embedding vector of the neighboring node in the previous layer. The function represents taking the average of the vectors of all neighboring nodes as the aggregated neighborhood feature. , The operation will node vector of the previous layer It is then combined with the aggregated neighborhood features. It is the first The trainable weight matrix of a layer network, It is the activation function, the initial embedding vector of layer 0. The initial feature vector is constructed from the original attribute encodings of the nodes. For example, the attributes of a doctor node include its department (unique thermal encoding), professional title (unique thermal encoding), and number of historical operations (numerical value). These attributes are concatenated into an initial feature vector. The entire calculation process includes the atlas data and the intermediate layer activation features calculated from each layer (i.e., the activation features of each layer). All local node embedding vectors and intermediate layer activation features are kept within their respective gateway nodes and are not transmitted across domains. They are only generated at preset update cycles (e.g., every 24 hours) or when there is a real-time access request.

[0063] To calculate the topological similarity between the requesting doctor and the target patient while protecting data privacy, a standardized medical terminology set jointly maintained by both parties is first used as a knowledge anchor. Terms whose IDs can be mapped to both the requesting doctor's node neighborhood (e.g., diseases the doctor frequently treats) and the target patient's node neighborhood (e.g., the patient's past medical history) are selected. A knowledge anchor correspondence sequence is established according to these shared term IDs to align the vector dimensions in subsequent calculations. A scheme based on the Paillier homomorphic encryption algorithm is used to securely calculate the dot product of the embedding vectors of the two local nodes. Let requesting institution A possess the embedding vector of the requesting doctor. The target organization B possesses the embedding vector of the target patient. Organization B first generates its Paillier public-private key pair. and public key Send it to organization A, which uses the public key. Encrypt each component of its own vector to obtain the encrypted vector. It is then sent to Agency B, which locally performs a dot product calculation on the ciphertext using the homomorphic addition and scalar multiplication properties of the Paillier encryption algorithm:

[0064] ;

[0065] in Indicates the ciphertext Scalar multiplication performed This represents a homomorphic addition between encrypted values. After the calculation is complete, organization B uses its private key. Decrypt the result to obtain the dot product sum. That is, the initial topological similarity, and the vector of mechanism A throughout the process. It remains encrypted and invisible to organization B. Before calculation, the vector components need to be scaled based on the number of nodes associated with the knowledge anchor. For example, each component value is divided by the degree of its corresponding anchor in the local graph to avoid high-popularity nodes dominating the similarity calculation. Then, the obtained initial topological similarity is... By using the inverse function (e.g., Perform numerical conversion and map to The interval is used to obtain the initial topological dissimilarity. The higher the similarity, the lower the difference.

[0066] The steps to obtain the cross-domain topology association distance are as follows:

[0067] Read the lower and upper bounds of the initial topology difference value, and read the lower and upper bounds of the intended difference penalty factor value. Calculate the relative positions of the initial topology difference value and the intended difference penalty factor within their respective value boundaries. Transform the two relative positions to the same numerical range. Read the preset topology difference weight and intended penalty weight, and verify the sum of the topology difference weight and intended penalty weight. If the sum of the weights is not equal to the unit value, perform weight correction according to the proportion of each weight to the total weight. Multiply the converted initial topology difference value by the corrected topology difference weight, and multiply the converted intended difference penalty factor by the corrected intended penalty weight. Summate the two products. Values ​​exceeding the upper boundary of the summation are defined as the upper boundary, and values ​​below the lower boundary are defined as the lower boundary. Generate the cross-domain topology association distance.

[0068] Specifically, obtain the initial topological difference degree Intention Difference Punishment Factor Then, these two heterogeneous risk metrics are merged into a unified cross-domain topological association distance. First, read the valid value range of each of these two indicators, which are both... Since they are already in the same numerical range, no additional scaling or transformation is needed; the preset topology difference weights are read. and intentional punishment weight These two weights reflect the relative importance of the doctor-patient relationship and the legitimacy of the visit intent in the final risk assessment. The weights are set based on the experience of security strategy experts and statistical analysis of historical abnormal visit events. For example, for research-oriented visits, the match of intent is more important and can be set accordingly. For routine consultations, both are equally important and can be set up. Before calculation, verify the sum of the two weights. Then, normalize and correct the weights. , To ensure the weights sum to 1, the initial association distance is calculated using a weighted summation method, as shown in the following formula:

[0069] ;

[0070] by , and the corrected weights For example, the calculated initial distance is Furthermore, to ensure that the final output distance value strictly falls within the standard range... Within this scope, boundary constraints are applied to the weighted summation result to exclude any values ​​that might exceed the limits due to floating-point calculation errors or other reasons. The numerical range is subject to mandatory constraints, i.e., if Then let ,like Then let Generate the final cross-domain topological association distance. .

[0071] The steps for obtaining the feature vector of multidimensional access behavior are as follows:

[0072] The historical operation records corresponding to the requesting doctor node are read from the electronic medical record access request. The start timestamp, end timestamp, access event identifier, and target patient identifier of each historical operation record are extracted. The duration of each historical operation is calculated in the order of end timestamp minus start timestamp. The number of historical operation records with different access event identifiers within a preset statistical period is counted to form a historical access frequency parameter. The historical minimum and historical maximum values ​​of the historical operation duration and historical access frequency parameter are read respectively. The difference between the historical minimum value and the historical maximum value is obtained by subtracting the historical minimum value from the historical operation duration. The historical access frequency parameter is converted to a unified numerical range in the same order. The feature array is concatenated according to the fixed arrangement order of the historical operation duration normalization result, the historical access frequency parameter normalization result, and the cross-domain topological association distance to generate a multi-dimensional access behavior feature vector.

[0073] Specifically, behavioral pattern features are extracted from the historical operation logs of the requesting doctor nodes associated with electronic medical record access requests. Specifically, all historical operation records within the most recent statistical period (e.g., the past 30 days) are filtered out. Each record contains a start timestamp, end timestamp, access event identifier, and target patient identifier. The historical operation duration is calculated by subtracting the start timestamp from the end timestamp of each record. Then, all historical operation durations are aggregated, and their average is calculated. As a characteristic parameter of this visit, the total number of times the doctor visited different patients (i.e., those with different target patient identifiers) within this period is also counted and recorded as the historical visit frequency parameter. To eliminate the influence of different features on their dimensions, it is necessary to normalize these original features and the cross-domain topological association distances obtained from previous steps, and map them uniformly to... The interval, relative to the historical average operation duration. and historical access frequency parameter The minimum-maximum normalization method is used, and its calculation formula is as follows:

[0074] ;

[0075] in, It is the original eigenvalue ( or ), and These are the minimum and maximum values ​​of this feature recorded throughout the system's historical data. These values ​​are updated periodically as system configurations. For example, if the average historical operation time for all doctors in the system ranges from [5 seconds to 300 seconds], the calculated value... If the time is 65 seconds, then its normalized result is: Cross-domain topological association distance Because its own calculation results are already in For the interval, no additional processing is required. Then, the three values—the normalized result of historical operation duration, the normalized result of historical access frequency parameter, and the cross-domain topological association distance—are concatenated in this fixed order to form a three-dimensional feature array, i.e. This yields a multidimensional access behavior feature vector for subsequent risk assessment.

[0076] The steps to obtain the real-time risk quantification score are as follows:

[0077] The system reads the normalized results of historical operation duration, the normalized results of historical access frequency parameters, and the cross-domain topological association distance in a fixed order from the multidimensional access behavior feature vector. It then inputs the multidimensional access behavior feature vector into a pre-trained risk decision tree. The system reads the feature dimension number, quantization segmentation limit, left branch pointer, and right branch pointer bound to the root node. It compares the corresponding feature dimension value with the quantization segmentation limit. If the value is less than the quantization segmentation limit, it enters the left branch pointer node; if it is greater than or equal to the quantization segmentation limit, it enters the right branch pointer node. The system repeats reading the feature dimension number and quantization segmentation limit bound to the current node until the branch pointer corresponds to the leaf node. It records the feature dimension value, quantization segmentation limit, branch pointer, and leaf node number used in each quantization comparison to form a risk determination path.

[0078] Read the corresponding risk weights according to the leaf node numbers recorded in the risk assessment path, verify whether all feature dimension values ​​in the risk assessment path have been quantified and compared, verify whether the adjacent branch pointing indicators are continuously connected according to the previous node's pointing result, delete invalid risk weights corresponding to missing feature dimension values, missing quantification segmentation limits, and interrupted branch pointing indicators, accumulate the remaining risk weights one by one according to the recording order of the leaf node numbers, retain the risk weights corresponding to duplicate leaf node numbers only once, limit the values ​​of the accumulated result below the preset risk score lower boundary to the preset risk score lower boundary, limit the values ​​of the accumulated result above the preset risk score upper boundary to the preset risk score upper boundary, and output the real-time risk quantification score for the current electronic medical record access request.

[0079] Specifically, the multi-dimensional access behavior feature vector generated in the previous step... The input is fed into a pre-trained risk decision tree model for path reasoning in risk assessment. This decision tree model, such as the C4.5 or CART algorithm, is trained based on historical labeled access logs (labeled "compliant," "risky," and "high-risk"). Each non-leaf node of the tree is associated with a feature dimension and a splitting limit. Starting from the root node, the associated feature dimension number is read (e.g., dimension 3 corresponds to...). ) and quantization segmentation limit (e.g., 0.4), which will determine the numerical values ​​of the corresponding dimensions in the multidimensional access behavior feature vector (i.e., The value of the feature is compared with the limit. If the feature value is less than the limit, the node proceeds to the next node along the left branch; if it is greater than or equal to the limit, the node proceeds to the next node along the right branch. For example, if the input... If the value is 0.325, which is less than the splitting threshold of 0.4 for the root node, the decision path shifts to the left child node. This process is then repeated on the new node, which may be bound to another feature dimension, such as dimension 1 (corresponding to...). The system continuously compares features and selects branches, using the root node and its segmentation limit (e.g., 0.8), until a leaf node is reached. During this process, it is necessary to record the identifier of each non-leaf node traversed from the root node to the leaf node, its bound feature dimension number, the quantized segmentation limit used for comparison, the actual input feature value, and the selected branch direction (left or right). Finally, the leaf node number reached is added. This ordered set of records constitutes the unique decision trajectory of this access request in the decision tree model, resulting in the risk determination path.

[0080] Based on the risk assessment path generated in the previous step, the final real-time risk quantification score is calculated. First, according to the final leaf node number recorded in the risk assessment path, the corresponding risk weight value is read from the leaf node risk weight table of the risk decision tree model. This weight table is generated during the model training phase. Each leaf node is assigned a risk weight based on the risk purity of the sample it represents (i.e., the proportion of "risk" and "high-risk" samples contained in the node). Risk weights are assigned between nodes. For example, if a leaf node contains 80% high-risk samples, its risk weight can be set to 0.8. Before reading the weights, the completeness and continuity of the risk determination path must be verified. This involves checking whether the feature dimension values ​​of each decision node recorded in the path have been successfully used for quantification and comparison, and checking whether the branch pointing indicators between adjacent nodes are continuous. If a missing feature value or interrupted branch jump is found in the path (e.g., the previous node points to the left branch, but the recorded next node is not its left child node), the path is deemed invalid, and its corresponding risk weight will be discarded. For valid risk determination paths, the risk weight corresponding to its leaf node is used as the base score for this calculation. In some complex decision models (such as random forests), multiple trees may generate multiple decision paths and risk weights. In this case, the risk weights generated by all valid paths need to be accumulated or averaged. To avoid double calculation, if multiple paths point to the same leaf node, its weight is only counted once. Then, the total risk score after accumulation or averaging is subject to boundary constraints, limiting it to the preset upper and lower boundaries of the risk score. For example, the lower boundary of the risk score is set to 0 and the upper boundary to 100. If the calculation result is -10, it is corrected to 0; if it is 110, it is corrected to 100. A standardized real-time risk quantification score for the current electronic medical record access request is output.

[0081] The steps to obtain the source log summary are as follows:

[0082] The real-time risk quantification score is read level by level from the preset multi-level access control rule base, and the risk stratification limit is arranged in intervals according to the numerical order of the risk stratification limit. The real-time risk quantification score is compared with each risk stratification limit in turn. When the real-time risk quantification score is in the pass interval, a pass control flag is written. When the real-time risk quantification score is in the verification interval, a second authentication invocation control flag is written. When the real-time risk quantification score is in the blocking interval, a blocking control flag is written. The control flags along with the electronic medical record access request are sent to the corresponding institutional gateway node. The institutional gateway node opens the access session according to the pass control flag, terminates the access link according to the blocking control flag, and suspends the access session and initiates identity credential verification according to the second authentication invocation control flag, thus forming the access control operation result.

[0083] Based on the access control operation results, the feature dimension numbers, quantization segmentation limits, branch pointing identifiers, and leaf node numbers traversed sequentially from the root node to the hit leaf node during the risk assessment process are read. A control rule assessment path is formed according to the node traversal order. The digital signature public keys used by the requesting institution gateway node and the target institution gateway node in this calculation are read separately and arranged according to the institution code order. The intermediate layer activation features generated by the requesting doctor node and the target patient node at each propagation level are read and arranged according to the node identifier, propagation level, and feature component position. The control rule assessment path, digital signature public keys, and intermediate layer activation features are converted into character sequences of uniform length and concatenated according to a fixed field order. The concatenated character sequences are compressed and mapped segment by segment. The compression result of the previous segment is merged with the next segment of the character sequence until all character sequences are processed, generating a traceability log summary with tamper-proof characteristics.

[0084] The system reads the request identifier, requesting doctor node identifier, target patient node identifier, requesting institution code, target institution code, and access timestamp corresponding to the electronic medical record access request. It also reads the control identifier, identity verification status, and access link execution status from the access control operation results. The system sets the request identifier as the unique index of the access event node, binds the requesting doctor node identifier and target patient node identifier to the access event node, writes the requesting institution code and target institution code to the institution association field, writes the access timestamp to the time sequence field, combines the control identifier, identity verification status, and access link execution status into a judgment result, writes the source traceability log summary to the summary field, establishes the association edges between the access event node and the requesting doctor node, target patient node, and institution gateway node, and writes the unique indexes to storage records after verifying that there are no duplicates, thus forming the source traceability graph database archive record.

[0085] Specifically, the calculated real-time risk quantification score is compared with the risk stratification limits in a multi-level access control rule base to determine the specific access control operation. This rule base predefines the control levels corresponding to different risk score ranges, and is set according to the organization's security policy and compliance requirements (such as HIPAA). For example, a typical three-level control rule base is set as follows: the risk stratification limits are set to 30 and 70, thus forming three risk ranges and a permission range. This indicates low risk; the access behavior is considered compliant, and the verification range is [not specified]. This indicates medium risk; access behavior is uncertain and requires additional verification; blocking range. This indicates high risk; the access behavior is considered abnormal or unauthorized. Real-time risk quantification scores (e.g., 55) are sequentially compared to these limits. If the score falls within the verification range, the system assigns a "remind two-factor authentication" control flag to this access request. If the score is 25, it falls within the allow range and a "allow" control flag is written. If the score is 85, it falls within the block range and a "block" control flag is written. This control flag, along with the original electronic medical record access request, is then sent back to the requesting institution's gateway node. The gateway node acts as the policy enforcement point to execute the corresponding access control operations. Specifically, upon receiving the "allow" flag, the gateway node establishes or maintains the access session, allowing the data flow to pass normally. Upon receiving the "block" flag, the gateway node immediately terminates the TCP connection with the target system and records the blocking event. Upon receiving the "remind two-factor authentication" flag, the gateway node suspends the current access session and pushes an authentication challenge to the requesting doctor's client, such as requiring a dynamic password or fingerprint / facial recognition. The session can only be resumed after successful verification. The final execution status (allow, block, successful verification, failed verification) is recorded, forming the access control operation result.

[0086] Based on the access control operation results determined in the previous step, the process begins to generate a source log summary for post-event auditing and accountability. First, the complete control rule decision path is extracted from the risk assessment process, i.e., the sequence of all nodes traversed from the root node of the decision tree to the hit leaf node, including the feature dimension number, quantization segmentation limit, and branch direction of each node. This information is then organized into a standard format string. Next, the digital signature public keys used for this cross-domain computation are read from the local certificate storage of both the requesting and target institution gateway nodes, and these two public keys are then sorted alphabetically according to their institution codes. The process involves arranging and concatenating the vectors. Next, when generating local node embedding vectors, the intermediate layer activation features generated by the requesting doctor node and the target patient node at each propagation level of the graph neural network (e.g., layer 1 and layer 2) are extracted. These high-dimensional vectors are expanded into a long string in a fixed order of "node identifier - propagation level - feature component position". The three parts—the control rule decision path string, the digital signature public key string, and the intermediate layer activation feature string—are then combined using a cascaded hashing mechanism to generate a tamper-proof digest. Specifically, the SHA-256 algorithm is used, and the process is as follows:

[0087] ;

[0088] ;

[0089] ;

[0090] The "+" sign represents string concatenation. The hash result of the first part is used as salt in the hash calculation of the second part, forming a hash chain. This ensures that even a small change in any part of the content will result in a significant change in the final digest value, thus guaranteeing the integrity and non-repudiation of the log and obtaining the final source log digest.

[0091] After generating the source log summary, it is integrated with other relevant information of the access event and archived in a graph database (such as Neo4j) designed for source analysis. First, key metadata is extracted from the original electronic medical record access requests and access control operation results, including the globally unique identifier of the request, the requesting doctor node identifier, the target patient node identifier, the requesting institution code, the target institution code, and the access timestamp accurate to milliseconds. Simultaneously, the final control identifier (allow, block, verify), the specific status of the two-factor authentication (e.g., "not triggered," "successful," "failed"), and the final execution status of the access link (e.g., "connected," "blocked") are extracted from the access control operation results. Then, a new access event node of type "AccessEvent" is created in the source graph database, using the globally unique identifier of the request as the node name. To ensure unique index attributes and prevent record duplication, the requesting doctor node identifier and the target patient node identifier are used as attributes of this node. Based on these two identifiers, "REQUESTED_BY" and "TARGETED" relationship edges are established from the access event node to the existing "Doctor" and "Patient" nodes in the database. The requester and target organization codes, access timestamps, combined judgment results (control identifier, verification status, link status), and the long string format traceability log summary generated in the previous step are all stored as attributes in this access event node. Then, a database write operation is performed, which is encapsulated in a transaction to ensure data consistency. After verifying that there are no conflicts in the unique index, the transaction is committed, the archiving is completed, and a complete traceability graph database archive record with rich contextual relationships is formed.

[0092] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A method for controlling and tracing access to electronic medical record data, characterized in that, Includes the following steps: The system receives the electronic medical record access request and associated access intent text transmitted by the gateway node of the requesting institution, constructs an intent subgraph, extracts the current chief complaint and past medical history records of the target patient pointed to by the electronic medical record access request, and generates a comprehensive medical history subgraph. Calculate the subgraph isomorphic matching degree of the intent subgraph in the comprehensive medical history subgraph, and map the subgraph isomorphic matching degree to an intent difference penalty factor; Locate the requesting doctor node and the target patient node in the local knowledge graph, and generate local node embedding vectors and intermediate layer activation features; Based on the knowledge anchors mapped from the standardized medical terminology set, cross-domain feature alignment is performed on the local node embedding vectors of the requesting doctor node and the target patient node to calculate an initial topological similarity; the initial topological similarity is converted into an initial topological difference, and the cross-domain topological association distance is generated by combining the initial topological difference with the intent difference penalty factor. Extract the historical operation duration and historical access frequency parameters of the requesting doctor node, and concatenate them with the cross-domain topology association distance to generate a multi-dimensional access behavior feature vector; The multidimensional access behavior feature vector is input into the risk decision tree model, which outputs a real-time risk quantification score for the electronic medical record access request.

2. The method for accessing and tracing medical electronic medical record data according to claim 1, characterized in that, The method further includes: The real-time risk quantification score is compared with the risk stratification limit in the multi-level access control rule base, and a control identifier corresponding to the electronic medical record access request is assigned and the corresponding access control operation is triggered. Extract the control rule determination path, the digital signature public key of the agency gateway node participating in the calculation, and the activation feature of the intermediate layer to generate a traceability log digest, and archive the traceability log digest into the traceability graph database.

3. The method for accessing and tracing medical electronic medical record data according to claim 1, characterized in that, The steps for obtaining the intent sub-map and the comprehensive medical history sub-map are as follows: The system receives an electronic medical record access request and associated access intent text transmitted by the requesting institution's gateway node. It reads the request identifier, target patient identifier, access field range, and access timestamp from the electronic medical record access request. It binds the access intent text according to the target patient identifier, locating the start and end positions of characters corresponding to disease names, symptom names, examination items, treatment behaviors, drug names, and body parts sentence by sentence. It unifies the medical terminology names corresponding to synonyms, determines semantic relationships based on the word order association, modification direction, and treatment orientation between adjacent medical entities, sets medical entities as nodes, and sets semantic relationships as node connection edges to establish an intent subgraph. Simultaneously, it retrieves the chief complaint text and past medical history record text from the current consultation triage system according to the target patient identifier, locating the medical history entities in the current consultation chief complaint and past medical history record text sentence by sentence. It establishes node connection edges according to the diagnostic, medication, examination, and symptom relationships between medical history entities, obtaining an intent subgraph and a comprehensive medical history subgraph.

4. The method for accessing and tracing medical electronic medical record data according to claim 1, characterized in that, The steps for obtaining the intent difference penalty factor are as follows: The medical term name, medical entity category, and attribute label of each node are read separately. The semantic relationship type, connection direction, and associated node number of each node connection edge are read separately. The nodes in the intent subgraph are mapped one by one to the candidate nodes in the comprehensive medical history subgraph that have the same medical term name or the same synonym unification result. The candidate node mappings that do not have the same medical entity category are deleted. The mapping is extended layer by layer to the adjacent nodes along the candidate node mapping. The semantic relationship type, connection direction, and node adjacency order in the extension path are verified. The mapping combination that holds both the node correspondence and the connection edge correspondence is retained. The number of matching nodes, the number of matching connection edges, and the continuous matching level in all mapping combinations are compared. The mapping combination with the most matching nodes and the largest number of matching connection edges is selected. The subgraph isomorphic matching degree is obtained by quantifying the selected mapping combination according to the proportion of the selected mapping combination to all nodes and all connection edges in the intent subgraph. Read the preset matching degree segment boundaries, the penalty benchmark value corresponding to each matching degree segment, and the decreasing change range corresponding to each matching degree segment. Compare the subgraph isomorphic matching degree with the matching degree segment boundaries in descending order to determine the unique matching degree segment corresponding to the subgraph isomorphic matching degree. Call the penalty benchmark value corresponding to the matching degree segment. Subtract the subgraph isomorphic matching degree from the upper boundary of the matching degree segment to obtain the matching difference value. Amplify the matching difference value according to the decreasing change range corresponding to the matching degree segment. Superimpose the amplified matching difference value onto the penalty benchmark value. The superimposed result exceeding the preset penalty upper boundary is defined as the penalty upper boundary, and the superimposed result below the preset penalty lower boundary is defined as the penalty lower boundary. Complete the numerical mapping according to the correspondence that the superimposed result increases when the subgraph isomorphic matching degree decreases to obtain the intent difference penalty factor.

5. The method for accessing and tracing medical electronic medical record data according to claim 1, characterized in that, The steps for obtaining the initial topological difference are as follows: The requesting doctor identifier, requesting institution code, target patient identifier, and target institution code in the electronic medical record access request are parsed. Personnel nodes are retrieved in the local knowledge graph of the requesting institution gateway node according to the requesting doctor identifier, and patient nodes are retrieved in the local knowledge graph of the target institution gateway node according to the target patient identifier. The occupational type marker of the personnel node and the identity type marker of the patient node are verified to determine the requesting doctor node and the target patient node. The first-order adjacent nodes, progressively expanded adjacent nodes, connection edge directions, connection edge categories, node degrees, and adjacent path lengths of the requesting doctor node and the target patient node are read respectively. The node's own attribute components, adjacent node attribute components, and connection edge category codes are aggregated layer by layer according to the corresponding levels. Dimension compression, component weighting, and nonlinear mapping are performed on the aggregation results of each layer. The local plaintext graph data is always kept inside the corresponding institution gateway node. Local node embedding vectors and intermediate layer activation features are generated in real time or cached according to a preset update cycle. The standardized medical terminology numbers jointly stored by the requesting institution gateway node and the target institution gateway node are read. Standardized medical terminology numbers that can be mapped to adjacent nodes of the requesting doctor node and the target patient node are selected. A knowledge anchor correspondence sequence is established according to the standardized medical terminology numbers. The local node embedding vectors of the requesting doctor node and the target patient node are split into multiple random numerical shares. An encrypted communication tunnel is established between the requesting institution gateway node and the target institution gateway node. Only the random numerical shares that cannot be independently restored from the local node embedding vectors are exchanged. The cross-product of the random numerical shares is calculated in the ciphertext state using a homomorphic encryption algorithm. The vector component arrangement position is unified according to the knowledge anchor correspondence sequence. The scale of the corresponding vector component is corrected according to the number of associated nodes of each knowledge anchor. The vector components that have completed the position unification and scale correction are multiplied and accumulated one by one to obtain the initial topological similarity. Then, according to the preset mapping order that the difference value decreases as the initial topological similarity increases, interval positioning, boundary constraints, and reverse numerical transformation are performed on the initial topological similarity to obtain the initial topological difference.

6. The method for accessing and tracing medical electronic medical record data according to claim 1, characterized in that, The steps for obtaining the cross-domain topological association distance are as follows: Read the lower and upper bounds of the initial topology difference degree and the lower and upper bounds of the intention difference penalty factor. Calculate the relative positions of the initial topology difference degree and the intention difference penalty factor within their respective value boundaries. Convert the two relative positions to the same numerical range. Read the preset topology difference weight and intention penalty weight. Verify the sum of the topology difference weight and intention penalty weight. If the sum of weights is not equal to a unit value, perform weight correction according to the proportion of each weight to the total weight. Multiply the converted initial topology difference degree by the corrected topology difference weight, and multiply the converted intention difference penalty factor by the corrected intention penalty weight. Summate the two products. Values ​​exceeding the upper boundary of the summation are defined as the upper boundary, and values ​​below the lower boundary are defined as the lower boundary, generating the cross-domain topology association distance.

7. The method for accessing and tracing medical electronic medical record data according to claim 1, characterized in that, The steps for obtaining the multidimensional access behavior feature vector are as follows: The historical operation records corresponding to the requesting doctor node are read from the electronic medical record access request. The start timestamp, end timestamp, access event identifier, and target patient identifier of each historical operation record are extracted. The duration of each historical operation is calculated in the order of end timestamp minus start timestamp. The number of historical operation records with different access event identifiers within a preset statistical period is counted to form a historical access frequency parameter. The historical operation duration, the historical minimum value, and the historical maximum value of the historical access frequency parameter are read respectively. The difference obtained by subtracting the historical minimum value from the historical operation duration and dividing by the historical maximum value minus the historical minimum value is calculated. The historical access frequency parameter is converted to a unified numerical range in the same order. The feature array is concatenated according to the fixed arrangement order of the historical operation duration normalization result, the historical access frequency parameter normalization result, and the cross-domain topological association distance to generate a multi-dimensional access behavior feature vector.

8. The method for accessing and tracing medical electronic medical record data according to claim 1, characterized in that, The steps for obtaining the real-time risk quantification score are as follows: The normalized results of historical operation duration, the normalized results of historical access frequency parameters, and the cross-domain topological association distance are read according to the fixed arrangement order in the multidimensional access behavior feature vector. The multidimensional access behavior feature vector is input into the pre-trained risk decision tree. The feature dimension number, quantization segmentation limit, left branch pointing identifier, and right branch pointing identifier bound to the root node are read. The corresponding feature dimension value is compared with the quantization segmentation limit. If it is less than the quantization segmentation limit, the left branch pointing node is entered. If it is greater than or equal to the quantization segmentation limit, the right branch pointing node is entered. The feature dimension number and quantization segmentation limit bound to the current node are read repeatedly until the leaf node corresponding to the branch pointing identifier is reached. The feature dimension value, quantization segmentation limit, branch pointing identifier, and leaf node number used in each quantization comparison are recorded to form a risk judgment path. The risk weights are read according to the leaf node numbers recorded in the risk assessment path. It is verified whether the feature dimension values ​​in the risk assessment path have been quantified and compared. It is verified whether the adjacent branch pointing indicators are continuously connected according to the previous node pointing result. Invalid risk weights corresponding to missing feature dimension values, missing quantization segmentation limits, and interrupted branch pointing indicators are deleted. The remaining risk weights are accumulated item by item according to the recording order of the leaf node numbers. The risk weights corresponding to duplicate leaf node numbers are only retained once. The values ​​of the accumulated result below the lower boundary of the preset risk score are limited to the lower boundary of the preset risk score. The values ​​of the accumulated result above the upper boundary of the preset risk score are limited to the upper boundary of the preset risk score. The real-time risk quantification score for the current electronic medical record access request is output.

9. The method for accessing and tracing medical electronic medical record data according to claim 2, characterized in that, The steps for obtaining the source tracing log summary are as follows: The real-time risk quantification score is read level by level according to the preset multi-level access control rule base, and the risk stratification limit is arranged in intervals according to the numerical order of the risk stratification limit. The real-time risk quantification score is compared with each of the risk stratification limits in turn. When the real-time risk quantification score is in the release interval, a release control flag is written. When the real-time risk quantification score is in the verification interval, a second authentication activation control flag is written. When the real-time risk quantification score is in the blocking interval, a blocking control flag is written. The control flags along with the electronic medical record access request are sent to the corresponding institutional gateway node. The institutional gateway node opens the access session according to the release control flag, terminates the access link according to the blocking control flag, and suspends the access session and initiates identity credential verification according to the second authentication activation control flag, thus forming the access control operation result. Based on the access control operation results, the feature dimension numbers, quantization segmentation limits, branch pointing identifiers, and leaf node numbers traversed sequentially from the root node to the hit leaf node during the risk determination process are read. A control rule determination path is formed according to the node traversal order. The digital signature public keys used by the requesting institution gateway node and the target institution gateway node in this calculation are read respectively. The digital signature public keys are arranged according to the institution coding order. The intermediate layer activation features generated by the requesting doctor node and the target patient node at each propagation level are read and arranged according to the node identifier, propagation level, and feature component position. The control rule determination path, digital signature public keys, and intermediate layer activation features are converted into character sequences of uniform length. The first and last parts are concatenated according to the fixed field order. The concatenated character sequences are compressed and mapped segment by segment. The compression result of the previous segment is merged with the next segment of the character sequence until all character sequences are processed, generating a traceability log summary with anti-tampering characteristics. The system reads the request identifier, requesting doctor node identifier, target patient node identifier, requesting institution code, target institution code, and access timestamp corresponding to the electronic medical record access request. It also reads the control identifier, identity verification status, and access link execution status from the access control operation result. The request identifier is set as the unique index of the access event node. The requesting doctor node identifier and target patient node identifier are bound to the access event node respectively. The requesting institution code and target institution code are written to the institution association field. The access timestamp is written to the time sequence field. The control identifier, identity verification status, and access link execution status are combined into a judgment result. The source tracing log summary is written to the summary field. An association edge is established between the access event node and the requesting doctor node, target patient node, and institution gateway node. After verifying that the unique index does not have duplicates, the record is written to the storage record, forming an archived record in the source tracing graph database.