Medical data trusted circulation method and system based on multi-level authority dynamic management

By constructing a multi-dimensional user identity and permission model and applying multi-pattern matching algorithms and graph algorithms, the problems of coarse-grained permission management, static settings, and low retrieval efficiency in existing medical data management systems are solved, achieving refined management and efficient data sharing.

CN121583436BActive Publication Date: 2026-05-15GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
Filing Date
2026-01-26
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing medical data management systems suffer from problems such as coarse-grained access control, static access settings, low retrieval efficiency, and a lack of real-time monitoring and value assessment, making it difficult to meet the complex and ever-changing business processes and refined data sharing needs of the medical industry.

Method used

We construct a multi-dimensional user identity and permission model, apply a multi-pattern matching algorithm for efficient retrieval, use graph algorithms for dynamic permission evaluation, and build a security monitoring network for anomaly detection and risk warning.

Benefits of technology

It enables refined management of medical data permissions, improves data retrieval efficiency, prevents abuse of permissions, and conducts anomaly detection and risk warning during the circulation process, ensuring data security and sharing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583436B_ABST
    Figure CN121583436B_ABST
Patent Text Reader

Abstract

The application provides a medical data trusted circulation method and system based on multi-level permission dynamic management, which comprises the following steps: constructing a multi-mode matching index, performing feature extraction and vectorization on medical text data, structured data and image data; establishing a three-layer index architecture and an intelligent routing layer; applying a graph algorithm to construct a user-permission-data relationship graph, real-time assess the permission usage mode, and detect abnormal behaviors; combining with the situation factors such as medical emergency degree to dynamically adjust the abnormal judgment threshold, and grading processing and automatically collecting the evidence information of abnormal behaviors. The application also comprises a data classification method based on the secrecy level and time effectiveness, and a fine-grained permission control mechanism using the principle of least privilege, which supports dynamic permission granting and recycling. The application balances the security and availability, realizes the safe sharing and effective utilization of medical data, and meets the diversified data access requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information, and in particular to a method and system for the trusted circulation of medical data based on multi-level permission dynamic management, which is used to achieve efficient sharing and value mining of medical data while ensuring the security of medical data and patient privacy. Background Technology

[0002] As a highly valuable and sensitive data resource, the secure circulation and efficient utilization of medical data have always been core challenges in the development of healthcare informatization. With the explosive growth in the scale of medical data and the increasing prominence of its value, how to achieve efficient sharing and value mining of medical data while ensuring its security and patient privacy has become a critical issue that urgently needs to be addressed in the field of healthcare informatization.

[0003] Traditional medical data management systems typically employ role-based access control (RBAC) or attribute-based access control (ABAC) models for permission management. For example, some hospital information systems use pre-defined role divisions, such as doctors, nurses, and administrators, assigning fixed data access permissions to different roles; other systems dynamically calculate access permissions based on user attributes such as department and job level to achieve more flexible permission management.

[0004] Currently, more advanced medical data circulation platforms employ technologies such as blockchain and federated learning to establish centralized management and auditing mechanisms for data permissions. These systems achieve traceability of data flow by recording the entire data access process and ensure compliance to a certain extent by setting permission rules through technologies such as smart contracts.

[0005] However, existing technologies still have significant shortcomings: First, existing permission management models are coarse-grained, making it difficult to adapt to the complex and ever-changing business processes and refined data sharing needs of the healthcare industry; second, permission settings in existing systems are often static, lacking dynamic management of permission lifecycles, leading to permission creep and security risks; third, existing systems perform poorly in data retrieval efficiency, especially in the context of massive medical data, making it difficult to achieve accurate and rapid data matching and retrieval; finally, the lack of real-time monitoring and value assessment mechanisms for the entire data flow process makes it impossible to effectively quantify the release effect of data element value. Summary of the Invention

[0006] To address the aforementioned technical issues, this invention provides a method and system for the trusted circulation of medical data based on multi-level dynamic permission management. It achieves precise permission control of medical data by constructing a multi-dimensional user identity and permission model; it builds an efficient medical data indexing and retrieval mechanism by applying an improved multi-pattern matching algorithm; it prevents permission abuse by applying graph algorithms for dynamic permission evaluation; and it enables anomaly detection and risk warning during the medical data circulation process by constructing a security monitoring network.

[0007] The technical solution of the present invention is as follows:

[0008] This invention provides a method for trusted flow of medical data based on multi-level dynamic access control, comprising:

[0009] Based on multidimensional features including medical institutions, user roles, data attributes and access scenarios, medical data is classified and a data tagging system is established. User permission feature vectors are constructed, and dynamic permission determination is carried out in combination with contextual factors to obtain a multidimensional user identity and permission model.

[0010] Based on the multi-dimensional user identity and permission model, the medical data is feature extracted and vectorized. A distributed index structure is constructed using a multi-pattern matching algorithm. The user permission mapping relationship in the multi-dimensional user identity and permission model is extracted to filter and sort the search results, thereby obtaining the medical data search results with matching permissions.

[0011] Collect user access behavior data for medical data retrieval results matching the permissions, construct a user-permission-data relationship graph, apply graph algorithms to evaluate the user-permission-data relationship graph in real time, detect abnormal permission usage patterns, generate permission adjustment suggestions and execute dynamic adjustment strategies to obtain dynamic permission evaluation and adjustment results;

[0012] Based on the dynamic evaluation and adjustment results of the permissions, a medical data circulation security threat model is constructed, the layout of monitoring points is optimized, network traffic data, system logs, user behavior data and environmental data are collected from each monitoring point and integrated into multi-source heterogeneous data, security event correlation analysis is performed on the multi-source heterogeneous data, a hierarchical early warning and automatic response mechanism is established, and anomaly detection and risk warning data of the security monitoring network are obtained.

[0013] Furthermore, based on multi-dimensional features including medical institutions, user roles, data attributes, and access scenarios, medical data is classified and a data tagging system is established. A user permission feature vector is constructed, and dynamic permission determination is performed in conjunction with contextual factors, resulting in a multi-dimensional user identity and permission model, including:

[0014] The data attributes and access scenarios in the multidimensional features are obtained. Based on the data attributes, the medical data is classified into sensitivity levels and business value levels. Based on the access scenarios, the medical data is classified into usage scenarios. A multi-level medical data classification standard is constructed, a unified data tag system is established, and medical data classification tags are obtained.

[0015] Combining the medical data classification labels and the medical institutions and user roles in the multidimensional features, the user's affiliated medical institution, user professional role, scope of responsibilities and qualification level are extracted to construct a user permission feature vector. Based on the matching relationship between the medical data classification labels and the user permission feature vector, permission mapping is performed to obtain the user permission feature vector and permission mapping relationship.

[0016] The user permission feature vector and permission mapping relationship are obtained. User access time, access location, device type, and network environment are collected as contextual factors. Combined with historical access behavior patterns, machine learning algorithms are applied to construct a context-permission association model, calculate the permission credibility score in the current context, make dynamic permission judgment decisions, and obtain context-aware permission judgment results.

[0017] Based on the context-aware permission determination results, a full lifecycle management framework including permission application, approval, use, supervision, and revocation is established to achieve traceable permission management and automatic expiration mechanism, thereby obtaining the multi-dimensional user identity and permission model.

[0018] Furthermore, the construction of a multi-level medical data classification standard and the establishment of a unified data labeling system result in medical data classification labels, including:

[0019] Medical data is acquired and classified into public, internal, sensitive and confidential levels based on its sensitivity; basic data, business data and core data based on its business value; and diagnosis and treatment data, scientific research data and management data based on its usage scenario, resulting in multi-level data classification results.

[0020] Based on the multi-level data classification results, standardized labels containing data type, sensitivity level, business attributes, and access requirements are established for each type of data to obtain the medical data classification labels.

[0021] Further, the construction of the user permission feature vector, based on the matching relationship between the medical data classification labels and the user permission feature vector, performs permission mapping to obtain the user permission feature vector and the permission mapping relationship, including:

[0022] The medical institution and user role in the multidimensional features are obtained, and the user's medical institution, professional role, scope of responsibility, qualification level and department level are extracted. The features of each dimension are quantified and encoded, and the feature fusion technology is used to obtain the user permission feature vector.

[0023] Based on the user permission feature vector and the medical data classification label, the matching degree between the user permission feature vector and the medical data classification label is calculated, and a permission mapping matrix between users and data is established to obtain the user permission feature vector and the permission mapping relationship.

[0024] Furthermore, the application of machine learning algorithms constructs a context-permission association model, calculates the permission credibility score in the current context, performs dynamic permission determination decisions, and obtains context-aware permission determination results, including:

[0025] Obtain the user permission feature vector and permission mapping relationship, collect user access time, access location, device type, network environment, extract context features and perform standardization processing to obtain the context feature vector;

[0026] Based on the context feature vector and historical access behavior data, machine learning algorithms are applied to establish the correlation between context factors and the rationality of permission use, calculate the permission credibility score in the current context, and obtain the context-adaptive permission score.

[0027] Based on the context-adaptive permission score and the preset permission judgment threshold, a dynamic permission judgment decision is made to generate a judgment result of granting, denying, or requiring secondary verification, thus obtaining the context-aware permission judgment result.

[0028] Furthermore, the process of extracting and vectorizing features from medical data, constructing a distributed index structure using a multi-pattern matching algorithm, and extracting user permission mapping relationships from the multi-dimensional user identity and permission model to filter and sort the search results yields permission-matched medical data search results, including:

[0029] Obtain structured and unstructured medical data, apply natural language processing and image recognition technologies to extract key features of the medical data, convert them into standardized feature vectors, and obtain medical data feature vectors;

[0030] Based on the medical data feature vector, a multi-pattern matching algorithm is applied to index the data, resulting in a multi-pattern matching index.

[0031] Based on the multi-pattern matching index, a distributed index structure containing a primary index, secondary indexes, and temporary indexes is constructed. An index routing layer is established, and the index path is determined based on query characteristics, user roles, and system load to obtain a distributed index system.

[0032] By combining the distributed indexing system and the multi-dimensional user identity and permission model, the user permission mapping relationship in the multi-dimensional user identity and permission model is extracted. The search results are dynamically filtered and security-de-identified according to the user permission mapping relationship. Personalized sorting is performed based on user preferences and historical behavior to obtain the medical data search results with the matching permissions.

[0033] Furthermore, the application of a multi-pattern matching algorithm for data indexing to obtain a multi-pattern matching index includes:

[0034] Based on the medical data feature vectors and query complexity, a sliding window with dynamically adjustable size is maintained. A window suffix tree is constructed only for the medical data feature vectors within the sliding window to obtain a dynamic window suffix tree index.

[0035] The medical data feature vector is divided into multiple data shards, and a suffix tree index is constructed in parallel for each data shard. A work-stealing load balancing strategy is adopted to obtain a set of parallel suffix tree indexes.

[0036] Based on the dynamic window suffix tree index and the parallel suffix tree index set, a two-layer index structure is constructed using the Aho-Corasick automaton. The outer layer uses the parallel suffix tree index set to index the dataset, while the inner layer uses the Aho-Corasick automaton to simultaneously match multiple patterns, resulting in the multi-pattern matching index.

[0037] Furthermore, the construction of a distributed index structure including primary indexes, secondary indexes, and temporary indexes, the establishment of an index routing layer, and the determination of index paths based on query characteristics, user roles, and system load result in a distributed index system, including:

[0038] Based on the multi-pattern matching index, a master index is constructed for the disease type, diagnosis and treatment methods and drug treatment characteristics of medical data. The master index is stored in a distributed database cluster using a consistent hashing sharding strategy to obtain a global master index.

[0039] Based on the multi-pattern matching index and business scenario requirements, secondary indexes are constructed for data subsets of predetermined departments, predetermined disease areas, or predetermined research directions, and deployed on edge nodes to obtain scenario-based secondary indexes.

[0040] Based on the needs of hot queries or temporary tasks, a short-term index is built using in-memory data structures and dynamic window features, and an automatic expiration and clearing mechanism is set to obtain a temporary index.

[0041] By integrating the global primary index, the scenario-based secondary index, and the temporary index, an index routing layer is established. Based on query features, user roles, and system load, a machine learning model is applied to evaluate the index path, resulting in the distributed index system.

[0042] Furthermore, the application graph algorithm performs real-time evaluation of the user-permission-data relationship graph, detects abnormal permission usage patterns, generates permission adjustment suggestions, and executes dynamic adjustment strategies to obtain dynamic permission evaluation and adjustment results, including:

[0043] Obtain the medical data retrieval results matching the permissions, collect the user's access behavior data on the medical data retrieval results matching the permissions in real time, construct the user-permission-data relationship graph containing user nodes, permission nodes and data nodes, and obtain the permission usage behavior relationship graph;

[0044] Based on the permission usage behavior relationship graph, a graph algorithm is applied to calculate node weights and path weights, and the permission usage behavior relationship graph is evaluated in real time to obtain permission usage evaluation results.

[0045] Based on the permission usage evaluation results and the permission usage behavior relationship diagram, a normal permission usage pattern feature library is constructed. The deviation between the current user's permission usage pattern and the normal permission usage pattern feature library is calculated. The degree of medical urgency, business stage characteristics, and organizational policy changes are introduced as contextual factors to dynamically adjust the anomaly judgment threshold. The deviation is compared with the anomaly judgment threshold, and the anomaly is classified according to the comparison results to obtain the permission usage anomaly detection results.

[0046] Based on the abnormal permission usage detection results, the permission usage evaluation results are integrated to generate permission adjustment suggestions. Based on the permission adjustment suggestions, an approval process is automatically generated and the permission adjustment is executed to obtain the dynamic evaluation and adjustment results of the permissions.

[0047] Furthermore, the application graph algorithm calculates node weights and path weights, performs real-time evaluation of the permission usage behavior relationship graph, and obtains permission usage evaluation results, including:

[0048] Based on the permission usage behavior relationship graph, key nodes are identified using node centrality and feature vector centrality, and the subgraph formed by the key nodes is extracted to obtain the key node subgraph.

[0049] Based on the key node subgraph, user-permission granting weight, permission-data coverage weight, user-data access weight and time decay factor are defined. The proportion of each type of weight in the total weight is dynamically adjusted according to the characteristics of the medical business process to obtain adaptive weight parameters.

[0050] Based on the adaptive weight parameters, the relevant path weights are recalculated only for the newly added or changed behavior data in the permission use behavior relationship graph. The weight distribution is estimated by using the random walk approximation technique to obtain the approximate weight calculation results.

[0051] Based on the approximate weight calculation results, a distributed computing model is used to distribute the computing tasks to multiple computing nodes, and the results are aggregated through a message passing mechanism to obtain the permission usage evaluation results.

[0052] Furthermore, the process involves constructing a normal permission usage pattern feature library, calculating the deviation between the current user's permission usage pattern and the feature library, introducing factors such as medical urgency, business stage characteristics, and organizational policy changes as contextual factors to dynamically adjust the anomaly detection threshold, comparing the deviation with the anomaly detection threshold, and classifying the anomaly based on the comparison results to obtain anomaly detection results for permission usage, including:

[0053] Based on the permission usage evaluation results and the permission usage behavior relationship diagram, historical behavior data is extracted, and the permission usage normal mode feature library is constructed for different roles, different departments and different business scenarios. The permission usage normal mode feature library includes statistical features, sequence features, correlation features and content features. The normal behavior distribution is represented by a Gaussian mixture model to obtain the permission usage benchmark model.

[0054] Based on the permission usage benchmark model and the current user permission usage pattern, calculate the deviation between the current user permission usage pattern and the permission usage benchmark model, evaluate the deviation of single behavior, the deviation of sequential behavior and the deviation of pattern, and obtain the permission usage deviation score.

[0055] Based on the permission usage deviation score, the severity of medical urgency, business stage characteristics, and organizational policy changes are introduced as contextual factors. The anomaly judgment threshold is dynamically adjusted through preset rules and machine learning models. The permission usage deviation score is compared with the anomaly judgment threshold to obtain the context-corrected anomaly score.

[0056] Based on the anomaly score after context correction, abnormal behaviors are categorized into observation level, alert level, intervention level, and blocking level. For anomalies at the intervention level and above, evidence information including operation context, system status, and user response is automatically collected to obtain the permission usage anomaly detection results.

[0057] Furthermore, the process involves constructing a medical data flow security threat model, optimizing the layout of monitoring points, collecting network traffic data, system logs, user behavior data, and environmental data from each monitoring point and integrating them into multi-source heterogeneous data. Security event correlation analysis is then performed on this multi-source heterogeneous data to establish a tiered early warning and automatic response mechanism, resulting in anomaly detection and risk warning data for the security monitoring network, including:

[0058] Based on the dynamic evaluation and adjustment results of permissions, user nodes, data nodes, access paths and permission relationships in the medical data circulation process are extracted, a medical data circulation network topology is constructed, the types of security threats in the medical data circulation network topology are identified, and attack tree and attack graph methods are used to simulate attack paths and quantify risk levels to obtain a medical data circulation security threat model.

[0059] Based on the attack paths and risk levels in the medical data circulation security threat model, the medical data circulation network topology is abstracted into a directed weighted graph. The cut set algorithm is applied to calculate and optimize the location of monitoring points to obtain a monitoring point layout scheme.

[0060] According to the monitoring point layout scheme, network traffic data, system logs, user behavior data and environmental data are collected from each monitoring point, and standardized processing is performed to obtain the multi-source heterogeneous data.

[0061] Based on the aforementioned multi-source heterogeneous data, time-series correlation analysis and causal reasoning are applied to perform security event correlation analysis, assess the severity of threats and the urgency of responses, classify security events into observation level, alert level, warning level and emergency level, and execute automated response measures for different levels to obtain anomaly detection and risk warning data of the security monitoring network.

[0062] This invention also provides a trusted medical data circulation system based on multi-level permission dynamic management, comprising:

[0063] The multi-dimensional user identity and permission model construction module is used to classify medical data and establish a data tag system based on multi-dimensional features including medical institutions, user roles, data attributes and access scenarios, construct user permission feature vectors, and perform dynamic permission determination in combination with contextual factors to obtain a multi-dimensional user identity and permission model.

[0064] The permission-matching medical data retrieval module is used to extract and vectorize features from medical data based on the multi-dimensional user identity and permission model, construct a distributed index structure using a multi-pattern matching algorithm, extract user permission mapping relationships from the multi-dimensional user identity and permission model, filter and sort the retrieval results, and obtain permission-matched medical data retrieval results.

[0065] The permission dynamic evaluation and adjustment module is used to collect user access behavior data of medical data retrieval results matching the permissions, construct a user-permission-data relationship graph, apply graph algorithms to evaluate the user-permission-data relationship graph in real time, detect abnormal permission usage patterns, generate permission adjustment suggestions and execute dynamic adjustment strategies to obtain the permission dynamic evaluation and adjustment results.

[0066] The security monitoring network anomaly detection and risk warning module is used to construct a medical data circulation security threat model based on the dynamic evaluation and adjustment results of the permissions, optimize the layout of monitoring points, collect network traffic data, system logs, user behavior data and environmental data from each monitoring point and integrate them into multi-source heterogeneous data, perform security event correlation analysis on the multi-source heterogeneous data, establish a hierarchical early warning and automatic response mechanism, and obtain anomaly detection and risk warning data of the security monitoring network.

[0067] The beneficial technical effects of this invention include:

[0068] By constructing a multi-dimensional user identity and permission model, this invention achieves refined management of medical data permissions; by applying a multi-pattern matching algorithm, it improves the efficiency of medical data retrieval; by applying a graph algorithm for dynamic permission evaluation, it prevents permission abuse; and by constructing a security monitoring network, it enables anomaly detection and risk warning during the flow of medical data. This invention not only ensures the security of medical data but also improves the efficiency of medical data sharing and utilization, providing strong support for the value mining of medical data. Attached Figure Description

[0069] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0070] Figure 1 This is a flowchart of the trusted medical data circulation method based on multi-level permission dynamic management according to the present invention.

[0071] Figure 2 This is a flowchart of the medical data retrieval process based on the multi-pattern matching algorithm of the present invention;

[0072] Figure 3 This is a structural diagram of the trusted medical data circulation system based on multi-level permission dynamic management according to the present invention. Detailed Implementation

[0073] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0074] Example 1

[0075] like Figure 1 As shown, the trusted flow method for medical data based on multi-level permission dynamic management provided by the present invention includes:

[0076] Step S1: Based on the multi-dimensional features including medical institutions, user roles, data attributes and access scenarios, medical data is classified and a data tagging system is established. User permission feature vectors are constructed, and dynamic permission determination is carried out in combination with contextual factors to obtain a multi-dimensional user identity and permission model.

[0077] Step S2: Based on the multi-dimensional user identity and permission model, extract and vectorize the features of the medical data, apply a multi-pattern matching algorithm to construct a distributed index structure, extract the user permission mapping relationship in the multi-dimensional user identity and permission model, filter and sort the search results, and obtain the medical data search results with permission matching.

[0078] Step S3: Collect user access behavior data of medical data retrieval results matching the permissions, construct user-permission-data relationship graph, apply graph algorithm to evaluate the user-permission-data relationship graph in real time, detect abnormal permission usage patterns, generate permission adjustment suggestions and execute dynamic adjustment strategies to obtain permission dynamic evaluation and adjustment results;

[0079] Step S4: Based on the dynamic evaluation and adjustment results of the permissions, construct a medical data circulation security threat model, optimize the layout of monitoring points, collect network traffic data, system logs, user behavior data and environmental data from each monitoring point and integrate them into multi-source heterogeneous data, perform security event correlation analysis on the multi-source heterogeneous data, establish a hierarchical early warning and automatic response mechanism, and obtain anomaly detection and risk warning data of the security monitoring network.

[0080] Example 2

[0081] In this embodiment, based on multi-dimensional features including medical institutions, user roles, data attributes, and access scenarios, medical data is classified and a data tagging system is established. A user permission feature vector is constructed, and dynamic permission determination is performed in conjunction with contextual factors to obtain a multi-dimensional user identity and permission model, including:

[0082] The data attributes and access scenarios in the multidimensional features are obtained. Based on the data attributes, the medical data is classified into sensitivity levels and business value levels. Based on the access scenarios, the medical data is classified into usage scenarios. A multi-level medical data classification standard is constructed, a unified data tag system is established, and medical data classification tags are obtained.

[0083] Combining the medical data classification labels and the medical institutions and user roles in the multidimensional features, the user's affiliated medical institution, user professional role, scope of responsibilities and qualification level are extracted to construct a user permission feature vector. Based on the matching relationship between the medical data classification labels and the user permission feature vector, permission mapping is performed to obtain the user permission feature vector and permission mapping relationship.

[0084] The user permission feature vector and permission mapping relationship are obtained. User access time, access location, device type, and network environment are collected as contextual factors. Combined with historical access behavior patterns, machine learning algorithms are applied to construct a context-permission association model, calculate the permission credibility score in the current context, make dynamic permission judgment decisions, and obtain context-aware permission judgment results.

[0085] Based on the context-aware permission determination results, a full lifecycle management framework including permission application, approval, use, supervision, and revocation is established to achieve traceable permission management and automatic expiration mechanism, thereby obtaining the multi-dimensional user identity and permission model.

[0086] Specifically, firstly, data attributes and access scenario information from the medical data system are obtained through database connections or API calls. Data attributes cover metadata such as the categories, formats, and sources of patient personal information and medical records; access scenarios include different usage environments such as routine medical treatment and clinical research.

[0087] A four-level sensitivity classification system is implemented based on data attributes. Using a rule engine containing approximately 200 pre-defined rules, data is divided into public, internal, sensitive, and confidential levels. Data containing direct identifiers is automatically classified as confidential, while data containing specific disease codes is classified as sensitive. Simultaneously, data is categorized by business value into basic data, business data, and core data, using a multi-factor weighted average scoring mechanism to determine value levels. Combining access scenarios, semantic analysis and context recognition technologies are used to classify data into clinical data, research data, and management data, forming a three-dimensional classification system based on sensitivity, value, and scenario. Furthermore, standardized tags containing data type and sensitivity level information are created for each data category, stored in JSON format, including main tags and extended tags, facilitating system parsing and matching.

[0088] Second, extract multidimensional user characteristics from the user management system, including information such as the name and type of the medical institution to which the user belongs, the user's professional role and detailed departmental functions, the direct and extended responsibilities determined by analyzing job descriptions and actual work content, and the four levels of qualification (primary, intermediate, advanced, and expert) assessed based on factors such as professional qualification certificates.

[0089] The extracted user features are quantitatively encoded: hierarchical encoding is used for institutions, one-hot encoding for roles, multi-label encoding for responsibilities, and qualification levels are directly mapped to values ​​from 1 to 4. Through feature fusion technology (including feature standardization, selection, and combination), these features are integrated into a 128-dimensional or 256-dimensional user permission feature vector. Permission mapping is then performed based on the matching relationship between this vector and medical data classification labels. A two-layer mechanism combining rules and models is employed: the rule layer implements basic permission allocation, and the model layer handles complex scenario judgments, ultimately constructing a permission mapping matrix that records the user's specific operational permissions.

[0090] Third, after obtaining the user permission feature vector and permission mapping relationship, contextual factors such as access time, location, device type, and network environment are collected. Access time is obtained from the system clock, location is determined through technologies such as IP positioning, device type is identified using user agent strings, and network environment is determined by analyzing connection types.

[0091] The raw contextual data, after preprocessing including outlier detection, missing value imputation, and noise filtering, is standardized and converted into contextual feature vectors. Combining this with historical user access behavior data, an ensemble learning framework is applied to construct a context-permission association model, comprising three sub-models: temporal pattern, spatial pattern, and environmental security. These are then weighted and fused to form a comprehensive evaluation model. This model is trained through supervised learning. When a user requests access, they input the current contextual features and calculate a permission credibility score between 0 and 1. The score considers three factors: consistency between the current context and the regular pattern, consistency within the context, and similarity to historical similar contexts. A judgment threshold is dynamically set based on data sensitivity, and a judgment result of permission, denial, or requiring secondary verification is given based on the score. "Gray zone" requests require secondary verification such as CAPTCHA verification.

[0092] Fourth, based on context-aware permission determination results, construct a permission lifecycle management framework covering five stages: application, approval, use, supervision, and revocation, forming a closed-loop management system.

[0093] The permission application stage provides a self-service interface, with intelligent form recommendations for permission sets and compliance checks. The approval stage implements multi-level hierarchical approval based on data sensitivity levels, supporting conditional approval. During usage, operation information is recorded in real time, and logs are stored using blockchain; access to highly sensitive data is recorded using a "dual-active" system. The supervision stage establishes an automated, management-level, and cross-level supervision mechanism, triggering alerts for abnormal situations. The revocation stage terminates permissions through three trigger methods: time, event, and exception, clearing configurations and notifying the user after execution. The framework has full lifecycle traceability and includes an automatic expiration mechanism, setting a dynamic validity period of 7-90 days for each permission, providing renewal reminders upon expiration, and automatically revoking permissions for long-term inactivity.

[0094] This framework, combining user permission feature vectors, permission mapping relationships, and context-aware judgment mechanisms, forms a multi-dimensional user identity and permission model that integrates user attributes, data attributes, context, and lifecycle, enabling refined and dynamic management of medical data access permissions.

[0095] Example 3

[0096] In this embodiment, the construction of a multi-level medical data classification standard and the establishment of a unified data labeling system to obtain medical data classification labels include:

[0097] Medical data is acquired and classified into public, internal, sensitive and confidential levels based on its sensitivity; basic data, business data and core data based on its business value; and diagnosis and treatment data, scientific research data and management data based on its usage scenario, resulting in multi-level data classification results.

[0098] Based on the multi-level data classification results, standardized labels containing data type, sensitivity level, business attributes, and access requirements are established for each type of data to obtain the medical data classification labels.

[0099] Specifically, firstly, various types of medical data are acquired through interfaces from multiple medical data sources, including Hospital Information System (HIS), Electronic Medical Record System (EMR), Laboratory Information System (LIS), and Picture Archiving and Communication System (PACS). The acquisition process employs an incremental synchronization mechanism, periodically pulling updated data from the source systems, and using data fingerprinting technology (such as hash value comparison) to avoid duplicate acquisitions. After acquisition, the data undergoes preliminary cleaning to remove obvious outliers and records with format errors, ensuring the quality of data for subsequent classification.

[0100] Data is categorized into four levels based on its sensitivity. Sensitivity assessment employs a dual "content + structure" analysis method, examining both the sensitivity of the data content and analyzing its structure and relationships. Public-level data refers to non-sensitive information that can be publicly disclosed, such as basic hospital information, departmental setup, and physician professional backgrounds. Internal-level data refers to information that can be shared within the institution but is not suitable for public disclosure, such as de-identified medical statistics and general treatment procedures for non-specific patients. Sensitive-level data includes patient treatment information containing indirect identifiers, such as patient data aggregated by age group, region, and disease type; although processed, it still poses a potential risk of re-identification. Confidential-level data contains patient treatment information with direct identifiers, such as complete electronic medical records, prescription records, and laboratory reports for specific patients.

[0101] Sensitivity assessment employs a combination of automated analysis and manual review. Automated analysis uses natural language processing technology to detect sensitive words, personal identifiers, and specific disease information in the data, calculating an initial sensitivity level based on pre-defined sensitivity scoring rules. Manual review involves medical data security experts sampling and verifying the automated classification results, paying particular attention to boundary cases and emerging data types. The sensitivity assessment process places special emphasis on identifying particularly sensitive information, such as mental illness records, sexually transmitted disease data, and HIV test results; this type of information is automatically classified as confidential.

[0102] Business value grading is a classification of medical data from a value creation perspective, and this is one of the innovative aspects of this system. Basic data consists of essential but relatively standardized information necessary for medical activities, such as basic patient information, routine physical examination data, and standard treatment records. This type of data has stable value, but the marginal value of a single data point is relatively low. Business data is professional information supporting daily medical operations, such as diagnostic conclusions, treatment plans, medication records, and surgical records. This type of data has high clinical reference value. Core data is data with special research value or innovative potential, such as rare case records, innovative treatment plans, and long-term follow-up data. This type of data is usually scarce and difficult to replicate, possessing the highest business value.

[0103] The business value assessment employs a multi-dimensional quantitative scoring model, primarily considering four factors: data scarcity (the difficulty and uniqueness of data acquisition), application breadth (the range of scenarios in which the data can be applied), analytical depth (the level of insight that can be mined from the data), and decision-making influence (the degree to which the data supports medical decisions). The scoring process uses a weighted average method, allowing different types of medical institutions to adjust the weights of each factor according to their own characteristics. For example, research hospitals may place greater emphasis on data scarcity and analytical depth, while general hospitals may focus more on application breadth and decision-making influence.

[0104] Use case classification is a functional division based on the intended purpose of the data. Clinical data is directly used in clinical medical activities, such as patient diagnosis, treatment plan development, and nursing records. This type of data typically requires real-time access and is closely linked to specific patients. Research data is used for medical research and teaching, and is usually anonymized or aggregated, such as clinical trial data, disease trend analysis, and medical outcome evaluation. Management data is used for hospital operations and administration, such as resource allocation, cost analysis, and quality control.

[0105] The use case classification employs a goal-oriented analysis method, determining the most suitable use case category by analyzing the data's structural characteristics, access patterns, and correlations. Three key indicators are examined: data granularity (individual-level or aggregate-level), timeliness requirements (level of real-time requirement), and correlation scope (single-case correlation or multi-case correlation). The combination of these three indicators can accurately determine the main use case of the data. For example, high-granularity, high-timeliness, single-case correlation data is typically classified as clinical data; medium-to-low granularity, low-timeliness, multi-case correlation data is typically classified as research data.

[0106] It is important to note that the same data may belong to different scenario categories at different processing stages. For example, a patient's original test report data belongs to medical data, but after de-identification processing, it may become part of research data, and after statistical aggregation, it may be transformed into management data for quality control. The system tracks the flow of data between different scenarios through data processing tagging and version control.

[0107] By classifying data along these three dimensions, a three-dimensional classification matrix is ​​constructed, allowing each piece of medical data to be precisely located as a point in the matrix, such as "Sensitive Level - Business Data - Medical Data". This multi-level classification result forms the basis for the subsequent construction of the tagging system.

[0108] Secondly, based on the results of multi-level data classification, a unified medical data tagging system was established. Data tags are metadata attached to data, used to describe data characteristics and control data access, and are the infrastructure for achieving fine-grained access control. The tagging system of this system adopts a hierarchical structure design, including a basic tag layer and an extended tag layer.

[0109] The basic tag layer contains four core tags: data type tags, sensitivity level tags, business attribute tags, and access requirement tags. Data type tags describe the basic category and format characteristics of the data, such as "test report," "CT image," and "prescription record." Sensitivity level tags directly correspond to the sensitivity classification results, marked as "public," "internal," "sensitive," or "confidential." Business attribute tags contain business value classification and usage scenario categorization information, such as combined tags like "basic data - medical treatment" and "core data - scientific research." Access requirement tags specify the control requirements for data access, such as "visible to the entire hospital," "shared within departments," "access by designated personnel," and "access requiring approval."

[0110] The extended tag layer contains finer-grained feature descriptions and control rules, customized for specific data types or special application scenarios. Common extended tags include: timeliness tags (marking the validity period and update cycle of data), association tags (marking the relationships between data), processing tags (marking the allowed processing operations on data), and propagation tags (marking the propagation range of data), etc. Extended tags can be flexibly configured according to the specific needs of medical institutions, enhancing the adaptability of the tagging system.

[0111] The tagging process is a crucial step in transforming classification results into standardized labels. An automatic tag generation engine is employed for efficient tagging. This engine includes a classification-tag mapping rule base, capable of automatically generating corresponding standard labels based on the data's classification attributes. For example, for patient surgical records classified as "Confidential - Core Data - Medical Data," the system will automatically generate the following tag combination: {Data Type="Surgical Record", Sensitivity Level="Confidential", Business Attribute="Core - Medical", Access Requirement="Access by Designated Medical Personnel + Department Head Approval"}.

[0112] After tags are generated, a tag consistency verification mechanism is applied to ensure the integrity and logical consistency of the tags. Verification rules include: necessary tag integrity checks (ensuring all core tags exist), tag value validity checks (ensuring tag values ​​are within predefined ranges), and checks on logical relationships between tags (ensuring there are no logical conflicts between different tags). For tags that fail verification, the system generates warning messages and guides the data administrator to manually correct them.

[0113] Tag attachment is the process of actually applying generated tags to data. Three tag attachment modes are supported: metadata embedding (tags are directly embedded as metadata into data files or database records), reference linking (tags are stored in a separate tag library and linked to data through unique identifiers), and hybrid (core tag embedding with extended tag references). For different storage systems and data formats, the system automatically selects the most suitable tag attachment mode to ensure tight binding between tags and data and efficient access.

[0114] A key feature of the tagging system is its dynamic update mechanism. As data usage changes and security policies adjust, data tags need to be updated accordingly. Three tag update triggering mechanisms are implemented: rule-triggered updates (automatically updating tags for affected data when security rules change), time-triggered updates (periodically reassessing data classifications and updating tags), and event-triggered updates (revising tags when specific events such as data breaches or security audit findings occur). The tag update process maintains a complete change log, supporting tag version management and rollback.

[0115] By establishing this comprehensive and standardized data tagging system, we have achieved fine-grained classification and characteristic description of medical data, providing a solid foundation for subsequent tag-based access control. Tagged medical data is easier to manage and protect, and it is also easier to realize value mining and reasonable sharing while meeting security requirements.

[0116] Medical data classification labels are a fundamental component of this system for achieving refined access control. They describe data characteristics and access rules in a structured manner, supporting automated permission determination and access control. Through three-dimensional classification and standardized labels, the system can accurately identify the characteristics and protection requirements of each type of data, achieving refined management based on data characteristics, and significantly improving the efficiency and accuracy of medical data governance.

[0117] Example 4

[0118] In this embodiment, the construction of the user permission feature vector, based on the matching relationship between the medical data classification labels and the user permission feature vector, performs permission mapping to obtain the user permission feature vector and the permission mapping relationship, including:

[0119] The medical institution and user role in the multidimensional features are obtained, and the user's medical institution, professional role, scope of responsibility, qualification level and department level are extracted. The features of each dimension are quantified and encoded, and the feature fusion technology is used to obtain the user permission feature vector.

[0120] Based on the user permission feature vector and the medical data classification label, the matching degree between the user permission feature vector and the medical data classification label is calculated, and a permission mapping matrix between users and data is established to obtain the user permission feature vector and the permission mapping relationship.

[0121] Specifically, firstly, the system obtains user-related medical institution and user role information through the API interfaces of the hospital's human resources system and organizational structure management system. Distributed data integration technology is employed to synchronously acquire user attribute information from multiple data sources, ensuring data integrity and real-time performance. Data consistency verification is implemented during the acquisition process to resolve potential data redundancy and conflict issues, laying a reliable data foundation for subsequent processing.

[0122] For information about the medical institution to which a user belongs, the system not only obtains the specific name of the institution the user is currently attending, but also extracts multi-level attributes of the institution, including institution type (general hospital, specialized hospital, community health center, etc.), institution level (Grade III Class A, Grade II Class A, etc.), institution nature (public, private, non-profit, etc.), and the administrative region where the institution is located (provincial, municipal, district / county level). Furthermore, it identifies the affiliation and cooperative relationships between medical institutions, such as members of medical consortia or affiliated teaching hospitals. This multi-dimensional institutional information enables the system to understand the user's institutional background within the healthcare system, providing institutional-level judgment criteria for access control.

[0123] Extracting user professional role information involves multi-level role identification. First, the user's main functional categories are identified, such as basic role classifications like doctors, nurses, medical technicians, researchers, and administrators. For medical professionals, their department affiliation (e.g., internal medicine, surgery, radiology) and specific professional direction (e.g., cardiology, general surgery, CT diagnosis) are further extracted. Special role identifiers are also identified, such as department head, project leader, and ethics committee member; these special roles typically have specific data access permissions.

[0124] Scope of responsibility extraction is a crucial step in identifying which types of data a user should access in their work. By analyzing a user's job description, work assignment records, and actual business operation records, a scope of responsibility profile is constructed. The scope of responsibility is divided into three levels: core responsibilities (data types that must be accessed in daily work), extended responsibilities (data types that need to be accessed for specific projects or temporary tasks), and management responsibilities (data types that need to be accessed for management or supervision). Different weights are assigned to each type of responsibility, with core responsibilities having the highest priority in access control decisions.

[0125] Qualification level extraction is a crucial dimension for assessing a user's professional competence and credibility. The system collects information such as professional qualification certificates (e.g., practicing physician license, specialist physician certificate), professional title evaluation results (e.g., chief physician, associate chief physician), years of service, continuing education credits, and significant awards and honors. Based on this information, the system uses a comprehensive scoring model to calculate the user's qualification level score and categorizes users into four levels: primary (Level 1), intermediate (Level 2), advanced (Level 3), and expert (Level 4). The qualification level directly affects the user's available access permissions, particularly access to sensitive and core data.

[0126] Departmental hierarchy extraction is a crucial step in understanding a user's position within the organizational structure. It involves identifying the user's department's hierarchical position within the organization, including direct reports, superiors, and top management. This departmental hierarchy information, combined with the user's position within the department (e.g., ordinary member, team leader, department head), forms the user's organizational position characteristic. This characteristic is particularly important for determining whether a user has permission to access cross-departmental or organizational management data.

[0127] After extracting these multidimensional user features, quantitative encoding is performed to convert various textual and categorical features into numerical representations to facilitate subsequent feature fusion and model processing. Quantitative encoding employs multiple encoding strategies: hierarchical encoding is used for medical institutions, encoding features such as the institution's type, level, and nature into multi-level numerical codes; multi-hot encoding is used for user roles, converting multiple roles a user may simultaneously possess into binary feature vectors; responsibility scope is represented using weighted vectors, assigning different weights to different responsibility types; qualification levels are directly mapped to numerical representations of 1-4; and departmental levels use structured encoding to reflect the hierarchical relationships between departments.

[0128] Feature fusion technology is a set of methods that integrate features from multiple different dimensions into a unified representation. The feature fusion process employed is a three-stage process. The first stage is feature standardization, using Z-score standardization and Min-Max standardization to transform features with different dimensions and distribution ranges to the same scale, eliminating the impact of dimensional differences on subsequent processing. The second stage is feature selection and dimensionality reduction, using principal component analysis (PCA) and feature importance assessment methods to select the most discriminative and informative feature subsets, reducing data redundancy. The third stage is feature combination, using deep neural network models to learn the nonlinear relationships between features, generating a comprehensive feature representation with strong expressive power.

[0129] Through the above process, a unified format user permission feature vector is finally generated. This vector is typically a 256-dimensional floating-point array, which can compactly and comprehensively represent the user's characteristics in various dimensions of the permission system. The user permission feature vector is the basic data structure for subsequent permission mapping and determination. It is cached in a high-speed in-memory database, and a reasonable update strategy is set to ensure that the vector representation is updated in a timely manner when user characteristics change.

[0130] Secondly, a matching mechanism is established between user permission feature vectors and medical data classification labels to achieve accurate permission mapping. This process uses a two-layer mapping architecture, including a rule-based mapping layer and a model-based mapping layer.

[0131] The rule-based mapping layer implements basic permission determination, handling clear and universally applicable permission rules. The system has a built-in library of commonly used permission rules in the healthcare industry, containing approximately 500 core rules covering common user roles and data type combinations. Examples include "cardiologists can access the clinical data of cardiology patients" and "clinical researchers can access anonymized research data from ethically approved projects." Rules are stored using a decision tree structure, supporting fast querying and rule inheritance. Each rule includes a condition part (user characteristic conditions and data characteristic conditions) and a result part (allowed operation types and operation conditions). The rule-based mapping employs forward chain reasoning, evaluating applicable rules one by one based on user feature vectors and data labels to obtain preliminary permission determination results.

[0132] Matching degree calculation is the core step in the permission mapping process, used to quantify the degree of fit between user characteristics and data tags. The system adopts a multi-dimensional similarity calculation model, decomposing the matching degree into four key dimensions: role matching degree (the professional relevance between the user's role and the data type), institution matching degree (the relationship between the user's institution and the data source institution), responsibility matching degree (the degree of alignment between the user's responsibility scope and the data's purpose), and qualification matching degree (the suitability between the user's qualification level and the data sensitivity level). The system designs a dedicated similarity function for each dimension. For example, role matching uses a predefined professional domain association matrix, institution matching uses the shortest path distance of the institution relationship graph, responsibility matching uses a semantic similarity algorithm, and qualification matching uses direct level comparison. The matching degrees of each dimension are integrated into a comprehensive matching score through a weighted average method. The score ranges from 0 to 1, representing the user's suitability for accessing specific data.

[0133] A model-based mapping layer handles permission determination in complex scenarios, overcoming the limitations of rule-based systems. This layer uses machine learning models, particularly random forests and gradient boosting decision trees (GBDT) trained based on historical permission decisions, to predict permission levels for new user-data combinations. Model input includes user permission feature vectors, data classification labels, and contextual information (such as access purpose, business process stage, etc.), and outputs the probability of permission for various operations (view, modify, export, etc.). The model is retrained monthly using the latest permission approval records and access logs to ensure adaptation to constantly changing permission policies and organizational structures.

[0134] Based on the matching degree calculation and two-layer mapping architecture described above, a complete permission mapping matrix is ​​constructed. This matrix is ​​a multi-dimensional data structure that records the specific operation permissions of users on various types of data. Rows represent users (indexed by user ID or permission feature vector), columns represent data categories (indexed by data category labels), and matrix elements are permission descriptors, containing the allowed set of operations (view, modify, download, delete, etc.) and operation conditions (time limits, number of attempts, approval required, etc.). The permission mapping matrix adopts a sparse matrix storage format to save storage space and uses a block storage strategy to support efficient concurrent access and partial updates.

[0135] It also implements a permission inference mechanism to handle user-data combinations that are not explicitly defined in the matrix. When a user requests access to a data type that does not have a direct counterpart in the permission mapping matrix, the system infers unknown permissions from known permissions through similarity inference and the principle of permission inheritance. The inference process adopts the "principle of least privilege" to ensure that the inferred permissions do not exceed the scope of known permissions, and the inference process is recorded for subsequent review.

[0136] To ensure the security and compliance of permission mapping, a multi-layered verification mechanism was implemented. First, a compliance check ensures that the generated permission mappings conform to the organization's basic security policies and industry regulatory requirements. Second, conflict detection identifies and resolves potential permission conflicts; for example, when users obtain conflicting permissions through different roles, the system adopts the "strictest principle" to select the more stringent permission setting. Finally, an unauthorized access risk assessment analyzes the potential for excessive data exposure caused by permission combinations, marking and restricting high-risk combinations.

[0137] The permission mapping results are stored using a version control mechanism, recording the time, reason, and changes for each mapping update, supporting the tracing and rollback of permission changes. Simultaneously, a periodic review mechanism for permission mappings is established, regularly (usually quarterly) evaluating the overall permission distribution to identify and correct potential issues such as excessive permission granting or permission blind spots.

[0138] Through the detailed permission mapping process described above, a complete representation of the user permission feature vector and the permission mapping relationship is obtained. This mapping relationship forms the basic architecture for permission determination, enabling the system to quickly and accurately determine the user's permission level and permitted operation scope when the user requests access to medical data. Furthermore, this feature vector-based mapping method is highly scalable and can adapt to the ever-changing organizational structure and business needs of medical institutions.

[0139] Example 5

[0140] In this embodiment, the application of machine learning algorithms to construct a context-permission association model, calculate the permission credibility score in the current context, perform dynamic permission determination decisions, and obtain context-aware permission determination results, including:

[0141] Obtain the user permission feature vector and permission mapping relationship, collect user access time, access location, device type, network environment, extract context features and perform standardization processing to obtain the context feature vector;

[0142] Based on the context feature vector and historical access behavior data, machine learning algorithms are applied to establish the correlation between context factors and the rationality of permission use, calculate the permission credibility score in the current context, and obtain the context-adaptive permission score.

[0143] Based on the context-adaptive permission score and the preset permission judgment threshold, a dynamic permission judgment decision is made to generate a judgment result of granting, denying, or requiring secondary verification, thus obtaining the context-aware permission judgment result.

[0144] Specifically, firstly, user permission feature vectors and permission mapping relationships are obtained from the permission management center, serving as the foundational data for context-aware permission determination. The acquisition process employs a secure communication protocol to ensure transmission security, stores active user permission data in a local cache, and receives real-time updates via a message queue to ensure cache timeliness.

[0145] Contextual information collection covers multiple dimensions: access time is accurate to the second, marked as weekdays / rest days and time periods, with key monitoring of access during irregular periods; access location is obtained through multi-level positioning technologies such as IP resolution and GPS, and compared with the user's regular activity areas; device type is identified through technologies such as client fingerprinting, recording characteristics such as device model and system version, with medical-specific devices given higher credibility; network environment analysis includes connection type and security features, with access to highly sensitive data requiring a secure intranet or compliant VPN; simultaneously, extended contextual information such as medical business status and system load is collected to comprehensively assess the access background.

[0146] Contextual features undergo multi-stage engineering processing: decomposing composite features, generating derived features, and calculating interaction terms; short-term behavioral patterns are extracted using a sliding window approach. During standardization, outlier data is detected first, and then corresponding standardization methods are applied to different feature types, dynamically adjusting parameters to adapt to different institutional characteristics. This ultimately results in a 128-512 dimensional structured contextual feature vector, providing input for authorization credibility assessment.

[0147] Secondly, historical access behavior data is extracted from the access log database and used as training data after cleaning, deduplication, and balancing. The context-permission association model adopts a multi-model fusion architecture, including three basic models: a GBDT-based temporal pattern model to analyze access time patterns, a GMM-based spatial behavior model to identify location anomalies, and a random forest-based environmental security model to assess device and network security. The weights of each model are dynamically adjusted through an attention mechanism; access to highly sensitive data increases the weight of the environmental security model, while access at unusual times increases the weight of the temporal pattern model.

[0148] The model employs supervised learning training, utilizing cross-validation and hyperparameter optimization to improve performance, and addresses the label imbalance problem through comprehensive sampling and a weighted loss function. Deployment utilizes an online learning architecture, updating parameters weekly through incremental learning, and triggering retraining when model accuracy falls below a threshold. Permission credibility scores range from 0 to 1, comprehensively considering consistency with the current context and conventional patterns (40%), internal context consistency (25%), and behavioral similarity with historically similar contexts (35%), and decomposed into risk scores for each dimension, generating an interpretable report. Context-adaptive permission scoring incorporates data sensitivity and other contextual factors to ensure accurate and flexible judgment.

[0149] Finally, a 12-cell threshold matrix is ​​constructed based on data sensitivity (level four) and operational risk (level three). For example, the low-risk operation threshold for public-level data is 0.5, and the high-risk operation threshold for confidential data is 0.85. The thresholds support context-aware dynamic adjustment; they can be temporarily lowered in medical emergencies and raised when the system detects an abnormal attack. The entire adjustment process is recorded.

[0150] Access control employs a three-tiered mechanism: Adaptability scores above the high threshold (threshold + 0.1) grant direct permission; scores below the low threshold (threshold - 0.1) reject direct permission; and scores in the "grey area" require secondary verification. Secondary verification dynamically selects the method based on risk level, including CAPTCHA and biometric authentication. Successful results can be cached briefly, while caching is disabled for access to highly sensitive data.

[0151] The system generates a corresponding response based on the judgment result. If permission is granted, an access token is generated; if permission is denied, the reason is explained; if secondary verification is required, the system guides the user through the process. Detailed logs are recorded to support audit traceability. Results are communicated to users and administrators through multiple channels, with additional notification channels added for high-risk operations or anomaly detections.

[0152] After authorization is granted, user actions are continuously monitored. Abnormal patterns can trigger a reassessment or revoke authorization. Through a closed-loop learning mechanism, the system records the assessment results and feedback data, regularly optimizing algorithms and threshold settings to continuously adapt to changing user behavior and organizational policies. This ensures the security of medical data while minimizing interference with legitimate medical work.

[0153] Example 6

[0154] like Figure 2 As shown, the process involves feature extraction and vectorization of medical data, the application of a multi-pattern matching algorithm to construct a distributed index structure, and the extraction of user permission mapping relationships from the multi-dimensional user identity and permission model to filter and sort the search results, resulting in permission-matched medical data retrieval results, including:

[0155] Step S21: Obtain structured and unstructured medical data, apply natural language processing and image recognition technology to extract key features of the medical data, convert them into standardized feature vectors, and obtain medical data feature vectors.

[0156] Step S22: Based on the medical data feature vector, apply a multi-pattern matching algorithm to perform data indexing to obtain a multi-pattern matching index;

[0157] Step S23: Based on the multi-pattern matching index, construct a distributed index structure containing a primary index, secondary indexes, and temporary indexes; establish an index routing layer; determine the index path based on query characteristics, user roles, and system load; and obtain the distributed index system.

[0158] Step S24: Combining the distributed indexing system and the multi-dimensional user identity and permission model, extract the user permission mapping relationship in the multi-dimensional user identity and permission model, dynamically filter and de-identify the search results according to the user permission mapping relationship, and perform personalized sorting based on user preferences and historical behavior to obtain the medical data search results with the matching permissions.

[0159] Specifically, firstly, structured and unstructured medical data from healthcare institutions are acquired through multi-source data integration interfaces. Structured medical data mainly comes from Hospital Information Systems (HIS), Electronic Medical Record Systems (EMR), and Laboratory Information Systems (LIS), including patient basic information, diagnostic records, prescription information, test results, and other data with clearly defined fields and formats. Unstructured medical data comes from Picture Archiving and Communication Systems (PACS), pathology image systems, scanned copies of doctors' handwritten medical records, and surgical videos, etc. This type of data usually does not have a predefined structure and requires special techniques for processing and feature extraction.

[0160] For structured medical data, the first steps are data cleaning and standardization. Data cleaning includes handling missing values, identifying outliers, and resolving data consistency issues. Standardization transforms heterogeneous data from different sources into a unified data model, using standard medical terminology sets (such as ICD-10 and SNOMED CT) to map diagnostic names, drug names, etc., ensuring semantic consistency. Specific feature extraction strategies are applied to different types of structured data: statistical feature extraction and trend analysis are used for numerical data (such as blood pressure and blood glucose indicators); label coding and semantic association analysis are used for categorical data (such as disease categories and drug categories); and time series feature extraction is used for time-series data (such as continuous monitoring indicators) to capture the temporal dimension of the data.

[0161] For unstructured medical data, domain-specific deep learning models are used for feature extraction. Medical text data (such as doctor's records and medical history descriptions) are processed using pre-trained language models in the medical field (such as BioBERT and ClinicalBERT). These models have been trained on large-scale medical literature and clinical texts and are capable of understanding medical terminology and contextual semantics. A multi-level text analysis workflow is implemented: first, text preprocessing is performed, including word segmentation, stop word removal, and professional terminology recognition; then, semantic features are extracted, including key entity recognition (such as disease names, symptom descriptions, and drug names), relationship extraction (such as the relationship between symptoms and diseases, and the relationship between drugs and treatment effects), and semantic classification; finally, document-level features are extracted through an attention mechanism to capture overall semantic information.

[0162] Medical image data (such as CT, MRI, and X-ray images) is processed using convolutional neural networks specifically designed for medical imaging. The system employs transfer learning, using deep learning models (such as DenseNet and ResNet) pre-trained on the medical image dataset to extract image features. The feature extraction process involves multi-level visual feature representation: low-level features capture basic visual elements such as edges and textures; mid-level features represent intermediate semantics such as organ structure and lesion morphology; and high-level features correspond to advanced semantic concepts such as disease patterns and abnormal regions. The system also applies a region proposal network to automatically identify key regions in medical images and extract detailed features from these regions.

[0163] For multimodal medical data (such as examination results that simultaneously contain text reports and images), a cross-modal feature fusion technique was implemented. This technique first extracts features independently from each modality of data, and then integrates the features from different modalities into a unified representation through an attention fusion mechanism and a multimodal encoder. This fused representation can simultaneously capture complementary information from textual descriptions and visual features, forming a more comprehensive medical data representation.

[0164] After feature extraction, all features are vectorized into a standardized feature vector format. The vectorization process employs dimensionality unification technology, mapping features from different sources and types to a vector space of the same dimension (typically 256 or 512 dimensions). Feature normalization methods are also applied to ensure the comparability of different features across dimensions. The final generated medical data feature vectors possess three key characteristics: semantic preservation (semantically similar medical data are close in vector space), discriminative power (different types or content of medical data show significant vector differences), and retrieval friendliness (supporting efficient vector similarity calculation).

[0165] Medical data feature vectors serve as the system's foundational data representation, supporting not only subsequent index building and retrieval but also facilitating advanced applications such as cluster analysis, similar case recommendation, and knowledge graph construction. A feature vector update mechanism is established, automatically triggering feature re-extraction and vector updates when the original medical data is updated, ensuring consistency between the feature vectors and the source data.

[0166] Second, based on the medical data feature vectors generated in the previous step, an improved multi-pattern matching algorithm is applied to construct the data index. Multi-pattern matching algorithms are a class of algorithms capable of efficiently matching multiple patterns (query conditions) simultaneously. They are suitable for common complex query scenarios in medical data retrieval, such as "finding patient cases that simultaneously meet specific symptoms, medication records, and test results."

[0167] The multi-pattern matching algorithm employed is based on the traditional Aho-Corasick automata and suffix tree algorithm, but with specific optimizations for the medical scenario. First, feature decomposition is performed on the medical data feature vectors, breaking down the high-dimensional feature vectors into multiple semantic sub-features, such as disease features, symptom features, and treatment features, facilitating the construction of specialized index structures for different dimensions. Then, for the decomposed feature subsets, feature hashing technology is applied to map the continuous feature value space to discrete hash buckets, reducing index space complexity while maintaining the locality of similar features.

[0168] The core of the multi-pattern matching index is a compact data structure built on feature vectors, enabling efficient similarity queries and composite condition matching. It employs a multi-level index architecture: the first level is a feature classification index, grouping medical data according to its main categories (e.g., disease categories, department classifications); the second level is a semantic feature index, building a similarity index based on feature vectors for data within each category; and the third level is a keyword index, establishing direct access indexes for medical terminology and key entities. This multi-level index structure can quickly locate the most relevant subset of data based on the specific query content, avoiding a full scan.

[0169] To support fuzzy and semantic matching requirements, vector quantization is applied during index construction. This technique divides the continuous feature vector space into a finite codebook, mapping each feature vector to the nearest codeword, achieving a discrete representation of features. This significantly reduces index storage space and query computation complexity while maintaining acceptable matching accuracy. A hierarchical quantization strategy is employed, using fine-grained quantization for important feature dimensions and coarse-grained quantization for secondary dimensions, balancing accuracy and efficiency.

[0170] The time-sensitive nature of medical data was also given special consideration in index construction. A time-aware indexing strategy was implemented, creating independent index structures for data from different time windows. The most recent data (e.g., data generated in the past week) used a high-precision full index; older data (e.g., data from the past month to a year) used a medium-precision compressed index; and historical data (e.g., data from a year ago) used a low-precision archived index. This strategy reflects the common time preference patterns in medical data retrieval, where users typically focus more on recent data, while also optimizing system resource utilization.

[0171] Index updates employ an incremental update mechanism to avoid the performance overhead of a full rebuild. When new medical data is added to the system, only the affected index portions are updated, typically the latest data index and related category indexes. A background index optimization program is implemented, performing index defragmentation and structural optimization when system load is low (e.g., at night) to ensure that index performance does not significantly degrade over long-term operation.

[0172] The construction process of the multi-pattern matching index is optimized for the sparsity and high dimensionality unique to medical data. Medical feature vectors typically have high dimensionality, but most dimensions have values ​​of zero or close to zero. Leveraging this characteristic, a sparse matrix storage technique is employed to record only non-zero feature dimensions, significantly reducing index storage space. Simultaneously, a dimensional importance assessment is implemented, assigning higher weights to dimensions with strong retrieval discriminative power and prioritizing the construction of refined indexes for these dimensions.

[0173] Through the above optimization strategies, a highly efficient multi-pattern matching index was finally obtained, which can support multi-condition compound queries and similar case retrieval of medical data, laying the foundation for the subsequent construction of a distributed index system.

[0174] Third, based on the multi-pattern matching index built in the previous step, a hierarchical distributed index structure was further established to adapt to the characteristics of large-scale medical data and complex access patterns. The distributed index system includes a three-layer index structure: primary index, secondary index, and temporary index, with each layer optimized for different data access needs.

[0175] The primary index is the system's core permanent index, covering all medical data and ensuring data integrity and accessibility. The system constructs the primary index based on key characteristics of the medical data (such as disease type, treatment methods, and drug characteristics), and employs a sharding strategy to distribute the index across a distributed database cluster. Index sharding uses a consistent hashing algorithm to ensure even data distribution while minimizing data migration during node additions and deletions. To improve data locality, a relevance-based sharding strategy is used, allocating semantically related or frequently accessed data to the same or adjacent shards. The primary index uses an incremental update mechanism, supporting real-time updates to index content without affecting query services.

[0176] Secondary indexes are specialized indexes designed for specific business scenarios, built for frequently accessed subsets of data or specific analytical needs. Based on the business characteristics and user access patterns of medical institutions, the system predefines a series of secondary index templates, such as department-specific indexes (containing only cases related to a specific department), disease-specific indexes (e.g., cardiovascular disease indexes), and research-specific indexes (e.g., clinical trial data indexes). Secondary indexes are typically deployed on edge nodes or near business application servers to reduce network transmission overhead and improve response speed. The selection and optimization of secondary indexes is a dynamic process; by analyzing query logs and access patterns, the efficiency of existing secondary indexes is regularly evaluated, and index configurations are adjusted to adapt to changing business needs.

[0177] Temporary indexes are short-term indexes dynamically created by the system for frequently accessed queries or temporary task requirements. The system monitors query traffic patterns, identifies similar queries that are frequently executed within a short period, and automatically builds dedicated in-memory indexes for these hot queries, significantly improving query response speed. Temporary indexes employ an automatic expiration mechanism based on time and access frequency; when query activity decreases or a predetermined time window is exceeded, the temporary index is automatically cleared, releasing resources. Temporary indexes are primarily designed for sudden data access needs, such as frequent queries for cases with specific symptoms during a sudden epidemic or concentrated access to relevant research data during medical conferences.

[0178] The index routing layer is a crucial component connecting user queries with the underlying index structure. It is responsible for determining the optimal index access path based on query characteristics, user roles, and system load. The index routing process first analyzes the query content, extracting key features and constraints; then it matches user roles and permissions to determine the range of data accessible to the user; finally, it considers the current system load and index status to select the most suitable index combination to execute the query. Routing decisions employ a combination of rule-based and machine learning approaches: predefined rules handle explicit routing scenarios, such as directly routing specific types of queries to their corresponding dedicated indexes; machine learning models handle complex decision-making scenarios, such as selecting the optimal combination from multiple available indexes, or making trade-offs between query decomposition and direct execution.

[0179] A dynamic index optimization strategy was implemented to continuously improve the index structure based on actual operational data. The optimization strategy included index coverage analysis (assessing the index's support for queries), index usage monitoring (identifying inefficient or redundant indexes), and access pattern mining (discovering potential index optimization opportunities). Based on these analyses, index optimization suggestions were automatically generated, including creating new indexes, adjusting existing index structures, or merging redundant indexes.

[0180] To handle complex queries across indexes, a distributed query execution engine was implemented. This engine can decompose complex queries into multiple subqueries, execute them in parallel on different indexes, and then merge the result sets. Query decomposition and execution plan generation employ a cost-based optimization method, considering data distribution, network topology, and index characteristics to generate the optimal execution path. A query result caching mechanism was also implemented to cache the results of frequently executed similar queries, further improving response speed.

[0181] By constructing this multi-layered distributed indexing system, the basic requirement of comprehensive indexing of medical data is met, while optimized access paths are provided for different application scenarios, achieving a balance between retrieval performance and system resource utilization.

[0182] Fourth, the distributed indexing system is combined with a multi-dimensional user identity and permission model to achieve a permission-aware medical data retrieval process. First, user permission mapping relationships are extracted from the permission management center. This is a structured representation describing the range of data a user can access and the types of operations they can perform. The permission mapping relationship includes the data types a user can access, access levels (e.g., read-only, modify, download), and access conditions (e.g., time restrictions, location restrictions). The system converts the permission mapping relationship into query filtering conditions, serving as security constraints for the retrieval process.

[0183] At the start of the retrieval process, the user's input query undergoes semantic parsing and intent recognition. Semantic parsing transforms the natural language query into a structured query representation, identifying medical entities, relationships, and constraints. Intent recognition determines the type of query purpose, such as finding similar cases, statistically analyzing specific indicators, or searching for treatment plans. Based on the parsing results and the user's permission mapping, a permission-aware query plan is generated. This plan includes two sets of constraints: business constraints (from the user's original query) and security constraints (from the permission mapping).

[0184] During the query execution phase, the optimal index path is determined through the index routing layer, and medical data that meets the criteria is retrieved from the distributed index system. The query results first undergo an access control filtering process, removing data items that the user is not authorized to access based on user permission mappings. This filtering employs a multi-level strategy: first, coarse-grained filtering quickly excludes results that clearly exceed the user's permission scope based on macro-level characteristics such as data category and department affiliation; then, fine-grained filtering precisely determines permissions based on specific data tags and access control rules.

[0185] For data that passes access control filtering, further security de-identification processing is implemented. Medical data de-identification is the process of removing or replacing sensitive information while preserving the data's analytical value. Appropriate de-identification strategies are automatically determined based on user access levels and data sensitivity. De-identification techniques include direct identifier removal (e.g., deleting names and ID numbers), indirect identifier generalization (e.g., replacing precise age with age ranges), data masking (e.g., hiding some sensitive diagnostic information), and data substitution (e.g., replacing original data with similar but not authentic values). For image data, region blurring technology is also applied to automatically identify and blur areas that may contain patient identification information (e.g., name markers on images). The de-identification process employs a reversible encryption mechanism, allowing authorized users to access the original data through an approval process when necessary.

[0186] Results ranking is a crucial aspect of improving the search experience. A multi-factor weighted ranking model is employed, comprehensively considering factors such as query relevance, data freshness, data completeness, and user preferences. Query relevance is calculated based on feature vector similarity, measuring the semantic match between search results and the user's query. Data freshness considers the data's creation time, with newer data typically receiving higher ranking weight. Data completeness assesses the fillability and quality of data fields, with more complete results tending to rank higher. User preferences are based on historical interaction data, analyzing users' professional interests and habits, such as preferences for specific disease data or particular formats of test results.

[0187] Personalized sorting is a key feature of the system, capable of adjusting sorting strategies based on users' professional background, historical behavior, and current tasks. It constructs a user interest model, recording users' query patterns, click behaviors, and data usage to identify users' professional preferences and areas of interest. Personalized sorting employs a hybrid strategy: on the one hand, it ensures the objective relevance of results, avoiding the "information cocoon" effect caused by excessive personalization; on the other hand, it appropriately adjusts sorting weights, displaying results that are more likely to align with user interests earlier. It also supports explicit sorting preference settings, allowing users to specify particular attributes (such as time order, relevance, etc.) as the primary sorting criteria.

[0188] Finally, the results after filtering, anonymization, and sorting are integrated to form a set of medical data retrieval results with matching permissions. The result set is returned in a structured format, containing basic data content, anonymization tags (indicating which fields have been anonymized), and permission information (such as the types of operations that can be performed on the data). A complete retrieval process log is also recorded, including query content, execution path, permission determination, and result processing, supporting subsequent security audits and retrieval optimization analysis.

[0189] Example 7

[0190] In this embodiment, the application of a multi-pattern matching algorithm for data indexing to obtain a multi-pattern matching index includes:

[0191] Based on the medical data feature vectors and query complexity, a sliding window with dynamically adjustable size is maintained. A window suffix tree is constructed only for the medical data feature vectors within the sliding window to obtain a dynamic window suffix tree index.

[0192] The medical data feature vector is divided into multiple data shards, and a suffix tree index is constructed in parallel for each data shard. A work-stealing load balancing strategy is adopted to obtain a set of parallel suffix tree indexes.

[0193] Based on the dynamic window suffix tree index and the parallel suffix tree index set, a two-layer index structure is constructed using the Aho-Corasick automaton. The outer layer uses the parallel suffix tree index set to index the dataset, while the inner layer uses the Aho-Corasick automaton to simultaneously match multiple patterns, resulting in the multi-pattern matching index.

[0194] Specifically, firstly, an adaptive index building strategy is designed by combining the distribution characteristics of medical data feature vectors and the complexity of query patterns. The core of this strategy is a dynamic sliding window mechanism, which solves the problems of high memory consumption and low tree building efficiency when the traditional suffix tree algorithm processes large-scale medical data.

[0195] The sliding window mechanism leverages the locality of access in data access, building a suffix tree index only on data within the window. The window boundaries slide with the query focus, ensuring that active data remains within the window. The window size dynamically adjusts between 4KB and 64KB (corresponding to 1000-15000 feature vectors), based on query complexity, system resource status, and data access patterns, employing a gradual strategy to avoid drastic fluctuations. Sliding trigger modes include query-driven, timed sliding, and predictive sliding. The sliding process is optimized using incremental indexing techniques, with new data being indexed and removed data saved as compressed versions.

[0196] The dynamic window suffix tree is specifically designed for high-dimensional feature vectors. It reduces the tree branching factor through feature dimensionality reduction and clustering techniques, and optimizes memory usage by employing compact storage and lazy loading mechanisms for nodes. During construction, medical feature vectors are first quantized and discretized into codeword sequences. Key dimensions are finely quantized to ensure discriminability. It supports exact matching, range matching, and similarity matching. Similarity matching integrates approximate nearest neighbor search and pruning strategies to achieve efficient retrieval.

[0197] Secondly, a data sharding and parallel processing strategy is implemented to improve the efficiency of suffix tree index construction. Data sharding is first performed by natural classification such as department and disease type, and then further subdivided by feature similarity to ensure that the data in each shard is balanced and semantically complete. The sharding granularity is dynamically adjusted to 10,000-50,000 feature vectors. Shard boundaries are determined through hierarchical clustering. For complex cases, a data replication strategy is used to ensure retrieval integrity, while maintaining shard metadata summaries to facilitate subsequent routing.

[0198] Parallel building employs a task decomposition model, creating independent tasks for each shard. Dynamic worker pool scheduling, combined with a work-stealing load balancing strategy, utilizes double-ended queues to reduce thread contention and prioritizes tasks based on workload and complexity. Task dependencies are handled, and dynamic task splitting is supported to address load imbalances. The build process employs segmented commits and optimistic concurrency control, with checkpointing mechanisms for fault recovery. Lightweight directories are built through index merging, supporting both full merge and federated modes. Large-scale deployments prioritize the federated mode for enhanced flexibility.

[0199] Finally, based on the aforementioned index foundation, a two-tiered index structure integrating suffix trees and Aho-Corasick automata is designed to adapt to the complex query requirements of medical data with multiple conditions. The outer index is a set of parallel suffix tree indexes, responsible for fast location and range reduction. A skip list structure is introduced to enhance range query capabilities, and heuristic pruning is applied to reduce invalid computations. The outer index adopts a hierarchical design, containing a global directory, shard-level indexes, and leaf node data pointers. During queries, relevant shards are retrieved in parallel, and the results are merged.

[0200] The inner index is an improved Aho-Corasick automaton, supporting feature vector sequence matching, fuzzy matching, and dynamic pattern loading. Memory is optimized through state compression and sparse matrix storage. The two-layer index works collaboratively: first, the query is parsed into a set of feature vectors and patterns; the outer layer filters a subset of data; and the inner layer performs parallel matching of multiple patterns. Query optimization strategies such as condition reordering and result caching are implemented, supporting incremental result return. The update mechanism is designed with separation: the outer index is updated incrementally, while the inner index updates according to changes in the query pattern, maintaining consistency through version control. This design improves query performance by 8-12 times, supports diverse retrieval scenarios, and meets the data access needs of medical professionals.

[0201] Example 8

[0202] In this embodiment, the construction of a distributed index structure including a primary index, secondary indexes, and temporary indexes, the establishment of an index routing layer, and the determination of index paths based on query characteristics, user roles, and system load to obtain a distributed index system include:

[0203] Based on the multi-pattern matching index, a master index is constructed for the disease type, diagnosis and treatment methods and drug treatment characteristics of medical data. The master index is stored in a distributed database cluster using a consistent hashing sharding strategy to obtain a global master index.

[0204] Based on the multi-pattern matching index and business scenario requirements, secondary indexes are constructed for data subsets of predetermined departments, predetermined disease areas, or predetermined research directions, and deployed on edge nodes to obtain scenario-based secondary indexes.

[0205] Based on the needs of hot queries or temporary tasks, a short-term index is built using in-memory data structures and dynamic window features, and an automatic expiration and clearing mechanism is set to obtain a temporary index.

[0206] By integrating the global primary index, the scenario-based secondary index, and the temporary index, an index routing layer is established. Based on query features, user roles, and system load, a machine learning model is applied to evaluate the index path, resulting in the distributed index system.

[0207] Specifically, firstly, based on the multi-pattern matching index built in the previous stage, a global master index covering all medical data will be further constructed. As the permanent core index of the system, the master index is the infrastructure that ensures the integrity and accessibility of data, supports comprehensive retrieval of medical data, and is mainly constructed around three core features: disease type, diagnosis and treatment methods, and drug treatment. These features are commonly used query dimensions for medical retrieval and the main organizational methods of the medical knowledge system.

[0208] The disease type index constructs a multi-level classification system based on the International Classification of Diseases (ICD-10 or ICD-11), extracting primary and secondary diagnostic information from medical data. It standardizes free text diagnoses into standard disease codes using medical entity recognition technology. For complex multiple diagnoses, disease correlation analysis is applied to identify primary and secondary relationships and comorbidity patterns, constructing a disease association graph. The index structure adopts a multi-level tree design, with the upper level representing broad disease categories and the lower levels progressively refining to specific disease entities. It also integrates mapping relationships between different disease classification systems such as ICD and SNOMED CT to ensure the accuracy of cross-terminology queries.

[0209] The index of diagnostic and treatment methods is built around medical interventions such as surgical procedures and examinations. It uses a medical service operation coding system (such as ICD-9-CM-3 or ICD-10-PCS) as its basic classification framework to standardize the operation descriptions in medical records. For diagnostic and treatment descriptions with lower levels of standardization, natural language processing technology is used to parse the text, identify intent, and extract key diagnostic and treatment information. The index adopts a two-layer structure: the outer layer is organized by intervention type, and the inner layer is subdivided by anatomical location and treatment purpose, forming a detailed classification tree. It also pays attention to the temporal relationships and combination patterns of diagnostic and treatment methods, supporting complex queries based on treatment processes.

[0210] The drug therapy feature index constructs a drug classification system based on drug ontology (such as the ATC classification system), standardizing information such as drug names and dosages in medical records into structured data. The index structure comprehensively considers the chemical classification, therapeutic uses, and mechanisms of action of drugs, supporting multi-dimensional drug information retrieval, with particular attention to drug combination patterns. It constructs drug association networks by analyzing co-use patterns, supporting queries based on drug interactions. For new drugs and non-standard drug names, a drug name standardization module maps various expressions to a unified identifier, ensuring search consistency.

[0211] The three core feature indexes are connected through a feature association mechanism to form an integrated main index structure. The system establishes a three-dimensional association model of disease-treatment-drug, recording the clinical correlations between different features to support query expansion and result association. The main index adopts a distributed storage architecture, implementing sharded storage based on a consistent hashing algorithm, and introducing a virtual node mechanism to improve the uniformity of data distribution. The sharding strategy also includes affinity sharding enhancement, allocating data that is frequently queried together to nearby shards, while employing differentiated sharding methods for different types of medical data. Small, high-frequency data is replicated across all nodes, while large, low-frequency data is distributed across multiple shards. The main index updates use a two-phase commit protocol, supporting configurable consistency levels to balance write consistency and system performance.

[0212] Second, based on the business needs of medical institutions, a series of dedicated secondary indexes are constructed. These indexes have more targeted coverage and a more professional structure, aiming to provide an optimized search experience for specific scenarios. The construction of secondary indexes is based on business scenario analysis, identifying high-value application scenarios and data access patterns by analyzing user query logs, business process documents, and expert interviews.

[0213] Department-specific indexes provide customized independent index structures for each major department (such as cardiology and neurosurgery), focusing on common disease profiles, treatment protocols, and medication usage patterns. Index fields and weights are tailored to the department's specific professional concerns. For example, the cardiology index enhances the index depth of cardiovascular indicators and interventional treatment records, while the neurosurgery index optimizes the organization of neurological imaging and surgical records. Furthermore, the time dimension is optimized based on the department's workflow characteristics; outpatient departments strengthen the chronological association of patients' historical visit records, while emergency departments prioritize rapid access to the latest data. The indexes are also enhanced through departmental knowledge graphs, supporting intelligent query understanding and expansion.

[0214] Disease domain indexes are built around specific diseases or disease groups, serving scenarios such as specialized research and outpatient clinics. High-value disease domains such as diabetes and oncology are predefined and dedicated indexes are constructed. These indexes integrate data from the entire disease management process, covering information at each stage, including screening, diagnosis, and treatment. The index structure reflects the natural stages of disease progression, facilitating longitudinal queries of a patient's disease development. They also integrate special fields such as biomarkers and risk factors to enhance the analytical capabilities for specialized disease research. For chronic diseases, time-series data organization is strengthened, and for complex diseases, multidisciplinary diagnostic and treatment information is integrated to construct comprehensive view indexes.

[0215] The research focus index is designed for fields such as clinical research and medical education, including specialized indexes such as clinical trial indexes and rare disease research indexes. These indexes have strict requirements for data quality and completeness. During the index construction phase, high-quality records that meet research standards are rigorously screened. The index structure takes into account the specific characteristics of research data, supports precise retrieval based on research methodologies, and integrates metadata such as ethical approval status and research stage to facilitate research management and compliance verification.

[0216] Secondary indexes are deployed at the network edge, close to users, using an edge computing architecture to reduce network latency and improve query response speed. A central-edge collaborative management strategy is employed, with the central node responsible for metadata maintenance and update strategy formulation, and edge nodes responsible for local index storage and query services. Data updates use a targeted push model, pushing necessary updates only to affected secondary indexes. Secondary indexes are loosely coupled to the primary index and can evolve independently. The system implements secondary index lifecycle management, regularly evaluating index effectiveness, optimizing or eliminating inefficient indexes, and ensuring efficient utilization of system resources.

[0217] Third, implement a dynamic response mechanism to optimize the handling of short-term hot query patterns and temporary task requirements through temporary indexes. Temporary indexes are dedicated index structures dynamically created for specific short-term access patterns, with a lifecycle matching the duration of the access pattern, aiming to improve the response performance of frequently similar queries within a short period.

[0218] Hotspot query identification is the trigger mechanism for temporary index creation. The system deploys a real-time query monitoring module to analyze query traffic from three dimensions: query pattern similarity, query frequency, and query source distribution. When a certain type of query pattern exceeds a preset threshold within 10-30 minutes, a hotspot event is triggered. A multi-level threshold strategy is adopted: moderate hotspots trigger query result caching, while high hotspots trigger temporary index creation.

[0219] The hot query analysis module delves into the selectivity, complexity, and result set characteristics of hot queries, automatically determining the most suitable temporary index type and structure. High-selectivity simple queries correspond to lightweight hash indexes, range queries and fuzzy matching correspond to tree index structures, and complex multi-condition combined queries correspond to composite indexes or bitmap indexes, optimizing performance for multi-condition evaluation.

[0220] Temporary task requirement identification targets periodic tasks such as clinical data quality review and infectious disease surveillance. It receives task forecast information through the task management interface to understand the characteristics of data requirements. For foreseeable temporary tasks, a pre-creation strategy is adopted to build necessary temporary indexes before the task starts, ensuring optimized data access performance when the task starts.

[0221] Temporary indexes employ a memory-first data structure design, selecting appropriate structures such as hash tables and skip lists based on index size and available resources. Large temporary indexes exceeding single-machine memory capacity utilize a distributed in-memory computing framework, storing them in shards across multiple nodes. In-memory indexes employ a delayed write strategy, batch caching update operations, and periodically flushing to persistent storage to reduce the impact of I / O on query performance.

[0222] Dynamic windows are a core feature of temporary indexes. The window boundaries are determined by the time window, spatial window, and access window. The system continuously monitors changes in query patterns and adjusts the window position and size in real time to ensure that high-frequency access data is always within the window, and a smooth transition strategy is adopted to avoid performance fluctuations.

[0223] An automatic expiration and cleanup mechanism ensures efficient use of system resources. The expiration strategy is based on multiple dimensions of indicators, including time, activity, utility assessment, and resource pressure, to assign an initial lifespan to each temporary index and dynamically adjust it. When an index is close to expiration, its value is assessed. If the usage rate is high, the lifespan is extended; if the usage rate is low, the cleanup process is triggered. The cleanup process adopts a tiered degradation strategy, first migrating the index to disk and then completely deleting it. At the same time, high-value temporary index patterns are recorded to facilitate rapid reconstruction in the future.

[0224] Temporary indexes play a crucial role in medical emergencies. Emergency response modes are preset for sudden events such as major accidents and infectious disease outbreaks, automatically triggering the creation of corresponding temporary indexes to support needs such as epidemic monitoring and case tracing. Temporary indexes in emergency mode enjoy higher resource priority, ensuring unrestricted query response during critical moments.

[0225] Fourth, the global primary index, scenario-based secondary indexes, and temporary indexes are integrated into a unified distributed index system, and an intelligent index routing layer is constructed to find the optimal execution path for different queries. As the nerve center of the distributed index system, the index routing layer is responsible for receiving query requests, analyzing features, selecting index combinations, coordinating query execution, and returning result sets.

[0226] The index metadata repository is the infrastructure of the routing layer. It records the coverage, structural characteristics, performance features, and physical deployment information of all indexes, and implements hierarchical metadata management. The global layer stores summary information, while the local layer stores detailed structural descriptions. Metadata updates adopt a publish-subscribe model to ensure that routing decisions are based on the latest state.

[0227] Query parsing and feature extraction are the first steps in routing. The user query is subjected to syntactic analysis, semantic analysis and feature extraction to form a query feature vector containing content, structure, constraints and priority features. Based on this vector, the set of applicable indexes is initially selected.

[0228] User role analysis is an important basis for routing decisions. The system obtains role information such as user function category, professional direction and usage scenario, maintains role-index affinity matrix, and combines user historical query patterns to identify personal preferences and habits, providing a reference for routing decisions.

[0229] System load monitoring provides a resource constraint perspective, the routing layer maintains a real-time system health status map, monitors indicators such as computing load and memory usage of each index node, considers load balancing when making routing decisions to avoid assigning queries to high-load nodes, and reserves computing resources for high-priority queries to ensure timely response to critical tasks.

[0230] The routing decision core adopts a machine learning model based on a deep reinforcement learning framework. It models the routing problem as a sequential decision-making process and learns the optimal routing strategy by observing the actual query execution results, thus adapting to complex and ever-changing medical data access patterns.

[0231] The routing execution strategy includes three core modes: direct routing, decomposed routing, and hierarchical routing. The system selects the appropriate mode based on the query characteristics and available indexes. Complex queries will generate detailed execution plans that specify the order of subquery distribution, intermediate result processing, and final merging strategy.

[0232] Result aggregation and consistency assurance are the final stages of routing execution. When merging result sets from multiple sources, data redundancy, freshness, and integrity are considered, and a result verification mechanism is implemented to ensure the accuracy of distributed execution results. For critical medical data queries, transactional consistency guarantees are provided to avoid reading inconsistent data.

[0233] Adaptive optimization is the continuous evolution mechanism of the distributed index system. The system records the execution status of each query, periodically evaluates the effectiveness of the routing strategy, and optimizes performance through measures such as retraining the routing model, adjusting the index structure, and reallocating resources. It also supports A / B testing of routing strategies and selects the optimal solution for promotion.

[0234] Fault tolerance is an essential characteristic of distributed systems. The index routing layer implements a multi-layered fault tolerance mechanism at the node, path, and degradation service levels. When a node fails, it automatically routes to a backup node; when a path execution fails, it quickly switches to an alternative path; and when some indexes are unavailable, it provides limited functional degradation services to ensure that basic query capabilities are not interrupted. Through multi-layered and intelligent index routing design, efficient management and utilization of distributed index resources are achieved, providing users with a unified, high-performance query interface while maintaining the system's scalability and robustness to adapt to the growth in the scale of medical data and changes in access patterns.

[0235] Example 9

[0236] In this embodiment, the application graph algorithm performs real-time evaluation of the user-permission-data relationship graph, detects abnormal permission usage patterns, generates permission adjustment suggestions, and executes a dynamic adjustment strategy to obtain the dynamic evaluation and adjustment results of permissions, including:

[0237] Obtain the medical data retrieval results matching the permissions, collect the user's access behavior data on the medical data retrieval results matching the permissions in real time, construct the user-permission-data relationship graph containing user nodes, permission nodes and data nodes, and obtain the permission usage behavior relationship graph;

[0238] Based on the permission usage behavior relationship graph, a graph algorithm is applied to calculate node weights and path weights, and the permission usage behavior relationship graph is evaluated in real time to obtain permission usage evaluation results.

[0239] Based on the permission usage evaluation results and the permission usage behavior relationship diagram, a normal permission usage pattern feature library is constructed. The deviation between the current user's permission usage pattern and the normal permission usage pattern feature library is calculated. The degree of medical urgency, business stage characteristics, and organizational policy changes are introduced as contextual factors to dynamically adjust the anomaly judgment threshold. The deviation is compared with the anomaly judgment threshold, and the anomaly is classified according to the comparison results to obtain the permission usage anomaly detection results.

[0240] Based on the abnormal permission usage detection results, the permission usage evaluation results are integrated to generate permission adjustment suggestions. Based on the permission adjustment suggestions, an approval process is automatically generated and the permission adjustment is executed to obtain the dynamic evaluation and adjustment results of the permissions.

[0241] Specifically, firstly, the system retrieves medical data search results that match user permissions, serving as the foundational data for permission usage assessment. These search results include the user's query request, the data set returned by the system, and relevant permission determination information, comprehensively recording the initial stages of the user's interaction with the data. Based on these search results, the system establishes a complete behavior monitoring mechanism to gain a deeper understanding of how users utilize authorized data.

[0242] Access behavior data collection forms the data foundation for constructing the relationship graph. A comprehensive behavior monitoring framework was implemented to collect detailed information on user interactions with medical data. The collected content includes five key dimensions: access operation type (e.g., viewing, editing, downloading, deleting), access data range (the data set accessed and its characteristics), access time characteristics (the specific time and duration of the access), access environment information (device, location, network environment), and subsequent access behaviors (e.g., data export, secondary editing). Non-intrusive monitoring technology is employed to achieve accurate capture of behavioral data without affecting the smoothness of user operations.

[0243] The data acquisition architecture adopts a layered design: the front-end acquisition layer is deployed in the client application, responsible for capturing user interface operations and local interaction behaviors; the service layer acquisition points are deployed on the application server, recording API calls and service requests; and the storage layer acquisition points monitor database operations and track underlying data access patterns. This multi-layered acquisition ensures the integrity and consistency of behavioral data, preventing data omissions that might occur with single-point monitoring.

[0244] Preprocessing behavioral data is a necessary step in constructing a high-quality relationship graph. First, data cleaning is performed to handle missing values, outliers, and duplicate records. Then, session reconstruction is conducted, organizing discrete action events into meaningful action sequences that reflect the user's complete workflow. Finally, feature extraction is performed to extract high-level behavioral features from the raw behavioral data, such as access patterns, action sequences, and temporal distributions. The preprocessed behavioral data is stored in a structured format, providing standardized input for relationship graph construction.

[0245] A user-permission-data relationship graph (or simply permission relationship graph) is a multidimensional graph structure used to represent the complex interactions between users, permissions, and data in a healthcare information system. This graph contains three main types of nodes: user nodes (representing user entities in the system), permission nodes (representing various permission definitions), and data nodes (representing accessed healthcare data entities). Nodes are connected by various types of edges, reflecting the relationships and interactions between different entities.

[0246] User nodes contain user identity information and role characteristics, such as user ID, department, responsibilities, and professional background. User nodes also record user activity metrics, such as login frequency, average session duration, and system usage patterns. These characteristics help establish baseline user behavior patterns.

[0247] Permission nodes represent various permission items defined in the system, including permission type (such as read, write, delete, etc.), permission scope (applicable data categories), and authorization conditions (time restrictions, location restrictions, etc.). Permission nodes also contain metadata, such as permission creation time, authorization approval records, and intended use, providing contextual information for permission management.

[0248] Data nodes represent medical data entities within the system, recording characteristics such as data type, sensitivity level, department, and patient association. Data nodes also contain usage statistics, such as access frequency, common access patterns, and typical use cases, reflecting the data's value and importance within the organization.

[0249] Edges in a relationship graph connect nodes of different types, representing the relationships and interactions between entities. The main edge types include: user-permission edges (representing the permissions granted to a user), permission-data edges (representing the data scope to which permissions apply), and user-data edges (representing the user's actual access to data). Each edge carries attribute information, such as relationship establishment time, relationship strength, and interaction frequency, providing a quantitative representation of relationship characteristics.

[0250] The permission relationship graph is constructed using an incremental update mechanism, with new behavioral data continuously integrated into the existing graph structure to form a dynamically evolving relationship network. The graph construction process considers the time dimension, recording the temporal characteristics of nodes and edges, and supporting relationship analysis based on time windows. The system also implements graph compression technology to aggregate and abstract historical data, balancing storage overhead and analytical accuracy.

[0251] By constructing this multidimensional relationship graph, we can comprehensively capture the complex permission usage patterns in the medical environment, providing rich structured information for subsequent permission assessment and anomaly detection.

[0252] Second, based on the permission usage behavior relationship graph constructed in the previous step, a series of professional graph algorithms are applied for in-depth analysis and evaluation. Graph algorithms can uncover hidden patterns and structural features in the relationship network, thereby providing a comprehensive evaluation of permission usage. The system's evaluation process comprehensively considers node attributes, connection structure, and interaction patterns, forming a multi-dimensional evaluation of permission usage.

[0253] Node importance assessment is the first step in the evaluation process. It involves calculating the weights and centrality metrics of each node in the graph to identify key entities. For user nodes, an improved eigenvector centrality algorithm is applied to calculate the user's influence within the permission network. This algorithm considers not only the number of nodes directly connected to the user but also the importance of those nodes themselves, thus identifying users closely associated with important data and critical permissions. For user nodes with high centrality, more detailed behavioral pattern analysis is performed because these users typically have broader system access capabilities, and their abnormal behavior may pose a greater security risk.

[0254] For permission nodes, the betweenness centrality algorithm is applied to identify critical permissions that act as a bridge between users and data. Permission nodes with high betweenness centrality are typically common channels for multiple users to access various types of data; the rationality and usage of these permissions directly impact the overall security posture of the system. Stricter monitoring and more frequent auditing are applied to these nodes to ensure that critical permissions are not abused.

[0255] For data nodes, a variant of the PageRank algorithm is used to calculate a sensitivity score, taking into account the frequency of data access, the diversity of users accessing the data, and the importance of connection permissions. Data nodes with high scores typically represent widely used data resources within the organization that contain important information, and changes in access patterns to these nodes may indicate potential data breach risks.

[0256] Path analysis is a core technology for evaluating permission usage patterns. It identifies and evaluates user access paths to data within a relationship graph, reflecting how permissions are actually used. Path weight calculation comprehensively considers path length (number of nodes traversed), the types of permissions included in the path, and the frequency of path usage. The system pays particular attention to low-frequency and emerging paths, which may represent rare operational patterns or newly emerging usage needs.

[0257] Path critical point analysis identifies decisive links in the access path, typically specific permission nodes or decision points. The system calculates the impact of deleting each link on path feasibility, identifying single points crucial to access control. These critical points become key monitoring targets, receiving special attention to their status changes and usage patterns.

[0258] Community detection algorithms are used to discover natural clusters in relationship graphs, identifying tightly connected user-permission-data subnetworks. A modularity maximization method is applied to segment the relationship graph into multiple functionally independent communities. These communities typically correspond to business units or functional teams within an organization, reflecting collaboration and data sharing patterns in actual work. Community segmentation helps establish context-based behavioral baselines because different communities may have their own unique normal access patterns.

[0259] Time series graph analysis provides a dynamic dimension to the evaluation process. It constructs a sequence of time-window relationship graphs to capture the evolution of permission usage patterns over time. By comparing the graph structure characteristics of different time windows, it is possible to identify sudden changes (such as anomalous connection patterns appearing in the short term) and gradual trends (such as a sustained increase in the frequency of a certain type of permission usage). Time series analysis is particularly suitable for identifying seasonal behavioral patterns and normal fluctuations caused by cyclical business activities, avoiding the misjudgment of these patterns as anomalies.

[0260] Anomaly subgraph detection techniques are used to identify anomalous structures in relational graphs. Density-based and entropy-based anomaly detection methods are employed to identify subgraph structures with connection patterns that differ from the main network. These anomalous subgraphs may represent unauthorized access attempts, unauthorized operations, or unusual data access patterns. In-depth analysis of the identified anomalous subgraphs determines their causes and potential risks.

[0261] Knowledge graph augmentation is an innovative technology for improving assessment accuracy. It integrates medical domain knowledge graphs with permission relationship graphs, introducing professional knowledge constraints. For example, the system knows that cardiologists typically need to access electrocardiogram (ECG) data but rarely access bone density test results; this domain knowledge helps the system more accurately judge the rationality of specific access patterns. Knowledge graphs also provide semantic relationships between medical data, enabling the system to understand seemingly unrelated but actually relevant data access behaviors.

[0262] The final step in this process is generating the assessment results. The multi-dimensional analysis results from the integrated graph algorithm are used to generate a structured access control assessment report. The assessment results comprise four core parts: access control health score (quantifying the overall rationality of access control usage), key node assessment (identifying high-risk users, sensitive data, and critical permissions), anomaly pattern summary (summarizing discovered abnormal behavior patterns), and trend analysis (predicting the future direction of access control usage). The assessment results are presented visually, allowing security managers to intuitively understand the system's access control status.

[0263] By applying these advanced graph algorithms, we have achieved deep insights into permission usage behavior, providing a solid analytical foundation for subsequent anomaly detection and permission adjustment.

[0264] Third, based on the previously obtained permission usage assessment results and behavior relationship diagram, a baseline of normal permission usage patterns is constructed, and an anomaly detection mechanism is developed to identify behaviors that deviate from the normal pattern. A context-aware anomaly detection method is adopted to consider the special and dynamic nature of the medical environment and avoid false positives and false negatives that may be caused by simple rule-based judgments.

[0265] Building a normal behavior feature library is fundamental to anomaly detection. First, validated normal behavior samples are selected from historical data, including standard workflows of typical users, data access patterns in routine medical scenarios, and verified permission usage cases. These samples undergo feature extraction and pattern mining to form a multi-dimensional normal behavior feature set. Feature dimensions include temporal patterns (e.g., distribution of access time periods, rhythm of operation sequences), spatial patterns (e.g., distribution of data access scope, cross-departmental access patterns), and relational patterns (e.g., typical structures of user-data associations, common forms of permission combinations).

[0266] The feature library is organized in a hierarchical structure, progressing from a global model to a role-specific model and then to a user-specific model, forming a progressively refined description. The global model describes general behavioral constraints applicable to all users, such as basic rules of system usage; the role-specific model defines more specific behavioral characteristics for specific professional roles (such as cardiologists and radiology technicians); and the user-specific model captures the unique habits and work patterns of individual users. This hierarchical structure enables the system to evaluate the rationality of behavior at different granularities, balancing general rules and personalized judgments.

[0267] Continuous feature library updates are a key mechanism for maintaining detection accuracy. A pattern learning loop is implemented to continuously incorporate new, validated, and normal behaviors into the feature library, enabling the model to adapt to the evolution of organizational business and user habits. The update process includes new pattern discovery (identifying stable behavioral patterns that have not yet been included in the library), pattern validity verification (confirming the universality and persistence of new patterns), and pattern integration (merging new patterns with existing knowledge). The feature library updates employ a gradual strategy to avoid detection instability that may result from abrupt updates.

[0268] Deviation calculation is the core step in quantifying the difference between the current behavior and the normal pattern. The user's current permission usage pattern is converted into a feature vector of the same dimensions as the feature library, and then the distance between this vector and the set of normal patterns is calculated. Distance calculation employs a multi-metric fusion method, including Euclidean distance (measuring absolute deviation in the feature space), Mahalanobis distance (considering deviation based on correlation between features), and cosine similarity (measuring consistency in pattern direction). The system assigns differentiated weights to different feature dimensions, highlighting the importance of key behavioral features. The final deviation is a weighted combination of these distance metrics, ranging from 0 to 1, with larger values ​​indicating greater deviation from the normal pattern.

[0269] Context awareness is an innovative feature of this system's anomaly detection, enabling the detection process to adapt to the complexity of the medical environment. Three key contextual factors are introduced to dynamically adjust the criteria for anomaly judgment:

[0270] The degree of medical urgency is the primary situational factor considered. The urgency of the situation is assessed by analyzing the nature of current medical activities (e.g., routine outpatient care, emergency treatment, emergency resuscitation). In high-urgency situations (e.g., intensive care, disaster response), the criteria for anomaly assessment are appropriately relaxed, allowing medical personnel more flexible access to necessary information to ensure timely medical treatment. The urgency assessment is based on multiple signals, such as emergency event markers in the system, surges in activity during unusual work periods, and the mobilization of specific emergency resources.

[0271] Business phase characteristics reflect the cyclical changes in an organization's business activities. Identify and adapt to the specific needs of different business phases, such as end-of-month statistical analysis periods, annual medical quality reviews, and peak periods for infectious disease surveillance. During specific business phases, certain access patterns that are typically considered anomalous may be reasonable. The system maintains a business calendar knowledge base, recording the temporal characteristics and data requirements of cyclical business activities, to adjust threshold judgments.

[0272] Organizational policy changes are institutional factors that influence access usage patterns. Tracking policy updates in healthcare institutions (such as adjustments to data management regulations and the implementation of new privacy protection measures) helps predict potential shifts in behavioral patterns. When policy changes are identified, an adaptation period is initiated, during which more conservative anomaly criteria are used to avoid misjudging policy-induced changes in compliance behavior as abnormal.

[0273] The dynamic adjustment of the anomaly detection threshold directly reflects the influence of the aforementioned situational factors. Based on the current situational combination, a situational adjustment coefficient is calculated, which is used to correct the base threshold. The adjustment algorithm considers the independent influence and interaction of situational factors to generate the final dynamic threshold. In standard situations, the system uses an empirically set baseline threshold; in special situations, the threshold will be adjusted upwards or downwards according to the characteristics of the situation, but the range of variation is constrained by the system configuration to ensure that the safety baseline is not breached.

[0274] Anomaly grading classifies detected anomalies based on a comparison of deviation from a threshold, categorizing them by severity. A four-level classification framework is employed: Warning (minor deviation, requiring logging but usually no intervention), Attention (significant deviation, requiring security personnel review), Alert (serious deviation, potentially triggering automated intervention), and Emergency (extreme deviation, indicating a high probability of a security incident, triggering immediate response). The grading process considers not only the absolute value of the deviation but also its persistence, the sensitivity of the data involved, and the potential scope of impact, resulting in a comprehensive rating.

[0275] The presentation of anomaly detection results adopts a multi-layered design to meet the needs of different roles. For security managers, it provides a detailed anomaly analysis report, including anomaly characteristic descriptions, deviation calculation details, and contextual considerations; for department heads, it generates summary reports, highlighting anomalies within their department and users requiring attention; for end users, it provides behavioral reminders when necessary, guiding users to avoid operating patterns that may be considered abnormal.

[0276] This context-aware anomaly detection mechanism can accurately identify genuine risks of privilege abuse while avoiding unnecessary interference with normal medical work, thus achieving a balance between safety and efficiency.

[0277] Fourth, based on the previously obtained abnormal permission usage detection results, further analysis of the causes is conducted, and corresponding permission adjustment suggestions are proposed. Finally, permission changes are implemented through the approval process, completing the closed loop of permission management. The system adjustment strategy balances security and business continuity, avoiding work interruptions that may result from overreaction, while ensuring timely response to genuine security threats.

[0278] The generation of permission adjustment suggestions begins with anomaly analysis, which delves into detected anomaly patterns to identify potential causes. Analysis dimensions include over-granting permissions (users possess permissions beyond their actual needs), insufficient permissions (users lack the necessary permissions to perform normal tasks, leading to attempts to bypass permissions), flawed permission management (insufficiently defined permissions or lack of necessary constraints), and user behavior issues (users failing to follow established data access guidelines). For each cause type, a corresponding solution template is matched to generate preliminary adjustment suggestions.

[0279] The specifics of the proposed adjustments take into account multiple factors. First, the severity and urgency of the anomaly are assessed to determine the priority and scope of the adjustments. Then, the impact of the adjustments is analyzed, assessing potentially affected business processes and user groups. Finally, implementation costs and technical feasibility are considered to ensure the recommendations are actionable. The recommendations include specific adjustment actions (such as revoking specific permissions, adding approval steps, or adjusting access control rules), expected effects, and a potential risk assessment.

[0280] The recommendations for adjusting permissions are mainly divided into four categories: permission reduction recommendations (when it is found that a user has excessive permissions that are not used or may be abused, it is recommended to revoke or restrict these permissions), permission supplementation recommendations (when it is found that a user's normal work is restricted by unnecessary permissions, it is recommended to appropriately expand the scope of permissions), condition strengthening recommendations (when the permission definition is reasonable but lacks sufficient usage conditions, it is recommended to add time limits, scenario limits or approval requirements), and monitoring strengthening recommendations (when the permission settings are basically reasonable but require stricter supervision, it is recommended to increase the audit frequency or real-time monitoring measures).

[0281] Prioritization serves as a guiding principle for implementation recommendations. Adjustment recommendations are prioritized based on risk assessments to ensure the most critical security risks receive the fastest response. Risk assessments comprehensively consider data sensitivity (the confidentiality level of the data involved), the degree of anomaly (the distance of behavior from normal patterns), the scope of impact (the amount of data and the number of users affected), and the duration of the anomaly (the length of time the problem persists). High-priority adjustment recommendations trigger expedited approval processes to ensure rapid response.

[0282] Automated approval process generation ensures smooth access control adjustments. Based on the type, scope, and impact of the adjustment, a matching approval process is automatically generated. For minor adjustments with low impact (such as adjusting non-sensitive permissions for a single user), only department head approval may be required. For adjustments with medium impact (such as team-level permission policy changes), a dual approval process involving both the department head and the data security officer is generated. For major adjustments with high impact (such as changes affecting core business processes or a large number of users), a multi-level approval process including senior management is created. The approval process automatically includes relevant background information, evidence of anomalies, and risk assessments to facilitate informed decisions by approvers.

[0283] The emergency response mechanism is a special channel for handling high-risk anomalies. For severe anomalies deemed urgent (such as obvious data breach attempts or large-scale unauthorized access), the emergency response protocol is activated, which may include immediate measures such as temporary privilege freezes, session termination, or system area isolation. These measures begin to control the spread of risk while awaiting formal approval. Emergency measures typically have a preset expiration date, after which formal approval is required for extension.

[0284] The execution of permission adjustments is the final step in the entire process. Once the necessary approvals are obtained, the permission changes are automatically implemented, updating relevant access control lists, permission rules, and user configurations. The adjustment process employs atomic operation principles to ensure the consistency and integrity of permission changes. The system records detailed change logs, including the adjustment content, execution time, approvers, and triggering reasons, supporting subsequent auditing and traceability. For adjustments affecting the current active session, the system implements a graceful transition strategy, such as applying the new permissions after the user's operation is completed, avoiding work interruption.

[0285] Effectiveness evaluation is a crucial step after permission adjustments. Continuous monitoring of permission usage after the adjustments is essential to assess the effectiveness of the measures. Evaluation metrics include the elimination of abnormal behavior, the impact on work efficiency under the new permission configuration, and user satisfaction feedback. If the adjustments fail to achieve the expected results or have negative impacts, the system will propose further optimization suggestions, forming a closed-loop management system for continuous improvement.

[0286] Knowledge accumulation and experience feedback serve as the system's self-improvement mechanism. A complete case study of each permission adjustment is recorded, including the problem context, the measures taken, and the final results, to build a permission management knowledge base. This accumulated experience is used to optimize the anomaly detection model, improve the adjustment suggestion generation algorithm, and refine the approval process design, enabling the system to continuously adapt to changing medical environments and security requirements.

[0287] User communication and training are complementary measures to permission adjustments. When implementing permission changes, a notification is sent to affected users, explaining the reasons for the adjustment and the new permission rules. For users or departments that repeatedly exhibit similar anomalies, targeted security awareness training or operational guidance will be recommended to help users understand security regulations and adjust their work habits, reducing the occurrence of abnormal behavior from the outset.

[0288] This comprehensive, dynamic process of permission assessment and adjustment enables intelligent management of medical data access permissions. It allows for timely response to security risks and adapts to the dynamic needs of medical work, ensuring a balance between data security and business efficiency. Permission management is no longer a static configuration process but a continuously optimized, dynamic process, providing a solid guarantee for the secure sharing and effective utilization of medical data.

[0289] Example 10

[0290] In this embodiment, the application graph algorithm calculates node weights and path weights, performs real-time evaluation of the permission usage behavior relationship graph, and obtains permission usage evaluation results, including:

[0291] Based on the permission usage behavior relationship graph, key nodes are identified using node centrality and feature vector centrality, and the subgraph formed by the key nodes is extracted to obtain the key node subgraph.

[0292] Based on the key node subgraph, user-permission granting weight, permission-data coverage weight, user-data access weight and time decay factor are defined. The proportion of each type of weight in the total weight is dynamically adjusted according to the characteristics of the medical business process to obtain adaptive weight parameters.

[0293] Based on the adaptive weight parameters, the relevant path weights are recalculated only for the newly added or changed behavior data in the permission use behavior relationship graph. The weight distribution is estimated by using the random walk approximation technique to obtain the approximate weight calculation results.

[0294] Based on the approximate weight calculation results, a distributed computing model is used to distribute the computing tasks to multiple computing nodes, and the results are aggregated through a message passing mechanism to obtain the permission usage evaluation results.

[0295] Specifically, firstly, a node importance analysis algorithm is applied to the permission usage behavior relationship graph. The core objective is to reduce computational complexity and focus on the core network structure. Two complementary methods, node centrality and eigenvector centrality, are used to capture node importance.

[0296] Node centrality, as a fundamental metric, employs multiple calculation methods tailored to the characteristics of the permission relationship graph: Degree centrality identifies active nodes by counting the number of direct connections between nodes; Betweenness centrality assesses the role of nodes as bridges in information flow, identifying key permission nodes controlling data access paths; Proximity centrality calculates the average distance from a node to all other nodes, filtering nodes capable of efficiently accessing network resources. The system adopts a differentiated centrality combination strategy for different types of nodes: user nodes are primarily evaluated based on degree centrality and betweenness centrality, permission nodes emphasize betweenness centrality, and data nodes focus on proximity centrality and degree centrality.

[0297] Eigenvector centrality, as an advanced metric, is based on the idea that "nodes connected to important nodes are more important." It can identify users with a small number of connections but who are associated with key data or high-privilege nodes. This study uses a power-law iteration method to calculate the centrality and introduces medical-specific adjustments, assigning higher base weights to patient-sensitive data nodes, thus giving higher scores to users associated with such nodes.

[0298] The centrality calculation adopts an incremental update mechanism, which only recalculates the centrality of the changed regions of the relationship graph, rather than recalculating the entire network, thereby improving real-time analysis capabilities. At the same time, it incorporates a time decay factor, which gives higher weight to recent connection behavior in the calculation, accurately reflecting the current permission usage pattern.

[0299] Key node identification employs a multi-threshold screening based on centrality metrics: user nodes are selected from the top 15% of centrality or those exceeding a preset threshold; permission nodes are selected from the top 20% of betweenness centrality; and data nodes are selected from the top 10% of feature vector centrality. The thresholds utilize an adaptive mechanism, dynamically adjusting based on network size and connection density to balance computational burden with coverage of critical nodes.

[0300] Key node subgraph extraction must retain key nodes and their direct connections, as well as indirect connections with shortest paths not exceeding two hops, to ensure the integrity of the subgraph structure. After extraction, enhancement processing is required, including edge weight recalculation, implicit relationship inference, and contextual information appending, providing a high-quality semantic data foundation for subsequent analysis. Through these steps, a large-scale permission usage behavior graph is simplified into a concise subgraph, reducing computational complexity while maintaining analytical value and supporting real-time permission assessment.

[0301] Second, we define refined weight calculation methods for different relationships in key node subgraphs and establish a dynamic adjustment mechanism based on medical business characteristics to improve the business relevance of the evaluation results.

[0302] User-permission grant weight measures the formality and compliance of permission granting. It considers four factors: the source of permission granting, the completeness of the authorization process, the timeliness of permission, and the sufficiency of the authorization reason. The comprehensive weight is calculated by analyzing the authorization records and approval documents of the permission management system. The value ranges from 0.1 to 1.0. The higher the value, the stronger the compliance and formality.

[0303] Permissions - Data Coverage Weight reflects the accuracy of permission coverage of target data. The weight is determined by calculating the range matching degree, granularity matching degree and purpose consistency between permissions and data classification. The value range is 0.2-1.0. The higher the value, the more precise the permission design and the more it conforms to the principle of least privilege.

[0304] User-Data Access Weight assesses the rationality and necessity of data access behavior. Based on the user's historical access records and current business role, it is calculated from four dimensions: reasonableness of access frequency, relevance of access content, reasonableness of access timing, and reasonableness of access operation. The value ranges from 0 to 1.0, with higher values ​​indicating that the access behavior is more in line with business expectations.

[0305] The time decay factor is used to adjust the influence of historical data. It adopts an exponential decay function. In scenarios with rapid changes, such as clinical operations, a fast decay rate of 7 days half-life is used, while in stable scenarios such as access to scientific research data, a slow decay rate of 30 days half-life is used to balance the reference value of historical patterns with the indicative significance of current behavior.

[0306] Dynamic weight adjustment is the core mechanism for adapting to different medical business processes: in routine medical service scenarios, the weights of the three categories are relatively evenly distributed (30%-40%-30%); in emergency and critical care scenarios, the weight of user-data access is increased to over 50%; in medical audit scenarios, the weight of permission-data coverage is increased to 60%; and in scientific research and analysis scenarios, the weight of user-permission granting is increased to 55%. The system identifies business scenarios by recognizing active user types, operating patterns, and system load characteristics, and adjusts weights using a smooth transition mechanism of 15-30 minutes to avoid abrupt changes affecting assessment stability.

[0307] The weight parameters adopt a hierarchical storage design: system-level default parameters are the basic configuration, department-level custom parameters are adapted to the specific needs of departments, and scenario-level dynamic parameters respond to real-time business changes. Parameter updates are recorded through a version control mechanism, supporting backtracking and auditing.

[0308] Third, a strategy combining incremental computation and approximate algorithms is adopted to efficiently process dynamic permission usage behavior data, reduce the computational complexity of real-time evaluation, and ensure high responsiveness in large-scale medical data environments.

[0309] The incremental calculation strategy only processes the newly added or changed parts of the relationship graph. The key is to identify the scope of influence through the change propagation model. This scope is determined by three factors: topological distance, relationship strength, and dependency. Each update only recalculates on the subgraph of the affected region, which greatly reduces the amount of computation.

[0310] Path weight calculation is the core of the evaluation. It comprehensively considers the characteristics of nodes on the path (appropriateness of user roles, accuracy of permission scope, and level of data sensitivity), edge weights (user-permission granting weight, permission-data coverage weight, and user-data access weight), and overall path characteristics (length, uniqueness, and frequency of use). An adaptive weight parameter is applied to calculate a comprehensive weight value in the range of 0-1. The higher the value, the more compliant and reasonable the use of permissions.

[0311] Random walk approximation is used to efficiently estimate the connection strength between nodes. Starting from the source node, nodes move randomly according to the edge weight transition probability until the target node is reached or the maximum number of steps is reached. A large number of samples are used to obtain the reachability distribution, which approximately reflects the probability of permission usage. Key optimizations include biased random walk, restarted random walk, and parallel random walk. The system dynamically selects a combination of strategies based on the graph structure and resource conditions.

[0312] The approximate accuracy control adopts an adaptive sampling strategy. The high-sensitivity evaluation increases the number of samplings, while the routine evaluation uses the standard number of samplings. At the same time, convergence monitoring is implemented. When the change in the sampling result is lower than the threshold, the calculation is terminated early to ensure that the error is controlled within the acceptable range of 3%-5%.

[0313] Incremental computation combined with approximation techniques forms an efficient evaluation framework: incremental computation is directly applied to routine data updates; important updates are precisely calculated in the core area and approximated by random walks in the peripheral area; at the same time, the pre-calculation results of the critical path are cached to reduce redundant calculations, thereby achieving efficient evaluation of large-scale dynamic permission usage behavior and balancing accuracy and real-time performance.

[0314] Fourth, implement a distributed computing architecture to distribute the computing load, improve system throughput and response speed through parallel processing, and at the same time ensure the consistency and reliability of results.

[0315] Task decomposition and scheduling employ a domain decomposition strategy, splitting permission evaluation tasks according to three dimensions: data partitioning, user grouping, and functional division, and maintaining a global task dependency graph to ensure correct execution order. The scheduler comprehensively considers load balancing, data locality, and task priority to formulate execution plans, assigning higher priority to time-critical tasks such as emergency medical scenarios to ensure timely response.

[0316] The compute node management adopts a flexible framework, deploying a three-tiered compute node system: core nodes, standard nodes, and edge nodes. Core nodes are always running and responsible for critical computing and coordination; standard nodes dynamically start and stop based on load to handle routine tasks; and edge nodes are deployed at the departmental level to handle localized, simple assessments. A node health monitoring mechanism tracks status in real time, supports heterogeneous computing environments, and maximizes the utilization of existing IT infrastructure.

[0317] Data distribution and access optimization employs a partitioning and replication strategy, placing frequently queried data in the same partition, backing up high-frequency core data across multiple nodes, and prioritizing local data processing tasks on each node to reduce network transmission overhead. The system implements a distributed caching mechanism, maintaining recent permission data and intermediate result caches on each computing node, and ensuring cache consistency through version tagging and proactive expiration notifications.

[0318] The message passing mechanism adopts an asynchronous model and supports four types of messages: task allocation, status update, intermediate result transmission, and result aggregation. Through optimization measures such as batch processing, compression, priority queues, and persistent retransmission, it ensures efficient and reliable communication, while implementing message flow control to maintain system stability.

[0319] The aggregation and integration of results adopts a layered strategy, first summarizing locally and then integrating at the upper layer to reduce the amount of data transmission; different consistency levels are required according to the assessment type, with strong consistency required for key permission audits and eventual consistency for routine monitoring; the result verification mechanism checks consistency and reasonableness to ensure the reliability of assessment results.

[0320] Error handling and fault tolerance are implemented with a multi-layered strategy: at the node level, single points of failure are handled through task retries and backup switching; at the data level, integrity is ensured through redundant storage and checksums; at the system level, service availability is maintained through monitoring and automatic recovery, while graceful degradation strategies are implemented to retain core evaluation functions when resources are limited or components fail.

[0321] Performance optimization identifies bottlenecks through monitoring and analysis, implementing targeted optimizations across four dimensions: computation, communication, storage, and load. It supports incremental updates of computation results, further improving response speed. This distributed architecture efficiently handles permission assessment tasks for large-scale medical institutions, generating comprehensive assessment results reflecting the organization's permission usage, providing a solid foundation for subsequent anomaly detection and permission optimization.

[0322] Example 11

[0323] In this embodiment, the construction of a normal permission usage pattern feature library is performed, the deviation between the current user's permission usage pattern and the normal permission usage pattern feature library is calculated, and medical urgency, business stage characteristics, and organizational policy changes are introduced as contextual factors to dynamically adjust the anomaly judgment threshold. The deviation is compared with the anomaly judgment threshold, and anomalies are classified according to the comparison results to obtain permission usage anomaly detection results, including:

[0324] Based on the permission usage evaluation results and the permission usage behavior relationship diagram, historical behavior data is extracted, and the permission usage normal mode feature library is constructed for different roles, different departments and different business scenarios. The permission usage normal mode feature library includes statistical features, sequence features, correlation features and content features. The normal behavior distribution is represented by a Gaussian mixture model to obtain the permission usage benchmark model.

[0325] Based on the permission usage benchmark model and the current user permission usage pattern, calculate the deviation between the current user permission usage pattern and the permission usage benchmark model, evaluate the deviation of single behavior, the deviation of sequential behavior and the deviation of pattern, and obtain the permission usage deviation score.

[0326] Based on the permission usage deviation score, the severity of medical urgency, business stage characteristics, and organizational policy changes are introduced as contextual factors. The anomaly judgment threshold is dynamically adjusted through preset rules and machine learning models. The permission usage deviation score is compared with the anomaly judgment threshold to obtain the context-corrected anomaly score.

[0327] Based on the anomaly score after context correction, abnormal behaviors are categorized into observation level, alert level, intervention level, and blocking level. For anomalies at the intervention level and above, evidence information including operation context, system status, and user response is automatically collected to obtain the permission usage anomaly detection results.

[0328] Specifically, firstly, the core of abnormal permission usage detection is to build a feature library of normal permission usage patterns, and to create differentiated benchmark models for different user groups and business scenarios through a hierarchical classification method, so as to ensure the accuracy and applicability of detection.

[0329] Historical behavior data is derived from permission usage assessment results and permission usage behavior relationship diagrams, and undergoes multi-dimensional screening to ensure quality. Screening criteria cover a time range of the most recent 3-6 months, normal behavior confirmed by security audits, complete record data, and representative samples covering various user roles and business scenarios. In the data preprocessing stage, outliers and missing values ​​are handled through cleaning algorithms, and normalization and standardization techniques are used to unify feature scales, providing standardized data for model construction.

[0330] The feature library employs a layered design to adapt to the diversity of medical organizations: the role layer captures specific behavioral patterns for professional roles such as cardiologists and radiology technicians; the department layer reflects the workflow and data access characteristics of different departments such as the emergency department and internal medicine wards; and the business scenario layer focuses on permission usage patterns for specific activities such as outpatient and inpatient treatment. The feature library contains four core feature categories: statistical features capture quantitative distribution characteristics such as access frequency and operation duration, calculated using a sliding time window; sequence features focus on the temporal relationships and rhythm of operations, extracting micro, meso, and macro multi-scale patterns; association features describe the relationship structure between users, data, and permissions, identifying frequent substructures through graph pattern mining; and content features focus on the specific content of operations, establishing topic models by combining domain knowledge and natural language processing techniques.

[0331] A Gaussian mixture model (MoG) is used to represent the distribution of normal behavior, adapting complex behavioral patterns by weighting multiple Gaussian distributions. The system constructs an independent model for each feature class, with parameters learned from historical data using the Expectation-Maximization (EM) algorithm. The number of Gaussian components is adaptively determined based on the Bayesian Information Criterion (BIC). Model performance is evaluated through cross-validation, and optimization techniques such as feature selection and parameter tuning are employed. Confidence intervals are set to define the boundaries of normal behavior. The feature library is maintained using a strategy combining regular updates and event-driven updates. Regular monthly updates are implemented, with special updates triggered when significant changes occur within the medical organization. Version control records the evolution history.

[0332] Second, the deviation of current user behavior is calculated based on the benchmark model, using a multi-level, multi-dimensional evaluation method. The system collects user permission usage behavior data in real time, ensuring that the format is consistent with the benchmark model, complex operations are broken down into basic units, and the preprocessing process is consistent with the model construction.

[0333] The deviation calculation adopts a three-layer architecture: single behavior deviation converts the current operation into a feature vector and compares it with the benchmark model, giving higher weight to key feature deviations and focusing on rare and high-risk operations; sequence behavior deviation constructs a time window to extract temporal features, and applies Markov models, dynamic time warping and other techniques to detect anomalies, focusing on rate and rhythm anomalies; pattern deviation constructs a user behavior profile based on a long time window, and uses concept drift detection and multi-dimensional comparative analysis to analyze overall pattern changes, focusing on progressive anomalies.

[0334] Feature domain weighting assigns weights based on the security sensitivity and discriminative ability of different features, and the weight configuration supports dynamic adjustment. Context-dependent deviation calculation considers the business environment in which the operation occurs, adjusting the benchmark and sensitivity to reduce false alarms caused by reasonable business changes. Deviation score fusion integrates multi-level and multi-dimensional deviation calculation results into a standardized score of 0-100. The fusion algorithm uses a non-linear model, giving disproportionate weights to high-risk combinations. An interpretability mechanism generates an explanation report detailing the specific reasons for high deviations, providing a basis for communication between security analysts and users.

[0335] Third, a context-aware mechanism is introduced to improve the adaptability of anomaly detection, and three types of core context information are collected in real time: the degree of medical urgency adopts a five-level quantitative standard and is evaluated through indicators such as the status of the emergency system and the scheduling of medical resources; the characteristics of business stages are identified based on the business calendar to identify current periodic activities and extract key features such as data access frequency; and organizational policy changes are tracked and management system adjustments are made to analyze the scope and depth of their impact on operational behavior.

[0336] The context integration model transforms multidimensional contextual information into a unified numerical representation, considering the interactions between factors and temporal continuity. The anomaly detection threshold employs a dual-engine dynamic adjustment: a rule engine defines the explicit relationship between contextual factors and threshold adjustment based on expert knowledge; a machine learning model (gradient boosting decision tree) captures the impact of complex contexts and is periodically retrained using the latest data. Threshold adjustments set safety boundaries to ensure that basic safety limits are not exceeded, with stricter restrictions applied to specific highly sensitive operations.

[0337] The context-adjusted anomaly score is calculated using normalized variance. A positive score indicates that the threshold is exceeded, while a negative score indicates that the threshold is within the threshold. The absolute value represents the degree of deviation. Contextual explanation tags are attached to the anomaly score, recording the contextual factors that led to the threshold adjustment, providing a reference for anomaly evaluation and pattern optimization.

[0338] Fourth, implement a four-level anomaly management strategy based on context-corrected anomaly scores: Observation level (anomaly score slightly above the threshold up to 30%) adopts a silent recording strategy, only recording in the security log; Alert level (30%-60%) sends a notification to security personnel, and may display operation prompts to users; Intervention level (60%-90%) initiates proactive intervention, requiring additional identity verification or supervisor approval, sending a high-priority alert to the security team and initiating automatic evidence collection; Blocking level (above 90%) immediately blocks operations and temporarily freezes accounts, triggers emergency response procedures, and initiates comprehensive evidence collection and security isolation.

[0339] The anomaly grading criteria support dynamic adjustment, taking into account data sensitivity, the scope of operational impact, and the credibility of user history. Automated forensics, targeting anomalies at the intervention level and above, collects three key types of information: operational context, system status, and user response, ensuring evidence integrity through technologies such as digital signatures and encrypted storage. Anomaly detection results are presented in a layered design to meet the needs of different roles. The result feedback mechanism records security analyst markings and feedback for model optimization and tracks the entire anomaly handling process to form a knowledge base.

[0340] This mechanism enables accurate identification and effective response to abnormal permission usage, balancing the needs of medical data security with normal business operations.

[0341] Example 12

[0342] In this embodiment, the construction of a medical data circulation security threat model, optimization of monitoring point layout, collection of network traffic data, system logs, user behavior data, and environmental data from each monitoring point and integration into multi-source heterogeneous data, and the performance of security event correlation analysis on the multi-source heterogeneous data, establishing a hierarchical early warning and automatic response mechanism, and obtaining anomaly detection and risk warning data of the security monitoring network, including:

[0343] Based on the dynamic evaluation and adjustment results of permissions, user nodes, data nodes, access paths and permission relationships in the medical data circulation process are extracted, a medical data circulation network topology is constructed, the types of security threats in the medical data circulation network topology are identified, and attack tree and attack graph methods are used to simulate attack paths and quantify risk levels to obtain a medical data circulation security threat model.

[0344] Based on the attack paths and risk levels in the medical data circulation security threat model, the medical data circulation network topology is abstracted into a directed weighted graph. The cut set algorithm is applied to calculate and optimize the location of monitoring points to obtain a monitoring point layout scheme.

[0345] According to the monitoring point layout scheme, network traffic data, system logs, user behavior data and environmental data are collected from each monitoring point, and standardized processing is performed to obtain the multi-source heterogeneous data.

[0346] Based on the aforementioned multi-source heterogeneous data, time-series correlation analysis and causal reasoning are applied to perform security event correlation analysis, assess the severity of threats and the urgency of responses, classify security events into observation level, alert level, warning level and emergency level, and execute automated response measures for different levels to obtain anomaly detection and risk warning data of the security monitoring network.

[0347] Specifically, firstly, a network topology diagram of medical data flow is constructed, and a comprehensive security threat model is performed based on this diagram. The network topology construction uses the results of dynamic permission assessment and adjustment as the basic data source, extracts key elements and relationships in the data flow process, and forms a complete view of medical data flow.

[0348] User node extraction is the first step in topology construction. Active user information, including medical staff, technicians, administrators, and system service accounts, is extracted from the permission assessment database. User node information includes identity identifiers, role classifications, permission sets, and behavioral characteristics. User node classification employs a two-tier design: the outer layer is categorized by function (e.g., clinicians, nurses, administrators), and the inner layer is categorized by risk level (based on historical behavior and permission scope). The system specifically marks user nodes with high-level permissions, such as system administrators, database administrators, and security administrators; these nodes are typically high-value attack targets.

[0349] Data node extraction focuses on medical data assets within the system. Data node information is extracted from the data catalog and asset management system, including various types of medical data such as electronic medical records, medical images, test results, and prescription information. Data node attributes include data type, sensitivity level, storage location, and access frequency. Data sensitivity levels are divided into four levels: general information (such as publicly available hospital information), basic medical information (such as non-identifiable statistical data), sensitive medical information (such as specific medical records), and highly sensitive information (such as mental illness records, HIV test results, etc.). The system also records the data lifecycle stages, such as data generation, use, archiving, and destruction, with different security risks at each stage.

[0350] Access path extraction reconstructs the access channel between users and data by analyzing system logs and network traffic. An access path defines the flow of data from storage location to the user, including the system components, network devices, and security control points it traverses. Access paths are represented as directed edges, with attributes including transmission protocol, frequency statistics, and time patterns. Access path analysis pays particular attention to non-standard paths, i.e., data access routes that deviate from conventional processes; these paths typically carry higher security risks. Technical and business constraints for each access path are also marked, such as the requirement for encrypted data transmission or dual approval for cross-departmental data access.

[0351] The extraction of permission relationships focuses on the configuration of user operation permissions for data. Permission rules are extracted from the access control matrix and permission management database to establish a permission mapping between user nodes and data nodes. Permission relationship attributes include operation type (read, modify, delete, etc.), authorization source, and validity period. The system pays particular attention to permission propagation paths, i.e., through which intermediate links a user can indirectly obtain data access rights; these indirect paths are often key channels for permission escape.

[0352] The medical data flow network topology is a complete network diagram integrating the above elements. It employs a multi-layered structure: the physical layer describes hardware devices and network connections; the application layer describes software systems and service components; the data layer describes data assets and storage locations; the user layer describes personnel roles and organizational structure; and the access control and authorization layer describes access control and authorization rules. These layers are interconnected to form a three-dimensional network, comprehensively expressing the technical and business views of data flow. Topology generation utilizes automated tools combined with manual verification to ensure accuracy and completeness.

[0353] Security threat type identification is based on the constructed network topology, systematically analyzing potential threats. A threat intelligence database and a healthcare industry security knowledge base are applied to identify threat types applicable to the current network environment. Major threat types include: unauthorized internal access (e.g., doctors viewing patient data outside their department), data breaches (e.g., copying patient information to unauthorized devices), man-in-the-middle attacks (e.g., network traffic hijacking), authentication bypass (e.g., credential theft), and denial-of-service attacks (e.g., system resource exhaustion). For each threat type, the applicable conditions, triggering mechanisms, and scope of impact are analyzed, assessing its feasibility and severity in the current environment.

[0354] An attack tree is a structured threat modeling method used to represent the logical path to achieving an attack objective. The system constructs an attack tree for each major threat, where the root node represents the final attack objective (e.g., obtaining specific sensitive data), intermediate nodes represent attack sub-objectives or intermediate states, and leaf nodes represent basic attack behaviors. Nodes are connected logically using AND (meaning multiple conditions must be met simultaneously) and OR (meaning any one of multiple choices is sufficient). Based on threat intelligence and historical event data, the system assigns attributes such as success probability, detection difficulty, and resource requirements to each leaf node in the attack tree. These attributes propagate upwards through the tree structure, ultimately forming a comprehensive assessment of the root node (the attack target).

[0355] The Attack Graph extends the concept of the Attack Tree by incorporating network topology information into attack path analysis. Based on the aforementioned network topology and attack tree, the system constructs an attack graph model for the medical data environment. Attack graph nodes include attack preconditions (system state or conditions), attack behaviors (specific operations), and attack results (state changes). Edges represent dependencies between attack behaviors and state transitions. Attack graph generation employs an automated reasoning engine, starting from the initial state and enumerating possible attack paths until a specific attack target is achieved. The system pays particular attention to multi-stage attack paths, i.e., complex attack scenarios where attackers need to gradually breach multiple defense layers to achieve their objective; these are often the most threatening attack paths.

[0356] Risk level quantification is a key output of threat modeling. Based on the constructed attack tree and attack graph, a risk score is assigned to each attack path. The score considers multiple factors: technical feasibility (difficulty of attack execution), required resources (time, expertise, and tools), detection probability (probability of detection by existing security measures), and potential impact (the degree of damage to data integrity, availability, and confidentiality). These factors are combined into a standardized risk score, typically ranging from 1 to 100, and attack paths are categorized into low-risk (1-40), medium-risk (41-70), and high-risk (71-100) classes. An automatic risk score update mechanism ensures that the risk assessment remains up-to-date as the environment changes and new threats emerge.

[0357] Through the above steps, a comprehensive medical data flow security threat model was constructed, providing a theoretical foundation and technical basis for subsequent monitoring point deployment and defense strategy formulation. This model not only identified potential threats but also clarified the direction for prioritizing defense resource allocation through attack path simulation and risk quantification.

[0358] Second, based on the aforementioned security threat model, the optimal layout of security monitoring points should be scientifically determined. The layout of monitoring points directly affects the effectiveness of security monitoring and the efficiency of resource utilization, and is a key link in building an efficient security monitoring network.

[0359] Network topology abstraction is the first step in monitoring point deployment. The medical data flow network topology is converted into a mathematical representation, namely a directed weighted graph. In this graph, nodes represent network entities (such as servers, network devices, application systems, etc.), and edges represent data flow paths. The edge weight is a combination of multiple factors, including data traffic volume, data sensitivity, risk level in the threat model, and path importance. Weight calculation uses a risk-weighted method, giving higher weights to high-risk attack paths to ensure these critical paths are prioritized for monitoring. The graph construction considers the network's hierarchical structure, unifying physical and logical connections while retaining node attribute information (such as device type, system function, etc.) as a reference for deployment decisions.

[0360] The cut set algorithm is a mathematical method in graph theory used to determine critical monitoring points. In the context of security monitoring, a cut set refers to a set of edges or nodes whose removal would break a specific path in the graph. This system applies the cut set algorithm to identify critical locations that can cover the most high-risk attack paths. Specifically, it employs a minimum cut set approach, finding the smallest set of nodes or edges that can monitor all critical attack paths. The algorithm first performs path analysis on the attack graph, identifying all possible paths from the attack source to the target; then it calculates the intersections of these paths, which are typically ideal locations for monitoring; finally, it applies set coverage optimization to find the minimum set of monitoring points that can cover all critical paths.

[0361] Classifying monitoring points is a crucial step in layout optimization. Monitoring points are divided into four categories, each with different monitoring functions and deployment requirements: network monitoring points are deployed at critical network nodes (such as gateways and switches) to monitor network traffic and communication patterns; monitoring points are deployed on critical servers and application systems to monitor system status, process behavior, and resource usage; data monitoring points are deployed at data storage and processing locations to monitor data access and operation behavior; and user monitoring points are deployed on terminal devices and access points to monitor user authentication and operation behavior. Based on monitoring objectives and technical constraints, the system assigns the most suitable monitoring type to each calculated monitoring location.

[0362] Monitoring coverage assessment is a core metric for deployment optimization. A method for calculating monitoring coverage is defined to evaluate the extent to which a specific deployment scheme covers attack paths. Coverage calculation considers three key factors: path coverage (the percentage of attack paths observable by monitoring points), risk coverage (the proportion of the cumulative risk value of monitored paths to the total risk), and detection depth (the number of attack stages that monitoring points can capture). Coverage assessment employs a weighted scoring method, giving higher weight to coverage of high-risk paths to ensure that limited monitoring resources prioritize the security of critical assets.

[0363] Monitoring point optimization is a crucial step in improving monitoring efficiency under resource constraints. A multi-objective optimization algorithm is applied to balance various design goals: maximizing monitoring coverage, minimizing the number of monitoring points (reducing deployment and maintenance costs), minimizing data collection volume (reducing storage and processing burden), and minimizing network impact (reducing interference from monitoring activities to business systems). The optimization process employs a genetic algorithm framework, starting with an initial layout scheme and continuously improving the scheme's quality through multiple iterations. Each iteration includes mutation (randomly adjusting monitoring point locations), crossover (combining the advantages of different schemes), and selection (retaining the best-performing scheme). The optimization process also considers technical feasibility constraints, such as the possibility that monitoring points cannot be deployed in certain locations due to hardware limitations or network architecture.

[0364] Redundancy and fault tolerance design ensure the reliability of the monitoring network. The layout design considers the fault tolerance requirements of the monitoring system itself, implementing a redundancy strategy for monitoring points. Backup monitoring points are set up along critical paths to ensure that basic monitoring capabilities are maintained even if the primary monitoring point fails. The redundancy design adopts a heterogeneous deployment principle, meaning that primary and backup monitoring at the same location are implemented using different technologies to avoid systemic failures caused by defects in similar technologies. The system also defines a fault detection and recovery mechanism for monitoring points, including health check protocols, automatic fault reporting, and automatic reallocation of monitoring tasks.

[0365] Prioritizing monitoring point deployments is a key implementation-oriented output. Based on optimization results, deployment priorities are assigned to monitoring points to guide phased implementation. Prioritization is based on three factors: risk coverage benefit (the total amount of risk covered by deploying this monitoring point), technical complexity (the difficulty of deployment and configuration), and dependencies (functional dependencies between monitoring points). Monitoring points are divided into three deployment batches: Batch 1 (essential monitoring points, covering the highest-risk paths), Batch 2 (important monitoring points, covering medium-risk paths), and Batch 3 (auxiliary monitoring points, providing comprehensive coverage). This phased deployment strategy enables security teams to prioritize the security monitoring of critical assets with limited resources.

[0366] The final output monitoring point layout plan is a comprehensive document that includes the location, type, functional description, technical specifications, deployment priority, and expected coverage of each monitoring point. The layout plan also includes a monitoring network topology diagram, visually illustrating the relationships between monitoring points and data flow, facilitating understanding and implementation by the security team. The digital version of the layout plan can be directly imported into network planning tools and security management platforms, supporting automated deployment and configuration.

[0367] Through the above optimization process of monitoring point layout, the scientific allocation of security monitoring resources has been achieved, ensuring that limited monitoring resources can cover key attack paths to the maximum extent, and providing comprehensive and efficient monitoring protection for medical data security.

[0368] Third, based on the aforementioned monitoring point layout plan, deploy various monitoring sensors and implement a comprehensive data acquisition and preprocessing process. Multi-source data acquisition is a fundamental step in security monitoring, providing raw data for subsequent analysis and detection.

[0369] Network traffic data collection is deployed at critical network nodes to capture detailed information about data communication. Collection devices include network TAPs (Test Access Points), switch port mirroring, and dedicated traffic analyzers. Collection content is divided into three levels: packet level (capturing complete network packets, including headers and payloads), session level (extracting statistical characteristics of communication sessions, such as source and destination addresses, duration, and number of bytes transmitted), and application level (identifying application layer protocols and communication content types). For sensitive medical network segments, deep packet inspection technology is used to analyze the flow of sensitive information within the communication content. The traffic collection system is configured with an intelligent sampling mechanism that dynamically adjusts the sampling rate based on network load, maintaining data representativeness while avoiding excessive storage and processing burdens.

[0370] System log collection is deployed on critical servers and application systems to record system behavior and state changes. The collection scope includes operating system logs (such as Windows event logs and Linux system logs), application logs (such as web server logs and database logs), security device logs (such as firewall logs and IDS / IPS alarms), and middleware logs (such as message queues and API gateway logs). Log collection uses standardized agents and supports multiple log formats and transmission protocols. For legacy systems lacking logging capabilities, auxiliary monitoring tools are deployed to record behavior. Real-time transmission mechanisms are configured for log collection to ensure that critical security events are delivered to the analysis system immediately, reducing detection latency.

[0371] User behavior data collection focuses on terminal devices and user interactions, recording user operations and access patterns. The collected content includes authentication events (login, logout, permission changes), data access behaviors (query, modification, download, etc.), system interaction patterns (operation sequence, dwell time, operation frequency), and abnormal behavior markers (such as multiple authentication failures, access at unusual times). The collection method combines application instrumentation (embedding behavior monitoring code in critical applications) and terminal monitoring tools (capturing user interface operations and system calls). User behavior collection places particular emphasis on privacy protection, implementing data minimization and anonymization, collecting only the behavioral data necessary for security analysis, and avoiding excessive intrusion into user privacy.

[0372] Environmental data acquisition complements traditional security data sources, providing contextual information about the operational environment. The collected data includes physical environment data (such as server room temperature, humidity, and power status), network environment data (such as load status and latency changes), business environment data (such as current business phase and activity type), and external threat intelligence (such as newly discovered vulnerabilities and threat activity reports). Environmental data acquisition is achieved by integrating various existing monitoring systems and external data sources, such as hospital facility management systems, business operation monitoring systems, and threat intelligence subscription services. This environmental data provides crucial contextual references for security incidents, helping to accurately determine the severity and urgency of events.

[0373] Data standardization is a crucial step in integrating heterogeneous data from multiple sources. A three-tiered standardization process was implemented:

[0374] Format standardization transforms data from diverse sources and formats into a unified structured format. The system employs a JSON-based universal data model, defining core fields such as event type, timestamp, source, subject, operation, and attributes. For semi-structured data (such as log text), regular expressions and parsing rules are applied to extract key information; for unstructured data (such as network load content), feature extraction algorithms are used to generate structured feature vectors. The formatting process preserves the complete semantics of the original data while ensuring consistency in subsequent processing.

[0375] Semantic standardization addresses the problem of different systems using different terms to describe similar concepts. An ontology model for the medical security domain was constructed, defining a unified conceptual system and relational framework. The standardization process maps the terms and concepts in the original data to standard terms in the ontology model, achieving semantic uniformity. For example, descriptions such as "user login," "identity authentication," and "access verification" from different systems are uniformly mapped to the standard "authentication" event. Semantic standardization significantly improves the accuracy of cross-system data association.

[0376] Time standardization addresses time inconsistencies in distributed systems. An NTP-based time synchronization mechanism is implemented to ensure all monitoring points use a unified time base. For collected historical data, a time offset correction algorithm is applied to adjust timestamps based on known system time differences. Time standardization achieves millisecond-level accuracy, ensuring the accuracy of event sequence analysis. For isolated systems that cannot be directly synchronized, event tagging technology is used to establish indirect time correspondences.

[0377] Quality control is a crucial part of the data standardization process. Comprehensive data quality management measures were implemented, including integrity checks (ensuring no key fields are missing), consistency verification (checking internal logical relationships within the data), accuracy assessment (comparing with a reference source), and timeliness monitoring (detecting data delays). For data that does not meet quality standards, the system employs a tiered processing strategy: repairable issues are automatically corrected; issues that are unrepairable but do not affect core functionality are flagged with warnings; and data with severe quality problems is isolated to prevent contamination of analysis results.

[0378] Data enrichment is a standardized extension step that adds valuable contextual information to the raw data. The system obtains supplementary information from reference sources such as asset management databases, organizational structure databases, and threat intelligence databases to enrich event data. For example, device type and location information are added to IP addresses, role and department attributes are added to user IDs, and data sensitivity level tags are added to access events. Data enrichment significantly enhances the analytical value of event data, enabling security analysis to understand the meaning of events within a richer context.

[0379] Process automation is key to ensuring efficient data processing. The system constructs an automated data processing pipeline, enabling end-to-end automatic flow from data acquisition to standardization. The pipeline comprises a data access layer (supporting multiple data sources and transmission protocols), a preprocessing layer (performing data cleaning and format conversion), a standardization layer (implementing format, semantic, and temporal standardization), a quality control layer (performing data quality checks and repairs), and a data storage layer (saving processed data to a unified data platform). The automated pipeline supports both real-time and batch processing modes, adapting to different data types and urgency levels.

[0380] Through the above multi-source data collection and standardized processing, the system has constructed a unified and standardized security data foundation, providing high-quality input data for subsequent security incident analysis and correlation detection. This standardized multi-source heterogeneous data not only supports real-time security monitoring but also provides valuable resources for long-term security trend analysis and pattern mining.

[0381] Fourth, based on standardized multi-source heterogeneous data, advanced security analysis is conducted to identify security incidents and assess their severity, ultimately implementing tiered response strategies. This step embodies the core value of security monitoring, transforming raw data into actionable security insights and automated protective measures.

[0382] Temporal correlation analysis is a key technique for identifying hidden relationships between scattered events. It employs a sliding time window technique to collect potentially related events within a configurable timeframe (typically from seconds to hours). Temporal analysis utilizes multiple correlation patterns: sequence patterns identify combinations of events in a specific order, such as a privilege escalation attempt immediately following an authentication failure; frequency patterns identify anomalous changes in event frequency, such as a large number of database queries within a short period; and periodic patterns identify regular patterns in event occurrence, such as data export activities at specific times each day. The system also implements a time decay mechanism, giving more weight to recent events in correlation analysis and more accurately reflecting the current security status.

[0383] Causal reasoning is a technique for analyzing deep logical relationships between events. A causal knowledge graph in the field of healthcare security was constructed, describing the cause-and-effect chains of typical security events. Based on this knowledge graph, the system applies a Bayesian network model to calculate whether observed event sequences conform to the causal structure of known attack patterns. Causal reasoning pays particular attention to the attack chain perspective, integrating discrete events into coherent attack scenarios, such as a complete chain from initial reconnaissance and credential acquisition to privilege escalation and data theft. This causal-based analysis significantly improves the accuracy of security alerts and reduces false alarms caused by isolated events.

[0384] The anomaly detection engine is a core component for discovering unknown threats. It deploys multiple complementary anomaly detection algorithms: statistical anomaly detection, based on historical baselines, identifies behaviors that deviate from normal distributions; machine learning detection applies supervised and unsupervised learning techniques to automatically learn normal and anomalous patterns from data; and rule engine detection uses expert-defined rule sets to capture known dangerous behavioral characteristics. Anomaly detection employs an integrated strategy, combining the results of multiple detectors to improve the detection rate and reduce the false alarm rate. The system also implements detector performance monitoring and automatic tuning to continuously optimize anomaly detection performance.

[0385] Security incident correlation analysis integrates the results of temporal correlation, causal reasoning, and anomaly detection to generate advanced security incidents. First, correlation analysis reduces data redundancy through event aggregation, merging multiple low-level alerts expressing the same security incident into a single event representation. Then, impact analysis is applied to assess the scope of assets involved and the potential damage. Finally, context enhancement is performed, extracting relevant information from environmental data and threat intelligence to enrich the incident description. The output of correlation analysis is a structured security incident description, including information such as incident type, scope of impact, evidence links, and severity assessment.

[0386] Threat severity and response urgency assessments employ a multi-factor scoring model. The assessment considers six key dimensions: asset value (business importance and sensitivity of affected assets), threat level (malicious intent and technical complexity exhibited by the incident), vulnerability exposure (effectiveness and deficiencies in system safeguards), scope of impact (range of systems and data potentially affected), business impact (potential disruption to healthcare operations), and recovery difficulty (resources and time required to recover from the damage caused by the incident). The system applies a weighted scoring algorithm to generate standardized severity scores (typically 0-100) and urgency ratings (from low to very high).

[0387] Security incident classification is based on severity and urgency assessments, categorizing incidents into four levels: Observational incidents are potential issues with low severity and urgency, such as a single authentication error or a brief network anomaly; these incidents are documented but typically do not require immediate response. Alert incidents are issues of low to medium severity that require the security team's attention but do not constitute an urgent threat, such as data access at atypical times or minor security configuration deviations. Warning incidents are explicit threats of medium to high severity, requiring timely response and investigation, such as persistent authentication attacks or anomalous transmission of sensitive data. Critical incidents are serious security incidents with high severity and urgency, requiring immediate response and a comprehensive investigation, such as clear evidence of data breaches or breakthrough intrusions into critical systems. The classification criteria are customized based on the organization's risk tolerance and business needs, allowing the security team to adjust the thresholds according to actual circumstances.

[0388] Automated response measures implement differentiated strategies based on event level. Three types of response actions are supported: Protective actions directly block or restrict threat activity, such as blocking suspicious IPs, isolating infected devices, or locking suspicious accounts; Investigative actions gather more evidence to confirm the threat, such as capturing complete network traffic, initiating memory forensics, or extracting suspicious file samples; and Notification actions communicate event information to relevant personnel, such as generating security alerts, sending email notifications, or triggering security team calls. Automated response is executed based on pre-configured response policies, which define the mapping between event types and response actions. For observation-level events, only logging and summarizing are typically performed; alert-level events trigger low-priority notifications and basic investigations; warning-level events activate proactive protective measures and detailed investigations; and emergency-level events trigger a full response, including highest-priority notifications, immediate protective actions, and in-depth investigations.

[0389] Balancing human intervention with automation is a crucial design principle for response systems. Different levels of automation are set for different event levels: low-level events (observation and alert levels) are highly automated to minimize human intervention; high-level events (warning and emergency levels) employ a "human-in-the-loop" model, where the automated system executes the initial response, but key decision points require confirmation from security analysts to avoid potential business interruption risks from automated responses. Security constraints on automated responses are also implemented, defining the scope of impact and execution conditions for different response actions to ensure that automated actions do not cause excessive disruption.

[0390] Response effectiveness evaluation is a crucial step in continuous improvement. Track the execution results and business impact of each response action to assess its effectiveness and appropriateness. Evaluation metrics include response time (the delay from detection to action), threat control rate (the percentage of threats successfully blocked or limited), false alarm impact (business disruption caused by erroneous responses), and analyst satisfaction (the security team's evaluation of automated responses). Based on this evaluation data, continuously optimize response strategies and execution logic to improve response accuracy and efficiency.

[0391] The accumulation and feedback loop of security knowledge serves as the system's self-improvement mechanism. Each security incident and its response are recorded as a structured case, including incident characteristics, analysis process, response measures, and final results. These cases are integrated into a security knowledge base to improve detection rules, optimize response strategies, and train analysts. An active learning mechanism is also implemented, adjusting algorithm parameters and decision thresholds based on analyst feedback on automated analysis and responses, gradually bringing the system's behavior closer to expert judgment.

[0392] Through the aforementioned security incident correlation analysis and tiered response mechanism, a closed-loop process from raw security data to actionable protective measures is achieved, providing comprehensive and precise protection for medical data security. The anomaly detection and risk warning data continuously generated by the security monitoring network not only support daily security operations but also provide a solid data foundation for security posture assessment and strategy optimization.

[0393] Example 13

[0394] like Figure 3 As shown, the present invention also provides a trusted medical data circulation system based on multi-level permission dynamic management, comprising:

[0395] The multi-dimensional user identity and permission model construction module 10 is used to classify medical data and establish a data label system based on multi-dimensional features including medical institutions, user roles, data attributes and access scenarios, construct user permission feature vectors, and perform dynamic permission determination in combination with contextual factors to obtain a multi-dimensional user identity and permission model.

[0396] The permission-matching medical data retrieval module 20 is used to extract and vectorize features from medical data based on the multi-dimensional user identity and permission model, construct a distributed index structure by applying a multi-pattern matching algorithm, extract user permission mapping relationships in the multi-dimensional user identity and permission model, filter and sort the retrieval results, and obtain permission-matching medical data retrieval results.

[0397] The permission dynamic evaluation and adjustment module 30 is used to collect user access behavior data of medical data retrieval results matching the permissions, construct a user-permission-data relationship graph, apply graph algorithm to evaluate the user-permission-data relationship graph in real time, detect abnormal permission usage patterns, generate permission adjustment suggestions and execute dynamic adjustment strategies to obtain the permission dynamic evaluation and adjustment results.

[0398] The security monitoring network anomaly detection and risk warning module 40 is used to construct a medical data circulation security threat model based on the dynamic evaluation and adjustment results of the permissions, optimize the layout of monitoring points, collect network traffic data, system logs, user behavior data and environmental data from each monitoring point and integrate them into multi-source heterogeneous data, perform security event correlation analysis on the multi-source heterogeneous data, establish a hierarchical early warning and automatic response mechanism, and obtain anomaly detection and risk warning data of the security monitoring network.

[0399] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for trusted circulation of medical data based on multi-level dynamic access control, characterized in that, include: Based on multidimensional features including medical institutions, user roles, data attributes and access scenarios, medical data is classified and a data tagging system is established. User permission feature vectors are constructed, and dynamic permission determination is carried out in combination with contextual factors to obtain a multidimensional user identity and permission model. Based on the multi-dimensional user identity and permission model, medical data is feature-extracted and vectorized. A distributed index structure is constructed using a multi-pattern matching algorithm. The user permission mapping relationship in the multi-dimensional user identity and permission model is extracted to filter and sort the search results, resulting in permission-matched medical data search results. This includes: acquiring structured and unstructured medical data; extracting key features of the medical data using natural language processing and image recognition technologies; converting these features into standardized feature vectors to obtain medical data feature vectors; indexing the data using a multi-pattern matching algorithm based on the medical data feature vectors to obtain a multi-pattern matching index; constructing a distributed index structure containing a primary index, secondary indexes, and temporary indexes based on the multi-pattern matching index; establishing an index routing layer; determining the index path based on query features, user roles, and system load to obtain a distributed index system; combining the distributed index system and the multi-dimensional user identity and permission model, extracting the user permission mapping relationship in the multi-dimensional user identity and permission model; dynamically filtering and security-de-identifying the search results based on the user permission mapping relationship; and performing personalized sorting based on user preferences and historical behavior to obtain the permission-matched medical data search results. Collect user access behavior data for medical data retrieval results matching the permissions, construct a user-permission-data relationship graph, apply graph algorithms to evaluate the user-permission-data relationship graph in real time, detect abnormal permission usage patterns, generate permission adjustment suggestions and execute dynamic adjustment strategies to obtain dynamic permission evaluation and adjustment results; Based on the dynamic evaluation and adjustment results of the permissions, a medical data circulation security threat model is constructed, the layout of monitoring points is optimized, network traffic data, system logs, user behavior data and environmental data are collected from each monitoring point and integrated into multi-source heterogeneous data, security event correlation analysis is performed on the multi-source heterogeneous data, a hierarchical early warning and automatic response mechanism is established, and anomaly detection and risk warning data of the security monitoring network are obtained.

2. The method according to claim 1, characterized in that, The method classifies medical data and establishes a data tagging system based on multidimensional features including medical institutions, user roles, data attributes, and access scenarios. It then constructs a user permission feature vector and performs dynamic permission determination by combining contextual factors, resulting in a multidimensional user identity and permission model, including: The data attributes and access scenarios in the multidimensional features are obtained. Based on the data attributes, the medical data is classified into sensitivity levels and business value levels. Based on the access scenarios, the medical data is classified into usage scenarios. A multi-level medical data classification standard is constructed, a unified data tag system is established, and medical data classification tags are obtained. Combining the medical data classification labels and the medical institutions and user roles in the multidimensional features, the user's affiliated medical institution, user professional role, scope of responsibilities and qualification level are extracted to construct a user permission feature vector. Based on the matching relationship between the medical data classification labels and the user permission feature vector, permission mapping is performed to obtain the user permission feature vector and permission mapping relationship. The user permission feature vector and permission mapping relationship are obtained. User access time, access location, device type, and network environment are collected as contextual factors. Combined with historical access behavior patterns, machine learning algorithms are applied to construct a context-permission association model, calculate the permission credibility score in the current context, make dynamic permission judgment decisions, and obtain context-aware permission judgment results. Based on the context-aware permission determination results, a full lifecycle management framework including permission application, approval, use, supervision, and revocation is established to achieve traceable permission management and automatic expiration mechanism, thereby obtaining the multi-dimensional user identity and permission model.

3. The method according to claim 2, characterized in that, The construction of a multi-level medical data classification standard and the establishment of a unified data labeling system result in medical data classification labels, including: Medical data is acquired and classified into public, internal, sensitive and confidential levels based on its sensitivity; basic data, business data and core data based on its business value; and diagnosis and treatment data, scientific research data and management data based on its usage scenario, resulting in multi-level data classification results. Based on the multi-level data classification results, standardized labels containing data type, sensitivity level, business attributes, and access requirements are established for each type of data to obtain the medical data classification labels.

4. The method according to claim 2, characterized in that, The construction of the user permission feature vector involves mapping permissions based on the matching relationship between the medical data classification labels and the user permission feature vector, resulting in the user permission feature vector and the permission mapping relationship, including: The medical institution and user role in the multidimensional features are obtained, and the user's medical institution, professional role, scope of responsibility, qualification level and department level are extracted. The features of each dimension are quantified and encoded, and the feature fusion technology is used to obtain the user permission feature vector. Based on the user permission feature vector and the medical data classification label, the matching degree between the user permission feature vector and the medical data classification label is calculated, and a permission mapping matrix between users and data is established to obtain the user permission feature vector and the permission mapping relationship.

5. The method according to claim 2, characterized in that, The application uses machine learning algorithms to construct a context-permission association model, calculates the permission credibility score in the current context, performs dynamic permission determination decisions, and obtains context-aware permission determination results, including: Obtain the user permission feature vector and permission mapping relationship, collect user access time, access location, device type, network environment, extract context features and perform standardization processing to obtain the context feature vector; Based on the context feature vector and historical access behavior data, machine learning algorithms are applied to establish the correlation between context factors and the rationality of permission use, calculate the permission credibility score in the current context, and obtain the context-adaptive permission score. Based on the context-adaptive permission score and the preset permission judgment threshold, a dynamic permission judgment decision is made to generate a judgment result of granting, denying, or requiring secondary verification, thus obtaining the context-aware permission judgment result.

6. The method according to claim 1, characterized in that, The application of a multi-pattern matching algorithm for data indexing yields a multi-pattern matching index, including: Based on the medical data feature vectors and query complexity, a sliding window with dynamically adjustable size is maintained. A window suffix tree is constructed only for the medical data feature vectors within the sliding window to obtain a dynamic window suffix tree index. The medical data feature vector is divided into multiple data shards, and a suffix tree index is constructed in parallel for each data shard. A work-stealing load balancing strategy is adopted to obtain a set of parallel suffix tree indexes. Based on the dynamic window suffix tree index and the parallel suffix tree index set, a two-layer index structure is constructed using the Aho-Corasick automaton. The outer layer uses the parallel suffix tree index set to index the dataset, while the inner layer uses the Aho-Corasick automaton to simultaneously match multiple patterns, resulting in the multi-pattern matching index.

7. The method according to claim 1, characterized in that, The construction of a distributed index structure comprising primary indexes, secondary indexes, and temporary indexes, the establishment of an index routing layer, and the determination of index paths based on query characteristics, user roles, and system load result in a distributed index system, including: Based on the multi-pattern matching index, a master index is constructed for the disease type, diagnosis and treatment methods and drug treatment characteristics of medical data. The master index is stored in a distributed database cluster using a consistent hashing sharding strategy to obtain a global master index. Based on the multi-pattern matching index and business scenario requirements, secondary indexes are constructed for data subsets of predetermined departments, predetermined disease areas, or predetermined research directions, and deployed on edge nodes to obtain scenario-based secondary indexes. Based on the needs of hot queries or temporary tasks, a short-term index is built using in-memory data structures and dynamic window features, and an automatic expiration and clearing mechanism is set to obtain a temporary index. By integrating the global primary index, the scenario-based secondary index, and the temporary index, an index routing layer is established. Based on query features, user roles, and system load, a machine learning model is applied to evaluate the index path, resulting in the distributed index system.

8. The method according to claim 1, characterized in that, The application graph algorithm performs real-time evaluation of the user-permission-data relationship graph, detects abnormal permission usage patterns, generates permission adjustment suggestions, and executes dynamic adjustment strategies to obtain dynamic permission evaluation and adjustment results, including: Obtain the medical data retrieval results matching the permissions, collect the user's access behavior data on the medical data retrieval results matching the permissions in real time, construct the user-permission-data relationship graph containing user nodes, permission nodes and data nodes, and obtain the permission usage behavior relationship graph; Based on the permission usage behavior relationship graph, a graph algorithm is applied to calculate node weights and path weights, and the permission usage behavior relationship graph is evaluated in real time to obtain permission usage evaluation results. Based on the permission usage evaluation results and the permission usage behavior relationship diagram, a normal permission usage pattern feature library is constructed. The deviation between the current user's permission usage pattern and the normal permission usage pattern feature library is calculated. The degree of medical urgency, business stage characteristics, and organizational policy changes are introduced as contextual factors to dynamically adjust the anomaly judgment threshold. The deviation is compared with the anomaly judgment threshold, and the anomaly is classified according to the comparison results to obtain the permission usage anomaly detection results. Based on the abnormal permission usage detection results, the permission usage evaluation results are integrated to generate permission adjustment suggestions. Based on the permission adjustment suggestions, an approval process is automatically generated and the permission adjustment is executed to obtain the dynamic evaluation and adjustment results of the permissions.

9. The method according to claim 8, characterized in that, The application graph algorithm calculates node weights and path weights, performs real-time evaluation of the permission usage behavior graph, and obtains permission usage evaluation results, including: Based on the permission usage behavior relationship graph, key nodes are identified using node centrality and feature vector centrality, and the subgraph formed by the key nodes is extracted to obtain the key node subgraph. Based on the key node subgraph, user-permission granting weight, permission-data coverage weight, user-data access weight and time decay factor are defined. The proportion of each type of weight in the total weight is dynamically adjusted according to the characteristics of the medical business process to obtain adaptive weight parameters. Based on the adaptive weight parameters, the relevant path weights are recalculated only for the newly added or changed behavior data in the permission use behavior relationship graph. The weight distribution is estimated by using the random walk approximation technique to obtain the approximate weight calculation results. Based on the approximate weight calculation results, a distributed computing model is used to distribute the computing tasks to multiple computing nodes, and the results are aggregated through a message passing mechanism to obtain the permission usage evaluation results.

10. The method according to claim 8, characterized in that, The process involves constructing a normal permission usage pattern feature library, calculating the deviation between the current user's permission usage pattern and the feature library, and dynamically adjusting the anomaly detection threshold by incorporating factors such as medical urgency, business stage characteristics, and organizational policy changes as contextual factors. The deviation is then compared with the anomaly detection threshold, and anomalies are classified based on the comparison results to obtain permission usage anomaly detection results, including: Based on the permission usage evaluation results and the permission usage behavior relationship diagram, historical behavior data is extracted, and the permission usage normal mode feature library is constructed for different roles, different departments and different business scenarios. The permission usage normal mode feature library includes statistical features, sequence features, correlation features and content features. The normal behavior distribution is represented by a Gaussian mixture model to obtain the permission usage benchmark model. Based on the permission usage benchmark model and the current user permission usage pattern, calculate the deviation between the current user permission usage pattern and the permission usage benchmark model, evaluate the deviation of single behavior, the deviation of sequential behavior and the deviation of pattern, and obtain the permission usage deviation score. Based on the permission usage deviation score, the severity of medical urgency, business stage characteristics, and organizational policy changes are introduced as contextual factors. The anomaly judgment threshold is dynamically adjusted through preset rules and machine learning models. The permission usage deviation score is compared with the anomaly judgment threshold to obtain the context-corrected anomaly score. Based on the anomaly score after context correction, abnormal behaviors are categorized into observation level, alert level, intervention level, and blocking level. For anomalies at the intervention level and above, evidence information including operation context, system status, and user response is automatically collected to obtain the permission usage anomaly detection results.

11. The method according to claim 1, characterized in that, The process involves constructing a medical data circulation security threat model, optimizing the layout of monitoring points, collecting network traffic data, system logs, user behavior data, and environmental data from each monitoring point and integrating them into multi-source heterogeneous data. Security event correlation analysis is then performed on this multi-source heterogeneous data to establish a tiered early warning and automatic response mechanism, resulting in anomaly detection and risk warning data for the security monitoring network, including: Based on the dynamic evaluation and adjustment results of permissions, user nodes, data nodes, access paths and permission relationships in the medical data circulation process are extracted, a medical data circulation network topology is constructed, the types of security threats in the medical data circulation network topology are identified, and attack tree and attack graph methods are used to simulate attack paths and quantify risk levels to obtain a medical data circulation security threat model. Based on the attack paths and risk levels in the medical data circulation security threat model, the medical data circulation network topology is abstracted into a directed weighted graph. The cut set algorithm is applied to calculate and optimize the location of monitoring points to obtain a monitoring point layout scheme. According to the monitoring point layout scheme, network traffic data, system logs, user behavior data and environmental data are collected from each monitoring point, and standardized processing is performed to obtain the multi-source heterogeneous data. Based on the aforementioned multi-source heterogeneous data, time-series correlation analysis and causal reasoning are applied to perform security event correlation analysis, assess the severity of threats and the urgency of responses, classify security events into observation level, alert level, warning level and emergency level, and execute automated response measures for different levels to obtain anomaly detection and risk warning data of the security monitoring network.

12. A trusted medical data circulation system based on multi-level permission dynamic management, characterized in that, include: The multi-dimensional user identity and permission model construction module is used to classify medical data and establish a data tag system based on multi-dimensional features including medical institutions, user roles, data attributes and access scenarios, construct user permission feature vectors, and perform dynamic permission determination in combination with contextual factors to obtain a multi-dimensional user identity and permission model. The permission-matching medical data retrieval module is used to extract and vectorize features from medical data based on the multi-dimensional user identity and permission model, construct a distributed index structure using a multi-pattern matching algorithm, and filter and sort the retrieval results by extracting user permission mapping relationships from the multi-dimensional user identity and permission model to obtain permission-matched medical data retrieval results. This includes: acquiring structured and unstructured medical data; extracting key features from the medical data using natural language processing and image recognition technologies, converting them into standardized feature vectors to obtain medical data feature vectors; indexing the data using a multi-pattern matching algorithm based on the medical data feature vectors to obtain a multi-pattern matching index; constructing a distributed index structure containing a primary index, secondary indexes, and temporary indexes based on the multi-pattern matching index; establishing an index routing layer; determining the index path based on query features, user roles, and system load to obtain a distributed index system; and combining the distributed index system and the multi-dimensional user identity and permission model to extract user permission mapping relationships from the multi-dimensional user identity and permission model, dynamically filtering and security-de-identifying the retrieval results based on the user permission mapping relationships, and performing personalized sorting based on user preferences and historical behavior to obtain the permission-matched medical data retrieval results. The permission dynamic evaluation and adjustment module is used to collect user access behavior data of medical data retrieval results matching the permissions, construct a user-permission-data relationship graph, apply graph algorithms to evaluate the user-permission-data relationship graph in real time, detect abnormal permission usage patterns, generate permission adjustment suggestions and execute dynamic adjustment strategies to obtain the permission dynamic evaluation and adjustment results. The security monitoring network anomaly detection and risk warning module is used to construct a medical data circulation security threat model based on the dynamic evaluation and adjustment results of the permissions, optimize the layout of monitoring points, collect network traffic data, system logs, user behavior data and environmental data from each monitoring point and integrate them into multi-source heterogeneous data, perform security event correlation analysis on the multi-source heterogeneous data, establish a hierarchical early warning and automatic response mechanism, and obtain anomaly detection and risk warning data of the security monitoring network.