Clinical test resource scheduling collaborative management system based on risk assessment

By constructing a risk assessment-based collaborative management system for clinical trial resource scheduling, and utilizing digital twin technology and knowledge graphs for real-time risk assessment and resource scheduling, the problem of lagging resource scheduling in clinical trials has been solved, thereby improving trial quality and safety.

CN120954716APending Publication Date: 2025-11-14JIANGSU SHIYAN PHARMACEUTICAL TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511080095.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies lack the ability to perceive dynamic risks in real time throughout the entire clinical trial process. Resource scheduling lags behind risk evolution, leading to resource waste, hindered trial progress, and reduced data reliability.

Method used

A risk assessment-based collaborative management system for clinical trial resource scheduling is constructed. Through data integration, risk assessment analysis, and resource scheduling management modules, digital twin technology and knowledge graphs are used for real-time mapping and risk probability updates to generate accurate resource scheduling plans.

Benefits of technology

It enables timely and accurate assessment of clinical trial risks, identifies hidden risk patterns, improves trial quality and safety, optimizes resource allocation, and reduces the probability of risk occurrence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954716A_ABST
    Figure CN120954716A_ABST
Patent Text Reader

Abstract

The invention discloses a clinical test resource scheduling collaborative management system based on risk assessment, and relates to the technical field of clinical test management, the clinical test resource scheduling collaborative management system comprises a clinical test collaborative management platform, and the clinical test collaborative management platform is in communication connection with the following modules: a data integration module, a resource scheduling module and a resource scheduling module. The data collection module is used for collecting clinical test related data from different systems, preprocessing the collected data and integrating the data into a multi-source data set in a unified format. According to the method, the physical world is mapped to the virtual space in real time through the digital twinning technology, and the risk probability is dynamically updated in combination with the association rule of the knowledge graph, so that the system can accurately evaluate the possibility of occurrence of the current risk, a timely and accurate decision basis is provided for risk management of a clinical test, and the risk management efficiency is improved. The problems of hysteresis and subjectivity caused by the fact that a traditional risk assessment method depends on historical data or expert experience are effectively avoided, and therefore the constantly changing risk situation in a clinical test can be better coped with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of clinical trial management technology, and specifically to a risk assessment-based collaborative management system for clinical trial resource scheduling. Background Technology

[0002] Clinical trials involve numerous stages, including trial design, participant recruitment, investigational drug management, and data collection and analysis. Each stage faces different risks, such as participant safety risks, trial delay risks, and data quality risks. Furthermore, clinical trials require the integration of resources from multiple parties, including researchers, equipment, funding, and facilities. The rational allocation and coordinated operation of these resources are crucial for the success of the trial. However, in the actual implementation of clinical trials, resource allocation is often constrained by various factors, such as differences in resource needs at different trial stages and communication and coordination barriers between participants. Without an effective resource allocation and collaborative management mechanism, problems such as resource waste, trial delays, and reduced data reliability may occur, ultimately affecting the overall quality and results of the trial.

[0003] In existing technologies, traditional risk assessments often rely on historical data or expert experience, lacking the ability to perceive the dynamic risks throughout the entire clinical trial process in real time. Moreover, risk indicators are scattered across different systems, making it difficult to form a global relational view, resulting in resource scheduling lagging behind risk evolution. Therefore, how to construct a network of entity relationships in clinical trials, integrate multi-source heterogeneous data, realize semantic association analysis of risk factors, and use digital twin technology to map the physical world of the trial to the virtual space in real time, verify the scheduling strategy recommended by the knowledge graph, and realize the scheduling management of clinical trial resources is the problem to be solved by this invention. To this end, a collaborative management system for clinical trial resource scheduling based on risk assessment is proposed. Summary of the Invention

[0004] The purpose of this invention is to provide a risk assessment-based collaborative management system for clinical trial resource scheduling to address the problems mentioned in the background section.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0006] A risk-assessment-based clinical trial resource scheduling and collaborative management system includes a clinical trial collaborative management platform, which is communicatively connected to the following modules, wherein:

[0007] The data integration module is used to collect clinical trial-related data from different systems and preprocess the collected data, including cleaning, transformation and standardization, to eliminate noise and inconsistencies in the data and integrate the scattered data into a multi-source dataset in a unified format.

[0008] The risk assessment and analysis module, based on multi-source datasets and constructed knowledge graphs, identifies various entities in clinical trials and their relationships, constructs a network of relationships between clinical trial entities, and analyzes it to uncover hidden risk patterns.

[0009] The risk association analysis module is used to construct a digital twin model of clinical trials using digital twin technology, and analyze the current risk probability by combining the association rules of knowledge graphs.

[0010] The resource scheduling management module is used to generate specific clinical trial resource scheduling plans based on the current risk probability and the scheduling strategy recommended by the knowledge graph, and to verify the generated resource scheduling plans and evaluate their feasibility and effectiveness in real-world scenarios.

[0011] A further improvement to the technical solution of the present invention is that the data integration module includes:

[0012] It establishes connections with various clinical trial-related systems, including subject management systems, trial process monitoring systems, equipment operation monitoring systems, and environmental data acquisition systems. Through preset interface protocols, it extracts the required multi-source data from each system, including basic subject information, operation records at each stage of the trial, equipment operating parameters, and environmental temperature and humidity, while also recording the data source and acquisition time.

[0013] The collected multi-source data is preprocessed, including data cleaning, data transformation and standardization. The time of each system is unified to UTC time zone, time windows are divided according to trial stage, subject ID is used as the core, trial process timeline, equipment usage records and environmental data are linked, and the preprocessed multi-source data is integrated, partitioned according to trial stage, and the scattered data are linked together to build a complete clinical trial data view and form a multi-source dataset.

[0014] A further improvement to the technical solution of the present invention is that the risk assessment and analysis module includes a relationship network construction unit and a latent risk pattern mining unit;

[0015] The relationship network construction unit, based on multi-source datasets to centralize clinical trial-related data, identifies various entities in clinical trials, analyzes their relationships, and constructs a clinical trial entity relationship network using knowledge graph technology. This network links dispersed risk indicators to form a global view, intuitively displaying the mutual influence between various elements during the trial process.

[0016] The implicit risk pattern mining unit is used to mine the integrated data and entity relationship network using natural language processing and semantic analysis techniques, analyze the semantic relationships between risk indicators, and conduct in-depth mining of the clinical trial entity relationship network through graph neural network technology to discover implicit risk patterns.

[0017] A further improvement to the technical solution of the present invention is that the relational network construction unit includes:

[0018] Based on multi-source datasets, this study uses natural language processing and rule engine technology to automatically identify core entities in clinical trials, including subjects, researchers, equipment, drugs, test indicators, and adverse events. Through terminology standardization mapping (MedDRA coding), it unifies the naming differences of entities in different systems, establishes unique identifiers, ensures entity consistency across data sources, and defines entity attributes and classification levels.

[0019] This study employs a BERT-based machine learning model for relation classification combined with domain knowledge rules to extract entity relationships from structured and unstructured text in multi-source datasets. Utilizing the knowledge graph technology of the Neo4j graph database, it constructs entity-relation-entity triples to form a dynamically scalable graph model. Identified entities are treated as nodes in the graph database, each containing a unique identifier and attributes. Extracted relations are treated as edges, each connecting two entity nodes and including relation type and weight. Through Neo4j's graph structure, a complete clinical trial entity relationship network is constructed, supporting dynamic expansion and updates.

[0020] By using a graph-based path reasoning algorithm, potential risk transmission paths are uncovered, and scattered risk indicators, including equipment failure rate and adverse event incidence rate, are associated. An interactive global view is generated using visualization tools to display the hierarchical network of relationships between clinical trial entities.

[0021] A further improvement to the technical solution of this invention is that the latent risk pattern mining unit includes:

[0022] The integrated multi-source data and clinical trial entity relationship network are preprocessed. Natural language processing technology is used to perform semantic parsing on unstructured text. Combined with entity attributes and relationship tags in structured data, risk-related semantic features are extracted. Text is converted into semantic vectors through word vector embedding. At the same time, numerical risk-related semantic features are normalized and encoded. Semantic role labeling technology is used to identify the causal and temporal logical relationships in risk descriptions and construct a hybrid semantic feature set covering text and numerical data.

[0023] By using preprocessed entities as nodes and relationships between entities as edges, a dynamic weighted knowledge graph is constructed. A graph attention network is used to aggregate node features. The contribution weights of different relationship types to risk transmission are learned through the attention mechanism. Time dimension information is introduced, and the experimental stage is embedded into the model as an edge attribute to capture the pattern of risk evolution over time. Through multiple rounds of iterative training, the node embedding representation is optimized, and the graph neural network model is trained to form a distinguishable cluster of high-risk entities in the feature space.

[0024] Based on a trained graph neural network model, community detection and subgraph mining are performed on the entire graph to identify high-density risk association regions. The contribution of nodes to risk prediction is calculated through gradient backpropagation technology to locate key risk transmission nodes. The mining results are verified by combining the domain knowledge base to filter out false associations and generate a structured risk pattern report by mining the hidden risk patterns.

[0025] A further improvement to the technical solution of the present invention is that the risk association analysis module includes a digital twin mapping unit and a risk probability update unit;

[0026] The digital twin mapping unit is used to map the physical world of the clinical trial to the virtual space in real time using digital twin technology, and to construct a digital twin model of the clinical trial so that the virtual model can accurately reflect the actual operating status of the trial.

[0027] The risk probability update unit is used to calculate the current risk probability by combining real-time data fed back from the digital twin model and association rules in the knowledge graph.

[0028] A further improvement to the technical solution of the present invention is that the digital twin mapping unit includes:

[0029] By deploying various sensors, monitoring devices, and data interfaces in the physical setting of clinical trials, multi-source heterogeneous data is collected in real time, including real-time physiological data of subjects, operating parameters of experimental equipment, environmental monitoring data, and status information of the experimental process. The collected multi-source heterogeneous data is cleaned to remove noise and outliers, and data formats and coding standards are unified to ensure data quality. Data fusion technology is used to associate and integrate data from different sources and of different types.

[0030] Based on the fused multi-source heterogeneous data, combined with the business logic and rules of clinical trials, a digital twin model of the clinical trial is constructed using digital twin modeling technology, covering all aspects and elements of the trial, including the subject population, trial process, equipment and facilities, etc.

[0031] The constructed digital twin model of the clinical trial is mapped in real time to the actual operating status in the physical world. Through a data synchronization mechanism, real-time data from the physical world is dynamically updated to the virtual model, ensuring that the virtual model can reflect the latest status of the trial in real time.

[0032] A further improvement to the technical solution of the present invention is that the risk probability update unit includes:

[0033] Various types of real-time feedback data are extracted from the digital twin model of clinical trials, covering multi-source information such as the physiological state of subjects, equipment operating parameters, environmental conditions and trial process progress. At the same time, association rules related to risk assessment are obtained from the knowledge graph. The two are integrated to form a risk association dataset.

[0034] Based on the association rules obtained from the knowledge graph, we analyze the various data features fed back in real time in the digital twin model, match and analyze the data features with the association rules, and then fuse the data features involved in the successfully matched rules to extract the feature information that has a key impact on the risk probability calculation and form an effective feature vector.

[0035] Using a pre-defined risk probability calculation model, the fused feature vector is input into the model, and combined with the association rule logic of the knowledge graph, the current risk is quantitatively calculated to obtain the current risk probability.

[0036] A further improvement to the technical solution of the present invention is that the resource scheduling and management module includes:

[0037] The system obtains the current risk probability value and extracts scheduling strategy recommendations that match the risk pattern from the knowledge graph. Combined with the current resource status of the clinical trial, such as the available quantity and location of personnel, equipment, and materials, the system transforms the risk response needs and strategy requirements into specific resource allocation instructions, and initially generates a resource scheduling scheme framework.

[0038] The preliminary resource scheduling plan was fully validated. From the perspective of resource availability, the actual occupancy of the required resources during the scheduling time was checked. From the perspective of time rationality, the scheduling process was evaluated to see if it met the time requirements of the clinical trial. From the perspective of spatial matching, the feasibility of the physical path of resource allocation was confirmed, thereby judging the practical operability of the plan.

[0039] Based on the feasibility verification results, the effectiveness of the resource scheduling scheme is further evaluated, the ability of the resource scheduling scheme to control risks is analyzed, and it is determined whether the possibility of risk occurrence or the degree of risk impact can be effectively reduced. Then, the resource scheduling scheme is optimized and adjusted to determine the final resource scheduling scheme that meets the needs of the actual scenario.

[0040] Due to the adoption of the above technical solution, the technical progress achieved by this invention compared to the prior art is as follows:

[0041] 1. This invention provides a risk assessment-based collaborative management system for clinical trial resource scheduling. By using digital twin technology to map the physical world to a virtual space in real time, and combining the association rules of knowledge graphs to dynamically update the risk probability, the system can accurately assess the likelihood of current risks occurring. This provides timely and accurate decision-making basis for the risk management of clinical trials, effectively avoiding the lag and subjectivity problems caused by traditional risk assessment methods that rely on historical data or expert experience, thus better addressing the ever-changing risk situations in clinical trials.

[0042] 2. This invention provides a risk assessment-based collaborative management system for clinical trial resource scheduling. By constructing a network of entity relationships in clinical trials, integrating multi-source data, and performing semantic association analysis of risk factors, it can link scattered risk indicators to form a global relational view. At the same time, it uses graph neural network technology to deeply mine the entity relationship network, discovering hidden risk patterns. This enables the system to identify risk factors that may have a significant impact on the trial, helping researchers and managers to fully understand the interactions between various elements in the trial process, take measures in advance to deal with potential risks, reduce the probability of risk occurrence, and improve the overall quality and safety of clinical trials. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0044] Figure 1 This is a schematic diagram of the workflow of the present invention;

[0045] Figure 2 This is a schematic diagram of the system functional modules of the present invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] Example 1, as Figure 1 , Figure 2 As shown, this invention provides a risk assessment-based clinical trial resource scheduling and collaborative management system, including a clinical trial collaborative management platform. The clinical trial collaborative management platform has the following communication modules, wherein:

[0048] The data integration module collects clinical trial-related data from various systems, including but not limited to subject information, trial process data, equipment operation data, and environmental data. It preprocesses the collected data, including cleaning, transformation, and standardization, to eliminate noise and inconsistencies. This integrates the scattered data into a unified format multi-source dataset, enabling centralized management and sharing, and improving data availability and integrity. The module establishes connections with various clinical trial-related systems, including subject management systems, trial process monitoring systems, equipment operation monitoring systems, and environmental data acquisition systems. Through pre-defined interface protocols, it extracts the necessary multi-source data from each system, including basic subject information. The system records operational data, equipment operating parameters, and environmental temperature and humidity at each stage of the trial, along with data sources and collection times. Specifically, for the subject management system, subject IDs, demographic information, enrollment / exit times, and group information are extracted via API. For the trial process monitoring system, it interfaces with trial phase logs (e.g., screening, dosing, follow-up) to record operation times, operators, and abnormal events. For the equipment operation monitoring system, it collects real-time instrument status (e.g., centrifuge speed, refrigerator temperature), calibration records, and fault alarms. For the environmental data acquisition system, it integrates data from temperature and humidity sensors and light intensity, labeling the collection location (laboratory, pharmacy) and timestamps, using RESTful APIs. The API enables real-time data transmission, supporting incremental updates and full synchronization. It defines the source system, business meaning, data type, and update frequency for each data field, forming a data dictionary that records data extraction time, success / failure status, and original data snapshots to ensure traceability. Preprocessing operations are performed on the collected multi-source data, including data cleaning, data transformation, and standardization. Data cleaning techniques remove duplicate and erroneous data and fill in missing values. Data transformation unifies data with different codes and units, and standardization eliminates the influence of units of measurement, improving data quality and ensuring consistency in logic and format. The time of each system is unified to the UTC time zone, and time windows are divided according to trial phases. Using the subject ID as the core, the trial process timeline, equipment usage records, and environmental data are linked. The preprocessed multi-source data is integrated, partitioned by trial phase, and the scattered data is linked to construct a complete clinical trial data view, forming a multi-source dataset.

[0049] The risk assessment and analysis module, based on multi-source datasets and constructed knowledge graphs, identifies various entities in clinical trials and their relationships, constructs a relationship network of clinical trial entities, and analyzes it to uncover latent risk patterns. The risk assessment and analysis module includes a relationship network construction unit and a latent risk pattern mining unit.

[0050] The relational network construction unit, based on centralized clinical trial data from multi-source datasets, identifies various entities within clinical trials and analyzes their relationships. It then uses knowledge graph technology to construct a relational network of clinical trial entities, linking dispersed risk indicators to form a global view that intuitively displays the mutual influence between various elements during the trial. Based on multi-source datasets, it employs natural language processing and rule engine technology to automatically identify core entities in clinical trials, including subjects, researchers, equipment, drugs, test indicators, and adverse events. Furthermore, through terminology standardization mapping (MedDRA coding), it unifies the naming differences of entities across different systems, establishing unique identifiers to ensure... Cross-data source entity consistency is ensured, and entity attributes and classification hierarchies are defined. Subjects include individuals participating in the trial, including basic information and participation status at each stage. Researchers include researchers responsible for the trial, including roles and responsibilities. Equipment includes instruments and equipment used in the trial, including equipment name, model, and purpose. Drugs include drugs used in the trial, including drug name, dosage, and method of administration. Examination indicators include various examination indicators collected during the trial, such as blood indicators and imaging indicators. Adverse events include adverse reactions or events occurring during the trial, including event description and severity. A relation-based BERT machine learning model is employed. This approach combines type and domain knowledge rules to extract entity relationships from structured and unstructured text in multi-source datasets. These relationships include subject-use-device, drug-cause-adverse event, and researcher-responsible-trial phase. Subject-use-device indicates that a subject used a certain device in the trial; drug-cause-adverse event indicates that a certain drug may cause a certain adverse event; and researcher-responsible-trial phase indicates that a researcher is responsible for a certain trial phase. Furthermore, using Neo4j graph database's knowledge graph technology, entity-relationship-entity triples are constructed to form a dynamically scalable graph model. The identified entities are used as nodes in the graph database. Each node contains a unique identifier and attributes of the entity. Extracted relationships are used as edges in the graph database. Each edge connects two entity nodes and includes the relationship type and weight. Through the graph structure of Neo4j, a complete clinical trial entity relationship network is constructed, supporting dynamic expansion and updates. Through the path reasoning algorithm of the graph, potential risk transmission paths are mined and scattered risk indicators, including equipment failure rate and adverse event incidence rate, are associated. Interactive global views are generated using visualization tools, displaying the clinical trial entity relationship network in layers. For example, the upper layer displays the main entities and key relationships, while the lower layer displays detailed relationships and risk indicators, and marks key risk nodes and their impact scope.

[0051] The implicit risk pattern mining unit utilizes natural language processing and semantic analysis techniques to mine integrated data and entity relationship networks, analyze the semantic relationships between risk indicators, and deeply mine the clinical trial entity relationship network using graph neural network technology to discover implicit risk patterns. It preprocesses the integrated multi-source data and clinical trial entity relationship network, performs semantic parsing of unstructured text using natural language processing techniques, extracts risk-related semantic features by combining entity attributes and relationship tags from structured data, transforms text into semantic vectors through word embedding, normalizes and encodes numerical risk-related semantic features, and uses semantic role labeling technology to identify causal and temporal logical relationships in risk descriptions. It constructs a hybrid semantic feature set encompassing text and numerical values, and uses the preprocessed entities as nodes. Nodes and relationships between entities are used as edges to construct a dynamic weighted knowledge graph. A graph attention network is used to aggregate node features. The contribution weight of different relationship types to risk transmission is learned through the attention mechanism. Time dimension information is introduced, and the experimental stage is embedded into the model as an edge attribute to capture the pattern of risk evolution over time. Through multiple rounds of iterative training, the node embedding representation is optimized, and the graph neural network model is trained to form distinguishable clusters of high-risk entities in the feature space. Based on the trained graph neural network model, community detection and subgraph mining are performed on the entire graph to identify high-density risk association areas. The contribution of nodes to risk prediction is calculated through gradient backpropagation technology to locate key risk transmission nodes. The mining results are verified by combining the domain knowledge base to filter false associations and generate a structured risk pattern report for the mined implicit risk patterns.

[0052] Community detection is performed based on the Louvain algorithm, and the modularity is calculated for full-graph community partitioning. The expression is as follows:

[0053] ;

[0054] ;

[0055] ;

[0056] In the formula, Modularity is used to measure the quality of community segmentation. Its value ranges from -1 to 1; a higher value indicates a more distinct community structure. This is the sum of the weights of all edges in the graph. For an undirected graph, each edge is calculated only once; for a directed graph, ... , Let be the elements of the adjacency matrix of the graph, representing the nodes. and nodes The edge weight between nodes, if nodes and nodes If there is an edge connecting them, then The weight of the edge; otherwise , For nodes The degree (for an undirected graph) or out-degree (for a directed graph), For nodes Community number to which it belongs Let Kronecker function be used when hour, ;otherwise By maximizing modularity To find the optimal community division, so that the connections within the community are tight and the connections between communities are sparse;

[0057] Subgraph mining is based on density. The subgraph density is defined as follows:

[0058] ;

[0059] In the formula, For subgraph The density of the subgraph ranges from [0,1]. A larger value indicates a tighter connection between nodes in the subgraph. For an undirected subgraph... , For subgraph The number of middle edges, For subgraph The number of nodes in the subgraph; used to measure the density of the subgraph. By setting a density threshold, high-density subgraphs can be identified, corresponding to high-density risk-related areas.

[0060] The contribution of a node to risk prediction is calculated using the gradient backpropagation technique, and its expression is as follows:

[0061] ;

[0062] In the formula, For nodes Features Risk prediction output Contribution As the output of risk prediction, The input features include feature information of nodes in the graph; the nodes are calculated through gradient backpropagation. Features For output The degree of contribution; For output For nodes feature The partial derivatives are calculated through gradient backpropagation. The contribution of each node to risk prediction is measured by calculating the absolute value of the gradient of each node's feature to the risk prediction output. The greater the contribution, the more critical the node is in risk transmission.

[0063] The key risk transmission nodes are located based on a comprehensive assessment of contribution, and the expression is as follows:

[0064] ;

[0065] In the formula, For nodes The overall contribution score is used to measure the node's contribution. The extent to which it serves as a key risk transmission node For the number of nodes, For nodes For nodes The influence weights satisfy the region , For nodes The contribution of each feature to risk prediction is considered by comprehensively taking into account the contribution of all nodes in the graph to the target node. The impact of computing nodes The overall contribution score indicates that the node is more likely to be a key risk transmission node;

[0066] By combining the domain knowledge base, a similarity metric is defined to validate the mining results, and its expression is as follows:

[0067] ;

[0068] In the formula, To uncover risk patterns Known risk patterns in the domain knowledge base The similarity score between the two values ​​ranges from [0,1], with a higher value indicating greater similarity. To uncover risk patterns, For known risk patterns in the domain knowledge base, Risk model and dot product, and Risk models and The norm of the mining is used to calculate the similarity between the mined risk patterns and known risk patterns in the domain knowledge base, thereby filtering out false associations with low similarity to known patterns and verifying the rationality of the mining results.

[0069] The risk association analysis module is used to construct a digital twin model of clinical trials using digital twin technology, and analyze the current risk probability by combining the association rules of knowledge graphs.

[0070] The resource scheduling management module is used to generate specific clinical trial resource scheduling plans based on the current risk probability and the scheduling strategy recommended by the knowledge graph, and to verify the generated resource scheduling plans and evaluate their feasibility and effectiveness in real-world scenarios.

[0071] Example 2, as Figure 1 , Figure 2 As shown, based on Embodiment 1, the present invention provides a technical solution: preferably, the risk association analysis module includes a digital twin mapping unit and a risk probability update unit;

[0072] The digital twin mapping unit utilizes digital twin technology to map the physical world of clinical trials onto a virtual space in real time, constructing a digital twin model of the clinical trial. This allows the virtual model to accurately reflect the actual operational status of the trial, enabling real-time monitoring and simulation of the entire clinical trial process. Through various sensors, monitoring devices, and data interfaces deployed in the physical scene of the clinical trial, it collects multi-source heterogeneous data in real time, including real-time physiological data of subjects, operating parameters of experimental equipment, environmental monitoring data, and status information of the experimental process. The collected multi-source heterogeneous data is cleaned to remove noise and outliers, and data formats and encoding standards are standardized to ensure data quality. Using data fusion technology, data from different sources and of different types are correlated and integrated. Based on the fused multi-source heterogeneous data, combined with the business logic and rules of the clinical trial, a digital twin model of the clinical trial is constructed using digital twin modeling technology. This model covers all aspects and elements of the trial, including the subject population, experimental process, equipment, and facilities. The constructed digital twin model of the clinical trial is mapped in real time to the actual operational status in the physical world. Through a data synchronization mechanism, real-time data from the physical world is dynamically updated to the virtual model, ensuring that the virtual model reflects the latest status of the trial in real time.

[0073] The specific working process of the digital twin mapping unit is as follows:

[0074] Physiological data of subjects is collected in real time through various physiological sensors (ECG sensors, blood pressure sensors, body temperature sensors, etc.) deployed in clinical trial settings. This includes dynamic changes in physiological indicators such as heart rate, blood pressure, and body temperature. Simultaneously, sensors on the experimental equipment (flow sensors, temperature sensors, etc. of drug delivery equipment) are used to collect equipment operating parameters, such as drug delivery flow rate and internal equipment temperature. Environmental monitoring equipment (temperature and humidity sensors, air quality sensors, etc.) is used to acquire relevant data about the experimental environment, including environmental parameters such as temperature, humidity, and air quality at the experimental site. This data is then connected to the clinical trial management system via a data interface to obtain status information of the trial process, such as trial stages, subject enrollment, and trial operation progress. The collected data is filtered to remove noise introduced by sensor errors, electromagnetic interference, and other factors. Statistical methods are used to detect and process outliers. A unified data format and coding standard are established to convert data from different sources and formats into a standard format. Based on the semantics and business logic of the data, correlations are established between data from different sources. The relationship between the physical world and the digital twin model is as follows: For example, the physiological data of the subjects is correlated with the status information of the experimental process to determine the changes in the subjects' physiological indicators at specific experimental stages; the operating parameters of the experimental equipment are correlated with environmental monitoring data to analyze the impact of environmental factors on equipment operation; data fusion algorithms are used to integrate the correlated data to generate a comprehensive dataset; a data synchronization mechanism is established between the physical world and the digital twin model to ensure that real-time data in the physical world can be updated in the virtual model; the fused multi-source heterogeneous data is input into the digital twin model, and the parameters and status of the model are updated according to data changes. For example, the physiological status parameters of individual subjects in the digital twin model are updated according to the real-time physiological data of the subjects; the operating status and performance indicators of the equipment in the model are updated according to the operating parameters of the experimental equipment; and the environmental parameters in the model are updated according to the environmental monitoring data. This allows the digital twin model to reflect the latest status of the experiment in the physical world in real time. Finally, 3D visualization technology is used to present the experimental scene, subjects, equipment, and other elements in the digital twin model, achieving real-time mapping between the physical world and the virtual model.

[0075] The risk probability update unit is used to combine real-time data from the digital twin model and association rules in the knowledge graph to calculate the current risk probability and accurately assess the likelihood of the current risk occurring. It extracts various types of real-time feedback data from the digital twin model of the clinical trial, covering multi-source information such as the physiological state of the subjects, equipment operating parameters, environmental conditions, and the progress of the trial process. At the same time, it obtains association rules related to risk assessment from the knowledge graph and integrates the two to form a risk association dataset. Based on the association rules obtained from the knowledge graph, it analyzes the characteristics of various types of data from the real-time feedback in the digital twin model, matches the data features with the association rules, and then fuses the data features involved in the successfully matched rules to extract the feature information that has a key impact on the risk probability calculation, forming an effective feature vector. Using a pre-set risk probability calculation model, the fused feature vector is input into the model, and combined with the association rule logic of the knowledge graph, the current risk is quantitatively calculated to obtain the current risk probability.

[0076] The specific working process of the risk probability update unit is as follows:

[0077] Through a data synchronization mechanism that maps digital twin models to the physical world in real time, various types of real-time feedback data are extracted from the digital twin model of clinical trials. These data include subjects' physiological states, equipment operating parameters, environmental conditions, and trial progress. Specifically, for subjects' physiological states, real-time data on parameters such as heart rate, blood pressure, body temperature, and blood glucose are extracted. For equipment operating parameters, real-time data on equipment temperature, pressure, power, and fault alarms are extracted. For environmental conditions, environmental monitoring data such as temperature, humidity, air quality, and light intensity are extracted. For trial progress data, trial phases, task completion status, and operation records are extracted. Furthermore, through a knowledge graph query interface, entity relationships and rules related to risk assessment are obtained. The extracted association rules are stored as a structured rule set. Then, based on the entity identifiers in the knowledge graph, the real-time data in the digital twin model is matched with the association rules. The matching process involves integrating the matched data and rules into a risk-related dataset. Based on the association rules obtained from the knowledge graph, the data features fed back in real time from the digital twin model are analyzed to identify risk-related data features. For example, for the rule of equipment failure → increased risk, outliers in equipment operating parameters are identified as risk features. Real-time data features are matched and analyzed with association rules to determine risk-related data features. According to the importance of the association rules, weights are assigned to each successfully matched feature. The weighted features are then fused to form an effective feature vector. A pre-defined risk probability calculation model is used to input the fused feature vector into the model. The risk probability calculation model is built based on a logistic regression model and then combined with the association rule logic of the knowledge graph to perform quantitative calculation of risk probability. Based on the model's output and the association rule logic of the knowledge graph, the current risk probability is calculated, representing the likelihood of the current risk occurring.

[0078] The formula for calculating the current risk probability is as follows:

[0079] ;

[0080] In the formula, This represents the probability of the current risk occurring, ranging from 0 to 1, where 0 indicates no risk and 1 indicates that the risk will definitely occur. The intercept term in the logistic regression model is typically a real number representing the baseline risk level, and represents the probability of baseline risk without any features influencing it. For the first The weight coefficient of the i-th feature represents the weight coefficient of the i-th feature. The degree of influence of each feature on the probability of risk For the first The values ​​of these features represent real-time data features extracted from the digital twin model. The total number of features, For the first The weight coefficient of the i-th association rule represents the weight coefficient of the i-th association rule. The degree of influence of each association rule on the probability of risk For the first The activation value of an association rule indicates whether a certain association rule in the knowledge graph has been triggered (0 indicates that it has not been triggered, and 1 indicates that it has been triggered). This represents the total number of association rules.

[0081] The resource scheduling and management module includes:

[0082] The system acquires the current risk probability value and extracts scheduling strategy recommendations matching the risk pattern from the knowledge graph. Combined with the current resource status of the clinical trial, such as the available quantity and location of personnel, equipment, and materials, it transforms risk response needs and strategy requirements into specific resource allocation instructions, initially generating a resource scheduling scheme framework. This framework is then comprehensively validated. From a resource availability perspective, the actual occupancy of required resources within the scheduling time is examined. From a time rationality perspective, the scheduling process is assessed to determine if it meets the timeline requirements of the clinical trial. From a spatial matching perspective, the feasibility of the physical path for resource allocation is confirmed, thus judging the practical operability of the scheme. Based on the feasibility validation results, the effectiveness of the resource scheduling scheme is further evaluated, and its ability to control risks is analyzed. It is determined whether the scheme can effectively reduce the probability of risk occurrence or mitigate the degree of risk impact. Finally, the resource scheduling scheme is optimized and adjusted to determine the final resource scheduling scheme that meets the needs of the actual scenario.

[0083] The specific working process of the resource scheduling and management module is as follows:

[0084] The system obtains the current risk probability value from the risk probability calculation model and extracts matching scheduling strategy recommendations from a knowledge graph based on the current risk pattern. The knowledge graph stores scheduling strategies including resource allocation suggestions and countermeasures. Then, it retrieves current resource status information from the clinical trial management system, including personnel, equipment, and supplies. Personnel information includes the number of available medical staff, their professional skills, and current task allocation; equipment information includes the number, type, location, and status (whether in use) of available equipment; and supplies information includes the number, type, and storage location of available supplies. This resource status information is stored as a structured dataset. Based on the scheduling strategy recommendations in the knowledge graph, risk response needs are transformed into specific resource allocation instructions, generating a preliminary resource scheduling scheme framework. If the risk of equipment failure is high, instructions for allocating available equipment are generated, including equipment type, quantity, and allocation path. If the risk of physiological abnormalities in subjects is high, instructions for increasing medical staff are generated, including the professional skills of medical staff and allocation path. The preliminary resource scheduling scheme framework includes the specific content, time schedule, and allocation path of resource allocation. The system verifies the practical operability of the resource scheduling scheme from the perspectives of resource availability, time rationality, and spatial matching. For resource availability verification, the actual occupancy of required resources within the scheduling time is checked to ensure resource availability. For time rationality verification, the scheduling process is evaluated to ensure it meets the time requirements of the clinical trial and that the scheduling time is reasonable. For spatial matching verification, the feasibility of the physical path for resource allocation is confirmed to ensure that resources can arrive at the designated location on time. Based on the feasibility verification results, the risk control capability of the resource scheduling plan is evaluated to determine whether it can effectively reduce the probability of risk occurrence or mitigate the degree of risk impact. For example, whether allocating the allocated equipment can effectively reduce the risk of equipment failure, and whether increasing the number of medical staff can effectively cope with the risk of physiological abnormalities in subjects. Based on the evaluation results, the resource scheduling plan is optimized and adjusted to ensure its effectiveness. Based on the optimized and adjusted resource scheduling plan, the final resource scheduling instruction is generated and fed back to the clinical trial management system to execute the resource scheduling operation. Real-time monitoring and adjustment functions for plan execution are provided to ensure smooth resource scheduling. Through a real-time data synchronization mechanism, the execution status of the resource scheduling plan is monitored, including resource allocation progress and task completion status. If deviations or problems are found during the execution process, the resource scheduling plan is adjusted to ensure smooth resource scheduling.

[0085] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A risk assessment-based clinical trial resource scheduling and collaborative management system, comprising a clinical trial collaborative management platform, characterized in that, The clinical trial collaborative management platform has the following communication modules, including: The data integration module is used to collect clinical trial-related data from different systems, preprocess the collected data, and integrate it into a multi-source dataset in a unified format. The risk assessment and analysis module, based on multi-source datasets and constructed knowledge graphs, identifies various entities in clinical trials and their relationships, constructs a network of relationships between clinical trial entities, and uncovers hidden risk patterns. The risk association analysis module is used to construct a digital twin model of clinical trials using digital twin technology, and to analyze the current risk probability by combining the association rules of knowledge graphs. The resource scheduling management module is used to generate specific clinical trial resource scheduling plans based on the current risk probability and the scheduling strategy recommended by the knowledge graph.

2. The clinical trial resource scheduling and collaborative management system based on risk assessment according to claim 1, characterized in that: The data integration module includes: It establishes connections with various clinical trial-related systems, including subject management systems, trial process monitoring systems, equipment operation monitoring systems, and environmental data acquisition systems. Through preset interface protocols, it extracts the required multi-source data from each system, including basic subject information, operation records at each stage of the trial, equipment operating parameters, and environmental temperature and humidity, while also recording the data source and acquisition time. The collected multi-source data is preprocessed, including data cleaning, data transformation and standardization. The time of each system is unified to UTC time zone, time windows are divided according to trial stage, subject ID is used as the core, trial process timeline, equipment usage records and environmental data are linked, and the preprocessed multi-source data is integrated, partitioned according to trial stage, and the scattered data are linked together to build a complete clinical trial data view and form a multi-source dataset.

3. The clinical trial resource scheduling and collaborative management system based on risk assessment according to claim 1, characterized in that: The risk assessment and analysis module includes a relationship network construction unit and a hidden risk pattern mining unit; The relationship network construction unit, based on multi-source datasets that centralize clinical trial-related data, identifies various entities in clinical trials, analyzes their relationships, and constructs a clinical trial entity relationship network using knowledge graph technology, associating dispersed risk indicators to form a global view. The implicit risk pattern mining unit is used to mine the integrated data and entity relationship network using natural language processing and semantic analysis techniques, analyze the semantic relationships between risk indicators, and conduct in-depth mining of the clinical trial entity relationship network through graph neural network technology to discover implicit risk patterns.

4. The clinical trial resource scheduling and collaborative management system based on risk assessment according to claim 3, characterized in that: The relational network construction unit includes: Based on multi-source datasets, this study uses natural language processing and rule engine technology to automatically identify core entities in clinical trials, including subjects, researchers, equipment, drugs, test indicators, and adverse events. Through terminology standardization mapping, it unifies the naming differences of entities in different systems, establishes unique identifiers, and defines entity attributes and classification levels. This study employs a BERT-based machine learning model for relation classification combined with domain knowledge rules to extract entity relationships from structured and unstructured text in multi-source datasets. It then utilizes Neo4j's knowledge graph technology to construct entity-relation-entity triples, forming a dynamically scalable graph model. Identified entities are treated as nodes in the graph database, each containing a unique identifier and attributes. Extracted relations are treated as edges, each connecting two entity nodes and including relation type and weight. Through Neo4j's graph structure, a complete clinical trial entity relationship network is constructed. By using a graph-based path reasoning algorithm, potential risk transmission paths are uncovered, and scattered risk indicators, including equipment failure rate and adverse event incidence rate, are associated. An interactive global view is generated using visualization tools to display the hierarchical network of relationships between clinical trial entities.

5. The clinical trial resource scheduling and collaborative management system based on risk assessment according to claim 4, characterized in that: The hidden risk pattern mining unit includes: The integrated multi-source data and clinical trial entity relationship network are preprocessed. Natural language processing technology is used to perform semantic parsing on unstructured text. Combined with entity attributes and relationship tags in structured data, risk-related semantic features are extracted. Text is converted into semantic vectors through word vector embedding. At the same time, numerical risk-related semantic features are normalized and encoded. Semantic role labeling technology is used to identify the causal and temporal logical relationships in risk descriptions and construct a hybrid semantic feature set covering text and numerical data. By using preprocessed entities as nodes and relationships between entities as edges, a dynamic weighted knowledge graph is constructed. A graph attention network is used to aggregate node features. The contribution weights of different relationship types to risk transmission are learned through the attention mechanism. Time dimension information is introduced, and the experimental stage is embedded into the model as an edge attribute to capture the pattern of risk evolution over time. Through multiple rounds of iterative training, the node embedding representation is optimized, and the graph neural network model is trained to form a distinguishable cluster of high-risk entities in the feature space. Based on a trained graph neural network model, community detection and subgraph mining are performed on the entire graph to identify high-density risk association regions. The contribution of nodes to risk prediction is calculated through gradient backpropagation technology to locate key risk transmission nodes. The mining results are verified by combining the domain knowledge base to filter out false associations and generate a structured risk pattern report by mining the hidden risk patterns.

6. The clinical trial resource scheduling and collaborative management system based on risk assessment according to claim 3, characterized in that: The risk association analysis module includes a digital twin mapping unit and a risk probability update unit; The digital twin mapping unit is used to map the physical world of the clinical trial to a virtual space in real time using digital twin technology, thereby constructing a digital twin model of the clinical trial. The risk probability update unit is used to calculate the current risk probability by combining real-time data fed back from the digital twin model and association rules in the knowledge graph.

7. The clinical trial resource scheduling and collaborative management system based on risk assessment according to claim 6, characterized in that: The digital twin mapping unit includes: By deploying various sensors, monitoring devices, and data interfaces in the physical setting of clinical trials, multi-source heterogeneous data is collected in real time. The collected multi-source heterogeneous data is cleaned, and data fusion technology is used to associate and integrate data from different sources and of different types. Based on the fused multi-source heterogeneous data, combined with the business logic and rules of clinical trials, a digital twin model of the clinical trial is constructed using digital twin modeling technology, covering all aspects and elements of the trial; The digital twin model of the constructed clinical trial is mapped in real time to the actual operating status in the physical world. Through a data synchronization mechanism, the real-time data in the physical world is dynamically updated to the virtual model.

8. The clinical trial resource scheduling and collaborative management system based on risk assessment according to claim 6, characterized in that: The risk probability update unit includes: Various types of real-time feedback data are extracted from the digital twin model of clinical trials. At the same time, association rules related to risk assessment are obtained from the knowledge graph. The two are integrated to form a risk association dataset. Based on the association rules obtained from the knowledge graph, we analyze the various data features fed back in real time in the digital twin model, match and analyze the data features with the association rules, and then fuse the data features involved in the successfully matched rules to extract the feature information that has a key impact on the risk probability calculation and form an effective feature vector. Using a pre-defined risk probability calculation model, the fused feature vector is input into the model, and combined with the association rule logic of the knowledge graph, the current risk is quantitatively calculated to obtain the current risk probability.

9. A risk assessment-based clinical trial resource scheduling and collaborative management system according to claim 8, characterized in that: The resource scheduling and management module includes: The system obtains the current risk probability value and extracts scheduling strategy recommendations that match the risk pattern from the knowledge graph. Combined with the current resource status of the clinical trial, the risk response needs and strategy requirements are transformed into specific resource allocation instructions, and a preliminary resource scheduling scheme framework is generated. The preliminary resource scheduling plan was fully validated. From the perspective of resource availability, the actual occupancy of the required resources during the scheduling time was checked. From the perspective of time rationality, the scheduling process was evaluated to see if it met the time requirements of the clinical trial. From the perspective of spatial matching, the feasibility of the physical path of resource allocation was confirmed, thereby judging the practical operability of the plan. Based on the feasibility verification results, the effectiveness of the resource scheduling scheme is further evaluated, the ability of the resource scheduling scheme to control risks is analyzed, and it is determined whether the possibility of risk occurrence or the degree of risk impact can be effectively reduced. Then, the resource scheduling scheme is optimized and adjusted to determine the final resource scheduling scheme that meets the needs of the actual scenario.

Citation Information

Cited By

  • Device for assessing cerebral hemorrhage risk in perioperative period of cardiovascular surgery and storage medium

    CN121687497A

  • Clinical nursing data management method and system for clinical big data

    CN121983297A