Systems and methods to produce clinical trial site distribution model

A MINLP model optimizes clinical trial site and country selection, addressing inefficiencies in traditional methods by providing rapid, objective, and compliant site distribution for clinical trials.

JP2025155873APending Publication Date: 2025-10-14MEDIDATA SOLUTIONS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025018910
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2025-02-07
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Traditional methods for selecting clinical trial sites and countries rely on manual processes, subjective decision-making, and unsophisticated tools like spreadsheets, leading to inefficiencies, errors, and suboptimal site selection due to fragmented data, human error, and inability to handle complex data analysis, which can result in delays and compliance issues.

Method used

A mathematical mixed integer nonlinear program (MINLP) is used to optimize the selection of clinical trial sites and countries, minimizing the number of sites and countries while meeting enrollment targets and regulatory requirements, using open-source frameworks to solve the model in seconds.

Benefits of technology

The MINLP approach provides efficient, objective site and country selection, reducing the time and cost of clinical trial planning by generating optimal scenarios in seconds, compared to traditional methods that take days or weeks, and ensures compliance with regulatory and demographic considerations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025155873000001_ABST
    Figure 2025155873000001_ABST
Patent Text Reader

Abstract

To provide a computer-implemented method, an institution distribution model generation system, and a computer-readable medium for providing a clinical trial distribution model for identifying a set of clinical trial sites for satisfying operational requirements of a clinical trial.SOLUTION: A method for generating a site distribution model includes defining an objective function as a first element indicating whether a site is included in a clinical trial and a second element indicating whether a country is included in the clinical trial, receiving a first condition group that must be satisfied by the site distribution model, implementing the site distribution model based on the objective function and the first condition group by using an optimization modeling language, analyzing the site distribution model to generate values for a site decision variable and a country decision variable, notifying a user of the fact when the analysis fails, and generating a list of clinical trial sites from the site decision variable if the site distribution model can be generated.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technique for generating a facility distribution model for identifying a group of facilities that meets the operational requirements of a clinical trial. [Background technology]

[0002] Traditional approaches to selecting specific countries and sites for clinical trials often involve manual processes, reliance on limited data, and subjective decision-making, which can lead to several problems. In some cases, existing solutions for clinical trial planning require multiple iterative edits using unsophisticated tools such as spreadsheets to arrive at an acceptable distribution of clinical trial locations and countries. Relying on heuristics or past experience to select clinical trial sites can introduce subjectivity and potential bias into the process. In some cases, decisions are driven by personal preferences or anecdotal evidence rather than objective and comprehensive data analysis, leading to suboptimal site selection. Summary of the Invention [Problem to be solved by the invention]

[0003] The use of spreadsheets and similar unsophisticated or informal data management and calculation tools to manage complex data sets can lead to fragmentation and inconsistencies. Information may be spread across multiple files and formats, making it difficult to obtain a comprehensive view of the data. This fragmentation can lead to errors in data interpretation and decision-making. Spreadsheets, especially when numerous and complex, pose challenges to effective collaboration and sharing among stakeholders involved in the site selection process. Version control issues arise, with different team members working on different versions of documents, potentially leading to confusion and discrepancies.

[0004] As clinical trials grow in size, managing data and processes in spreadsheets becomes increasingly unwieldy. The manual effort required to update, analyze, and maintain spreadsheets for numerous sites and countries is substantial and inefficient. Manual data entry and analysis in spreadsheets is also prone to human error, and simple mistakes such as incorrect data entry, applying the wrong formula, or copy-paste errors can have significant impacts and lead to flawed analyses and decisions.

[0005] Additionally, spreadsheets are typically static and do not easily integrate real-time data updates. This means that decision-making processes may be based on outdated information, which is particularly problematic in the dynamic field of clinical trials, where conditions and regulations change rapidly. Furthermore, while spreadsheets offer analytical capabilities, they are limited in their ability to handle the complex, multidimensional analysis required for optimal site selection. They also do not easily allow for advanced statistical analysis, predictive modeling, or scenario simulation. Clinical trials It does not provide greater insight into the site's likelihood of success.

[0006] Traditional approaches may analyze historical clinical trial site and country recruitment data and attempt to come up with scenarios that reflect the median number of sites and countries. Study planners may then manually review the site list and select sites based on heuristics and past experience. This process can take weeks or even months, requiring hundreds of hours of iterating through countless potential scenarios. Thus, manual processes and traditional decision-making approaches are time-consuming and do not easily scale to the complexity and scope of large, multinational clinical trials. As the number of potential sites and countries increases, the ability of traditional methods to efficiently process and evaluate all the necessary information decreases. Traditional Clinical trialsThe manual nature of the site selection process can introduce significant inefficiencies and delays. Evaluating each site and coordinating among different stakeholders can take time, lengthening the overall timeline of the trial set-up phase. Such delays can have a significant impact on the success of a trial.

[0007] Furthermore, traditional methods may not be able to fully integrate or effectively analyze the vast amounts of data needed to make informed decisions, including epidemiological data, regulatory information, past facility performance, patient demographics, etc. Without comprehensive data integration, decisions may be made based on incomplete information, potentially resulting in suboptimal facility selection.

[0008] Traditional approaches often rely heavily on decision-makers' experience and intuition, which can introduce subjectivity and bias into the facility selection process. While experience is valuable, over-reliance on it without sufficient support from objective data can lead to inconsistent and biased results. Furthermore, human-generated scenarios are likely to be suboptimal, potentially having to choose from over 2,000 relevant facilities across more than 30 countries. In reality, there are simply too many potential site-country combinations for a human to accurately evaluate all possible scenarios.

[0009] Traditional methods make it difficult to keep up with the ever-changing regulatory landscape in various countries. There is a risk of overlooking important regulatory changes and requirements, which can result in compliance issues and Clinical trials The various stakeholders (sponsors, CROs, regulators, Clinical trials Coordinating information and decision-making among multiple stakeholders (including site coordinators) can be cumbersome using traditional methods. Ineffective communication can lead to misunderstandings, misaligned goals, and delays.

[0010] Traditional methods typically do not offer real-time decision support or adaptive planning capabilities, and in the dynamic environment of clinical trials, where conditions and information change rapidly, the inability to quickly adapt can be a significant drawback. [Means for solving the problem]

[0011] The disclosed embodiments relate to solving the problem of determining which sites and countries to use to conduct testing of a new drug within the context of a clinical trial. More specifically, the embodiments described herein output the number and specific names of sites and countries that need to be involved in order for the clinical trial to achieve its operating parameters.

[0012] Disclosed embodiments solve the problem of planning the distribution of sites and countries for a specific clinical trial by formulating a mathematical mixed integer nonlinear program (MINLP). The objective function of the MINLP is to minimize the number of sites and countries in a scenario. The MINLP's constraints include achieving a target enrollment, selecting a minimum number of sites per country, requiring specific countries to be present in the scenario, and maintaining a user-specified balance between top, middle, and bottom sites based on historical enrollment. In embodiments, the MINLP is solved using open-source frameworks such as COIN-OR Branch and Cut (CBC) or Basic Open-source Nonlinear Mixed Integer Programming (BONMIN), and the output is presented in various formats to allow for action. The output provides a list of sites and countries that meet all user criteria and minimize the clinical trial's geographic footprint.

[0013] The disclosed embodiments generate scenarios in seconds instead of days, a major improvement in the fast-paced world of clinical trial design.

[0014] In one aspect, the disclosed embodiments provide a method, system, and computer-readable medium for generating a site distribution model for identifying a set of sites to meet operational requirements for a clinical trial. The method includes accessing a database of clinical trial sites, where each site is designated by a site identifier and an associated country identifier. The database includes a site estimate of the site's estimated cumulative trial enrollment, e i Contains data from.

[0015] The method further includes defining an objective function, the objective function including finding the minimum value of the sum of at least a first element and a second element, the first element being a facility decision variable z i The facility decision variable z is the iterative sum of the i Each of the variables corresponds to a facility identifier and has a discrete value indicating whether the facility specified by the corresponding facility identifier is included in the clinical trial. The second element is a country determination variable c, where the index value j ranges from 1 to the total number of countries, C. j The iterative sum of the two variables is multiplied by the second weighting factor β. j Each of the , corresponds to a country identifier and has a discrete value that indicates whether the country specified by the corresponding country identifier is included in the clinical trial.

[0016] The method further includes receiving a first set of conditions that must be met by the site distribution model, the first set of conditions including (i) an estimated cumulative trial enrollment, e i (ii) that the estimated total enrollment, determined at least in part based on the first set of criteria, reaches a defined target enrollment, and (ii) for each facility designated as included, an associated country is designated as included. The method further includes generating, using an optimization modeling language, computer code for implementing the facility distribution model based at least in part on the objective function and the first set of criteria.

[0017] The method further includes solving the facility distribution model to generate values ​​for the facility decision variables and the country decision variables, if possible, and otherwise indicating to the user that a solution is not possible. The method further includes generating a list of clinical trial sites from the generated values ​​of the site decision variables if the site distribution model is solvable.

[0018] Embodiments can include one or more of the following features, either alone or in combination. In receiving the first set of conditions, the first set of conditions may further include (i) not exceeding a maximum number of defined facilities, and (ii) not exceeding a maximum number of defined countries. The method may further include receiving a second set of conditions that the facility distribution model seeks to satisfy, and generating computer code for implementing the facility distribution model based at least in part on the objective function, the first set of conditions, and the set of second set of conditions.

[0019] The second set of criteria may include a defined ratio between the number of sites in at least the first tier and the number of sites in at least the second tier. The at least the first tier and the at least the second tier may be defined based at least in part on historical enrollment data of the sites and may include sites ranked lower according to enrollment volume and sites ranked higher according to enrollment volume, respectively. The second set of criteria may include a defined set of countries to be included in the clinical trial. The second set of criteria may include a defined minimum number of sites per country that a solution must satisfy. The second set of criteria may include a defined maximum number of sites per country that a solution must satisfy. Analyzing the site distribution model may include: (i) instantiating a model object including the objective function, the first set of criteria, the site decision variables, and the country decision variables, and (ii) passing the model object to a analyzer.

[0020] The facility distribution model may include a Mixed Integer Non-Linear Program (MINLP). The optimization modeling language generates the MINLP in the form of a Python program. Indicating to the user that a solution is not possible may include indicating to the user to modify one or more constraints in the first set of conditions and the second set of conditions. The method may further include, if the facility distribution model is solvable, generating a predicted enrollment timeline based at least in part on the list of clinical trial sites from the generated values ​​of the facility decision variables and the estimated facility cumulative enrollment data. [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 illustrates a system for generating a site distribution model for identifying a group of sites to meet operational requirements for a clinical trial, according to a disclosed embodiment. [Figure 2] FIG. 10 plots a predicted enrollment timeline based on an example facility distribution model solution. [Figure 3] FIG. 10 plots multiple iterations of an example solution of a facility distribution model for different target registration times. [Figure 4] FIG. 10 is a plot of the optimal number of facilities for multiple iterations of the solution of the facility distribution model for different target registration times. [Figure 5] FIG. 1 illustrates a method for generating a site distribution model to identify a group of sites that meet the operational requirements of a clinical trial. DETAILED DESCRIPTION OF THE INVENTION

[0022] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the present invention. However, it will be understood by those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail.

[0023] Selecting specific countries for a clinical trial and determining the number of sites within each country is a multifaceted technical problem, requiring sophisticated technical solutions due to the complex and significant factors involved. For example, a Phase III clinical trial in breast cancer requires the recruitment of 200 patients within 24 months. Furthermore, it requires the use of up to 180 sites across 20 countries, with at least five and a maximum of 30 sites within each country. Furthermore, it may require the exclusion of sites in China and the Czech Republic, but the inclusion of sites in Brazil and Japan. Furthermore, it requires the selection of sites based on historical and / or estimated enrollment data, including those with a high 25%, a mid-50%, and a low 25% enrollment rate. Thus, the site selection process is not arbitrary but requires a careful balance of scientific, regulatory, logistical, and demographic considerations, each of which presents its own technical challenges, as discussed in more detail below.

[0024] Clinical trial regulatory requirements, administered by the U.S. Food and Drug Administration (FDA), Europe's European Medicines Agency (EMA), and other agencies around the world, vary from country to country. Ensuring compliance with these regulations requires in-depth technical knowledge, and failure to do so can result in significant delays, increased costs, or even invalidation of trial results. The more countries participating in a clinical trial, the greater the regulatory burden, which can result in inefficiencies and delays in the clinical trial timeline.

[0025] Identifying and accessing appropriate patient populations is critical for successful clinical trials. This requires understanding the epidemiology of the disease or condition being studied, including its prevalence and demographic characteristics in various regions and countries. Technological tools, data processing and storage infrastructure, and statistical methods are used to analyze demographic data and disease incidence, and to ensure that selected sites have sufficient eligible participants.

[0026] The ability of a particular site in a particular country to effectively conduct a clinical trial depends on the site's infrastructure, personnel, and clinical trial experience. Site evaluation and selection involves a technical assessment of their operational capabilities, past performance, and the quality of the data they generate. This includes quantitative assessments and site audits that are technical in nature.

[0027] Clinical trials Drug transport, supply chain management, multiple international Clinical trials Spanning facilities Clinical trials Logistics such as ensuring data integrity are complex technical challenges that require advanced planning and coordination, often supported by specialized software and systems designed for clinical trial management.

[0028] Addressing these challenges typically requires a combination of advanced software tools, data analytics, simulation models, and decision support systems. These technological solutions help:

[0029] ·Identification of appropriate patient populations through epidemiological data analysis.

[0030] ·Assessment of regulatory landscape and compliance requirements in various jurisdictions.

[0031] ·Evaluation and selection of facilities based on various performance and capacity indicators.

[0032] · Planning and management of logistics and supply chains with the help of specialized software.

[0033] ·Ensuring data integrity and standardization through EDC systems and data management protocols.

[0034] As such, clinical trial country and site selection involves a complex web of technical considerations that require specialized knowledge, tools, and systems, making it a technical problem that requires a comprehensive technical solution to ensure the success of a clinical trial in terms of regulatory compliance, patient recruitment, data integrity, and overall operational efficiency.

[0035] It is a type of mathematical optimization or decision-making problem that involves both continuous and discrete variables and has at least one nonlinear element in the objective function or constraints. MINLP problems include:

[0036] Objective function: The function to be minimized or maximized. Objective functions can be nonlinear and can include both continuous and integer variables.

[0037] Variables: A problem can contain two types of variables (called "decision variables" here): Continuous variables can take on any value within a given range, while integer variables can only take on integer values ​​within a specified range. The values ​​of the variables that optimize the objective function (and satisfy the conditions and bounds) are the "solution" of the MINLP.

[0038] Constraints: Conditions that the solution must satisfy. This includes linear and nonlinear equations or inequalities involving variables.

[0039] Bounds: Each variable typically has upper and lower bounds that define the feasible range of values ​​that the variable can take on.

[0040] Solving a MINLP problem requires finding values ​​of continuous and integer variables that optimize an objective function while satisfying all constraints. Nonlinearities and the discrete nature of the variables can result in complex solution landscapes with multiple local optima, making traditional optimization methods less effective or inapplicable. Therefore, a variety of specialized algorithms and heuristics are used to tackle MINLP problems, including branch-and-bound methods, cutting plane methods, and metaheuristic approaches such as genetic algorithms and simulated annealing.

[0041] 1 illustrates a system 100 for generating a site distribution model 145 for identifying a set of sites to meet the operational requirements of a clinical trial. The system includes one or more computer systems or subsystems, such as a model construction subsystem 110 that includes a processor and associated memory and storage. The system 100 further includes a clinical trial site database 120, a portion of which may be selected for a particular clinical trial.

[0042] Clinical trials capable of communicating with the facility data processor 125 Clinical Trial Facility The database 120 may include, for example, Clinical trials Name of the facility and Clinical trials Facility identifier, associated country identifier, and Clinical trials Each facility includes facility-level data. Clinical trials In an embodiment, the information includes facility information. Clinical Trial Facility The site-level enrollment data stored by the database 120 includes, for each study site, an estimated cumulative number of study enrollments over the study period, e i where i is an index value ranging from 1 to the total number of sites N. The estimated cumulative study enrollment e i can be expressed as the estimated cumulative number of patients enrolled by month X, where month X is the target trial duration. Site-level enrollment data may be provided by various predictive models based on historical enrollment data and / or may be provided by the user based on external models and datasets (see, e.g., US2023 / 0026758A1).

[0043] In an embodiment, Clinical Trial Facility The site and country identifiers for each site stored in database 120 may be numeric or alphanumeric codes uniquely associated with a particular physical site located in a particular country. In some cases, separate identifiers may not be necessary if site and country names can be standardized within a set of clinical trial sites to ensure they are uniquely and accurately associated with a particular site.

[0044] Clinical Trial Facility Based on data retrieved from database 120, Clinical trials The facility data processor 125 calculates facility decision variables z with index i having values ​​ranging from 1 to the total number of facilities N. i Generate facility decision variables z i Each of the values ​​corresponds to a site identifier and has a discrete value, eg, a value of 0 or 1, indicating whether the site specified by the corresponding site identifier is included in the clinical trial site distribution. Clinical trials Facility Data Processor 1 25 is also the country decision variable c j where index j ranges from 1 to the total number of countries C. j Each of the corresponds to a country identifier, and the country specified by the corresponding country identifier is Clinical trials Whether it is included in the facility distribution, i.e., Clinical trials It has a discrete value, eg, a value of 0 or 1, that indicates whether any of the facilities in the facility distribution are located in a particular country identified by the country identifier.

[0045] The objective function processor 130 calculates the facility decision variables z i and country decision variable c j and applies an objective function that finds the minimum value of the sum of at least the first and second elements.

number

[0046] Therefore, the objective function can be written as the minimum of the sum of two factors: (i) the total number of sites used multiplied by a user-controlled weight α, and (ii) the total number of countries used multiplied by a user-controlled weight β. In an embodiment, the weights α and β are the clinical trial weights. operational Parameter database 160 These specific parameters, α and β, allow the user to control how “lean” or “spread out” the geographic distribution should be. Specifically, the facility decision variable z i Increasing the weight of (e.g., the first weighting coefficient α) tends to result in a solution in which enrollments are distributed among more facilities, while the country decision variable c j Increasing the weights (e.g., the second weighting factor s) tends to result in solutions in which enrollments are distributed across more countries. The first weighting factor α and the second weighting factor β can be determined based on the relative value of these two objectives in terms of the overall objectives of the study.

[0047] The optimization modeler 140 receives data defining the objective function from the objective function processor 130 and receives a set of conditions based on user input from the condition processor 150. In an embodiment, the user-entered set of conditions may be retrieved from a database 160 of clinical trial operational parameters.

[0048] The condition processor 150 provides a first set of conditions (called "constraints") that must be satisfied by the facility distribution model 145. These conditions are effectively mandatory constraints for generating the facility distribution model 145. Constraints are standard conditions in MINLP problems that must always be satisfied by a feasible solution. These conditions are part of the problem definition and ensure that the solution satisfies the necessary conditions imposed by the problem. The conditions can be linear or nonlinear and include any combination of continuous and integer variables in the problem. Constraints can represent physical laws, capacity limits, etc., and define the feasible solution space by eliminating values ​​or combinations of values ​​that do not satisfy these requirements.

[0049] In an embodiment, the condition processor 150 can further provide a second set of conditions that the facility distribution model 145 attempts to satisfy. Such conditions are, in effect, optional constraints when generating the facility distribution model 145. In general, optional constraints introduce a level of conditional logic or flexibility into the problem. These conditions do not need to be strictly satisfied by all feasible solutions, but can be "turned on," i.e., enabled, under certain conditions, often governed by additional binary or integer variables introduced specifically for this purpose. Optional constraints can be used to model scenarios, decision-dependent situations, and conditional requirements that apply only when a particular decision is made or a condition is met.

[0050] The optimization modeler 140 generates computer code for formulating the facility distribution model 145 based at least in part on the objective function and the first set of criteria using an optimization modeling language. In an embodiment, the mathematical formulation (e.g., the objective function and the first set of criteria) may be computer coded in the form of a Python program using an optimization modeling language (e.g., Pyomo). In an embodiment, the implementation of the facility distribution model 145 is based at least in part on the objective function, the first set of criteria, and the second set of criteria.

[0051] The first set of conditions includes the condition that a defined target enrollment (target_enr) must be reached. To determine whether the target enrollment condition is met, the estimated total enrollment e T is the estimated cumulative study enrollment for the selected sites, e i In an embodiment, the condition processor 150 calculates the value of the condition based at least in part on clinical Estimated cumulative study enrollment numbers from 125 study site data processors i receive the data and, in turn, optimize it - 140. Specifically, the facility decision variable z i For each index value i, where i has a discrete value indicating that the corresponding facility is included in the facility distribution, the estimated cumulative study enrollment e i is the estimated total number of registrations e T is included in the sum of all included facilities to determine the target enrollment condition.

number

number

number

[0052] In embodiments, the second set of conditions may include a predetermined country group to be included in the clinical trial site distribution (i.e., the condition requires that at least one of the sites in the site distribution be included in the predetermined country group), a defined minimum number of sites per country, and / or a defined maximum number of sites per country. These conditions may be expressed, respectively, as follows:

number

number

[0053] The mapper 180 calculates the facility decision variables z generated when solving the facility distribution model 145. i and country decision variable c j The value of the facility decision variable z i Each of the variables c corresponds to a facility identifier and has a discrete value, e.g., a value of 0 or 1, indicating whether the facility specified by the corresponding facility identifier is included in the clinical trial facility distribution. j Each of the values ​​corresponds to a country identifier and has a discrete value, for example, 0 or 1, and the country specified by the corresponding country identifier is Clinical trials Indicates whether the facility is included in the distribution of facilities.

[0054] A list of clinical trial sites is generated from the generated values ​​of the site decision variables. The list is displayed and / or output to the user in various forms, for example, via user interface 185. In an embodiment, the output may include a summary of the criteria used to generate the solution, and statistics regarding the number and location of sites and countries.

[0055] The condition is the facility decision variable z that solves the facility distribution model 145. i and country decision variable c j If a solution is not possible, e.g., due to the value of , being too restrictive to allow, e.g., the user interface 18 5 The user is notified via a notification message. In an embodiment, such notification, i.e., indicating to the user that a solution is not possible, may include indicating to the user to change one or more conditions in the first set of conditions and the second set of conditions. The user can then adjust the conditions and repeat the implementation and solution of facility distribution model 145.

[0056] Figure 2 shows the predicted,probability based on the solution of a hypothetical example facility,distribution model. Registration 14. As described above, once the facility distribution model is solved, the predicted enrollment timeline is calculated by solving the facility decision variables z i These values ​​correspond to the site identifiers of the sites selected for inclusion in the clinical trial site distribution. T , i.e., the total cumulative enrollment is the estimated cumulative study enrollment at the selected facilities, e i may be calculated based at least in part on

[0057] In the example shown in Figure 2, the cumulative number of enrolled patients is plotted against the number of months since the "first patient in" (FPI), i.e., the time when the first patient was enrolled, marking the start of the enrollment period. The dashed line corresponds to a user-specified target enrollment of 200 patients. In this example, the predicted enrollment reaches the target enrollment approximately 12 months after the FPI. Other criteria were also applied, such as a maximum of 150 sites and a maximum of 30 countries. Furthermore, the selection of sites required a certain percentage of sites ranked high, medium, and low in enrollment rates based on historical and estimated enrollment data. Another criterion required at least three sites per selected country.

[0058] FIG. 3 is a plot of the number of facilities based on multiple iterations of the solution of an example facility distribution model for different target enrollment times compared to a conventional example. Continuous lines interconnect data points generated using the techniques described herein for a particular objective function, set of criteria, etc. Each data point is for a different value of the time to target enrollment. The plot shows that as the time to target enrollment increases, the number of facilities required decreases.

[0059] This plot includes two data points for a conventional site selection approach for the same baseline scenario. The first conventional data point shows that with a target enrollment period of approximately 32 months, 630 sites would be selected compared to a solution with 535 sites obtained using the techniques disclosed herein. The second conventional data point shows that with a target enrollment period of approximately 50 months, 500 sites would be selected compared to a solution with 400 sites obtained using the techniques disclosed herein. Thus, the disclosed approach provides for the use of substantially fewer sites than the conventional approach, resulting in significant savings in cost, time, and study duration.

[0060] Figure 4 plots the optimal number of sites based on multiple iterations of an example solution of a site distribution model for different target enrollment periods compared to a specific target enrollment period. The solid line plots the number of sites versus months to target enrollment. Each connected dot on the solid line corresponds to a scenario generated using the techniques described herein. The dashed line corresponds to a user-specified 12-month study period. The same conditions were applied to this hypothetical example as in the example in Figure 2 above. This graph demonstrates that extending the period can significantly reduce the number of sites, serving as a sensitivity analysis to the user-specified period. In the depicted example, the 12-month scenario is feasible. However, if the user had selected 8 or 10 months, it is easy to see from the plot that extending the target enrollment period to 12 months would require significantly fewer sites. This provides valuable guidance to users when selecting clinical trial parameters.

[0061] FIG. 5 illustrates a method 500 for generating a site distribution model to identify a set of sites to meet the operational requirements of a clinical trial. Generally, this method 500 involves converting a user's clinical trial parameters into a scenario that includes optimal sites and countries. In an initial step, the user provides the clinical trial parameters that the output scenario must satisfy. For example, the user may input target patient enrollment, regional preferences, minimum number of sites per country, etc. In an embodiment, the user-entered parameters are converted into a mixed integral nonlinear program (MINLP).

[0062] The user input is supplemented with site-level enrollment data, e.g., number of patients enrolled by month X (where X is the desired study duration). Thus, method 500 includes accessing a database of clinical trial sites, each site designated by a site identifier and associated country identifier (510). The database includes an estimated site's estimated cumulative study enrollment, e i Includes:

[0063] The method 500 further includes defining 520 an objective function that minimizes the sum of at least a first element and a second element, where the first element is a facility decision variable z, where i ranges from 1 to a total number N of facilities. i The facility decision variable z is the iterative sum of the i Each of the variables corresponds to a facility identifier and has a discrete value indicating whether the facility specified by the corresponding facility identifier is included in the clinical trial. The second element is a country decision variable c, with index value j ranging from 1 to the total number of countries, C. j The iterative sum of the two variables is multiplied by the second weighting factor β. j Each of the , corresponds to a country identifier and has a discrete value that indicates whether the country specified by the corresponding country identifier is included in the clinical trial.

[0064] In an embodiment, user-input parameters for the planned study are converted into MINLP conditions. Supported conditions are divided into two categories: "always" and "optional." Accordingly, method 500 further includes receiving 530 a first set of conditions that must be satisfied by the center distribution model. The first set of conditions includes estimated total enrollment (estimated cumulative study enrollment e i The first set of conditions may include: reaching a defined target number of enrollments; and for each facility designated as inclusive, designating an associated country as inclusive. In an embodiment, the first set of conditions may further include not exceeding a maximum number of defined facilities and not exceeding a maximum number of defined countries.

[0065] In embodiments, the site distribution model may receive a second set of criteria to be satisfied, where implementation of the site distribution model is based at least in part on the objective function, the first set of criteria, and the second set of criteria, including a definition of the countries to be included in the clinical trial, a defined minimum number of sites per country that a solution must satisfy, and a defined maximum number of sites per country that a solution must satisfy.

[0066] The method 500 further includes implementing 540 a facility distribution model based at least in part on the objective function and the first set of criteria using an optimization modeling language. For example, the mathematical formulation (e.g., the objective function and the first set of criteria) may be programmed in Python using an optimization modeling language (e.g., Pyomo).

[0067] Method 500 further includes step 550 of solving the site distribution model to generate values ​​for site and country decision variables, if possible. For example, an analyzer such as CBC may be used to solve the formulated MINLP program. If the analysis is not possible, a message indicating to the user that a solution is not possible (555) is displayed if a feasible solution does not exist. In such a case, the user can initiate further iterations and modify clinical trial parameters to find a feasible solution. In an embodiment, the method may include suggesting to the user that the conditions be relaxed and / or the trial duration be extended.

[0068] If a feasible solution exists, an optimal selection of sites and countries (i.e., "raw" results) is output. Thus, method 500 further includes step 560 of generating a list of clinical trial sites from the generated values ​​of the site decision variables if a solution to the site distribution model is possible. In an embodiment, the raw results may be summarized, for example, by determining the total number of sites and countries, the number of sites per country, and plotting a predicted enrollment curve for this scenario. The user can review the summary to assess whether there are additional parameters not accounted for or whether the solution has other shortcomings. In such cases, the user can modify the clinical trial parameters and initiate further iterations. In practice, such iterations are completed in a matter of seconds. This is in contrast to traditional approaches, which require days or weeks to perform iterations.

[0069] Aspects of the present invention may be embodied in the form of a system, a computer program product, or a method. Likewise, aspects of the present invention may be embodied in hardware, software, or a combination of both. Aspects of the present invention may also be embodied as a computer program product stored on one or more computer-readable medium(s) with computer-readable program code embodied thereon.

[0070] The computer-readable medium may be a computer-readable storage medium, which may be, for example, an electronic, optical, magnetic, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof.

[0071] The computer program code in embodiments of the present invention can be written in any suitable programming and / or scripting language. The program code can be executed on a single computer or on multiple computers. The computer can include a processing unit in communication with a computer-usable medium, the computer-usable medium including a set of instructions, and the processing unit designed to execute the set of instructions and / or a trained machine learning algorithm.

[0072] The above discussion is meant to be illustrative of the principles and various embodiments of the present invention. Numerous variations and modifications will become apparent to those skilled in the art once the above disclosure is fully appreciated. It is intended that the following claims be interpreted to embrace all such variations and modifications.

Claims

1. 1. A computer-implemented method for generating a site distribution model for identifying a group of sites that meet operational requirements for a clinical trial, comprising: (1) A database of clinical trial sites, each designated by a site identifier and associated country identifier, containing an estimated cumulative number of study registrations, e i accessing a database containing (2) defining an objective function for finding the minimum value of the sum of at least a first element and a second element, The first factor is a first weighting coefficient α and a facility decision variable z i where index value i ranges from 1 to the total number of facilities, N, The facility decision variable z i each of which corresponds to a facility identifier and has a discrete value indicating whether the facility specified by the corresponding facility identifier is included in the clinical trial; The second factor is a second weighting coefficient β and a country decision variable c j where the index value j ranges from 1 to the total number of countries, C, The country decision variable c j each of which corresponds to a country identifier and has a discrete value indicating whether the country designated by the corresponding country identifier is included in the clinical trial; (3) receiving a first set of conditions that must be satisfied by the facility distribution model, The first group of conditions is Estimated cumulative study enrollment i that the estimated total enrollment, determined at least in part based on the above, has reached a defined target enrollment, and that for each designated facility, an associated country has been designated; (4) generating, using an optimization modeling language, computer code for implementing the facility distribution model based at least in part on the objective function and the first set of criteria; (5) a model analysis step of analyzing the facility distribution model to generate values ​​for the facility decision variables and the country decision variables, if possible, and otherwise notifying a user that the facility distribution model cannot be analyzed; (6) generating a list of clinical trial sites from the generated values ​​of the site determination variables if the site distribution model can be analyzed; 10. A computer-implemented method comprising:

2. The computer-implemented method of claim 1 , wherein the first set of conditions further includes the conditions that a predetermined maximum number of facilities is not exceeded and a predetermined maximum number of countries is not exceeded.

3. receiving a second set of conditions that the facility distribution model is required to satisfy; The computer-implemented method of claim 1 , wherein the generating step is based at least in part on the objective function, the first set of conditions, and the second set of conditions.

4. The computer-implemented method according to claim 3 , wherein the second group of conditions includes a condition that the number of facilities belonging to the first layer and the number of facilities belonging to the second layer are a predetermined ratio.

5. 5. The computer-implemented method of claim 4, wherein the first tier and the second tier are defined based on facility registration history data, the first tier including facilities ranked lower according to the number of registrations, and the second tier including facilities ranked higher.

6. The computer-implemented method of claim 3 , wherein the second set of conditions includes a predetermined group of countries to be included in the clinical trial.

7. The computer-implemented method of claim 3 , wherein the second set of conditions includes a defined minimum number of facilities per country that a solution must meet.

8. The computer-implemented method of claim 3 , wherein the second set of conditions includes a defined maximum number of facilities per country that the solution must satisfy.

9. The model analysis step includes: instantiating a model object including an objective function, a first set of conditions, facility decision variables, and country decision variables; Passing the model object to a parser; The computer-implemented method of claim 1 , comprising:

10. The computer-implemented method of claim 1 , wherein the facility distribution model comprises a mixed integer nonlinear program (MINLP).

11. 11. The computer-implemented method of claim 10, wherein the optimization modeling language generates a MINLP in the form of a Python program.

12. The computer-implemented method according to claim 1 , wherein when notifying the user that the facility distribution model cannot be analyzed, the method instructs the user to change one or more conditions in the first condition group and the second condition group.

13. 10. The computer-implemented method of claim 1, further comprising: if the facility distribution model is solvable, generating a predicted enrollment timeline based at least in part on a list of clinical trial sites from the generated values ​​of the facility decision variables and the estimated facility cumulative enrollment data.

14. A facility distribution model generation system for generating a facility distribution model for identifying a group of facilities that meets operational requirements for a clinical trial, a computer having one or more processors and a memory storing instructions executable by the one or more processors; The processor: (1) A database of clinical trial sites, each designated by a site identifier and associated country identifier, containing an estimated cumulative number of trial registrations, e i accessing a database containing (2) defining an objective function for finding the minimum value of the sum of at least a first element and a second element, The first factor is a first weighting coefficient α and a facility decision variable z i where the index value i is a number between 1 and the total number of facilities, N, The facility decision variable z i each of which corresponds to a facility identifier and has a discrete value indicating whether the facility specified by the corresponding facility identifier is included in the clinical trial; The second factor is a second weighting coefficient β and a country decision variable c j where the index value j ranges from 1 to the total number of countries, C, The country decision variable c j each of which corresponds to a country identifier and has a discrete value indicating whether the country designated by the corresponding country identifier is included in the clinical trial; (3) receiving a first set of conditions that must be satisfied by the facility distribution model, The first group of conditions is Estimated cumulative study enrollment i that the estimated total enrollment determined at least in part based on the above has reached a defined target enrollment, and that for each facility designated as included, the relevant country has been designated as included; (4) generating, using an optimization modeling language, computer code for implementing the facility distribution model based at least in part on the objective function and a first set of criteria; (5) a model analysis step of analyzing the facility distribution model to generate values ​​for the facility decision variables and the country decision variables, if possible, and otherwise notifying a user that the facility distribution model cannot be analyzed; (6) generating a list of clinical trial sites from the generated values ​​of the site determination variables if the site distribution model can be analyzed; A facility distribution model generation system that causes the computer to execute the above.

15. The facility distribution model generation system according to claim 14 , wherein the first group of conditions further includes the conditions that the number of facilities does not exceed a predetermined maximum number, and that the number of countries does not exceed a predetermined maximum number.

16. receiving a second set of conditions that the facility distribution model is required to satisfy; The facility distribution model generation system of claim 14 , wherein the generating step is based at least in part on the objective function, the first set of conditions, and the second set of conditions.

17. 17. The facility distribution model generation system according to claim 16, wherein the second group of conditions includes a condition that the number of facilities belonging to the first hierarchical level and the number of facilities belonging to the second hierarchical level are a predetermined ratio.

18. 18. The facility distribution model generation system of claim 17, wherein the first tier and the second tier are defined based on facility registration history data, the first tier including facilities ranked lower according to the number of registrations, and the second tier including facilities ranked higher.

19. 1. A non-volatile computer-readable medium storing instructions that, when executed by one or more computer-implemented processors, cause the one or more processors to perform a method for generating a site distribution model for identifying a set of sites that meet operational requirements for a clinical trial, the method comprising: (1) A database of clinical trial sites, each designated by a site identifier and associated country identifier, containing an estimated cumulative number of trial registrations, e i accessing a database containing (2) defining an objective function for finding the minimum value of the sum of at least a first element and a second element, The first factor is a first weighting coefficient α and a facility decision variable z i where the index value i is a number between 1 and the total number of facilities, N, The facility decision variable z i each of which corresponds to a facility identifier and has a discrete value indicating whether the facility specified by the corresponding facility identifier is included in the clinical trial; The second factor is a second weighting coefficient β and a country decision variable c j where the index value j ranges from 1 to the total number of countries, C, The country decision variable c j each of which corresponds to a country identifier and has a discrete value indicating whether the country designated by the corresponding country identifier is included in the clinical trial; (3) receiving a first set of conditions that must be satisfied by the facility distribution model, The first group of conditions is Estimated cumulative study enrollment i that the estimated total enrollment determined at least in part based on the above has reached a defined target enrollment, and that for each facility designated as included, the relevant country has been designated as included; (4) generating, using an optimization modeling language, computer code for implementing the facility distribution model based at least in part on the objective function and a first set of criteria; (5) a model analysis step of analyzing the facility distribution model to generate values ​​for the facility decision variables and the country decision variables, if possible, and otherwise notifying a user that the facility distribution model cannot be analyzed; (6) generating a list of clinical trial sites from the generated values ​​of the site determination variables if the site distribution model can be analyzed; 1. A computer-readable medium comprising:

20. The computer-readable medium of claim 19 , wherein the first group of conditions further includes the conditions that a defined maximum number of facilities is not exceeded and a defined maximum number of countries is not exceeded.

21. receiving a second set of conditions that the facility distribution model is required to satisfy; 20. The computer-readable medium of claim 19, wherein the generating step is based at least in part on the objective function, the first set of conditions, and the second set of conditions.

22. 22. The computer-readable medium of claim 21, wherein the second group of conditions includes a predetermined ratio between the number of facilities belonging to the first tier and the number of facilities belonging to the second tier.

23. 23. The computer-readable medium of claim 22, wherein the first tier and the second tier are defined based on facility registration history data, the first tier including facilities ranked lower according to number of registrations, and the second tier including facilities ranked higher.