System and method for semantic polyhierarchical data organization and optimized discovery of clinical trials

TrialOptima addresses the limitations of existing clinical trial registries by using an expert-curated, polyhierarchical taxonomy and natural language processing to enhance data organization and search capabilities, enabling efficient and accurate retrieval of clinical trial information from multiple dimensions.

WO2025229692A1PCT designated stage Publication Date: 2025-11-06DIVATE GEETA PATHIK +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IN2025/050714
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-03
Filing Date
2025-05-03
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Existing clinical trial registries face challenges with variable data quality, inconsistent data reporting, inadequate categorization, non-intuitive user interfaces, and limited search capabilities, leading to inefficiencies in data retrieval and discoverability, particularly due to the lack of semantic or polyhierarchical tagging and expert-curated taxonomies.

Method used

The TrialOptima system employs an empirically derived, expert-curated taxonomy (EDGE) with polyhierarchical tagging and natural language processing to transform unstructured trial data into a richly structured format, enabling multi-dimensional search and retrieval capabilities, and incorporates an intelligent search interface for user-friendly exploration.

Benefits of technology

This approach enhances the accuracy and efficiency of clinical trial data discovery by allowing users to explore data from multiple facets, supports complex queries, and adapts to user preferences, improving the reliability and applicability of trial information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2025050714_06112025_PF_FP_ABST
    Figure IN2025050714_06112025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented system for semantic polyhierarchical organization and optimized discovery of clinical trials. Trial records from multiple registries are transformed into a structured format and annotated using an expert-curated Empirically Derived Ground-Truth Encoding (EDGE) taxonomy. Each trial is tagged across multiple semantic dimensions such as therapeutic area, intervention, phase, enabling multi-faceted classification. Users interact through a dynamic interface that supports both structured search filters and a conversational assistant powered by a generative AI model, allowing intuitive, multi-turn query refinement. A natural language processing module interprets user inputs and maps them to relevant taxonomy terms, including synonyms and local expressions. The system updates results in real time and uses semantic annotation for high-speed, context-aware retrieval. An integrated analytics module provides visual summaries and insights supporting decision-making. By combining expert-curated taxonomy with adaptive tagging and intelligent search, the system improves clinical trial discoverability, accuracy, and user experience beyond conventional single-hierarchy registries.
Need to check novelty before this filing date? Find Prior Art

Description

TITLE OF THE INVENTION: SYSTEM AND METHOD FOR SEMANTIC POLYHIERARCHICAL DATA ORGANIZATION AND OPTIMIZED DISCOVERY OF CLINICAL TRIALSFIELD OF INVENTION

[0001] The present invention generally relates to digital information management systems for clinical trial data. More particularly, it relates to a system and method for enhancing the search for clinical trial data by utilizing advanced semantic data organization techniques applied to legacy clinical trial databases and other clinical trial information sources by integrating expert-defined, empirically derived taxonomy with natural language processing, semantic polyhierarchical tagging techniques, and intelligent search mechanisms to enable optimized data retrieval.BACKGROUND OF INVENTION

[0002] Clinical trials are essential for advancing medical science by rigorously evaluating the safety and efficacy of new medical interventions. As per Bhaskar, S. Bala, “Clinical trial registration: A practical perspective.” Indian Journal of Anaesthesia 62(1 ) (2018): 10-15, “The clinical trial represents a prospective study comparing the outcomes of interventions in human participants in the allotted group / groups. Interventions can be in terms of drugs, cells and other biological products, surgical procedures, radiologic procedures, devices, behavioral treatments, process-of-care changes, preventive care including pharmaceuticals, medical devices, and treatment protocols.” These trials require meticulous planning and strict adherence to global health regulations and standards, often involving coordination across multiple sites and the recruitment of appropriate participants. The complexity of managing trials - which may span various regions and involve diverse patient demographics and clinical parameters - necessitates robust data management systems to facilitate intelligent study and analysis of historical clinical trial data.

[0003] According to the World Health Organization’s International Standards for Clinical Trial Registries (2018), “The registration of all interventional trials is ascientific, ethical and moral responsibility.” To promote harmonization in data collection and validation by clinical trial registries, baseline quality standards need to be established and implemented globally. Properly managed clinical trial data not only supports regulatory approvals but also aids in the continuous improvement of medical research and practice. As noted by Viergever, Roderik F., et al. in “The quality of registration of clinical trials: still a problem.” PLOS One 9(1 ) (2014): e84727, there are important advantages to increased transparency in clinical trial conduct and the use of registries. Transparent trial registries improve access to information for healthcare workers, researchers, and patients; allow steps to be taken against publication bias and selective reporting; increase the accountability of those conducting clinical research; and help identify gaps in the health research landscape, thus facilitating better priority-setting in research.

[0004] However, existing clinical trial registries face numerous challenges that significantly impact their utility and effectiveness. One major issue is variable data quality, stemming from differing standards of data entry and maintenance across trials and registries. Many registries suffer from incomplete or inconsistent data entries — since not all trial aspects are mandated for registration — and there is a lack of standardization in data reporting. These inconsistencies complicate comparisons and aggregation of data. The search capabilities of current registries are often rudimentary, lacking advanced filtering options for detailed criteria or complex queries. Inadequate categorization of trial information further hinders quick and efficient data retrieval. For example, a systematic review by Viergever, Roderik F., et al. in “The quality of registration of clinical trials: still a problem.” PLOS One 9(1 ) (2014): e84727 found that only 51.9% of registered clinical trials had complete information on intervention specifics (such as drug name, dose, duration, frequency, and administration route), and only 57.6% of records specified a primary outcome with a meaningful timeframe. Such gaps are crucial, as they affect the reliability and applicability of trial results.

[0005] Accessibility of trial data is another concern. In many cases, while data may technically be available in a registry, it is not easily accessible due to non- intuitive user interfaces or cumbersome navigation systems. Researchers and otherusers often struggle to locate specific trial information. Outdated information remains a frequent problem due to irregular updates, which poses the risk of misleading researchers or patients and can adversely affect ongoing studies. Ensuring data integrity over time is challenging when updates and corrections are not uniformly recorded. Moreover, compliance with global data standards and regulations is inconsistent, hampering international collaboration and limiting the global applicability of the collected trial data. These issues underscore the necessity for improved systems and standards to enhance the utility, accessibility, and efficiency of clinical trial registries in supporting medical research and regulatory assessments.

[0006] The Clinical Trials Registry of India (CTRI) is a specific example of a registry that, while crucial for enhancing transparency in clinical trials, exemplifies many of the aforementioned challenges regarding data quality, accessibility, and searchability. Data entries in the CTRI often lack completeness and accuracy, with frequent discrepancies even in fundamental details like trial registration dates. Such issues can lead to misclassification of trials as prospectively or retrospectively registered, as noted by Chakraborty, Indraneel, and Gayatri Saberwal, in "CTRI requirement of prospective trial registration: Not always consistent." Indian Journal of Medical Ethics 7(4) (2022): 312-314. Additionally, although the CTRI platform is freely accessible, it is hampered by a user interface that is not very intuitive. This complicates efforts by researchers, clinicians, and the public to easily locate specific trial information. The search functionality is notably limited, failing to support complex queries or advanced filtering, which in turn impedes comprehensive data retrieval necessary for tasks like systematic reviews or meta-analyses (as noted by Saberwal, Gayatri in "How to make Clinical Trials Registry-lndia world class." Current Science 124(7) (2023): 785). These deficiencies highlight the need for substantial improvements in the design of CTRI’s interface and search mechanisms to better serve a diverse user base.

[0007] One specific gap in existing trial registries is the lack of real-world, ‘ground-truth’ contextual tagging of trial records. Current registries do not provide semantic or polyhierarchical tags that capture the multiple contexts andrelationships of a trial beyond the basic fields in the database. In practice, this means trial entries are not annotated with expert-validated labels (such as related disease categories, interventions, or outcomes) that reflect real-world clinical context. Consequently, users cannot easily connect one trial to other related trials or concepts unless those relationships are explicitly in the text, and queries using different terminology or exploring multiple facets often fail to retrieve all relevant information. For example, a search using a local or colloquial disease name may miss trials entered under a formal medical term, because the system lacks any mapping of synonyms or contextual metadata. This absence of ‘ground-truth’ contextual tagging leads to limited cross-linking of information and hampers the ability to perform nuanced, multi-dimensional searches across the dataset. Existing registry systems require multiple discrete search operations or rely on manual data inspection to resolve connections across trials, thereby limiting the derivation of meaningful, actionable intelligence from the dataset.

[0008] A number of systems and methods have been disclosed in prior art to improve clinical trial data management and search, but none adequately address the aforementioned shortcomings, particularly the use of a ground-truth based semantic taxonomy for polyhierarchical tagging.

[0009] International Patent Application WO2011089568A1 (Nair et al.) for invention titled “Method for Organizing Clinical Trial Data” discloses a method for managing clinical trial data by gathering information from various sources, eliminating redundancies, and consolidating it into an organized dataset. It introduces a baseline tagging approach using non-indication parameters and creates a disease-specific list of indication parameters categorized into main and sub-indications. Through advanced tagging with these parameters, a multi-layered dataset is produced. However, the tagging system is essentially a traditional linear or flat tagging system. It defines hierarchical categories of indications, but it does not allow a given data element to exist under multiple hierarchies simultaneously. In other words, the system does not provide a semantic layer wherein a single trial entry can be contextually linked to multiple categories or dimensions (such as to adisease category and a treatment modality concurrently). It therefore fails to use any ground-truth expert taxonomy to flexibly tag trials in multiple contexts.

[0010] Korean Patent KR102521963B1 (Ung et al.) for invention titled “Data Classification System and Method for Clinical Trial Discovery” discloses a system and method for normalizing data from diverse sources to a standard format for search and retrieval. The classification approach employs unique code extraction for clinical trial data and normalization of various terminologies to a unified lexicon, coupled with a structured storage system to allow retrieval based on user-defined search terms. While such normalization can improve consistency, the method of converting diverse expressions into a single normalized form can oversimplify the data, potentially stripping away valuable contextual nuances. Moreover, it does not describe any dynamic or semantic tagging mechanism that considers the context in which terms appear. It lacks a flexible tagging system that assigns metadata based on both content and context. In particular, the system cannot capture complex, multi-faceted relationships (e.g., a trial could be related to multiple conditions or categories simultaneously, or local terminologies and synonyms may not be fully recognized in context). The absence of an expert-driven contextual taxonomy means this system cannot fully address real-world tagging needs.

[0011] U.S. Patent US11107559B1 (Wynden et a / .) for invention titled “System and Method for Identifying One or More Investigative Sites of a Clinical Trial and Centrally Managing Data Therefrom” discloses a system aimed at identifying and managing clinical trial sites by generating Health Ontology Mapper (HOM) Boolean logic statements based on codes from biomedical databases. These logic statements are used to search across site-specific and restricted databases to match patient data with trial requirements. Essentially, it uses ontology-driven search logic to find suitable trial sites given certain criteria (often patient-centric criteria or trial requirements). While this approach leverages biomedical ontologies for site matching, it is focused on matching patient or trial criteria to sites and managing site data. It does not provide advanced mechanisms for enhancing the user-driven discovery of trials across multiple facets. Notably, it lacks a rich, interactive search interface with adaptive filtering or multi-criteria refinement. Userscannot iteratively refine search parameters in a flexible, exploratory manner using this system. The emphasis is on back-end matching logic rather than a front-end discovery tool for broad trial data exploration. In short, US11107559B1 does not incorporate the kind of adaptive, user-guided filters that adjust to user preferences or allow exploration of trial data from different contextual starting points (e.g., starting with a condition, a sponsor, or an investigator name and then pivoting to related information). It does not use any polyhierarchical tagging of trial records; each query remains largely a predefined matching exercise rather than an open- ended search across semantically enriched data.

[0012] Indian Patent Application IN201821019402 (Tata Consultancy Services Ltd.) for invention titled “Method and System for Performing a Data Driven Cognitive Clinical Trial Feasibility Analysis” discloses a system that receives trial protocol requirements, identifies relevant metadata from these requirements, and leverages historical and third-party data to assess the feasibility of potential trial sites. It integrates various data sources and uses historical performance metrics to predict which sites would be suitable for a new trial, thereby optimizing site selection for clinical studies. However, its scope is primarily the internal process of site selection and trial feasibility analysis. It does not address the problem of discovering clinical trial information via a user-facing search across digital channels. In particular, it does not disclose a method for end-users (such as researchers or patients) to find clinical trials or trial sites using a flexible search interface. The system described does not incorporate real-time data filtering in a user query context, nor does it handle multi-category search parameters (like searching by disease, location, sponsor, etc. in a single interface). Essentially, while it improves how trial organizers might identify sites, it does not improve how a user might discover trials or trial information from a registry or database perspective. There is no semantic tagging of trial records for public search purposes, and no multi-faceted taxonomy is employed for organizing trial data.

[0013] U.S. Patent US11145390B2 (Merative US LP) for invention titled “Methods and Systems for Recommending Filters to Apply to Clinical Trial Search” discloses a method and system to leverage biomedical ontologies to enhanceclinical trial search by recommending relevant filters or search criteria to users. It uses a knowledge base (e.g., disease ontologies and possibly gene or condition synonyms) to interpret a user’s query or context and suggest additional filters that could narrow or improve the search results. For instance, if a query is related to a broad condition, the system might suggest more specific sub-condition filters, drawing from an ontology of diseases. While it introduces a degree of semantic understanding — recognizing synonyms and hierarchical disease relationships to guide users — its focus is on the recommendation of search filters in a primarily single-hierarchy framework. It does not describe tagging trial records with multiple categories simultaneously; instead, it operates on improving query refinement. The system lacks a comprehensive expert-curated ground-truth taxonomy that is applied directly to the data for multi-dimensional classification. Thus, US11145390B2, while improving search usability through ontology-based suggestions, still does not provide a platform where trial data is pre-organized into a multi-faceted semantic structure for discovery. It relies on filter suggestions rather than fundamentally restructuring the trial database with an expert-defined taxonomy spanning all trial facets.

[0014] U.S. Patent Application US20180046780A1 for an invention titled “Computer-Implemented Method for Determining Clinical Trial Suitability or Relevance” discloses a method that uses machine learning to evaluate patient data and determine the relevance or suitability of clinical trials for a given patient, effectively prioritizing clinical trials for patient recruitment. The system integrates patient-specific information (such as medical history or genomic data) with clinical trial criteria to automatically match and rank trials. While this approach can improve the efficiency of patient-to-trial matching by learning which trials are likely relevant to certain patient profiles, it is fundamentally geared toward a matching algorithm rather than open-ended trial exploration. The method treats the problem as one of filtering trials for a specific patient context, rather than broadly organizing trial data for any user query. It does not utilize an expert-defined taxonomy to tag each trial record with multiple attributes for general search; instead, it largely depends on comparing patient attributes to trial inclusion / exclusion criteria. The absence of an explicit taxonomy or ontology to categorize trial metadata means it doesn’t solvethe discoverability issues of registries for researchers or the public - it is focused on personalized matching and fails to incorporate a ground-truth contextual labeling of trial records that would benefit wider discovery use cases.

[0015] In summary, despite advancements in the prior art, there remains a significant gap in the field. None of the existing systems or methods adequately address the semantic polyhierarchical organization of clinical trial data combined with an expert-curated ground-truth taxonomy and natural language search capabilities. Existing solutions either rely on flat tagging schemas or simple normalization that loses context, or they focus narrowly on specific problems like site matching or patient-trial matching. They do not provide a unified platform where clinical trial records are enriched with multi-dimensional semantic tags (covering, for example, therapeutic area, disease / indication, intervention type, trial phase, site location, sponsor, etc.), derived from expert knowledge, that allows the data to be searched and navigated from any of those dimensions seamlessly. Moreover, they fail to accommodate local expressions and synonyms that real-world users might employ when querying data, and they lack adaptive, user-friendly interfaces for iterative query refinement. There is a need for a novel solution that overcomes these shortcomings by combining a robust expert-defined ontology / taxonomy for trials with a polyhierarchical tagging system and an intelligent search interface. The present invention fulfills this need by providing an integrated system that significantly improves the discovery of relevant clinical trials and associated data across multiple categories and contexts.OBJECTIVES OF THE INVENTION

[0016] It is the objective of the present invention to overcome the aforementioned shortcomings in existing clinical trial registry systems and search tools.

[0017] An objective of the invention is to transform unstructured clinical trial data into a richly tagged, structured format using advanced semantic polyhierarchical tagging, enabling multiple overlapping dimensions.

[0018] A further objective of the invention is to develop and apply an empirically derived, expert-curated taxonomy that incorporates ground-truth knowledge from clinical researchers and field audits. This taxonomy provides the semantic foundation for context-aware tagging and ensures accurate classification across heterogeneous data sources.

[0019] It is another objective of this invention to incorporate intelligent search capabilities that can interpret and execute natural language expressions in user queries.

[0020] Another objective of the invention is to optimize the data retrieval process such that searching for and retrieving clinical trial information becomes faster and more accurate.

[0021] A further objective is to design the system for seamless integration with various external systems and databases, making it adaptable to incorporate data from multiple sources.

[0022] It is an objective of the invention to provide advanced analytical tools and visualizations such as dashboards, charts and graphs, on top of the organized trial data to support complex decision-making.

[0023] A further objective is to ensure a user-centric design for the search and retrieval interface. The system’s user interface should be customizable or adaptive to individual user preferences and professional requirements.

[0024] Yet another objective is to incorporate adaptive learning within the system’s search functionality, enabling the system to learn from user interactions and feedback.

[0025] Another objective is to integrate disparate data sources, including not only multiple trial registries but also related data such as published clinical trialresults, regulatory approvals, or even real-world evidence data streams (for instance, feeds of newly registered trials globally).

[0026] By meeting these objectives, the present invention aims to provide a transformative solution for clinical trial data discovery and analysis that transcends the limitations of current registry systems and search tools.SUMMARY OF THE INVENTION

[0027] The invention forming the subject matter of the instant Application and Specification, herein referred to as “TrialOptima,” is a novel system and method for semantic polyhierarchical data organization and optimized discovery of clinical trials, their regulatory journey and stakeholders. In broad terms, TrialOptima provides a comprehensive platform that takes unstructured or semi-structured clinical trial data from existing trial registries like the CTRI, International Clinical Trials Registry Platform (“ICTRP”) and transforms it into a richly structured, searchable knowledge base. This is achieved through a combination of expert- driven taxonomy design, automated text processing, and an intelligent search interface.

[0028] At its core, TrialOptima integrates an ‘Empirically Derived Ground-Truth Encoding Taxonomy’ - referred to hereinafter as the “EDGE” taxonomy- with a polyhierarchical tagging mechanism powered by natural language processing and expert-defined heuristics. Unlike conventional systems that rely on static ontologies or flat tagging schemas, where tags are used to categorize information without a hierarchical structure or predefined categories, this hybrid approach leverages ‘ground-truth’ data, being data gathered through field-level insights and validated by clinical domain experts. The technical advancement made by TrialOptima lies in enabling each clinical trial record to be simultaneously classified across multiple, interlinked semantic dimensions — such as disease indication, trial phase, sponsor type, geographic site, and intervention strategy — thereby facilitating context-aware discovery and adaptive retrieval in ways not achievable through traditional registry architectures. This polyhierarchical tagging enables the processing architecture to support multi-dimensional indexing and query traversal across various facets (e.g.,disease, intervention, phase, location, sponsor), overcoming the limitations of prior systems that were constrained by single-hierarchy or flat classification models. The use of an empirically derived ground-truth taxonomy ensures that the tagging reflects actual usage and domain consensus, improving the relevance and accuracy of the organization.

[0029] TrialOptima provides a unified solution that significantly improves the discovery of clinical trial information. By leveraging an expert-curated semantic taxonomy (the EDGE taxonomy) and applying it in a polyhierarchical manner to legacy trial data, the system enables users to explore the data in ways that were previously impractical. A given trial can be found through multiple pathways: e.g., a trial might be discovered by searching its condition, or by searching the investigator’s name, or by searching the site location - each approach will surface the same trial due to the interconnected tagging. The natural language understanding capability means users can search in the terms they know (even lay terms or abbreviations), and the system will interpret their intent and fetch relevant results. Through iterative filtering, users can then hone in on exactly what they need, without being overwhelmed by too many results or frustrated by too few irrelevant ones.

[0030] This invention thus marries the depth of expert domain knowledge (through the EDGE taxonomy) with the breadth of machine-driven data processing (through NLP), delivered via a user-friendly interface that encourages exploration and insight. As a result, researchers, clinicians, sponsors, and even patients can rapidly find pertinent trials, identify patterns across trials, and make informed decisions, thereby advancing clinical research and its applications.DESCRIPTION OF FIGURES

[0031] Figure 1 illustrates an overview of the functional architecture of TrialOptima depicting the major modules of the system and how data flows between them.

[0032] Figure 2 illustrates the data processing and search workflow in TrialOptima for translating the query in natural language into a precise search on the tagged data, and refining results in a loop through user applied filters.

[0033] Figure 3 illustrates the hardware and network layout for deploying TrialOptima.

[0034] Figure 4 depicts an example of the user interface of the TrialOptima system, highlighting the multi-faceted search filters and the presentation of summary analytics.DETAILED DESCRIPTION OF THE INVENTION

[0035] The following is a detailed description of the invention as per one preferred embodiment. The embodiments are described in such a way that the disclosure is clearly communicated. The level of detail provided, however, is not meant to limit the expected variations of embodiments; rather, it is intended to include all modifications, equivalents, and alternatives that come within the spirit and scope of the current disclosure as defined by the attached claims. Unless the context indicates otherwise, the term “comprise” and its variants such as “comprises” and “comprising” throughout the specification are to be read in an open, inclusive manner, that is, “including, but not limited to.” When “embodiment” or “an embodiment” is used in this specification, it signifies that a particular feature, structure, or characteristic described in conjunction with the embodiment is present in at least one embodiment. As a result, the expressions “one embodiment” and “in an embodiment” that appear throughout this specification do not necessarily refer to the same embodiment. Furthermore, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner. Unless the content clearly requires otherwise, the singular terms “a,” “an,” and “the” include plural referents. Similarly, unless explicitly stated otherwise, the term “or” is generally used in its inclusive sense (i.e., “and / or”).

[0036] The use of examples or exemplary language (e.g., “such as”) provided with respect to certain embodiments herein is intended to better illuminate theinvention and does not impose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any element that is not explicitly claimed as being essential to the practice of the invention.

[0037] Headings and the abstract of the invention provided herein are for convenience only and do not interpret the scope or meaning of the embodiments. All publications and patent documents cited herein are incorporated by reference to the same extent as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference. Where a definition or use of a term in an incorporated reference is inconsistent or contrary to the definition of that term provided herein, the definition provided herein applies and the definition of that term in the reference does not apply.

[0038] Groupings of alternative elements or embodiments of the invention disclosed herein should not be construed as limitations. Each group member can be referred to and claimed individually or in any combination with other members of the group or other elements found herein. One or more members of a group can be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is herein deemed to contain the group as modified, thus fulfilling the written description requirement for all combinations of group members.

[0039] It should be appreciated that the present invention can be implemented in numerous ways, including as a system, a method, or a device. In this specification, these implementations, or any other form that the invention may take, may be referred to as processes. In general, the order of the steps of the disclosed processes may be altered within the scope of the invention. Various terms used herein are explained below. To the extent a term used in a claim is not defined below, it should be given the broadest definition persons of ordinary skill in the pertinent art have given that term as reflected in printed publications and issued patents at the time of filing.

[0040] The invention is described as per one non-limiting embodiment with reference to Figures. Referring to FIG. 1 , the TrialOptima system (100) is composed of distinct but interrelated modules deployed in a computing environment. These include the following modules:

[0041] Data Acquisition Module (102) is responsible for ingesting clinical trial data from one or more source repositories. In one embodiment, the module accesses static clinical trial records from publicly available registries such as the CTRI, ICTRP etc. Each record (often a trial registration dataset) contains detailed information about a single clinical trial, including fields such as trial title, scientific title, health conditions, interventions, sponsor, site locations, investigators, phase, outcomes, etc. The data acquisition process involves extracting these details from their raw or semi-structured format into a structured format suitable for further processing. Notably, the data acquisition pipeline is flexible and extensible - it can accommodate concurrent or future data sources beyond the initial one. For example, additional sources might include other national trial registries, databases of published trial results, drug approval databases, regulatory agency databases, or academic clinical trial data repositories. The output of this module is cleaned and structured dataset of trial records ready for tagging.

[0042] Polyhierarchical Tagging Module (103) is responsible for the application of its polyhierarchical tagging process to organize the information semantically from the extracted trial data. This module uses the expert-defined EDGE taxonomy to tag each trial record with a variety of labels across multiple categories. A “polyhierarchical” system means that a single data item (in this case, a trial or an attribute of a trial) can be classified into multiple hierarchies or categories simultaneously. The taxonomy provides several top-level categorical dimensions (for example: Therapeutic Area, Indication (Disease / Condition), Intervention Type, Phase of Trial, Sponsor Type, Geography / Location, etc.), each of which has its own internal hierarchy or network of terms. In operation, the tagging module first defines or loads the set of categorical tags corresponding to each top-level category. For instance, one top-level category could be Therapeutic Area, which might include values such as Cardiology, Pulmonology, Oncology, Neurology, etc. Thetherapeutic area can have multiple tags under the same category e.g. a trial on diabetic patients with heart disease is tagged with Cardiology as well as Endocrinology, allowing cross-disciplinary searches using these multiple tags under the same category. Another category might be Type of Study (e.g., Drug (Chemical), Behavioral, Biologies, Cosmetics, etc.), and another could be Sponsor Type (e.g., Industry, Academic, Government, etc.), and so on. These categories and their allowed values are curated by domain experts to reflect meaningful groupings in clinical research. The EDGE taxonomy is seeded from extensive field audits of clinical trial registry sites, coupled with rigorous manual validation of trial records by domain experts. Terms and synonyms identified during these audits are incorporated into the controlled vocabulary only after demonstrating high consistency and inter-annotator reliability. This process ensures the taxonomy reflects authentic, empirically validated terminologies and relationships, enhancing both accuracy and semantic depth. After establishing the categories and their principal values, the tagging module uses a specified set of keywords or phrases related to each value to perform the actual tagging. For example, consider the category Therapeutic Area with one value Pulmonary. Domain experts provide a list of diseases and conditions that fall under the pulmonary therapeutic area, such as Asthma, Chronic Bronchitis, Pulmonary Fibrosis, Pneumonia, Emphysema, and so forth. These terms including their synonyms or variants form the keyword set for the Pulmonary tag. The Polyhierarchical Tagging Module uses either Direct Keyword Matching or NLP-based Semantic Matching. During tagging, the system scans the content of each trial record (for instance, the title or the condition field) for occurrences of any of these keywords. If a match is found (e.g., the trial is about asthma treatment), the trial is tagged with the Pulmonary therapeutic area label. Importantly, because the taxonomy is polyhierarchical, the same trial might also be tagged under other relevant categories if applicable. For example, a single trial could receive multiple therapeutic area tags if it spans multiple organ systems (though in practice most trials have one primary therapeutic area), or it could receive multiple Intervention Type tags if a combination therapy is used. Similarly, for a category like Indication (Disease), which might itself be hierarchical (for example, Oncology as a broad area with sub-categories like Breast Cancer, Lung Cancer, etc.), the system tags the trial at the appropriate level(s) of the hierarchy. A trialinvestigating a drug for metastatic breast cancer might be tagged under the specific indication “Breast Cancer,” which in turn is linked to the broader category “Oncology.” Because of the polyhierarchical design, an indication could potentially link to multiple higher-level categories (some diseases span multiple physiological systems), though typically one disease maps to one therapeutic area. Nonetheless, the system’s structure allows for the flexibility of multiple parent tags if needed (for instance, a trial about a combination therapy could be tagged under multiple intervention categories). This robust system of classification supports complex queries and discovery. Users can filter through large datasets by these tags and find precisely what they are looking for. The tagging module not only looks for direct keyword matches but also employ NLP techniques known in the art such as language models such as BioBERT (Jinhyuk Lee et al., Bioinformatics, Volume 36, Issue 4, February 2020, Pages 1234-1240) to recognize linguistic variations, synonyms, or context. For example, if the keyword list includes “heart attack” under a cardiovascular category, and a trial description uses the term “myocardial infarction,” the system’s NLP component will recognize that as a synonym and still apply the appropriate tag. In this manner, local terminologies or variations in phrasing (including abbreviations or colloquial disease names) are mapped to the standardized taxonomy, ensuring that the tagging is comprehensive and consistent even if the source data uses inconsistent terminology. Furthermore, if the tagging engine encounters previously unseen or ambiguous terms not represented in the EDGE taxonomy, it flags these terms for human-in-the-loop validation. Domain experts review these flagged terms periodically. Once validated, new terms are integrated into the taxonomy store, triggering automatic re-tagging of affected trial records, ensuring continuous and consistent taxonomy evolution. All tags assigned to a trial are stored as metadata linked to that trial’s record. The output of the polyhierarchical tagging module is a richly annotated dataset of trials, where each trial record now carries multiple labels linking it into various taxonomical hierarchies (the EDGE taxonomy structure).

[0043] Search-Optimized Database (104) : After tagging, the processed data is stored in a specialized database that is optimized for search and retrieval. In one embodiment, the structured and tagged data is stored in a search-optimizeddatabase with appropriate tags and key fields. In another embodiment, the data could be stored in a graph database, where each trial is a node connected to tag nodes (allowing many-to-many relationships inherently) - this naturally represents the polyhierarchy as a graph of interconnected nodes. The database design establishes relationships between tags that reflect the complex structure of clinical trial information. For instance, the tags themselves might be interrelated (the hierarchical relations: e.g., “Asthma” is a child of “Pulmonary Therapeutic Area” as is “Upper Respiratory Infection”). These relationships are captured so that traversing the network of information is possible (allowing, for example, a search to be expanded to parent or child categories if needed). The database also undergoes search optimization to support efficient querying. The search optimization scheme supports compound queries as well (e.g., find all trials that have tag A and tag B). Additionally, data verification and cleaning steps are applied at this stage - ensuring consistency and accuracy of the stored data. Any errors introduced during text extraction (like OCR errors or formatting issues) are identified and corrected, if possible, to ensure the integrity of the database. The end result is a networked repository where each trial entry is connected to various tag indices, enabling linked results and multi-level drill-down queries. For example, the database can quickly provide the list of all trials for a given disease, or all trials sponsored by a particular organization, or all trials conducted in a specific state - and because of the interconnected tagging, these queries can be combined or navigated without performing full-text scans across all records, thereby reducing CPU cycles, memory consumption, and I / O overhead during query execution, and enabling faster response times even on large-scale datasets.

[0044] Search Module (105) : The search module is the user-facing query engine of TrialOptima. It provides an interface (e.g., a web-based application or dashboard or a chat interface) where users can enter search criteria and get results from the database. The search module interprets user inputs, which can range from simple keyword queries to structured, field-specific searches. A key feature of the search module is its ability to handle natural language queries. A user might input a query in plain language, such as “phase 3 trials for diabetes in Delhi,” and the search module will parse this into the relevant tagged criteria (e.g., Phase = 3,Condition tag = Diabetes, Location tag = Delhi) and retrieve matching trials. In one embodiment of the interface, the search module presents multiple fields corresponding to the key categories of interest: for example, text boxes or dropdowns for Condition / Disease, Sponsor, Site Location, Investigator Name, etc., and additional filter options for attributes like trial phase, recruitment status, date ranges, etc. A user can fill one or several of these fields to initiate a search. For instance, the user could enter a particular disease in the Condition field and a city name in the Site field if they are interested in trials for that disease in that city. In another embodiment, the search interface is implemented as a chat-based conversational assistant powered by a generative Al model (e.g., a large language model). In this interface, the user may enter a natural language query — such as “Show me interventional lung cancer trials in Mumbai recruiting this year” — and the system interprets the query contextually using semantic mapping to the underlying taxonomy. The Al assistant supports follow-up questions and interactive refinement, offering a guided discovery experience through conversational prompts.

[0045] Upon executing the search, the system (100) queries the Search- Optimized Database (104) for trials matching all the specified criteria (using the tags and structured fields to match, rather than a naive full-text scan). The Search Module (105) then returns an initial set of results - these could be listed as trial titles with brief details in a table or list format. In addition to the list of individual trials, the interface can also display aggregate summaries of the results (as part of the Analytical Module (106), described next). For example, it might show that the query returned 20 trials, with a breakdown by phase (e.g., X number of Phase 1 , Y of Phase 2, etc.) or other relevant aggregates. This gives the user a high-level overview along with the detailed results. Critically, the search module supports iterative filtering and refinement of results. Users can progressively narrow down the results by applying additional filters or adjusting their search criteria, without starting over from scratch. The system (100) might present facets or filter options alongside the results, derived from the taxonomy tags. For example, after an initial search by disease, the interface might show filters like checkboxes for trial phases, locations, sponsor types, etc., populated with the values present in the result set. The user can tick one or more of these filters to refine the result list. This dynamicfiltering is adaptive; the available filter options and their counts update based on the current result subset.

[0046] User Interface (107) : A user operates the system through the User Interface (107). The user starts with a broad query by entering "Diabetes" as the condition. The system (100) returns 100 trials related to diabetes. The User interface (107) reveals that within these results, trials are categorized under various phases (Phase 1 , Phase 2, etc.), various locations, sponsors, etc., each with counts. The user then clicks on “Phase 3” to restrict results to Phase 3 diabetes trials (now perhaps 30 trials), then further filters by location, e.g., “Delhi” (narrowing to 10 trials), and additionally by sponsor type “Industry” (perhaps narrowing to 6 trials). Each step filters the current set of results, narrowing down to those meeting all criteria, and the interface updates immediately. This multi-filter refinement process is done interactively and quickly, thanks to the search-optimized data structure. The approach ensures that the search is both comprehensive (initially casting a wide net) and precise (iteratively narrowing down). The ability to apply many filters in TrialOptima means users have fine-grained control to sift through large datasets and pinpoint studies matching exact requirements. For instance, users can drill down to the level of individual studies conducted by a particular investigator for specific sponsors and indications - something that would be extremely tedious or impossible in traditional registry search interfaces.

[0047] Analytical Module (106): The analytical module provides data analysis and visualization capabilities on top of the search results or the entire dataset. After performing a search or filter operation, users may want to see aggregated data, trends, or summaries rather than just individual trial listings. The Analytical Module (106) taps into the organized data to generate such insights in real-time. It can produce statistics like the number of studies matching certain criteria, the distribution of trials by phase or by year, the number of unique sponsors involved in a result set, etc., and display these in interactive charts or dashboards. For example, the system might generate a summary dashboard for a search query: if a user searches for all trials by a certain sponsor, the dashboard could show how many trials that sponsor has in each phase (Phase 1 , Phase 2, etc.), across whichtherapeutic areas those trials are distributed, and the total number of sites or investigators involved across those trials. FIG. 4 (the user interface example) further illustrates how such summary information might be presented, with charts and counts by categories like trial phase or trial type. A user might see, say, a bar chart of that sponsor’s trials by phase, or a pie chart of those trials by therapeutic area.

[0048] The analytical views are not static; they can be filtered in tandem with the search. Continuing the example, if the user then filters those sponsor’s trials to only Phase 3 trials, the charts will update to reflect only Phase 3 trials of that sponsor (perhaps breaking them down by region or outcome status, etc., if such data is available). This ability to seamlessly transition between raw data and aggregated views, and to have those views update based on context, helps users derive insights quickly. It supports decision-making such as assessing a sponsor’s experience in a particular phase of trials, finding which locations are most common for a certain kind of trial, or identifying trends like an increase in trials for a certain therapy area over time. Dashboards are interactive, users can click on specific items (like trial phases or therapeutic areas) directly within the dashboard. This action will take them to the study search page, where results are pre-filtered based on their selection (e.g., Phase 3 trials or a specific therapeutic area).

[0049] On the study search page, users can then drill down further to filter by additional parameters like treatment area or sponsor. This process allows for seamless navigation from high-level views to detailed trial data, enhancing decision-making by enabling users to quickly explore and refine their search

[0050] Now, the detailed working of the system (100) as per the instant nonlimiting embodiment is described with reference to Figure 1. At least one Data Source (101 ) provide raw input data. A primary data source is the CTRI database, and optionally other trial data repositories such as ICTRP can be included. These sources feed into the Data Acquisition Module (102), which systematically retrieves clinical trial data from external registries (e.g., CTRI or ICTRP) via APIs, XML feeds, or JSON downloads, parses these raw inputs using libraries such as Python’s xml. etree or BeautifulSoup, and standardizes data into a structured internal schemausing tools like OpenRefine or Pandas. It conducts automated validation checks, deduplication based on unique trial IDs, and version-control management, outputting structured records (e.g., JSON objects or relational database tables). The output of the Data Acquisition Module (102) is passed to the Polyhierarchical Tagging Module (103), which enriches the data with semantic tags from the EDGE taxonomy. The tagged data is then stored in a Search-Optimized Database (104). The Search-Optimized Database (104) is structured to support the operations of the Search Module (105) and the Analytical Module (106). Users interact with the system via a User Interface (107) (e.g., a web application) to further operate the Search Module (105) and Analytical Module (106). User queries go to the Search Module (105), which queries the Search-Optimized Database (104) and returns results to the user. The Analytical Module (106) can also query the Search- Optimized Database (104) (for aggregate data) and present visualization outputs to the user through the User Interface (107).

[0051] Referring now to Figure 2 which outlines the data processing flow (200) of TrialOptima (100). The operation of the system can be understood as a two- phase process: an offline data preparation phase (ingestion and tagging) and an online query phase (search and retrieval).

[0052] Step 1 : Data Ingestion (201 ) - A legacy clinical trial record in an unstructured or semi-structured format (201 ) is fetched from the source. For example, this could be an XML, HTML, or PDF document retrieved from the CTRI, ICTRP or other global registries. The Data Acquisition Module (102), optionally implemented using a Python-based scraping engine or an ETL pipeline, parses this input and converts it to a common intermediate format such as JSON or CSV.

[0053] Step 2: Text Extraction (202) - In the extraction step (202), the system parses the raw record to extract relevant Structured Data Fields (203). This involves identifying key sections of the record such as “Health Condition,” “Intervention,” “Phase,” “Location / Site,” etc., and converting them into a structured format (like keyvalue pairs in a JSON or columns in a relational database such as PostgreSQL orMySQL). Text mining libraries (e.g., spaCy or Apache OpenNLP) also extracts information from free-text fields like eligibility criteria or trial descriptions.

[0054] Step 3: Apply EDGE Taxonomy (204) - The system analyzes the structured data fields (203) using the EDGE taxonomy (204), which combines structured semantic labels with real-world linguistic and contextual variation. Unlike conventional systems that apply flat or expert-defined ontologies, the EDGE taxonomy (204) is continuously refined through a hybrid expert-sourced and field- validated process. This includes (a) verified investigator-site lists and regional terminologies collected from on-ground teams, (b) domain expert heuristics on clinical trial structuring, and (c) natural language processing techniques that capture variations in terminology and structure from large trial corpora. The EDGE taxonomy is composed of both explicit tags, derived from structured field analysis, and implicit tags, derived from latent signals observed in unstructured trial descriptions, eligibility criteria, and other narrative elements. Explicit tag classes (e.g., trial phase, intervention type, geography) originate from consistent field headings found in standardized registries such as CTRI or ClinicalTrials.gov. Implicit tag classes (e.g., trial purpose, therapeutic rationale, patient population type) are derived from observed linguistic patterns and encoded by domain experts through iterative corpus analysis and manual abstraction, often validated through cross-annotator agreement. In its preferred embodiment, the EDGE taxonomy tags are used for Polyhierarchical Tagging (205) of data derived from the extraction step (202), using a combination of two complementary techniques: (a) Explicit Tagging via Direct Keyword Matching: This technique involves scanning structured fields of the trial record (e.g., condition, intervention) against a predefined keyword dictionary derived from the EDGE taxonomy to assign explicit tags that correspond to clearly stated attributes in the structured input, (b) Implicit Tagging via NLP-based Semantic Matching: For unstructured or semi-structured textual content (such as trial summaries, eligibility criteria, or descriptions), the module employs advanced natural language processing (NLP) methods to extract domain-relevant implicit attributes that may not be directly mentioned in structured fields. This NLP engine preferably utilizes biomedical transformer models such as Bidirectional Encoder Representations from Transformers for Biomedical Text Mining ("BioBERT”), a pre-trained biomedical language representation model for biomedical text mining, or domain-specific NLP libraries such as ScispaCy, known in the art. This functionality is illustrated using a fictitious structured clinical trial record with the title and description “Phase 3 trial on Metformin for Type-2 Diabetes Mellitus and Coronary Artery Disease conducted in Mumbai by PharmaCo.” The text extraction step first extracts concepts such as “Metformin,” “Type-2 Diabetes Mellitus,” “Coronary Artery Disease,” “Phase 3,” “Mumbai,” and “PharmaCo.” These terms are then semantically matched against EDGE taxonomy terms. The system assigns tags accordingly: (i) Explicit Tags: Therapeutic Area = [Endocrinology, Cardiology]; Disease = [Type-2 Diabetes Mellitus, Coronary Artery Disease]; Intervention = [Pharmaceutical]; Phase = [Phase 3]; Geography = [Mumbai — Maharashtra — India]; Sponsor Type = [Industry]; (ii) Implicit Tags: Trial Purpose = [Post-Marketing Surveillance]; Study Context = [Urban India]. The system ensures every record receives multi-dimensional semantic tagging, supported by context-aware lookup tables, synonym expansions, and prior mappings from the taxonomy. In cases where new terminology is encountered, the system flags these entries and routes them to a curation queue, allowing domain experts to review, approve, and add them to the taxonomy. This human-in-the-loop mechanism ensures that the taxonomy evolves while preserving consistency and traceability.

[0055] Step 4: Store Tagged Data (206) - The now-tagged trial data is stored in the search-optimized database (104). This is preferably a NoSQL document database (e.g., MongoDB), a graph database (e.g., Neo4j), or a full-text search engine (e.g., Elasticsearch) configured to handle multi-faceted queries. The system also detects new or previously unseen terms and flags them for expert review (207), allowing updates to the EDGE taxonomy. This taxonomy expansion process is handled via a human-in-the-loop interface backed by a version-controlled taxonomy database.

[0056] Step 5: User Query Input (208) - A user initiates a search via a frontend User Interface (Ul) or Application Programming Interface (API) or an Artificial Intelligence (Al) assistant chat interface (207), entering either free-text or structuredfilters. The query could be a set of keywords, a phrase in natural language, or a selection of filters in the Ul.

[0057] Step 6: Natural Language Processing of Query (209) - The Search (210) performs NLP on the user query (208). The NLP (209) is implemented using language models like BioBERT (Jinhyuk Lee et al., Bioinformatics, Volume 36, Issue 4, February 2020, Pages 1234-1240), tokenizes and interprets the query to extract semantic intent. Named entity recognition and ontology mapping help match abbreviations and synonyms via the taxonomy. For example, it would identify “Phase 2” likely refers to a trial phase, “lung cancer” refers to a condition (which might map to tags like Indication = Lung Cancer, Therapeutic Area = Oncology in the taxonomy), and “Mumbai” refers to a location (Site Location tag = Mumbai). This interpretation step preferably involves tokenization, part-of-speech tagging, and named-entity recognition (to recognize that “Phase 2” is a trial phase entity, “lung cancer” is a disease entity, and “Mumbai” is a geographic location). The system leverages the taxonomy here as well: it knows that “lung cancer” corresponds to a specific indication tag in its database (and is under the Oncology therapeutic area). If the query included an abbreviation or synonym (e.g., “Ml” for myocardial infarction, or “CML” for Chronic Myeloid Leukemia), the system would be aware of it via synonyms in the taxonomy or a controlled vocabulary and would match it to the appropriate tag.

[0058] Step 7: Semantic Match & Filtering - Using the interpreted query, the system performs a search (210) against the search-optimized database (104). The interpreted query is executed against the search-optimized database (104), retrieving relevant records by matching structured tags. Results are ranked by relevance using scoring algorithms and presented along with filter options for refinement. It retrieves all trial records that satisfy the query criteria, effectively performing a semantic match rather than a simple text match. Because the data is tagged and structured, this is akin to executing a query like: find all records where Trial Phase = 2 AND Indication = Lung Cancer AND Site Location = Mumbai. The result is a set of relevant trials (211 ). If the user had used a broader query or provided partial information, the system might retrieve a broader set of results. Thesearch module (105) at this step (209) preferably also applies a ranking algorithm known in the art, such as text relevance scoring (e.g., BM25), semantic vector similarity (e.g., cosine similarity using Sentence-BERT or BioSentVec), weighted tag matching, and faceted dimension scoring — to evaluate the contextual relevance of each clinical trial record against the user query across both structured and unstructured fields, and order results by relevance, especially for queries with free- text components, to order results by relevance, especially for queries with free-text components. For example, if the user query (207) was not fully structured (for instance, they just typed “cancer trials in Mumbai”), the system might score results by how well they match the combination of tags and any text in the query. At this stage, the system also prepares information for faceted filtering: it knows the breakdown of the current results by various dimensions (e.g., how many of those results are Phase 1 vs Phase 2, how many are in each region, etc.), and this info is passed along to generate the filter Ul options for the next step (212).

[0059] Step 8: Refine Filters / Drilldown (212) - The user is presented with the results and may choose to refine results (211 ). Filters may include sponsor type, trial design, intervention type, or registration date. The underlying data structure leverages a faceted, search optimization mechanism (e.g., Elasticsearch or equivalent), enabling rapid re-querying across multi-dimensional tags while preserving user context and filter history for dynamic Ul rendering and query optimization. If the initial results are too many or not specific enough, the user can apply additional criteria using the interface filters. For instance, continuing the lung cancer example, suppose 50 trials were found across all phases in Mumbai. The user might then check a filter for “Interventional studies only” and “Phase 2” if not already applied or add “Last 5 years” as a date filter if such an option exists. The search module takes these refinements and essentially adds those conditions to the query, performing another search on the database (returning to Step 7, now with narrower criteria). The results (211 ) update, typically shrinking to a more focused list. This iterative process can repeat multiple times until the user is satisfied with the specificity of the results (211 ). Throughout the search process, the system maintains context - meaning it remembers the user’s current filter selections and applies new ones on top of existing ones unless the user removes them. Thedesign ensures that even as filters are applied, the system is using precomputed tags, so each refinement query is very fast (much faster than having to rescan all text each time in a traditional system).

[0060] Figure 3 shows a schematic of the hardware infrastructure on which TrialOptima is deployed (300). The system can be implemented in a cloud-based server environment accessible via the internet. In the figure, multiple types of users (Researchers, Patients, Sponsor / lnvestigator representatives) operate client devices (301) which could be standard PCs, laptops, or mobile devices with a web browser. These clients connect over a network (302) (which could be the Internet or a secure intranet) to the TrialOptima Application Server (303), which runs the core logic of the Data Acquisition (102), Tagging (103), Search (105), and Analytical module (106). The Application Server interacts with a PSQL Database (304) that stores the tagged trials database (as described earlier). It also interfaces with a Taxonomy Data Store (305) which holds the expert-defined taxonomy EDGE taxonomy, including all the categories, tags, and keyword lists. The taxonomy store (305) is used by the tagging module (103) at runtime and could be implemented as a database or even a configuration file repository that can be updated as the taxonomy evolves. For security and maintenance, the Data Acquisition part is preferably run as a scheduled back-end process that periodically pulls new or updated trial records from the external data sources (307). This is indicated by the dotted line in FIG. 3 showing a data acquisition connection from the external databases (306) into the Application Server (303). Such a connection might use APIs or automated scripts and would typically run on a schedule (e.g., nightly or weekly updates) to keep the TrialOptima database up-to-date with the latest trial registrations. Users access the system via a web interface served by the Application Server. Communication between client devices and the server is secured (e.g., via HTTPS, shown as “Secure Connection” in the figure). Thus, when a user runs a search or clicks a filter on their device, the request travels to the application server, which then queries the database server, and the results are sent back to the client to be displayed in their web browser as updated results or visualizations.

[0061] This modular deployment means the system is scalable: for instance, the Database Server can be scaled (horizontally or vertically) for performance, or the Taxonomy Data Store can be managed by experts to update the taxonomy without altering application code. It also separates concerns - the front-end deals with user interaction and presentation, while the back-end handles heavy data processing and storage.

[0062] The user interface of TrialOptima is designed to be intuitive despite the system’s complexity. Figure 4 provides a snapshot of what a user might see when using the system’s search functionality. This figure is an example; actual implementations can vary in style.

[0063] Figure 4A, shows the left side panel with the search input area containing multiple fields and filters: At the top, text input fields are provided for key search parameters: Condition / Disease, Sponsor, Site Name, Investigator Name. These allow users to directly type in values or terms. For example, a user could type “Acne” in the Condition box and “Assam” in the Location box to look for acne trials in the state of Assam. Below the text fields, additional filters such as Trial Phase are listed with checkboxes (Phase 1 , Phase 2, Phase 3, Phase 4, Post Marketing Surveillance, BA / BE (bioavailability / bioequivalence), etc.). The user can check one or multiple phases to filter trials by phase. Similar filters for other attributes could include checkboxes or dropdowns for recruitment status, trial type, study design, etc., depending on what taxonomy tags are available in the system.

[0064] Figure 4B shows two main sections of the search results. In the upper section, the interface displays an overview of the search results in aggregate form. In the example, various trial phases are listed (All, Phase 1 , Phase 1 / 2, Phase 2, Phase 2 / 3, Phase 3, Phase 3 / 4, Phase 4, PMS (post-marketing surveillance), RWE (real-world evidence), BA / BE, Validation / Pivotal, PoC / Pilot, PK / PD, etc.), with each category showing a count of trials (in the figure, zeros are shown as it might depict an initial state or a state with no query applied). This suggests that, upon running a search, this area would populate with the counts of results falling into each category. For instance, if the user searched “diabetes” and got 1463 trials, the summaryshows: Phase 1 : 31 trials, Phase 1 / 2 25, Phase 2: 133 trials, Phase 2 / 3: 58 trials, Phase 3: 329, etc. These summaries give a quick glance at the distribution of results. The summary panel might also include other aggregate info, like the number of unique sponsors, or a breakdown by region, depending on the query. The summary panel also has a search box that could allow entering an additional keyword to filter within the current results, or for quickly locating something in the results. The lower section is the Results Table / List. Here each row might represent a trial record, showing key fields like Trial ID, Title, Primary Condition, Sponsor, Registered On date, etc. The figure shows column headers such as “No. of Sites,” “Registered On,” “Health Condition,” etc. Users can scroll through this list or page through it if there are many results. Clicking on a particular trial in the list might bring up a detailed view of that trial’s full record.

[0065] Figure 4C. shows the dashboard view that allows users to view charts based on therapeutic area, phase-wise trial counts, and year-wise trial counts for a particular sponsor, investigator, or site profile and download them (e.g., as a CSV or Excel file) for offline analysis or record-keeping. The interface is interactive: clicking on any filter or summary category likely refreshes the results list and summary counts. For example, if a user clicks on the “Phase 2” bar in the summary chart, the system would filter the results to only Phase 2 trials (combining with any other current filters) and update both the list and the summary counts (which might then update other categories like showing only those relevant to Phase 2 trials). This interactivity allows the user to navigate the data in a nonlinear fashion.

[0066] To further illustrate the capabilities of TrialOptima, consider a few example scenarios of how a user might utilize the system:Exploratory Filtering: A researcher is interested in dermatology trials, particularly for acne treatment, and wants to see what studies have been done in Northeast India. They type “acne treatment” in the Condition field. The system returns (for example) 30 trials related to acne. The researcher notices many are multi-center studies across India. They then add a filter by typing “Assam” in the location field (or selecting Assam from a location filter if available). Now the results filter down totrials that have a site in Assam (suppose 5 trials). The researcher can click on a specific sponsor’s name among those to see all trials by that sponsor in Assam, even beyond dermatology if desired. They can also pivot in another way: for each of those trials in Assam, they might want to see what other locations those trials took place in without manually searching each trial. The system could allow clicking on the site or investigator to pivot the search to that entity. This approach enables discovering, for example, that a certain hospital in Assam is involved in many dermatology trials, and then exploring what other trials that hospital is involved in, even beyond dermatology, all through a few clicks. Traditional search would require separate searches for each such question, but TrialOptima’s semantic linking and interface allow it to happen fluidly.

[0067] Multiple Paths to the Same Trial: TrialOptima’s polyhierarchical nature means a particular trial can surface via different queries. For example, consider a specific clinical trial with Trial ID CTRI / 2021 / 08 / 035648 which is a post-marketing surveillance study of a vaccine (Covishield) conducted at Christian Medical College by Dr. Winsley Rose. A user interested in vaccines might start by searching the vaccine name “Covishield,” navigate through the results to find that trial, and then see that Dr. Winsley Rose is an investigator - perhaps clicking on the investigator’s name to see what other studies Dr. Rose has conducted. Meanwhile, another user might start by searching for an investigator’s name directly - if they search for “Dr. Winsley Rose,” they would find the same trial among others, and see that it involves the Covishield vaccine. Both users converge on the same trial from two different starting points (one from the product, one from the investigator). This demonstrates the system’s ability to let users approach the data from any angle and still discover relevant connections, thanks to the underlying multi-tagging of trials.

[0068] Non-Linear Filter Application: In many conventional systems, the order in which you apply filters might matter, or some filters might not be available until you pick a certain path (for instance, you might have to choose a disease area before seeing phase filters). In TrialOptima, users are not constrained to a specific sequence of filtering. For instance, one user could first filter by Sponsor (choose a particular pharmaceutical company) then filter by a Therapeutic Area (sayOncology) to see that sponsor’s oncology trials, whereas another user could start by filtering by Therapeutic Area = Oncology first, then within those results filter by that same sponsor. Both approaches yield the same final subset of trials. The interface and underlying data support this commutative property of filtering because the data is semantically tagged in all relevant dimensions from the start. This flexibility significantly improves user experience, as users can apply their thinking in whatever order or priority, they consider important, and the system will accommodate it.

[0069] Analytical Insights: Suppose a user (perhaps a policy maker or research analyst) wants to understand the landscape of clinical trials in India for a particular therapy area. They could use TrialOptima to get a quick dashboard overview. For example, they run a broad search with no specific condition but then apply filters: Sponsor Type = “Academic” and Therapeutic Area = “Oncology.” The system might return all oncology trials sponsored by academic institutions. The analytical module can then show that, for instance, there have been 100 such trials in total: a pie chart might show distribution by phase (Phase 1 : 10 trials, Phase 2: 30 trials, Phase 3: 40 trials, Phase 4: 20 trials), another chart might categorize by region (North, South, East, West India), and a list might detail the top 5 academic sponsors by number of oncology trials. Such insights can be gleaned in seconds and can inform decisions like where to allocate funding or identify potential gaps (maybe no Phase 3 oncology trials in a certain region, etc.).

[0070] These scenarios underscore how TrialOptima provides both precision (retrieving trial records that match a narrowly defined set of criteria) and contextual breadth (supporting multi-angle exploration and discovery beyond the initial query). The inventive combination of a multi-category semantic backbone with a user- friendly exploratory interface is what enables these outcomes. This is technically enabled by the underlying polyhierarchical semantic tagging model, which allows each trial record to be indexed across multiple interlinked dimensions — such as disease, phase, sponsor, intervention, investigator, and location — using the EDGE taxonomy. Since these tags are assigned during preprocessing and stored in a search-optimized database, the system does not need to perform computationallyexpensive full-text scans or dynamically construct filter trees at runtime. In the preferred embodiment, when a user enters a query (e.g., “Phase 3 acne trials in Assam”), the system parses the input using NLP, maps each term to one or more pre-tagged dimensions (e.g., Indication = Acne; Phase = Phase 3; Location = Assam), and retrieves only those trial nodes that already carry these indexed tags. This approach ensures high precision (only semantically tagged matches are returned) while also increasing recall (e.g., trials tagged under broader categories like “Dermatology” or synonyms such as “pimples” are also retrieved if the taxonomy permits hierarchical or synonym expansion). Because these semantic relationships are already encoded and traversable, the system requires fewer query iterations to resolve relationships between entities, thereby reducing total query volume and memory usage per session. In contrast to traditional systems that require repeated, sequential searches for each attribute or dimension, TrialOptima's semantic tagging enables compound, multi-dimensional queries to be resolved in a single, low- latency operation. This results in faster performance, reduced CPU and I / O load, and higher throughput even as the trial corpus scales. The effect is a demonstrable systems-level efficiency gained through architectural design.

[0071] The implementation of TrialOptima can employ various technologies to achieve the described functionality. For NLP tasks (query parsing, synonym recognition, etc.), natural language processing toolkits or machine learning models (such as transformer-based language models trained on biomedical text) are preferably deployed to enhance accuracy. The taxonomy itself can be stored in an ontology format (like OWL or RDF) to integrate with ontology reasoners if needed. The system preferably includes an automated feedback loop wherein frequently searched but unrecognized terms are automatically identified through query logs and incorporated into the taxonomy via adaptive learning algorithms — such as embedding-based similarity matching or contextual vector clustering — once a confidence threshold is met. Over time, as the database accumulates user interactions, a recommendation system could suggest popular searches or related searches to users (for instance, “users who searched X also looked at Y”), further leveraging the rich network of data.

[0072] In conclusion, the TrialOptima system introduces a technically synergistic improvement over prior clinical trial registry architectures by integrating expert-curated semantic models with a polyhierarchical tagging module and a search-optimized retrieval layer. This architecture enables pre-indexed, semantically enriched representations of trial data across multiple dimensions, allowing for compound query resolution and graph traversal without the need for runtime full-text scans or complex manual filter chaining. As a result, TrialOptima significantly reduces query latency, memory overhead, and computational load while maintaining high precision and recall, thus demonstrating a structural and algorithmic advance beyond traditional flat or single-hierarchy registry systems.

[0073] The present invention is industrially applicable to the fields of clinical research, healthcare information technology, and pharmaceutical data management. By deploying the TrialOptima system, clinical research organizations, regulatory agencies, hospitals, and pharmaceutical companies can vastly improve how they manage and search through large volumes of clinical trial data. The invention enables efficient identification of relevant clinical trials for purposes such as trial feasibility analysis, competitor trial surveillance, patient recruitment, and evidence synthesis for research or regulatory review.

[0074] The system’s ability to integrate multiple data sources and update dynamically makes it suitable for national or international trial registry platforms that require constant data curation and user-friendly access. Its use of adaptive learning means it can become more effective over time in responding to user needs, which is valuable in any large-scale deployment (such as a national clinical trial registry portal). Overall, the invention provides a practical solution that addresses real-world inefficiencies in clinical trial data retrieval, thus having broad applicability in improving the transparency, accessibility, and productivity of clinical research efforts worldwide.

Claims

WE CLAIM:

1. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) comprising at least one network connected computer (303) having at least one processor, being programmed with a Data Acquisition Module (102); a Polyhierarchical Tagging Module (103); a Empirically Derived Ground-Truth Encoding (EDGE) taxonomy store (304); a Search Optimized Database (104); a User Interface Module (107); wherein the Polyhierarchical Tagging Module (103) accepts inputs from the Empirically Derived Ground-Truth Encoding (EDGE) taxonomy store (304) to organize information extracted from at least one clinical trial record semantically in the Search Optimized Database (104) and simultaneously tag the clinical trial record with at least one label corresponding to at least one category.

2. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein a Data Acquisition Module (102) is configured to retrieve at least one clinical trial record from at least one data source and transform said record into a structured format.

3. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the Empirically Derived Ground-Truth Encoding (EDGE) taxonomy (204) comprises a plurality of semantically interlinked categories representing a plurality of facets of clinical trial data and a plurality of associated terms including empirically validated keywords, synonyms, and context-specific expressions derived from on-ground domain knowledge, site-level audits, and expert heuristics.

4. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 said Polyhierarchical Tagging Module (103) is configured to accept and analyze at least onestructured clinical trial record using natural language processing and assign at least one tag to at least one record and wherein a plurality of tags have parentchild relationships forming a polyhierarchy.

5. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the Search- Optimized Database (104) is configured to store tagged clinical trial records, the database comprising data structures or indices that link each clinical trial record to its assigned tags and enable retrieval of records by tag-based queries.

6. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the Search Module (105) is configured to accept user queries through the User Interface (107), interpret the queries using natural language processing to identify relevant taxonomy terms and filters, and execute searches on the search- optimized database (104) to retrieve clinical trial records that match the user query; said User Interface Module (107) further being configured to present retrieved clinical trial records to the user through interactive filtering options to refine search by selecting or adjusting at least one of the taxonomy-based filters such that the Search Module (105) iteratively updates the retrieved records in response to the user’s refinements without requiring a new query.

7. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the EDGE taxonomy (204) comprises a hierarchical ontology of medical terms and trial attributes, including categories for at least therapeutic area, disease / indication, intervention type, clinical trial phase, of trial sites, and sponsor type, and wherein each category in the taxonomy contains a predefined list of acceptable values or terms and defined relationships (hierarchical or associative) to other categories, such that a single clinical trial record can be tagged with multiple values across said categories.

8. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the Search- Optimized Database (104) is implemented as a graph database in which each clinical trial record is a node linked to tag nodes representing taxonomy terms, and wherein each tag node can have multiple parent nodes and child nodes reflecting the polyhierarchical structure of the taxonomy, such that queries for a given tag automatically encompass its descendant or related tags unless otherwise specified.

9. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the Search Module’s (104) natural language processing includes a query parser that detects and distinguishes between different types of query terms by matching them to categories in the taxonomy, and further comprises a relevance ranking component that ranks the retrieved clinical trial records based on the degree of semantic match to the user’s query, prioritizing records that satisfy more of the query facets or that have higher semantic similarity to free-text query components.

10. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the Search Module (104) further comprises a conversational assistant interface powered by a generative Al model, configured to receive natural language queries from the user in a chat-based interface, interpret multi-turn interactions using natural language understanding, map user expressions to taxonomy terms, and refine or expand search results iteratively based on user prompts.

11. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the User Interface (107) deploys the Analytical Module (106) to generate summary analytics alongside the list of retrieved clinical trial records, said summary analytics including aggregate counts or visualizations (charts or graphs) of the results broken down by one or more categories of the taxonomy, and whereinactivation of a filter option corresponding to a category value triggers the search module to refine the results to those pertaining to the selected category value.

12. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 said User Interface (107) comprising interactive filtering options including at least the ability to filter results by trial phase, recruitment status, location, sponsor, condition, and intervention type, and wherein filters are applied in any order and combination, using the tagged metadata, such that the resulting set of clinical trial records is the intersection of all selected filters enabling the allowing a user to narrow down results progressively from multiple starting points.

13. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the Data Acquisition Module (102) is further configured to periodically update the structured clinical trial data by fetching new and modified clinical trial entries from at least one source and trigger the Polyhierarchical Tagging Module (103) to perform tagging.

14. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 wherein the expert-defined EDGE taxonomy (204) can be updated by authorized users to introduce at least one new category and term wherein the system (100) is configured such that after taxonomy updates, the Polyhierarchical Tagging Module (103) can reprocess existing records to apply a tag for the newly added category and term.

15. A system for semantic polyhierarchical data organization and optimized discovery of clinical trials (100) as claimed in claim 1 further comprising an Analytics Module (106) that operates on the tagged clinical trial records to generate insights including statistical summaries, trends, or comparative analyses, wherein the analytics module produces output including number of trials by year for a given indication, geographic maps of trial site distributions for a given sponsor, or network graphs of collaborations between investigatorsand sites, and wherein these analytic outputs are accessible through the user interface to support strategic decision-making and discovery of patterns in clinical trial activity.

16. A computer-implemented method for semantic organization and discovery of clinical trial data, comprising the steps of: a. acquiring clinical trial data by retrieving trial records (201 ) from at least one source registry or database, the trial records comprising textual fields describing aspects of clinical trials; b. parsing and structuring each trial record (202) into a set of data fields according to a predetermined schema (203); c. providing an expert-curated Empirically Derived Ground-Truth Encoding (EDGE) taxonomy (204) that defines multiple categories of trial information and a controlled vocabulary of terms for each category, the taxonomy being established by domain experts; d. tagging the trial records by analyzing data fields of each structured trial record to detect occurrences of the controlled vocabulary terms, and assigning at least one category tag to the record for each detected term, wherein a single trial record receives at least one tag across at least one category, to output a semantically tagged dataset of trial records; e. storing the tagged trial record in a searchable data repository (104) that maps each tag to the record that have that tag, and maintaining relationships among tags as defined in the taxonomy’s hierarchy; f. receiving a user query containing at least one term or filter pertaining to clinical trial attributes; g. interpreting the query by using natural language processing (209) to identify which taxonomy categories and specific terms correspond to the query terms, including resolving synonyms or ambiguous terms to their standardized taxonomy term; h. retrieving matching trial records from the search-optimized database (104) by finding records whose tags satisfy the query such that for each condition in the query there is a corresponding tag on a record that matches a termor falls under a term identified from the query, thereby forming an initial result set; i. presenting the result set to the user along with available filter options derived from the taxonomy showing which tags within the result set could further refine the results; and j. refining the results iteratively by, in response to user interactions selecting one or more of the available filter options or adding additional query terms, updating the result set to include only those trial records that meet the refined criteria by intersecting the selected tag filters with the prior query conditions and updating the presented filter options to reflect the refined result set, wherein the method loops through steps (i)-(j) as the user continues to adjust filters to generate and store multi-criteria, context-aware searching of clinical trial records.

17. A computer-implemented method for semantic organization and discovery of clinical trial data as claimed in Claim 16, wherein the step of tagging the trial records (step d) further includes using a text-mining algorithm to analyze free- text portions of the trial record to identify implicit information such as therapeutic area or trial purpose, and assigning appropriate tags from the taxonomy even if those terms are not explicitly mentioned in structured fields, thereby enriching the record’s metadata beyond what is directly provided by the source.

18. A computer-implemented method for semantic organization and discovery of clinical trial data as claimed in Claim 16, wherein interpreting the query (step g) comprises detecting if the user query is a broad or incomplete phrase and, if so, automatically expanding or clarifying the query using the taxonomy; for example, if a user query contains only a high-level category name, the method expands the query to include at least one specific indication tag under the said category.

19. A computer-implemented method for semantic organization and discovery of clinical trial data as claimed in Claim 16wherein the receiving and interpreting steps (steps f and g) further comprise receiving a free-form natural languagequery through a generative Al-based chatbot interface, interpreting user intents through multi-turn dialogue processing, resolving ambiguities, and translating conversational inputs into structured taxonomy-aligned filters for semantic search.

20. A computer-implemented method for semantic organization and discovery of clinical trial data as claimed in Claim 16, further comprising ranking the matching trial records (step h) before presenting them, wherein records that have tags matching more of the user’s query terms or filters are ranked higher, and wherein records can also be ranked based on additional criteria such as recency of the trial (more recent trials ranked higher) or learned relevance from past user behaviors, thus providing the user with the most relevant results first.

21. A computer-implemented method for semantic organization and discovery of clinical trial data as claimed in Claim 16, wherein the presenting step (i) includes displaying a visual dashboard that summarizes the current result set, including visual elements like charts or graphs for distribution of the result set across different taxonomy categories, and wherein the user can interact with these visual elements to trigger the refining step (j) for that category value.

22. A computer-implemented method for semantic organization and discovery of clinical trial data as claimed in Claim 16, wherein during the refining step (j), the method maintains the state of applied filters and allows the user to remove or toggle filters already applied, re-expanding the result set if a filter is removed, and recalculates the result set and available filter options after each such change, thereby supporting both narrowing and broadening of the search in a seamless user session.

23. A computer-implemented method for semantic organization and discovery of clinical trial data as claimed in Claim 16, further comprising an updating step wherein the system periodically repeats steps (a)-(e) to incorporate new or updated clinical trial records from the data sources, and automatically refinesand makes them available for search, such that users querying after an update will have access to the latest data without manual reconfiguration of the system.

Citation Information

Patent Citations

  • Structure detection models

    US20210150346A1

  • Knowledge discovery using a neural network

    US20220101113A1