Personalized breast cancer treatment computer model
The system addresses the limitations of existing breast cancer treatment systems by integrating genomic data, EHRs, and imaging data with a custom BRCA gene panel and RAG LLM, providing precise, up-to-date, and user-friendly treatment recommendations.
Patent Information
- Application Number
- PCT/US2025/027942
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-14
- Filing Date
- 2025-05-06
- Publication Date
- 2025-11-13
AI Technical Summary
Current systems for personalized breast cancer treatment fail to adequately leverage advanced bioinformatics and AI technologies, such as machine learning and retrieval augmented generation, to provide precise, up-to-date, and user-friendly treatment recommendations, often leading to delays and suboptimal care due to data complexity and the need for continuous medical guideline updates.
A system utilizing a Retrieval Augmented Generation (RAG) enriched Large Language Model (LLM) integrates patient-specific genomic data, electronic health records, imaging data, and public health information to generate personalized treatment plans, incorporating a custom BRCA gene panel with 50 relevant genes, and provides intuitive interfaces for clinicians.
The system enhances the precision and efficiency of breast cancer treatment by offering timely, accurate, and tailored treatment strategies aligned with the latest medical knowledge, ensuring reliable and user-friendly decision-making support for healthcare providers.
Smart Images

Figure US2025027942_13112025_PF_FP_ABST
Abstract
Description
PERSONALIZED BREAST CANCER TREATMENT COMPUTER MODEL
[0001] The application claims priority to U.S. Patent Appl. Serial No. 63 / 643,196, entitled “System And Method For Personalized Breast Cancer Treatment Decision Support Using A Retrieval-Augmented Generation Enriched Large Language Model,” filed May 6, 2024, to Debarka Sengupta, et al..
[0002] The application claims priority to U.S. Patent Appl. Serial No. 63 / 643,215, entitled “Method And System For Precision Oncology Report Generation Utilizing Custom BRCA Gene Panel,” filed May 6, 2024, to Debarka Sengupta, et al..
[0003] The application claims priority to U.S. Patent Appl. Serial No. 63 / 643,227, entitled “Custom BRCA Gene Panel And Method Of Use,” filed May 6, 2024, to Debarka Sengupta, et al..
[0004] The application claims priority to U.S. Patent Appl. Serial No. 63 / 692,988, entitled “System And Method For Personalized Breast Cancer Treatment Decision Support Using A Retrieval-Augmented Generation Enriched Large Language Model,” filed September 10, 2024, to Debarka Sengupta, et al..
[0005] The application claims priority to U.S. Patent Appl. Serial No. 63 / 702,408, entitled “System And Method For Personalized Breast Cancer Treatment Decision Support Using A Retrieval-Augmented Generation Enriched Large Language Model,” filed October 2, 2024, to Debarka Sengupta, et al..
[0006] The application claims priority to U.S. Patent Appl. Serial No. 63 / 788,478, entitled “System And Method For Personalized Breast Cancer Treatment Decision Support Using A Retrieval-Augmented Generation Enriched Large Language Model,” filed April 14, 2025, to Debarka Sengupta, et al.
[0007] All of the above-referenced patent applications are commonly owned by the owner of the present invention and are incorporated herein in their entirety.TECHNICAL FIELD
[0008] The present disclosure pertains to the field of medical informatics and personalizedmedicine, and, more specifically, to systems for decision support in cancer treatment. The disclosure relates to a novel system for personalized breast cancer treatment decision support, which is designed for use in analyzing and integrating patient-specific genomic data, electronic health records (EHR), imaging data, and public health information.
[0009] The present disclosure relates to systems for advanced computational techniques and artificial intelligence, including in some embodiments a Retrieval Augmented Generation (RAG) enriched Large Language Model (LLM), to provide comprehensive and personalized treatment recommendations.
[0010] The present disclosure further pertains to the field of genomic medicine and precision oncology, and, more specifically, to methods and systems for enhancing clinical decisionmaking in breast cancer treatment. The disclosure relates to a novel method and associated system for generating a personalized cancer treatment report, which leverages a custom BRCA gene panel for analyzing and integrating patient-specific genomic data, electronic health records (EHR), imaging data, and public clinical trials information.
[0011] The present disclosure further relates to methods and systems employing advanced bioinformatics techniques, including in some embodiments the use of precision algorithms and data integration platforms, to deliver targeted and personalized treatment plans. This approach facilitates the identification of suitable therapeutic options, encompassing both approved and investigational drugs, based on a comprehensive analysis of the patient's unique genomic profile and relevant medical data.
[0012] The present disclosure further pertains to the field of genomic diagnostics and targeted therapy within oncology, and more specifically, to methods and systems designed for genomic analysis in the context of breast cancer. The disclosure is directed towards a novel approach and associated system for conducting genomic analysis, particularly by utilizing a custom BRCA gene panel. This method integrates patient-specific genomic data with a wealth ofgenetic information from various knowledge sources to identify genetic targets relevant to breast cancer treatment.
[0013] The disclosure further relates to methods and systems that apply computational biology techniques, including, in some embodiments, precision data handling and bioinformatics algorithms, to create comprehensive genomic profiles. By employing such advanced methodologies, the system aims to facilitate the identification of precise treatment strategies, including the utilization of approved breast cancer drugs and the alignment with ongoing clinical trials. This enables a personalized approach to breast cancer treatment, tailored to the distinct genetic composition of each patient’s cancer, thereby enhancing the efficacy of therapeutic interventions.STATEMENT OF FEDERALLY FUNDED RESEARCH
[0014] None.BACKGROUND
[0015] The field of personalized medicine, particularly in the treatment of breast cancer, is becoming increasingly significant in the healthcare sector. The advancements in genomics and bioinformatics have led to the development of more personalized approaches to cancer treatment, emphasizing the need to tailor therapeutic strategies to individual patient profiles. The National Cancer Institute highlights the importance of integrating genomic data, electronic health records (EHR), and imaging data to enhance the precision and effectiveness of cancer treatments. This integration is crucial for identifying the most effective treatment plans, considering the unique genetic makeup and health history of each patient.
[0016] The National Cancer Institute highlights the importance of integrating genomic data, electronic health records (EHR), and imaging data to enhance the precision and effectiveness of cancer treatments. This integration is crucial for identifying the most effective treatment plans, considering the unique genetic makeup and health history of each patient.
[0017] The evolution of genomic sequencing technologies and computational biology has catalyzed a shift towards more personalized cancer care strategies. These advancements underscore the importance of utilizing a patient's genomic profile, in conjunction with their electronic health records (EHR) and imaging data, to devise customized treatment approaches. Such personalized strategies are pivotal for identifying optimal treatment modalities tailored to the genetic and clinical nuances of each patient's cancer.
[0018] However, the complexity and sheer volume of genomic data, coupled with the critical need to analyze BRCA gene variants accurately, present formidable challenges. Clinicians and oncologists often grapple with interpreting this dense information expediently and accurately, leading to potential delays and less-than-ideal treatment outcomes.
[0019] One prevalent approach to achieving personalized treatment recommendations involves the analysis and integration of vast amounts of data, including patient-specific genomic information, EHRs, and imaging studies. However, the complexity and volume of data present significant challenges. Clinicians and oncologists often face difficulties in interpreting this data swiftly and accurately, leading to delays in treatment decisions and potentially suboptimal care outcomes.
[0020] Existing methodologies and systems for data integration and analysis in this arena are not without their shortcomings. Many fail to fully exploit advancements in bioinformatics, such as machine learning algorithms and custom gene panels, which could significantly refine the process of generating personalized treatment recommendations. Additionally, these conventional solutions may not be equipped to incorporate the latest oncological insights and clinical guidelines dynamically, which diminishes their utility over time. The user experience is another area where current systems may fall short, as they often present a steep learning curve for clinicians, further impeding the treatment planning process.
[0021] Current methodologies for integrating and analyzing data, including crucial BRCAgene information, often fail to harness the full potential of bioinformatics innovations. There is a palpable gap in systems that adeptly utilize custom BRCA gene panels and machine learning algorithms to streamline the creation of personalized treatment recommendations.
[0022] The field faces multiple hurdles, include the need for advanced decision support systems that can adeptly analyze and synthesize heterogeneous data, including BRCA gene variations; the necessity for these systems to dynamically incorporate up-to-the-minute medical research and clinical guidelines; and the essential requirement for user-friendly interfaces that enable healthcare providers to make informed decisions without being encumbered by complex systems or vast, raw datasets.
[0023] Moreover, current systems and methodologies for integrating and analyzing these data types have limitations. They may not fully leverage the potential of artificial intelligence (Al) and machine learning technologies, including large language models (LLMs) and retrieval augmented generation (RAG) architectures, which can significantly enhance the accuracy and efficiency of data analysis. Moreover, existing solutions may lack the capability to continuously update with new clinical knowledge and guidelines, limiting their effectiveness over time. Additionally, these systems often do not provide a seamless and intuitive interface for clinicians, further complicating the decision-making process.
[0024] The challenges in the field are manifold. There is a need for advanced decision support systems that can effectively analyze and integrate heterogeneous data types to provide personalized treatment recommendations. Such systems should be able to adapt to the latest medical research and clinical guidelines, ensuring that treatment recommendations remain relevant and evidence-based. Furthermore, there is a need for these systems to be user-friendly, enabling healthcare providers to make informed decisions quickly without navigating through complex interfaces or interpreting vast datasets manually.
[0025] Moreover, there are significant challenges that must be addressed, particularly the issueof data hallucination in generative Al systems. Hallucination, or the generation of inaccurate or fabricated information by Al models, poses a substantial risk in the medical field, where precision and reliability are paramount. This problem is especially concerning when Al is used to generate treatment recommendations based on complex medical data. The occurrence of hallucinations can lead to erroneous treatment plans, misdiagnoses, and potentially disastrous outcomes for patients. Therefore, there is an urgent need for the development of Al systems that are not only capable of handling large volumes of medical data but are also highly trustworthy and reliable.
[0026] Accordingly, there exists a need in the field of personalized breast cancer treatment for innovative systems that are capable of harnessing the power of Al and machine learning to offer precise, up-to-date, and easily accessible treatment recommendations. Additionally, there is a pressing need for advanced Al systems that incorporate robust mechanisms to verify and validate the information generated before it is used in clinical decision-making. Such systems must ensure the accuracy and relevance of all Al-generated content, particularly in applications involving critical health decisions. This would significantly improve the quality of care for patients with breast cancer, aligning treatment strategies with individual patient profiles and the latest medical knowledge.
[0027] Furthermore, there exists a need in the field of personalized breast cancer treatment to leverage genomics to deliver precise treatment recommendations. Addressing this need would markedly enhance patient care, aligning treatment interventions with each patient’s specific genomic landscape and the forefront of oncological research.
[0028] Further, there is a pressing need in the personalized breast cancer treatment arena to harness the specific capabilities of BRCA gene panels. Addressing this need will significantly enhance patient care by aligning treatment interventions with each patient’s distinct genetic landscape and the latest advancements in oncological research, thus revolutionizing theapproach to breast cancer therapy.SUMMARY OF THE DISCLOSURE
[0029] The present disclosure sets forth a comprehensive system for personalized cancer treatment decision support, designed to enhance the precision and efficiency of cancer care. In certain embodiments, the system includes integration of computational and artificial intelligence mechanisms, encapsulated within a suite of program instructions executed by a central processing unit (CPU). In such embodiments, through sets of instructions the system operates to analyze patient-specific genomic data, electronic health record (EHR) data, imaging data, and public health information, facilitating the generation of personalized treatment plans that are informed by each patient’s unique clinical profile.
[0030] To address the needs in the field, the system and methods of using such system leverage state-of-the-art technologies to automate and refine the process of data integration and analysis. By utilizing a dynamic and continuously updated database of genomic profiles, alongside advanced algorithms for data comparison and interpretation, the system offers an unprecedented level of decision support to clinicians. This enables the identification of optimal treatment options tailored to the specific genetic makeup and health history of individual patients. Furthermore, the system’s integration of EHR and imaging data enriches the contextual understanding of each patient’s condition, enhancing the relevance and applicability of treatment recommendations.
[0031] Accordingly, in certain embodiments, the system has the capacity to streamline the decision-making process for healthcare providers and its adaptability to evolving medical knowledge and treatment methodologies. The system’s architecture is designed for intuitive interaction, allowing clinicians to access and interpret complex data through a user-friendly interface. This not only expedites the formulation of treatment plans but also ensures that these plans are grounded in the most current and comprehensive patient information available.Moreover, the system’s incorporation of feedback mechanisms and interdisciplinary collaboration tools supports a holistic and collaborative approach to cancer treatment planning. By addressing the intricate challenges of personalized medicine, particularly in the context of breast cancer treatment, this invention represents a significant advancement in the field, promising to improve outcomes and the overall quality of care for patients.
[0032] Further, the present disclosure delineates an advanced method and system for the creation of personalized cancer treatment reports, as related to precision oncology for breast cancer. Specifically, in certain embodiments, the present disclosure can involve a process that commences with the analysis of genomic data utilizing a custom BRCA gene panel, integrates this data with the patient's electronic health records (EHR), imaging data, and public clinical trials information, and culminates in the generation of a comprehensive personalized cancer treatment report. This methodology addresses the needs articulated above in respect to the personalization of cancer treatment, offering a tailored approach that is reflective of the individual's unique genetic makeup, medical history, and the latest medical insights.
[0033] To address the needs in the field, the system and methods utilize a custom BRCA gene panel, designed to include no fewer than 50 genes directly relevant to the diagnosis, prognosis, and treatment of breast cancer. This specificity ensures the precision of genomic analysis, which, when combined with a holistic view of the patient's health data, allows for the identification of the most suitable therapeutic options. The system enhances user interaction through features like pop-up window interfaces that elucidate terminologies within the personalized cancer treatment report, thereby streamlining the clinical decision-making process.
[0034] Accordingly, in certain embodiments, the system and method allow for the selection of both approved and investigational drugs that align with the patient’s genetic profile, as well as the provision for allocating clinical trial opportunities and genetic counseling referrals. In someembodiments, the system and method can further incorporates data validation, removing unreliable data before the generation of the treatment report, thereby ensuring the accuracy and reliability of the recommendations. Moreover, the system is adept at analyzing a comprehensive range of data, including patient history, genomics, and imaging studies, making it a robust tool in the arsenal against breast cancer.
[0035] Further, the present disclosure pertains to a method and system for breast cancer genomic analysis, focusing specifically on the employment of a BRCA gene panel. This disclosure sets forth a genomic analysis method that facilitates the integration of expansive genetic data from knowledge sources with individual patient genomic data. By creating a custom BRCA gene panel based on this integrated data, the invention presents a solution to the previously identified problems of data complexity and the need for up-to-date, personalized treatment strategies.
[0036] The disclosed method leverages the specificity of the BRCA gene panel to streamline the decision-making process for breast cancer treatment, offering a targeted approach that is both time-efficient and precise. This solution not only simplifies the interpretation of complex genetic data but also enhances the accuracy of treatment recommendations. By storing and continuously updating genetic information from a variety of sources, including the most recent patient data, the invention ensures that the treatment recommendations remain relevant and are based on the most current oncological evidence.
[0037] Accordingly, the system and method of the present disclosure utilize custom BRCA gene panel to provide a personalized healthcare approach, allowing for the testing of patient samples for specific genetic markers associated with breast cancer. As such, in some embodiments, the present disclosure offers a dynamic platform that empowers healthcare providers with the ability to quickly adapt treatment plans as new genetic information becomes available, and to provide patients with the most effective treatment options based on theirunique genomic profile.
[0038] In general, in one embodiment, the disclosure includes a system for personalized breast cancer treatment decision support. The system can include a central processing unit (CPU). The system can include a computer-readable memory. The system can include a computer- readable storage media. The system can include a set of program instructions. The set of programming instructions can include first program instructions configured to analyze patientspecific genomic data and compare it with a database of genomic profiles to provide personalized treatment recommendations. The analysis can include constructing genomic data profiles for the patient and comparing these profiles to the database to identify matching treatment options. The set of programming instructions can also include second program instructions configured to integrate electronic health record (EHR) data with the patientspecific genomic data. The integration can include analyzing patient health history and current medical conditions. The set of programming instructions can also include third program instructions configured to analyze imaging data. The imaging data can include a repository comprising known treatment outcomes. The set of programming instructions can also include fourth program instructions configured to retrieve and utilize public data sources. The set of programming instructions can also include fifth program instructions to generate a treatment plan based on compiled data in the set program instructions. The set of program instructions can be stored on the computer-readable storage media for execution by the CPU via the computer-readable memory.
[0039] Implementations of the present disclosure can include one or more of the following features:
[0040] The treatment plan based on the analysis of integrated patient data can include one or more treatments selected from the group consisting of on-label prescriptions, clinical trial allocations, and investigational drug opportunities
[0041] The first program instructions can further include analyzing genomic alterations for FDA-approved therapy matching.
[0042] The first program instructions can further include providing targeted therapy recommendations based on genomic data analysis.
[0043] The second program instructions can further include processing and analyzing patient history data, including previous treatments and outcomes, to personalize the treatment recommendations.
[0044] The set of program instructions can further include sixth program instructions configured to provide genetic counseling referrals as part of the treatment plan based on genetic markers identified in the genomic data analysis.
[0045] The sixth program instructions can further include analyzing genetic counseling needs based on the comprehensive integration of genomics data, EHR, and patient history data.
[0046] The set of program instructions can further include seventh program instructions to coordinate multidisciplinary team meetings and tumor board discussions for case preparation and treatment plan validation;
[0047] The seventh program instructions can further include documenting the decisions made and recommendations provided during tumor board discussions for future reference and follow-up.
[0048] The set of program instructions can further include eighth program instructions to summarize the treatment plan in a conversational format that mimics the decision-making process of a human-led tumor board;
[0049] The set of program instructions can further include ninth program instructions to continuously update the system’s database with new clinical oncology knowledge, treatment guidelines, and emerging therapeutic agents and strategies.
[0050] The ninth program instructions can further include instructions for the system to learnand adapt to new information about drug-drug interactions and drug response predictions using artificial intelligence.
[0051] The set of program instructions can further include tenth program instructions to facilitate discussion and case preparation by a multidisciplinary team using an integrated webbased platform.
[0052] The tenth program instructions can further include providing an interface for clinicians to input, review, and discuss the treatment plan, facilitating real-time adjustments based on clinician expertise and patient feedback.
[0053] In general, in another embodiment, the disclosure includes a method for personalized breast cancer treatment decision support implemented in a computer infrastructure having computer executable code tangibly embodied on a computer-readable storage medium including programming instructions to provide a personalized treatment plan. The method can include analyzing patient-specific genomic data to identify targeted therapy options. The method can include integrating electronic health record (EHR) data with genomic data to contextualize the treatment recommendations. The method can include conducting imaging data analysis to refine the treatment recommendations. The method can include retrieving public data sources. The method can include generating a treatment plan based on aforementioned steps.
[0054] Implementations of the present disclosure can include one or more of the following features:
[0055] The step of retrieving public data sources can include identifying potential clinical trial participations and investigational drug opportunities.
[0056] The step of generating a treatment plan can include incorporating genetic counseling referrals.
[0057] The method can include coordinating multidisciplinary team meetings and tumor boarddiscussions for case preparation and treatment plan validation.
[0058] The method can include summarizing the treatment plan in a conversational format for clinician review.
[0059] The method can include continuously updating the system’s database with clinical oncology knowledge and treatment guidelines.
[0060] The method can include retrieving feedback from clinicians, wherein the feedback is based on the treatment plan.
[0061] In general, in another embodiment, the disclosure includes a method for generating a personalized cancer treatment report. The method can include collecting a sample from a patient. The method can include obtaining genomic data using a custom BRCA gene panel. The method can include performing quality control on the genomic data. The method can include, resultant from performing quality control on the genomic data, removing unreliable sequences from the genomic data to produce quality-controlled genomic data, analyzing the quality-controlled genomic data using a custom BRCA gene panel to perform variant calling to obtain called variants. The method can include annotating the called variants to obtain annotated variants, integrating the annotated variants with the patient’s electronic health records (EHR), imaging data, and public clinical trials information. The method can include generating a personalized cancer treatment report based on the integrating. The report can include treatment-related information, gene alterations, FDA-approved drugs, investigational drugs, and clinical trial opportunities.
[0062] Implementations of the present disclosure can include one or more of the following features:
[0063] The custom BRCA gene panel can include at least 50 genes relevant to breast cancer diagnosis, prognosis, and treatment.
[0064] The method can further include providing descriptions of terminologies used in thepersonalized cancer treatment report in a pop-up window interface.
[0065] The method can further include utilizing both formalin-fixed paraffin-embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA) for genomic data analysis.
[0066] The method can further include identifying approved drugs and investigational drugs suitable for the patient as indicated by the personalized cancer treatment report.
[0067] The method can further include allocating clinical trial opportunities based on the integrated data analysis.
[0068] The method can further include providing genetic counseling referrals based on the personalized cancer treatment report.
[0069] The personalized cancer treatment report can include drug-drug interaction information and adverse reaction information for prescribed medications.
[0070] The method can further include, prior to generating the personalized cancer treatment report, removing unreliable data in the quality control on the genomic data.
[0071] The integrated data analysis can include analysis of one or more of patient’s history data, genomics data, and PET scan data.
[0072] In general, in another embodiment, the disclosure includes a digital clinical decision support system (DCDSS) for precision oncology. The system can include a processor The system can include a memory storing instructions. The instructions, when executed by the processor, can cause the system to analyze genomic data using a custom BRCA gene panel. The instructions, when executed by the processor, can cause the system to integrate analyzed genomic data with patient's electronic health records (EHR), imaging data, and public clinical trials information. The instructions, when executed by the processor, can cause the system to generate a personalized cancer treatment report based on the integrated data analysis.
[0073] Implementations of the present disclosure can include one or more of the following features:
[0074] The custom BRCA gene panel can include at least 50 genes relevant to breast cancer diagnosis, prognosis, and treatment.
[0075] The system can be configured to provide descriptions of terminologies used in the personalized cancer treatment report in a pop-up window interface.
[0076] The system can be configured to utilize both formalin-fixed paraffin-embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA) for genomic data analysis.
[0077] The system can be configured to identify approved drugs and investigational drugs suitable for the patient as indicated by the personalized cancer treatment report.
[0078] The system can be configured to allocate clinical trial opportunities based on the integrated data analysis.
[0079] The system can be configured to provide genetic counseling referrals based on the personalized cancer treatment report.
[0080] The personalized cancer treatment report can include drug-drug interaction information and adverse reaction information for prescribed medications.
[0081] The system can be configured to, prior to generating the personalized cancer treatment report, remove unreliable data in the quality control on the genomic data.
[0082] The integrated data analysis can include analysis of one or more of patient’s history data, genomics data, and PET scan data.
[0083] In general, in another embodiment, the disclosure includes a method for breast cancer genomic analysis. The method can include consolidating genetic information from a plurality of knowledge sources to obtain consolidated genetic information. The plurality of knowledge sources can include databases and trial data on breast cancer treatment. The method can include utilizing a computational environment to process the consolidated genetic information to facilitate the ranking of genes. The ranking of genes can be based on criteria comprising alteration frequency, gene length, and the presence of genes across commercial panels. Themethod can include employing a robust rank aggregation algorithm to prioritize genes for inclusion based on their significance in breast cancer to obtain ranked genes. The method can include creating a custom BRCA gene panel from the ranked genes comprising genes most relevant to breast cancer treatment. The method can include storing the custom BRCA gene panel in a genomic database designed to facilitate ongoing updates and integrations.
[0084] Implementations of the present disclosure can include one or more of the following features:
[0085] The biological sample can be selected from the group consisting of formalin-fixed paraffin-embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA).
[0086] The genomic database can be configured to continuously update with new genetic information from the knowledge sources and new patient genomic data.
[0087] The method can include using the custom BRCA gene panel to analyze a patient’s biological sample for genetic markers associated with breast cancer, where the analysis can include assessing the presence of specific genetic alterations linked to breast cancer prognosis and treatment responsiveness.
[0088] The method can include testing the patient’s biological sample using the custom gene panel to identify mutations indicative of breast cancer.
[0089] The testing for genetic markers can include utilizing the custom gene panel to identify and interpret mutations in the patient’s biological sample, including the assessment of homologous recombination repair (HRR) genes and pharmacogenomic (PGx) genes.
[0090] In general, in another embodiment, the disclosure includes a system for breast cancer genomic analysis includes a processor. The system can also include a memory. The processor and the memory can be configured to receive and consolidate genetic information from multiple knowledge sources. The system can also include a computational module programmed to use a robust rank aggregation algorithm to analyze and rank genes based on their relevanceto breast cancer treatment. The system can also include a genomic database configured to store the ranked genes and facilitate the creation of a custom gene panel based on the ranked genes.
[0091] Implementations of the present disclosure can include one or more of the following features:
[0092] The biological sample can be obtained from FFPE tissue blocks or ctDNA of the patient.
[0093] The custom gene panel can include approximately 50 genes, which are categorized into groups based on their clinical relevance and therapy options, facilitating targeted and precise breast cancer treatment strategies.
[0094] The genomic database can be continuously updated to include new genetic information from ongoing research and trial data, enhancing the accuracy and relevance of the custom gene panelBRIEF DESCRIPTION OF THE DRAWINGS
[0095] Other advantages of the present disclosure will be apparent from the following detailed description of the disclosure in conjunction with embodiments as illustrated in the accompanying drawings, in which:
[0096] FIG. 1 depicts a block diagram of the input and output of a distributed processing system that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0097] FIG. 2 depicts a distributed processing system that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0098] FIG. 3 depicts a computer system in a communication network that provides distributed processing, in accordance with certain embodiments of the present disclosure.
[0099] FIG. 4 depicts a schematic of the computer architecture for a distributed processingsystem that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0100] FIG. 5 depicts a process flow diagram of the summarization pipeline for a distributed processing system that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0101] FIG. 6 depicts a process flow diagram of a retrieval augmented generation architecture in a large language model for a distributed processing system that can be configured to provide cancer treatment support and functions based on the large language model, in accordance with certain embodiments of the present disclosure.
[0102] FIG. 7 depicts a block diagram of a method for providing personalized breast cancer treatment decision support, in accordance with certain embodiments of the present disclosure.
[0103] FIG. 8A depicts a user interface of distributed processing system that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0104] FIG. 8B depicts a user interface of distributed processing system that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0105] FIG. 8C depicts a user interface of distributed processing system that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0106] FIG. 8D depicts a user interface of distributed processing system that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0107] FIG. 9 depicts a schematic workflow of an agentic framework for precision oncology,showing the ingestion of medical data, tool-based retrieval, ReAct-style reasoning, and generation of clinically relevant responses, in accordance with certain embodiments of the present disclosure.
[0108] FIG. 10 depicts an example interaction in which an oncologist provides a patient case study to the system, and the system performs structured reasoning to generate a personalized treatment recommendation, in accordance with certain embodiments of the present disclosure.
[0109] FIG. 11 depicts a visualization of the vector search space using a 3D U-Map, illustrating the clustering and overlap of topic-specific embeddings in the vector database, in accordance with certain embodiments of the present disclosure.
[0110] FIG. 12 depicts the distribution of token volumes across distinct information silos in the vector database, providing insight into the relative size and topical focus of each silo, in accordance with certain embodiments of the present disclosure.
[0111] FIGS. 13A-13C depict tool usage patterns across three types of evaluation tasks illustrating how often each tool is utilized, in accordance with certain embodiments of the present disclosure. FIG. 13A depicts usage patterns for Objective QA evaluation tasks. FIG. 13B depicts usage patterns for Subjective QA evaluation tasks. FIG. 13C depicts usage patterns for Precision Oncology QA evaluation tasks.
[0112] FIG. 14 depicts an evaluation of system performance across diverse question answering tasks, highlighting improvements in faithfulness, precision, recall, and relevancy, in accordance with certain embodiments of the present disclosure.
[0113] FIGS. 15A-15C depict the composition of the evaluation dataset and the comparative performance of the system across multiple question formats and difficulty levels, in accordance with certain embodiments of the present disclosure. FIG. 15A depicts composition for Objective QA evaluation tasks. FIG. 15B depicts composition for Subjective QA evaluation tasks. FIG. 15C depicts composition for Precision Oncology QA evaluation tasks.
[0114] FIGS. 16A-16C depict a step-by-step breakdown of tool coordination during response generation, with Sankey diagrams showing how tools are selected and reused across each reasoning step, in accordance with certain embodiments of the present disclosure. FIG. 16A depicts the breakdown for Objective QA evaluation tasks. FIG. 16B depicts the breakdown for Subjective QA evaluation tasks. FIG. 16C depicts the breakdown for Precision Oncology QA evaluation tasks.
[0115] FIG. 17 depicts a schematic illustration of the functional architecture of a hierarchical agentic framework, gSage, for precision oncology, in accordance with certain embodiments of the present disclosure.
[0116] FIG. 18 depicts a schematic detailing the technical implementation of the gSage iterative reasoning loop, highlighting structured feedback and multi-step inference, in accordance with certain embodiments of the present disclosure.
[0117] FIG. 19 depicts a U-Map visualization of the embedding distributions in the vector database used by gSage, illustrating topic clustering and semantic overlap, in accordance with certain embodiments of the present disclosure.
[0118] FIG. 20 depicts a token count distribution across data topics in the gSage vector database, highlighting content density and topical breadth, in accordance with certain embodiments of the present disclosure.
[0119] FIG. 21 depicts a pie chart showing the distribution of retrieved data sources used by gSage in response to oncology multiple-choice questions, in accordance with certain embodiments of the present disclosure.
[0120] FIG. 22 depicts a Sankey diagram illustrating the flow of information sources retrieved by gSage into structured decision-making pathways, in accordance with certain embodiments of the present disclosure.
[0121] FIG. 23 depicts a pie chart representing the categorical distribution of multiple-choiceoncology questions used for model evaluation, in accordance with certain embodiments of the present disclosure.
[0122] FIG. 24 depicts a performance comparison of gSage against baseline large language models based on accuracy, precision, recall, and Fl -score in multiple-choice oncology question answering, in accordance with certain embodiments of the present disclosure.
[0123] FIG. 25 depicts a three-step framework for building and evaluating Deep Clinical and Genomic Patient Dossiers (DCGPDs), including data curation, gold-standard question and answer generation, and scoring protocol development, in accordance with certain embodiments of the present disclosure.
[0124] FIG. 26 depicts a schematic overview of how the gSage agent processes patientspecific queries by integrating molecular and clinical data through a multi-agent reasoning framework, in accordance with certain embodiments of the present disclosure.
[0125] FIG. 27 depicts the distribution of 46 oncology-related questions derived from DCGPDs across clinical categories, in accordance with certain embodiments of the present disclosure.
[0126] FIG. 28 depicts a comparative evaluation of gSage and baseline models across three dimensions: Clinical Reasoning, Evidence-Based Accuracy, and Patient-Centricity, in accordance with certain embodiments of the present disclosure.
[0127] FIG. 29 depicts a block diagram of a process for utilizing a custom BRCA gene panel for the treatment of breast cancer, in accordance with certain embodiments of the present disclosure.
[0128] FIG. 30 depicts a schematic diagram of panel design process for preparing a custom BRCA gene panel, in accordance with certain embodiments of the present disclosure.
[0129] FIG. 31 depicts a schematic of the report generation process based on the use of the custom BRCA gene panel, in accordance with certain embodiments of the present disclosure.NOTATION AND NOMENCLATURE
[0130] Various terms are used to refer to particular system components. Different companies may refer to a component by different names - this document does not intend to distinguish between components that differ in name but not function. In the following discussion and in the claims, the terms “including” and “comprising” are used in an open-ended fashion, and thus should be interpreted to mean “including, but not limited to . . . .” Also, the term “couple” or “couples” is intended to mean either an indirect or a direct connection. Thus, if a first device couples to a second device, that connection may be through a direct connection or through an indirect connection via other devices and connections.
[0131] The terminology used herein is for the purpose of describing particular example embodiments only, and is not intended to be limiting. Following long-standing patent law convention, the terms “a” and “an” mean “one or more” when used in this application, including the claims.
[0132] As used herein, the singular forms “a,” “an,” and “the” may be intended to include the plural forms as well, unless the context clearly indicates otherwise. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order discussed or illustrated, unless specifically identified as an order of performance. It is also to be understood that additional or alternative steps may be employed.
[0133] The terms first, second, third, etc. may be used herein to describe various elements, components, regions, layers and / or sections; however, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms may be only used to distinguish one element, component, region, layer, or section from another region, layer, or section. Terms such as “first,” “second,” and other numerical terms, when used herein, do not imply a sequence or order unless clearly indicated by the context. Thus, a first element,component, region, layer, or section discussed below could be termed a second element, component, region, layer, or section without departing from the teachings of the example embodiments. The phrase “at least one of,” when used with a list of items, means that different combinations of one or more of the listed items may be used, and only one item in the list may be needed. As used herein, the term “and / or” when used in the context of a listing of entities, refers to the entities being present singly or in combination. Thus, for example, the phrase “A, B, C, and / or D” includes A, B, C, and D individually, but also includes any and all combinations and subcombinations of A, B, C, and D. Accordingly, as an example, “at least one of: A, B, and C” includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A and B and C. In another example, the phrase “one or more” when used with a list of items means there may be one item or any suitable number of items exceeding one.
[0134] Spatially relative terms, such as “inner,” “outer,” “beneath,” “below,” “lower,” “above,” “upper,” “top,” “bottom,” and the like, may be used herein. These spatially relative terms can be used for ease of description to describe one element’s or feature’s relationship to another element(s) or feature(s) as illustrated in the figures. The spatially relative terms may also be intended to encompass different orientations of the device in use, or operation, in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, elements described as “below” or “beneath” other elements or features would then be oriented “above” the other elements or features. Thus, the example term “below” can encompass both an orientation of above and below. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptions used herein interpreted accordingly.
[0135] Unless otherwise indicated, all numbers expressing quantities of ingredients, reaction conditions, and so forth used in the specification are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numericalparameters set forth in this specification are approximations that can vary depending upon the desired properties sought to be obtained by the presently disclosed subject matter.DETAILED DESCRIPTION OF THE DISCLOSURE
[0136] The present disclosure is directed to an innovative system for personalized breast cancer treatment decision support. This system encompasses a comprehensive suite of program instructions executed by a central processing unit (CPU), utilizing advanced computational techniques and artificial intelligence, including a Retrieval Augmented Generation (RAG) supported Large Language Model (LLM). The set of program instructions is designed to analyze patient-specific genomic data, integrate electronic health records (EHR) data, analyze imaging data, and retrieve and utilize public data sources to generate personalized treatment recommendations.
[0137] The disclosure provides efficient and accessible devices, methods, and systems for personalized medicine in the treatment of breast cancer that address the limitations of prior art by introducing a novel decision support system. This system can include first program instructions for analyzing genomic data against a comprehensive database of genomic profiles, second program instructions for the integration of EHR data to contextualize the genomic analysis, third program instructions for analyzing imaging data to provide additional insights into the patient’s condition, and fourth program instructions for retrieving pertinent public health information. The system operates to compile these diverse data sources into a coherent treatment plan through fifth program instructions, ensuring that each patient receives care that is tailored to their unique genetic makeup and health history. By leveraging artificial intelligence and machine learning technologies, the system offers an unprecedented level of accuracy and efficiency in the analysis and interpretation of complex medical data, thereby significantly enhancing the quality of care for patients with breast cancer.
[0138] Further, the present disclosure is directed to a novel digital clinical decision supportsystem (DCDSS) for precision oncology, specifically tailored for breast cancer treatment. This system, and the method for using such system, integrates a custom BRCA gene panel analysis with patient-specific electronic health records (EHR), imaging data, and public clinical trials information to generate a comprehensive personalized cancer treatment report. The innovation encompasses a methodical process for analyzing genomic data, leveraging a custom BRCA gene panel that includes approximately 50 genes associated with breast cancer diagnosis, prognosis, and treatment.
[0139] The disclosure provides patient-centered methods and systems for cancer treatment planning that surmount the shortcomings of previous approaches by introducing a holistic digital solution. This solution facilitates the integration of disparate data sources through an advanced computational framework, ensuring a seamless synthesis of genetic information with broader patient health records. The DCDSS, in some embodiments, can be equipped with functionalities for interpreting both formalin-fixed paraffin-embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA), enriching the genomic analysis with a dual data input system. The system can also provide intuitive access to complex medical terminologies through pop-up window interfaces, thereby enhancing the user experience for clinicians. Additionally, the DCDSS is capable of autonomously updating its database with the latest clinical trials and medical guidelines, powered by an adaptive algorithmic core that keeps the treatment recommendations at the forefront of medical innovation.
[0140] Through the integration of technology and medical science, the system and methods disclosed herein dramatically enhances the precision, efficiency, and personalization of cancer treatment strategies, marking a significant leap forward in the management and care of breast cancer patients.
[0141] Further, the present disclosure is directed to a novel method and system for genomic analysis in the field of breast cancer treatment. The method and system can include the use anddesign of a custom BRCA gene panel. The system embodies a method for analyzing genomic data with heightened precision by employing a custom BRCA gene panel, meticulously selected to include genes with known associations to breast cancer diagnosis, prognosis, and responsiveness to specific treatments.
[0142] The invention addresses the limitations of conventional data integration in oncology by providing a method and system that ensures accurate and comprehensive analysis of genetic information, focusing on the relevant information for breast cancer diagnosis and treatment.
[0143] Accordingly, in some embodiments, the disclosed system and method are directed to facilitate more accurate and individualized treatment plans for breast cancer care. Enhanced by genomic medicine and computational technology, the disclosed system and method offer healthcare professionals the tools to craft data-driven strategies. Such strategies are meticulously aligned with each patient’s genetic profile, contributing significantly to the advancement of precision oncology.
[0144] In certain embodiments, the system is capable of processing genetic data from various sample types, including both formalin-fixed paraffin-embedded (FFPE) tissue and circulating tumor DNA (ctDNA), thus providing a robust foundation for the genomic analysis.
[0145] As such, the system and method described herein significantly increasing the accuracy and personalization of treatment plans for breast cancer care. In some embodiments, through the incorporation of genomic medicine and computational technology, the system provides a powerful tool for healthcare professionals, enabling them to offer individualized, data-driven treatment strategies that align with the unique genetic makeup of each breast cancer patient, thus contributing to the evolution of precision oncology.BRCA GPT
[0146] FIG. 1 depicts a block diagram of the input and output of a distributed processing system that can be configured to provide cancer treatment support and functions based on alarge language model, in accordance with certain embodiments of the present disclosure.
[0147] FIG. 1 illustrates a block diagram representing the workflow of a distributed processing system, designated as BRCA GPT 100, which is structured to provide decision support for breast cancer treatment by employing a large language model. The diagram is divided into three primary sections corresponding to the information input 101, the processing core 102, and the information output 103.
[0148] The information input section 101 can include various types of medical and patientspecific data that serve as inputs to the system. In some embodiments, these inputs include, but are not limited to, pathology, radiology, genetics, electronic health records (EHR), treatment history, and current health condition data. These diverse data points provide a multidimensional view of the patient’s health status, which in certain embodiments of the present disclosure are then utilized to construct personalized treatment plans.
[0149] The processing core 102 represents the large language model that processes the input data. This model can utilize advanced algorithms and artificial intelligence to analyze the input data, discern patterns, and synthesize information, as described in more detail in respect to FIGS. 2-4
[0150] The information output section 103 details the types of recommendations and reports that the large language model can generate based on the analysis of the input data. In some embodiments, the outputs include, but are not limited to, suggested treatments, clinical trial options, dosage recommendations, supportive care options, second opinions, and potential drug interactions or reactions. These outputs provide clinicians with actionable insights, enabling them to make informed decisions regarding the patient’s breast cancer treatment plan.
[0151] FIG. 2 depicts a distributed processing system that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0152] FIG. 2 depicts an example of a distributed processing system 200 configured to provide decision support for breast cancer treatment based on a large language model. This system 200 comprises a large language model that is distributed across multiple servers 240, each containing a part of the model to manage computational load and efficiency. Additionally, the system includes multiple front ends 210 to manage input 201 from various data sources such as EHRs, genomic databases, and imaging systems, and to format output 202, which consists of personalized treatment recommendations.
[0153] In this system 200, the front ends 210 are designed to receive complex medical data as input 201 and organize it into structured segments suitable for analysis by the processing servers 230. A load balancer 220 may be implemented to distribute the processing tasks among various servers 230, ensuring optimal utilization of resources based on the current demand and server workloads. Each processing server 230 functions to access the large language model servers 240 to retrieve the necessary data for analysis and uses this data to process segments of the medical input 201.
[0154] The distributed processing system 200 is particularly adept at handling the nuanced requirements of breast cancer treatment decision support, where inputs may include detailed patient medical histories, genetic information, and diagnostic images. The processing servers 230, in this context, may specifically serve as decision support servers, utilizing the large language model to analyze the input data and generate comprehensive, personalized treatment plans as output 202. The system is thus capable of offering tailored recommendations, identifying potential clinical trial participation, suggesting suitable drug therapies, and providing a platform for genetic counseling — all integral components of modern precision oncology.
[0155] The disclosed and other embodiments and the functional operations described in this specification can be implemented in digital electronic circuitry, or in computer software,firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products. In some embodiments, one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more them. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to suitable receiver apparatus.
[0156] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on onecomputer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0157] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0158] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0159] To provide for interaction with a user, the disclosed embodiments can be implemented using a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device,e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0160] The components of a computing system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet. A communication network that can be used to implement the described distributed processing may use various communication links to transmit data and signals, such as electrically conductor cables, optic fiber links and wireless communication links (e.g., RF wireless links).
[0161] FIG. 3 illustrates a computer system within a communication network that facilitates distributed processing, particularly tailored for a personalized breast cancer treatment decision support system, in accordance with certain embodiments of the present disclosure.
[0162] FIG. 3 shows an example computer system in a communication network that provides distributed processing. This system includes a communication network 300 that enables communications for communication devices connected to the network 300, such as computers. For example, the communication network 300 can be a single computer network such as a computer network within an enterprise or a network of interconnected computer networks such as the Internet. The communication network 300 can, in some embodiments, be a complex of interconnected networks spanning across multiple medical facilities, potentially integrating with the Internet for broader connectivity. Multiple computer servers 310 are linked to the communication network 300 to establish a distributed processing system analogous to the system described in FIG. 2. Accordingly, as shown in FIG. 3, one or more computer servers310 can be connected to the communication network 300 to form a distributed processing system. These servers 310, which may be geographically dispersed or co-located, host the decision support system and manage data processing and storage.
[0163] In some embodiments of operation, one or more computers (e.g., the computers of doctors 301 and 302) can use the communication network 300 to remotely access the distributed processing system 310. In certain embodiments, various user terminals, such as the computers used by oncologists 301 and 302, harness the communication network 300 to remotely engage with the distributed processing system 310. In some embodiments, an oncologist 301 may transmit a request to the system 310 for analysis of patient-specific information and generation of a treatment plan. The oncologist 301 can upload patient data, including genomic profiles, EHRs, and imaging studies, to the system 310. Upon receipt, the system 310 processes this data to output personalized treatment recommendations. The results are then conveyed to the oncologist 301 or are made retrievable within the system 310. Moreover, system 310 is capable of concurrently supporting multiple oncologists, thereby providing a scalable solution to personalize treatment across a spectrum of patient cases.
[0164] While this specification contains many specifics, these should not be construed as limitations on the scope of what being claims or of what may be claimed, but rather as descriptions of features specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0165] Similarly, while operations are depicted in the drawings in a particular order, this should not be understand as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0166] FIG. 4 depicts a schematic of the computer architecture for a distributed processing system 400 that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0167] As shown in FIG. 4, at step 401, the system 400 receives manually curated gene-drug pairs. Following at step 402, the system 400 conducts a series of bulk downloads and API calls. In some embodiments, during this same step 402, the system 400 then filters and cleans the received information.
[0168] After the manually curated gene-drug pairs are sent to a summarization pipeline 410, as described in more detail in respect to FIG. 5. Once the gene-drug pairs have proceeded through the summarization pipeline 410, the system 400 constructs hierarchical summaries at step 420. In some embodiments, the hierarchical summaries can be constructed in mark-down format with headings and sub-headings that facilitate RAG processing. In some embodiments, the hierarchical summaries made at step 420 have an overarching brief document summary, which is then separated into sub-topic summaries. In certain embodiments, each sub-topic summary include is further broken down in step 420 into detailed information on the sub-topic.
[0169] The hierarchical summaries constructed at step 420 then proceed in system 400 to step 430, where embedding is generated. In some embodiments, step 430 proceeds through the useof BioMistral. In certain embodiments, the system 400 at step 430 generates embedding using a BisMistral 7B Q8 base model. Through the generation, in such an embodiment, the system 400 can preserve links across vectors through the metadata. Further, in such an embodiment, in step 430 the system 400 can store citation information in metadata.
[0170] Following, the system 400 sends the information to a retrieval-augmented generation architecture 440, as described in more detail below in respect to FIG. 6. In the retrieval- augmented generation architecture 440, custom large language models can be utilized. In some embodiments, the large language models can include BioMistral, OpenOrca, Instruct, or combinations thereof. In retrieval-augmented generation architecture 440, as shown in FIG. 4, the processed information, now enriched with context and additional information, can be forwarded to the large language model. In certain embodiments, a framework optimized for language model-powered applications can facilitate the delivery of a query to the large language model from a user interface. In retrieval-augmented generation architecture 440, the large language model can processes the query and information coming from the user interface in connection with the information stored regarding a particular patient and the vector store, to generate an accurate and relevant response.
[0171] In some embodiments, the system 400 allows for the large language model’s generated response to be returned. This response is the system’s answer to the user’s query, containing treatment support information derived from the extensive analysis performed by the large language model.
[0172] Finally, the system 400 can deliver the response back to the user. In some embodiments, the return response is delivered by a chat window on the user’s console.
[0173] FIG. 5 depicts a process flow diagram of the summarization pipeline 410 for a distributed processing system 400 that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of thepresent disclosure.
[0174] In FIG. 5, the process begins at step 501 with the reception of a base corpus from the summarization pipeline 410, as discussed in respect to FIG. 4. This corpus contains manually curated gene-drug pairs and additional data obtained through bulk downloads and API calls, which have been preprocessed to remove irrelevant or redundant information.
[0175] At step 502, the received corpus undergoes contextual chunking. Here, the system leverages Al technologies to segment the comprehensive corpus into logically coherent chunks. This step involves identifying and categorizing the data based on thematic and contextual relevance to streamline further processing.
[0176] Following contextual chunking, in step 503, the document is separated into an abstract and a plurality of sub-topics. Each piece is treated as an individual unit of content to be further refined. This separation facilitates focused summarization processes where each segment is handled according to its specific context and relevance.
[0177] In step 504, the broken-up content, now segmented into complete documents, just the abstract, and various subtopics, undergoes summarization. This step uses the abstract as a contextual guide to ensure that summaries are not only concise but also retain the essential information relevant to each subtopic, ensuring a high level of accuracy and relevance in the synthesized summaries.
[0178] The process then moves to step 505, where it enters a summarization and autoevaluation pipeline. In some embodiments, the system 410 employs content collection techniques, organic keyword extraction, and prompt template collections in conjunction with a large language model. In some embodiments, the large language model can be Gemini, openAI, Anthropic, or combinations thereof.
[0179] In certain embodiments, while still in step 505, the response is evaluated for coherence, consistency, fluency, and relevance. Successful summaries are then logged as templates forfuture summarization tasks, enhancing the system's efficiency and accuracy over time. In such embodiments, the response is logged as a template and can be used for future runs of the system410
[0180] Finally, at step 506, the system generates hierarchical summaries using the best prompt and large language model for a given document source and summary depth. In some embodiments, this involves selecting the most appropriate model and prompt combination based on the source document’s complexity and the required depth of summary. This step can ensure that the final output is tailored to the specific needs of the clinical context, providing clear, concise, and relevant information to support cancer treatment decisions.
[0181] FIG. 6 depicts a process flow diagram of retrieval augmented generation architecture in a large language model, in accordance with certain embodiments of the present disclosure.
[0182] FIG. 6 provides a schematic representation of the architecture and workflow for the RAG portion 420 of the system 400, utilizing Retrieval Augmented Generation (RAG) within a large language model to support clinicians, particularly doctors, in making informed cancer treatment decisions.
[0183] The RAG portion 420 of the system 400, as depicted in FIG. 6, represents a sophisticated integration of Al and domain-specific knowledge bases to provide a dynamic, interactive, and intelligent tool for clinicians in the treatment of cancer. By leveraging the RAG method, the system enhances the factual consistency and reliability of the generated responses, thereby supporting the decision-making process in the complex field of oncology.
[0184] The process of FIG. 6 begins at step 601, where the RAG portion 420 of the system 400 defines a large language model and a configuration bin. Here, different configurations are determined based on the specific requirements of the clinical query being processed, establishing the groundwork for a tailored Al response.
[0185] At step 602, the RAG portion 420 of the system 400 selects a base large language modelfrom a set of models available. The selection is tailored to the specifics of the clinical question, ensuring the base model is optimally aligned with the type of information being sought. Models such as BioMistral, OpenOrca, Instruct, and Medicine Chat or combinations thereof are included in the selection pool.
[0186] Proceeding to step 603, the RAG portion 420 of the system 400 contains the large language model bin, which houses a variety of models from the same family but tuned for different tasks. This step allows for a targeted approach where the most appropriate model is chosen based on its specialized capability to handle specific aspects of oncology-related queries. The large language bin at step 603 can include large language models from the same family, tuned for different tasks. In some embodiments, the large language models can include BioMIstral, OpenOrca, Instruct, Medicine Chat, or combinations thereof.
[0187] At step 604, the RAG portion 420 of the system 400 contains the configuration bin, which contains merging algorithms and the corresponding legal values of parameters. In some embodiments, configurations such as SLERP (Spherical Linear Interpolation), DARE (Dynamically Adjusted Retrieval Enhancement), TIES, Param, or combinations thereof can be utilized to optimize the response generation process.
[0188] With this information, the RAG portion 420 of the system 400, culminates at step 605 with the evolutionary merge generation. In some embodiments, at step 605, the RAG portion 420 of the system 400, selects a model from the bin and generates a valid merge configuration. The process includes merging with the base model, evaluating the newly formed model, and then setting this merged model as the new base model for subsequent operations. This cycle may include random selection of models and merge algorithms, allowing for the generation of multiple merged models in parallel. In some embodiments, the best-performing model is then selected, and the process is repeated until satisfactory performance is achieved, ensuring that the final ALgenerated response is both accurate and relevant to the clinician’s needs.
[0189] FIG. 7 depicts a block diagram of a method 700 for providing personalized breast cancer treatment decision support, implemented in a computer infrastructure equipped with computer executable code embodied on a computer-readable storage medium. This medium includes programming instructions to create a personalized treatment plan through a sequence of analytical and integrative steps, each represented as a block in the diagram of FIG. 7.
[0190] The method 700 initiates with the analysis of genomic data. In some embodiments, as shown in FIG. 7, step 702 involves analyzing patient-specific genomic data to identify potential targeted therapy options. This involves evaluating the patient’s genomic profile against a vast database of genomic markers and their known associations with therapeutic outcomes.
[0191] The method 700 continues with the integration of EHR data. In some embodiments, as shown in FIG. 7, in step 704, the method includes integrating electronic health record (EHR) data with the genomic data, where this comprehensive analysis aims to contextualize the genomic findings within the broader scope of the patient’s medical history and current health condition.
[0192] Further, the method 700 encompasses imaging data analysis. In some embodiments, as shown in FIG. 7, step 706 includes conducting detailed analysis of imaging data, which may reveal additional insights into the patient’s condition, thereby refining the treatment recommendations derived from genomic and EHR data.
[0193] The method 700 also involves retrieving public data sources. In some embodiments, as shown in FIG. 7, in step 708, the system fetches data from public health databases, which may provide information on the latest clinical trials, emerging therapies, and a broader understanding of the patient’s condition within the context of current medical research.
[0194] The method 700 culminates in the generation of a treatment plan. In some embodiments, as shown in FIG. 7, step 710 involves synthesizing the insights gained from thepreceding steps to generate a comprehensive, personalized treatment plan that considers all facets of the patient’s medical profile and the latest oncological data.
[0195] Additionally, the method 700 may include steps for identifying clinical trial participation and investigational drug opportunities, integrating genetic counseling referrals into the treatment plan, and coordinating with multidisciplinary teams to validate the treatment strategy.
[0196] FIG. 7 illustrates a systematic and patient-centric method 600 for formulating a treatment plan in personalized breast cancer care. This method, with its sequential and integrated approach, provides a dynamic framework for oncologists to deliver precision medicine, tailored to the individual needs of their patients.
[0197] FIGS. 8A through 8D depict a user interfaces of distributed processing system that can be configured to provide cancer treatment support and functions based on a large language model, in accordance with certain embodiments of the present disclosure.
[0198] In some embodiments, the interfaces depicted in FIGS. 8A through 8D are designed to deliver cancer treatment support utilizing the capabilities of a large language model. These figures showcase an interface environment where users, such as oncologists or medical staff, can interact with the system in a conversational manner akin to engaging with an intelligent digital assistant.
[0199] As shown in FIG. 8A, the user interface can be presented as a streamlined chat window where the user can type in queries or information relevant to patient care. The system can be configured to process these inputs through its underlying large language model, which then generates and displays responses in the same window. This interactive dialogue format allows the user to ask complex questions regarding patient data or potential treatments and receive synthesized information and recommendations in real-time.
[0200] FIG. 8B continues to detail the user interface, demonstrating how the system can handlea sequence of interactions within a single session. This may include follow-up questions, requests for clarification, or deeper dives into specific topics, such as the interpretation of genetic test results or the evaluation of treatment efficacy based on the latest medical research. The interface is designed to retain context throughout the conversation, providing coherent and contextually relevant information as the user navigates through different facets of the patient’s case.
[0201] FIG. 8C illustrates a user interface wherein a medical professional can interact directly with the system to inquire about patient-specific information and evaluate approved treatment options. The interface displays a query field where the user can enter questions regarding patient data or seek advice on treatment strategies. Upon submission of the query, the system, leveraging its architecture as described in FIGS. 4 through 6, processes the input to mitigate the risk of data hallucination and ensures the generation of accurate and trustworthy information. The system then presents personalized treatment recommendations, which include up-to-date, relevant clinical trials and treatment options tailored to the patient’s specific medical profile. This interface supports the medical professional in making informed decisions quickly, enhancing patient care with precision.
[0202] FIG. 8D extends the interactive capabilities shown in FIG. 8C by demonstrating the system’s response to additional, more complex inquiries that a doctor may pose during a consultation. This figure shows a sequence where the medical professional can ask follow-up questions based on earlier responses, and the system dynamically generates further detailed information. The interface in FIG. 8D also highlights that the system can, in certain embodiments, provide deep insights into the queried topics. In some embodiments, the conversation history is maintained on the interface, allowing the doctor to easily reference earlier interactions, which aids in maintaining the context of ongoing patient case discussions. This ensures that the responses are not only accurate but also contextually appropriate to thespecific needs of the patient and medical situation.
[0203] In some embodiments, the interfaces in FIGS. 8A through 8D can provide an intuitive layout, facilitating ease of use while minimizing potential input errors. In some embodiments, the design of the user interface ensures that all interaction history remains visible, allowing users to reference earlier parts of the conversation easily. These interfaces embody the principles of human-centered design, ensuring that the advanced analytical capabilities of the system are accessible through a familiar and user-friendly conversational format.
[0204] FIG. 9 depicts a schematic of an agent-based large language model architecture configured to support precision oncology decision-making in accordance with certain embodiments of the present disclosure.
[0205] As shown in FIG. 9, the system 900 includes an intelligent agent 903 that interfaces with a suite of specialized tools 901 and retrievers 902 to enhance the clinical accuracy, coherence, and utility of treatment recommendations generated by the large language model. In certain embodiments, the intelligent agent 903 operates not merely as a passive generator of responses, but instead assumes the role of a planner and orchestrator, actively managing retrieval, refinement, and synthesis tasks across multiple modalities and data sources.
[0206] In some embodiments, the agent 903 determines which tool or combination of tools 901 to invoke at each step based on the evolving content of the clinician’s inquiry. For example, when a user query seeks clinical trial options, the agent may prioritize use of a trial retriever tool 902 configured to identify relevant, up-to-date clinical trials based on patient phenotype and genomic markers. In other embodiments, the agent 903 may invoke a dosage information retriever when queries concern pharmacological guidance, or a guideline summarizer to provide evidence-based care pathways.
[0207] The tools 901 are configured to execute asynchronous retrievals and post-processing operations on curated vector stores, which are divided into discrete silos based on the type andsource of data. In some embodiments, the silos include, but are not limited to, treatment guidelines, genomics data, drug references, clinical trial registries, and disease-specific literature. Each silo can be associated with a dedicated retriever tool and tailored postprocessing pipeline.
[0208] In some embodiments, the vector search within each silo combines both sparse and dense embeddings to support robust hybrid retrieval. Dense embeddings may be generated using a proprietary model such as Voyage Al, while sparse embeddings may be created using a model such as SPLADE. The hybrid approach ensures high semantic recall even in the presence of terminological variation or limited keyword specificity.
[0209] The intelligent agent 903 leverages a ReAct-style architecture, which enables it to dynamically determine the optimal course of action through a reasoning-and-acting framework. The agent maintains transparency in its operations by exposing intermediate reasoning steps and tool selections, allowing users to inspect or validate the factual basis for the final recommendation. In some embodiments, this visibility is provided through the user interface described in FIGS. 8A-8D, offering an interpretable chain-of-thought behind the model’s output.
[0210] FIG. 10 depicts a representative interaction flow within the agent-based framework, illustrating how a user query is decomposed into sub-tasks and processed through various tools before being synthesized into a comprehensive response. As illustrated, a clinician's query regarding treatment strategy for a BRCA-positive breast cancer patient may initiate a multi- step workflow. The agent 903 may first retrieve guideline summaries, then request drugspecific data including dosage safety, and finally scan clinical trials for eligibility criteria. Outputs from each tool are routed back into the agent’s internal state before the large language model generates a final, cohesive recommendation.
[0211] In certain embodiments, and confirmed through case studies and testing, as referencedin Improved Precision Oncology Question- Answering Using Agentic LLM by Rangan Das et al. (September 2024), 10.1101 / 2024.09.20.24314076, incorporated in their entirety herein, the system described in FIGS. 9 and 10 — designated GeneSilico Copilot or GSCP in such publications — exceeds state-of-the-art performance benchmarks for breast cancer decision support. The system’s ability to contextualize complex medical queries using distributed, specialized retrieval allows for highly tailored oncotherapy recommendations. These capabilities render the system suitable for real-world clinical deployment in precision oncology environments.
[0212] FIG. 11 evidences challenges associated with semantic clustering in vector embeddingbased retrieval for personalized oncology decision support, in accordance with certain embodiments of the present disclosure.
[0213] As shown in FIG. 11, conventional vector-based document retrieval systems can experience improper partitioning of the search space when applied to narrow biomedical domains such as breast cancer. Specifically, semantically similar phrases across thematically distinct documents may result in close proximity within the vector space, thereby causing information from unrelated sources to cluster together. This phenomenon can result in the unintentional intermixing of documents from distinct medical guideline sources, diluting the precision of information retrieval.
[0214] As illustrated in FIG. 12, the volume of content across these silos may vary significantly. For example, breast cancer guidelines from the National Comprehensive Cancer Network (NCCN) comprise approximately 17,100 tokens, while the clinical trial data spans over 3.4 million tokens and the PubMed-derived content includes approximately 460,000 tokens. Despite the relatively small token count of the NCCN guidelines, this silo is disproportionately influential in synthesizing clinically relevant responses, as shown in FIGS.13A-13C
[0215] To address this issue and to manage the significant data imbalance across document types, the system can implement a Siloed Abstractive Vector Store (SVS). This architecture, in certain embodiments, partitions the vector database into thematically distinct silos, each of which corresponds to a specific domain or data source. In some embodiments, these silos include, but are not limited to, NCCN Guidelines, ASCO Guidelines, PubMed articles, PharmGKB data, and Clinical Trial records.
[0216] The SVS framework assigns a dedicated retrieval module to each silo. In some embodiments, each module applies customized filtering, summarization, and post-processing logic to ensure that only relevant content is accessed in response to a query. These retrieval modules operate asynchronously, allowing the agentic architecture described in FIGS. 9-10 to coordinate multi-silo retrieval without risk of cross-contamination. For instance, when the agent issues a query requiring dosage guidance, the response is synthesized solely from the PharmGKB or guideline silo, without inadvertently drawing from unrelated literature in other silos.
[0217] In some embodiments, each document stored within the SVS includes both chunked content and a corresponding abstractive summary. These summaries are tailored to the document’ s use case and the intended silo context. For example, a clinical trial document may be summarized differently in the Clinical Trials silo versus the PharmGKB silo, depending on whether the emphasis is on trial eligibility criteria or pharmacogenomic findings.
[0218] Each chunk is linked to its parent summary via metadata, enabling dual-layer retrieval. In some embodiments, when a chunk is retrieved from a silo, the corresponding summary is retrieved concurrently, allowing the agent to assess the broader semantic context of the document. This linkage enhances the agent’s ability to reason efficiently, reducing the number of iterative reasoning-and-action loops required to generate a coherent response.
[0219] Where the internal ordering of document content is clinically meaningful, such as instepwise treatment protocols or temporally sequenced trial data, the system may employ a large language model to re-order or reinterpret chunked segments based on summary guidance. In such embodiments, the LLM can restructure the retrieved content to ensure that the generated response retains the correct narrative and logical flow.
[0220] Further, prior to delivery to the agent, all retrieved content — including document chunks, summaries, and source metadata — may undergo further refinement using a designated LLM for consistency of tone, clarity of information, and clinical relevance across all output. The metadata retained with each result can also be used to cite the origin of the information, facilitating transparency and interpretability for the clinician.
[0221] FIG. 14 illustrates the construction and composition of a comprehensive evaluation suite used to assess the performance of an agentic retrieval-augmented generation (RAG) system in comparison to conventional RAG systems and standalone large language models (LLMs), in accordance with certain embodiments of the present disclosure.
[0222] As shown in FIG. 14, due to the absence of a publicly available breast cancer-specific question-answering dataset, a multi-source evaluation corpus was constructed. This evaluation suite was compiled from a combination of objective and subjective medical question-answering datasets, including MedMCQA, MedQA, PubMedQA, and internally generated case-specific queries. In some embodiments, the dataset includes over 220 objective questions and 110 open- ended questions, encompassing a range of complexity, format, and domain relevance.
[0223] For the objective subset, questions were extracted and processed through a Pythonbased script that identified all potential answer options within each dataset entry. In certain embodiments, this script compared the generated model response to the full set of plausible choices, accommodating cases where the correct answer was present only implicitly within the model’s output.
[0224] FIGS. 15A-15C present the results of performance benchmarking between agenticsystems, basic RAG systems, and standalone LLMs. In some embodiments, evaluations were conducted using multiple foundation models, including GPT-4 and Claude Opus 3, both of which are known for competitive benchmark performance.
[0225] In the context of objective question answering, as shown in FIG. 15A, four quantitative metrics — accuracy, precision, recall, and Fl -score — were used to assess each system’s performance across 223 questions. As depicted in FIGS. 15A-15C, agentic systems significantly outperformed both basic RAG and standalone LLMs across all metrics. In some embodiments, an agentic system leveraging GPT-4 achieved uniform scores of 0.83 for all four metrics, representing the highest observed performance in this configuration. An agentic system powered by Claude Opus 3 also achieved high performance, though slightly lower than GPT-4. In contrast, basic RAG systems showed mixed performance, with relatively strong precision but significantly lower recall and Fl -scores, while standalone LLMs scored lowest across all categories.
[0226] For subjective question answering, as shown in FIG. 15B, consisting of 113 open- ended questions, performance was evaluated using the DeepEval framework, which assesses both retrieval and generation quality. Evaluation metrics included context precision, context relevancy, faithfulness, and answer relevancy. Standalone LLMs were excluded from this phase due to their inability to cite retrieval sources. In some embodiments, agentic configurations again demonstrated superior performance. Agentic (Claude Opus 3) achieved the highest context precision (0.44), while Agentic (GPT-4) led in answer relevancy with a score of 0.96. Both agentic configurations significantly outperformed basic RAG setups in context relevancy (0.27 vs. <0.15) and demonstrated strong fidelity to source material.
[0227] In a further evaluation, as shown in FIG. 15C, a custom Precision Oncology QA dataset was constructed to simulate real-world medical complexity in breast cancer genetics and personalized treatment. This dataset, comprising 25 carefully designed questions, includedscenario-based queries involving genetic markers, disease staging, and personalized therapeutic interventions. In this evaluation, agentic systems again outperformed their baseline counterparts. Agentic (GPT-4) achieved a context precision of 0.52 and a relevancy score of 0.80, while Agentic (Claude Opus 3) scored 0.51 and 0.82, respectively. Basic RAG models scored notably lower in both metrics, with precision scores ranging from 0.15 to 0.20 and relevancy scores of 0.55 across both LLMs.
[0228] In addition to benchmark-style questions, synthetic patient case studies were developed to replicate real-world clinical decision-making. These cases were constructed through a collaborative process involving practicing oncologists and the LLM-based system. In some embodiments, the oncologists contributed domain-specific knowledge — including patient history patterns, mutation profiles, and prior treatments — while the LLMs generated coherent case narratives based on this input. Final validation of the cases was conducted by the oncologists to ensure clinical authenticity.
[0229] Using these synthetic cases, the agentic systems were tasked with formulating complete treatment plans based solely on the provided case information. In such embodiments, agentic systems successfully demonstrated the ability to synthesize appropriate treatment pathways across a range of breast cancer subtypes, with outputs reflecting alignment with current clinical guidelines and practice norms. The integration of case-specific retrieval, structured planning, and LLM reasoning confirmed the system’s readiness for clinical decision support applications in personalized oncology.
[0230] FIGS. 16A-16C depicts a representative sequence of tool utilization events by the GSCP agent during evaluation scenarios, in accordance with certain embodiments of the present disclosure.
[0231] As shown in FIGS. 16A-16C, one of the core strengths of the GSCP framework lies in the orchestration of a suite of specialized tools operating across distinct thematic silos. Thesetools — each aligned with a particular data source or medical domain — are invoked in sequence by the agent, depending on the evolving needs of a query. This coordination allows for highly targeted information retrieval and synthesis, while also exposing the agent’s reasoning steps to the end user. In some embodiments, this visibility enables clinicians to trace which sources contributed to each portion of the system’s final response.
[0232] Tool usage statistics derived from the evaluation datasets reveal distinct usage patterns. For instance, NCCN guidelines emerged as a primary information source, particularly in responses to objective case-based questions. In certain embodiments, PubMed was also widely utilized, often serving as a secondary source to supplement guideline-based responses. FIGS. 13A-13C, as previously discussed, provides a comparative view of the contribution of each tool across objective, subjective, and case-based scenarios.
[0233] In most queries, the agent completes response generation within three coordinated reasoning steps. In complex cases, up to five tool invocation cycles may occur. This progressive breakdown is shown in FIGS. 16A-16C, and in certain embodiments provides insight into the cognitive path taken by the agent during response construction. Such transparency in planning enhances trust and interpretability for the system’s end users.
[0234] Beyond architecture and tooling, the GSCP framework advances the state-of-the-art in clinical LLM application by improving both response quality and reasoning fidelity. As detailed above, the GSCP system achieved marked improvements in faithfulness (up to 15.29%), context precision (up to 200.83%), and context relevancy (up to 47.27%) over baseline RAG configurations in the domain of precision oncology. In subjective QA tasks, GSCP surpassed prior methods with up to 93.65% gain in context precision and 2600% improvement in relevancy, confirming the efficacy of the agent-based reasoning model.
[0235] In some embodiments, the GSCP system employs pre-processed documents enriched with markdown annotations and structured hierarchical chunking. This approach allows theagent to reason more effectively about topic structure within documents, leading to clearer and more contextually aligned responses. The documents — particularly guidelines and literature — are divided by subtopic to facilitate targeted access and more granular synthesis by the agent.
[0236] Each tool within the GSCP framework is optimized for different information volumes and embedding strategies. In some embodiments, dense and sparse embeddings are combined using hybrid retrieval models, such as Voyage Al for dense embeddings and SPLADE for sparse. This embedding heterogeneity allows the GSCP agent to balance semantic similarity with keyword specificity in real time, improving retrieval quality across diverse query types.
[0237] While GPT-4 and Claude Opus 3 both demonstrated strong performance within the GSCP framework, qualitative evaluations revealed differences in response style. Claude Opus 3 was noted for producing well-structured, highly readable responses appreciated by oncology experts, albeit with slightly slower output times (often up to two minutes per case). These responses included nuanced explanations, detailed dosage information, and structured breakdowns, highlighting the value of human-perceived clarity in high-stakes domains.
[0238] Although both foundation models achieved comparable levels of medical accuracy, the superior readability of Opus 3 responses — despite lower automated evaluation scores — underscores a limitation of current frameworks such as DeepEval. These tools, which rely on LLMs for model evaluation, may not fully capture critical human-centric aspects such as coherence, clinical utility, and professional readability. Accordingly, future evaluation designs may incorporate human expert panels and rubric-based assessment to supplement automated benchmarks.
[0239] In certain embodiments, the GSCP system’s transparency further extends to its internal preference patterns. For example, analysis revealed a tendency to rely more heavily on NCCN guidelines, which underwent manual paraphrasing to convert structured flowcharts into coherent plain text. ASCO and ESMO guidelines, by contrast, were primarily summarizedusing LLM-based methods and saw lower utilization rates. This disparity in usage may inform future optimization efforts, such as retraining or curation of underutilized silos.
[0240] Similarly, PubMed emerged as a consistent and valuable source, particularly when filtered and categorized to emulate a domain-specific collection. In some embodiments, this transformation was accomplished through initial topic modeling and summarization prior to indexing, allowing the agent to extract oncology-relevant material more efficiently.
[0241] In certain embodiments, the GSCP framework can incorporate evaluation on real-life patient data and expand its clinical decision support capabilities. These improvements will require the development of novel, reproducible, and transparent evaluation methodologies. Because current tools like DeepEval may yield variable outputs based on internal prompt drift or model bias, more robust frameworks are being developed to assess planning and reasoning steps explicitly.
[0242] System-level performance remains an area for optimization. In some embodiments, tool invocation latency and vector store responsiveness are the principal bottlenecks. Current deployments rely on standard hardware for vector retrieval, contributing to delays in processing complex queries. Planned enhancements include hardware acceleration and vector quantization techniques to significantly improve retrieval speeds.
[0243] To generate a clinically viable treatment plan, a physician weighs multiple patientspecific variables including genetic profiles, comorbidities, prior therapies, and risk factors. This decision-making process can involve referencing evidence-based guidelines such as those published by NCCN, ASCO, or ESMO, in addition to consulting current clinical trials. In some embodiments, the GSCP agent addresses this complexity by integrating patient history, genomic data, and real-time literature synthesis to provide tailored treatment recommendations in alignment with these guidelines.
[0244] The GSCP system streamlines physician workflow by aggregating high-quality data,performing guided retrieval, and synthesizing personalized recommendations — all while preserving visibility into the agent’s reasoning path. This framework offers a promising solution for augmenting clinical judgment in oncology and beyond.
[0245] The present disclosure, including but not limited to the GSCP architecture, demonstrates the feasibility and utility of domain-specific, agent-based retrieval systems tailored for complex medical environments. By integrating structured planning, transparent reasoning, and optimized tooling, the system represents a significant advancement in the clinical deployment of LLM-based technologies.
[0246] FIG. 17 depicts a hierarchical multi-agent framework, designated as gSage, configured to provide transparent and clinically structured decision support in precision oncology, in accordance with certain embodiments of the present disclosure.
[0247] As shown in FIG. 17, the gSage architecture is structured as a multi-layered reasoning pipeline, designed to mimic a medical oncologist’s approach to therapeutic decision-making. The system receives patient-specific data inputs — including clinical history, pathological findings, demographic characteristics, and next-generation sequencing (NGS) results — and generates an evidence-based treatment recommendation accompanied by an interpretable reasoning trace.
[0248] At the core of the system is a primary agent, responsible for planning the decisionmaking process. In some embodiments, this primary agent analyzes the patient profile and formulates a case-specific reasoning plan. If required information is missing or insufficient, the agent may flag diagnostic knowledge gaps and suggest additional steps such as germline testing or immunohistochemical assays.
[0249] The execution of the plan is delegated to a set of five domain-specific research agents, each associated with a distinct clinical function. These agents include: (1) a guideline interpretation agent, (2) a drug synthesis agent, (3) a clinical trial matching agent, (4) aliterature retrieval agent, and (5) an evidence aggregation agent. As shown in FIG. 18, the primary agent decomposes a complex clinical question into structured sub-tasks, each of which is assigned to one or more research agents. These sub-tasks are processed in iterative loops that allow for feedback refinement and agent coordination.
[0250] Each research agent interacts with a curated vector database containing siloed biomedical content including treatment guidelines, drug databases, peer-reviewed literature, clinical trial records, and internal expert-reviewed materials. In some embodiments, the agents utilize hybrid semantic search algorithms and re-ranking mechanisms to retrieve the most relevant information.
[0251] Upon retrieval, each research agent synthesizes its findings by identifying clinically meaningful associations between key variables. For example, the system may link the presence of a TP53 mutation with lymphovascular invasion in a HER2-negative breast cancer case and use this relationship to inform treatment escalation strategies. In some embodiments, this synthesis includes structured outputs that are returned to the primary agent for integration.
[0252] The gSage system compiles a final recommendation that includes both a proposed clinical action, such as for example, but not limited to, recommended therapy, diagnostic follow-up, trial enrollment) and a multi-step reasoning trace. In some embodiments, this trace may include, in initial case assessment and planning outline, details on which agents were invoked and what tools were used, intermediate reasoning outputs and refinements across agent cycles, final synthesis with inline citations linking directly to authoritative sources, or combinations thereof.
[0253] By preserving a complete, auditable reasoning pathway, gSage enhances transparency and clinical interpretability. In certain embodiments, this structured trace can be displayed to end users — such as oncologists or tumor boards — to support validation, peer review, or regulatory documentation.Working Example Testing and Results
[0254] To benchmark system performance, a series of evaluation experiments were conducted using synthetic and real -world clinical data. In one such evaluation, 15 anonymized breast cancer patient profiles were curated in collaboration with medical oncologists. Each case included comprehensive phenotypic characterization and annotated NGS data derived using the GeneSilico Gene Panel. A total of 66 unique somatic alterations were identified and cataloged with clinical relevance annotations.
[0255] Using this dataset, a structured suite of 46 oncology-specific queries was developed and paired with ground-truth answers agreed upon by domain experts. These queries were designed to probe different facets of clinical reasoning, including therapeutic selection, biomarker interpretation, and guideline adherence. Performance was measured across multiple dimensions including accuracy, clinical soundness, and patient-centered relevance.
[0256] In comparative evaluations across five open-source LLMs and three proprietary foundation models, the gSage system achieved the highest accuracy, scoring 83.4%, surpassing the next-best model — Anthropic Sonnet 3.5 — which scored 81.8%. gSage also demonstrated superior performance in clinical reasoning quality and alignment with patient-specific variables.
[0257] In some embodiments, the evaluation framework incorporates an LLM-enabled, oncologist-driven jury system. In this configuration, language models are used to assist in scoring the correctness and completeness of the responses, while final adjudication is performed by human experts. This hybrid scoring model is intended to reduce bias and improve consistency across evaluation cycles.
[0258] By combining modular reasoning, transparent planning, and curated retrieval, gSage provides a clinically oriented alternative to black-box LLM tools. Its reproducible architecture and evaluation strategy offer a template for future Al copilots in medicine, with the potentialto accelerate safe and effective integration of large language models into oncology workflows.
[0259] FIG. 18 depicts a task delegation and coordination architecture for a hierarchical multiagent system configured to generate clinically validated oncology treatment recommendations, in accordance with certain embodiments of the present disclosure.
[0260] As shown in FIG. 18, the primary agent can coordinate five domain-specific research agents, each responsible for performing targeted information retrieval and synthesis aligned with oncology subdomains. The structure enables complex queries to be decomposed, evaluated, and resolved through iterative tool invocation and feedback loops, culminating in a consolidated, explainable output that integrates evidence across multiple trusted sources.
[0261] FIG. 19 depicts a schematic representation of a multi-vector embedding space used to organize oncology knowledge by clinical domain, in accordance with certain embodiments of the present disclosure.
[0262] To ensure the trustworthiness and clinical relevance of its outputs, the gSage system integrates a curated biomedical knowledge corpus, developed with direct contributions from medical oncologists and domain experts. This corpus includes both publicly available resources — such as OpenFDA, PubMed, OpenTargets, and ClinicalTrials.gov — and internally authored content designed to align with recognized international oncology standards, including those promulgated by the NCCN, ASCO, and ESMO.
[0263] In certain embodiments, every internally created guideline is linked to peer-reviewed publications or clinical study results, providing an auditable chain of evidence for each recommendation. This curated content is housed alongside external data sources in a centralized knowledge base, enabling a unified framework for semantic search, retrieval, and evidence ranking.
[0264] To facilitate high-precision information retrieval, gSage employs a vector database constructed with multi-embedding representations. In some embodiments, embedding modelssuch as Snowflake and MedCPT are utilized to generate both dense and sparse representations of source documents. The system supports query-adaptive retrieval strategies — meaning that the method of evidence selection varies depending on the clinical objective of the query.
[0265] For example, in certain embodiments, if the query pertains to clinical trial matching, the retrieval engine emphasizes eligibility criteria and patient-specific biomarkers. Conversely, when the objective is to retrieve landmark study evidence, the system prioritizes high -impact trials, drug efficacy data, and regulatory approvals. The multi-vector embedding architecture allows the system to dynamically navigate between granular patient data and broad populationlevel findings.
[0266] As shown in FIG. 19, this multi-dimensional embedding space partitions oncology knowledge into clinically coherent regions. These partitions are updated dynamically to reflect evolving evidence standards and maintain alignment with current clinical practice. This design helps mitigate the retrieval of semantically similar but contextually irrelevant documents, preserving fidelity across diverse clinical use cases.
[0267] FIG. 20 depicts a structured knowledge trace and source mapping for the gSage system’s decision-making process, in accordance with certain embodiments of the present disclosure.
[0268] In some embodiments, every response generated by gSage includes a multi-step reasoning trace. This trace explicitly identifies which knowledge sources were consulted, what retrieval strategies were employed, and how individual research agents contributed to the synthesis. As shown in FIG. 20, this structured representation includes source citations, logical inference steps, and an audit trail that links final recommendations to their supporting evidence.
[0269] By exposing its full reasoning path, gSage transforms from a black-box LLM system into an auditable decision-support tool. In certain embodiments, the transparency mechanisms described herein enable physicians to verify, contest, or further investigate specific componentsof the system’s recommendation. This is especially valuable in high-stakes clinical contexts, where interpretability and accountability are essential to responsible Al deployment.
[0270] In combining expert curation, structured retrieval, and evidence-linked reasoning, the gSage system delivers clinically rigorous recommendations while maintaining transparency at each step of the process. This design ensures not only trustworthiness but also adaptability, allowing the system to evolve with emerging medical literature and guideline updates.
[0271] As further discussed below, gSage achieved competitive performance on benchmarking tasks, demonstrating its potential as a deployable, evidence-driven Al system in real-world oncology workflows.
[0272] FIG. 21 depicts the distribution of data sources accessed by the gSage system during processing of objective oncology-related questions, in accordance with certain embodiments of the present disclosure.
[0273] FIG. 22 depicts a structured mapping of the gSage decision-making process in response to objective multiple-choice questions, illustrating the integration of literature, clinical guidelines, trial evidence, and drug information, in accordance with certain embodiments of the present disclosure.
[0274] FIG. 23 depicts a categorical distribution of oncology-focused multiple-choice questions used to evaluate the performance of various large language models, in accordance with certain embodiments of the present disclosure.
[0275] FIG. 24 depicts comparative model performance on a breast cancer multiple-choice question (MCQ) benchmark dataset across four key evaluation metrics — accuracy, precision, recall, and Fl -score — in accordance with certain embodiments of the present disclosure.
[0276] To assess the performance of the gSage system in clinically relevant question answering tasks, a structured evaluation was conducted using a breast cancer-specific benchmark composed of 265 multiple-choice questions. These questions were drawn from establisheddatasets, including MedQA, MedMCQA, and PubMedQA, as well as an internally curated dataset developed in collaboration with medical oncologists. Each question included a single ground truth answer verified by domain experts to ensure medical validity.
[0277] In some embodiments, the benchmark questions were categorized across key clinical domains, including diagnosis and staging, treatment modalities, complications and toxicities, tumor biology, and genetic profiling. As shown in FIG. 23, the largest share of questions focused on core oncology tasks such as therapeutic strategy and disease progression management, supporting a robust evaluation of model competency.
[0278] Performance metrics for the evaluation included accuracy, precision, recall, and Flscore, with results benchmarked against multiple leading foundation models. These included both proprietary LLM services such as GPT-4o and Anthropic Sonnet 3.5, as well as open- source biomedical models such as BioMistral, UltraMedical, OpenBioLLM, Phi-4-14B, and Llama 3.3-70B.
[0279] As shown in FIG. 24, the gSage system outperformed all baseline models across every evaluated metric. In certain embodiments, gSage achieved an accuracy of 0.92, precision of 0.84, recall of 0.85, and an Fl-score of 0.85. Its nearest proprietary competitors, GPT-4o and Sonnet 3.5, achieved accuracy scores of 0.84 and 0.89, respectively, but fell short in precision and recall — demonstrating less consistent alignment with oncologic standards and evidencebased guidelines.
[0280] Open-source models performed significantly below the proprietary and agentic systems. For example, BioMistral achieved an accuracy of only 0.28 and a precision of 0.18, indicating limited applicability in specialized oncology settings. These results highlight the importance of domain-specific model design and the limitations of broad-domain pretraining in high-precision clinical contexts.
[0281] As shown in FIG. 21, gSage's superior performance in multiple-choice benchmarks canbe attributed to its integrated use of diverse and clinically validated information sources. During response synthesis, the system frequently accessed literature databases, drug knowledge graphs, treatment guidelines, and trial registries. The contribution of each source varied depending on question type and complexity, as further detailed in FIG. 22, which maps the structured reasoning flow employed during evaluation.
[0282] In some embodiments, gSage's success in objective QA tasks is driven by its query- adaptive retrieval strategies, which enable differential prioritization of evidence based on user intent. For instance, when confronted with a drug interaction question, the system retrieves information from drug interaction modules and regulatory filings; whereas for treatment sequencing, it draws from NCCN pathways and historical clinical trials. This tailored retrieval behavior results in highly relevant and context-sensitive answers.
[0283] Unlike traditional RAG models that return static citations or generic content, gSage integrates these sources into a coherent, auditable reasoning trace. Each final answer is grounded in its supporting evidence and reflects clinical best practices. In certain embodiments, this interpretability allows for greater trust among oncologists and facilitates integration into multidisciplinary clinical workflows.
[0284] While these results demonstrate the depth and precision of gSage’ s knowledge base, objective QA tasks are limited in their ability to assess complex reasoning or individualized clinical interpretation. Accordingly, subsequent evaluations focus on the system’s performance in subjective question answering, where multi-step inference and patient-specific nuance are critical.
[0285] FIG. 25 depicts a clinical step-by-step timeline for a representative breast cancer patient case, illustrating diagnostic assessments, treatment progression, and follow-up evaluations, in accordance with certain embodiments of the present disclosure.
[0286] FIG. 26 depicts a schematic operational flow of the gSage system when applied topatient-specific case queries, showing how prognostic markers, molecular features, and guideline-driven evidence are integrated into a structured decision-making pathway, in accordance with certain embodiments of the present disclosure.
[0287] FIG. 27 depicts a categorical distribution of clinical questions used in subjective casebased evaluations, reflecting the relative focus on diagnostic, therapeutic, molecular, and surveillance-related topics, in accordance with certain embodiments of the present disclosure.
[0288] To evaluate gSage in realistic clinical contexts, a benchmark was developed using Deep Clinical and Genomic Patient Dossiers (DCGPDs) — structured case timelines derived from actual breast cancer patients who underwent comprehensive next-generation sequencing (NGS) profiling. The initial source cohort consisted of 42 patients sequenced using a custom 68-gene breast cancer panel. Clinical records were compiled from physician notes, pathology reports, imaging studies, treatment histories, and molecular test results. These data were deidentified and curated by practicing oncologists to ensure both clinical accuracy and privacy compliance.
[0289] From the 42-patient source pool, 15 dossiers were selected based on the completeness and clinical relevance of their underlying records. Each dossier was transformed into a time- sequenced narrative capturing the patient’s oncologic journey, with particular emphasis on molecular data integration and decision points in accordance with prevailing standards of care.
[0290] To assess gSage’ s capacity for applied clinical reasoning, a total of 46 scenario-based queries were developed from the selected dossiers. Each scenario simulated a real-world oncologic decision-making moment — such as interpreting biomarker significance, selecting therapy after recurrence, or evaluating trial eligibility — and was accompanied by three to four follow-up queries crafted by medical oncologists. These questions were designed to test not only factual knowledge but also diagnostic logic, treatment planning, and contextual awareness.
[0291] In some embodiments, evaluation criteria were defined across three primary dimensions: (1) Evidence-Based Accuracy, which measures the extent to which the model’s outputs align with published guidelines, trial data, and authoritative clinical literature; (2) Clinical Reasoning, which assesses the model’s ability to integrate diverse patient data and construct a logical treatment rationale; and (3) Patient-Centricity, which evaluates the personalization of the response in light of the patient’s comorbidities, treatment history, and genomic profile.
[0292] Ground truth responses for each scenario were independently authored and reviewed by board-certified oncologists. These answers served as the benchmark against which gSage and comparative models were assessed. Unlike binary MCQ scoring, the subjective nature of this evaluation necessitated a robust scoring rubric that reflected clinical nuance.
[0293] To mitigate potential bias in the evaluation process, a two-tiered review architecture was implemented. In the first tier, language model-based jurors conducted initial response assessments using structured prompts engineered by medical experts. In the second tier, oncologists independently reviewed the anonymized scoring outputs, confirming clinical validity and adjudicating disagreements. This multi-layered review structure, as illustrated in FIG. 25, can be utilized to enhance and ensure consistency, transparency, and domain alignment.
[0294] As shown in FIG. 26, gSage’ s response to a patient-specific query proceeds through a multi-step operational flow. Upon receiving a structured case query, the system identifies relevant prognostic and molecular markers, retrieves guideline-aligned recommendations from trusted knowledge silos, and composes a reasoning trace that connects clinical context with evidence-backed interventions.
[0295] In some embodiments, the diversity of clinical questions was distributed across key domains of oncology practice, as illustrated in FIG. 27. In certain embodiments, these caninclude therapeutic planning, disease monitoring, surgical timing, radiologic interpretation, genetic counseling, and trial enrollment eligibility.
[0296] As shown through the working examples, the framework as illustrated in FIGS. 25-27, evidences the robustness of gSage in simulated real-world use cases.
[0297] FIG. 28 depicts comparative scoring of large language models across three dimensions — clinical reasoning, evidence-based accuracy, and patient-centricity — during case-based evaluations derived from real-world breast cancer patient scenarios, in accordance with certain embodiments of the present disclosure.
[0298] Following generation of case-specific responses by each model, a structured scoring protocol was implemented to evaluate clinical fidelity. In certain embodiments, each model response was initially assessed by LLM-based evaluators configured with oncologist-designed prompts. These preliminary scores were then subjected to a blind review process conducted by board-certified oncologists, who validated, refined, or corrected the scoring outputs to ensure alignment with real-world clinical expectations.
[0299] As shown in FIG. 28, gSage achieved the highest aggregate scores across all three primary evaluation dimensions. For clinical reasoning, the gSage system achieved a mean score of 4.65, exceeding the performance of leading closed-source systems, including GPT-4o (3.61) and Sonnet 3.5 (4.30). Open-source models trailed further behind, with Llama 3-70B scoring 3.33 and OpenBioLLM scoring 2.80.
[0300] In evidence-based accuracy, gSage continued to lead, receiving a score of 4.17. This surpassed GPT-4o (3.17) and Sonnet 3.5 (4.09), with Llama 3-70B (3.43) and OpenBioLLM (2.70) again reflecting more limited ability to integrate guideline-backed content or reference authoritative oncology sources.
[0301] For patient-centricity, gSage demonstrated a strong capacity to incorporate individual clinical context into its responses, achieving a score of 4.39. This performance outpaced GPT-4o (3.30), Sonnet 3.5 (3.99), and the lower-scoring open-source models, including Llama 3- 70B (3.28) and OpenBioLLM (2.46).
[0302] In some embodiments, gSage’s superior performance reflects its multi-agent, silo- aware architecture, which enables dynamic retrieval and integration of heterogeneous clinical data points, including patient history, molecular biomarkers, and prior treatment response. This facilitates a structured synthesis process that aligns with complex oncologic decision-making pathways.
[0303] Further, in certain scenarios, gSage was observed to proactively identify information gaps — such as missing biomarker tests or unverified germline mutations — and recommended follow-up diagnostics or surveillance measures consistent with prevailing standards of care. This proactive reasoning behavior contributed positively to its scoring across all dimensions.
[0304] In contrast, baseline models — particularly open-source foundation models — frequently generated responses that lacked patient specificity, omitted essential clinical context, or failed to prioritize the most relevant sources of evidence. This resulted in lower clinical reasoning scores and diminished trustworthiness of outputs during blinded evaluation.
[0305] Accordingly, the evaluation results illustrate that gSage not only retrieves the correct data but also interprets and applies it through a clinician-like reasoning process. Its ability to produce personalized, guideline-concordant recommendations in complex patient scenarios reinforces its applicability as a reliable Al co-pilot in clinical oncology.Working Example Methodology for Cancer -Specific Corpus Development and Knowledge Representation
[0306] In certain embodiments, a domain-optimized corpus is utilized to support accurate, patient-tailored recommendations in precision oncology. The corpus, referred to herein as the “GeneSilico Corpus,” was constructed to provide an expert-aligned foundation for the gSage decision-support system. Unlike conventional biomedical corpora sourced indiscriminately from public datasets, the GeneSilico Corpus is manually curated by domain experts toemphasize clinical relevance, evidence traceability, and oncologic utility. The corpus integrates five primary content domains: (1) structured genomic insight, (2) pharmacological relationships, (3) treatment protocols and global guidelines, (4) peer-reviewed biomedical literature, and (5) clinical trial metadata.
[0307] In certain implementations, the corpus is anchored by a custom-designed 68-gene panel, selected based on therapeutic relevance, oncologic frequency, inclusion in existing diagnostic assays, normalized gene lengths, and representation of microsatellite instability hotspots. These genes include, for example, BRCA1, BRCA2, TP53, PIK3CA, ESRI, ATM, CHEK2, and PALB2, among others commonly implicated in homologous recombination repair (HRR) pathways or pharmacogenomic interactions.
[0308] Pharmacologic data is collected from multiple open-source and regulatory-aligned databases, including OpenTargets, the U.S. Food and Drug Administration (FDA), the European Medicines Agency (EMA), RxList, and CIViCDB. Drug-gene associations are recorded with metadata indicating approved uses, investigational status, biomarker linkage, and known contraindications. Each entry is cross-referenced to ensure consistency with source literature and label indications. For treatment guidelines, manually synthesized documents are developed to align with NCCN, ASCO, ESMO, and BCCA protocols, enabling interpretation across standard-of-care frameworks.
[0309] Literature sources are harvested from PubMed using targeted query filters specific to oncology, with documents selected, reviewed, and approved by clinicians for corpus inclusion. Unlike broad biomedical models, gSage excludes off-topic literature, maintaining focus on decision-relevant studies such as pivotal clinical trials, meta-analyses, and large cohort reviews. Clinical trial content is aggregated from ClinicalTrials.gov, HSMD, PubMed, and internal registries, and is enriched with trial metadata including eligibility criteria, phase, biomarker targets, geographic recruitment zones, and study design.
[0310] Working Example Methodology for Vector Store Design for Adaptive KnowledgeRetrieval
[0311] In certain embodiments, the gSage system employs a multi-tiered vector database tailored to precision oncology. The vector store is structured to support both dense and sparse retrieval strategies across semantically complex inputs, enabling responsive, query-adaptive search behavior. The documents are classified into two categories: (1) structured data sources (e.g., JSON-encoded trial metadata or drug-target mappings), and (2) plain-text oncology knowledge documents, such as for example, but not limited to, guidelines, literature, and clinical protocols /
[0312] All documents undergo preprocessing through a semantic summarization pipeline prior to embedding. Summarization is performed using an LLM model such as GPT-4o, and the generated summaries are linked to the source chunks via document metadata. For structured data, key -value pairs are selected and flattened into narrative representations to permit semantic interpretation. Embedding is performed using a combination of open-source models optimized for clinical and biomedical text, including Snowflake-xs (summaries), Snowflake-m and MedCPT (full texts and structured fields), and BM42 (sparse lexical matching). This hybrid strategy supports both broad context discovery and precision targeting within the oncology corpus.
[0313] Query resolution is further enhanced by enabling downstream tools and research agents to dynamically select from among the available embedding strategies. For example, a hypothesis-driven search may prioritize semantic similarity, while a regulatory lookup or drug interaction check may leverage sparse keyword matching. This flexibility mitigates issues such as embedding collapse or retrieval drift, which often affect static LLM embedding workflows in dense clinical domains.Working Example Methodology for Hierarchical Agentic Reasoning for Oncology Workflows
[0314] In certain implementations, the gSage system utilizes a hierarchical agentic structure developed using the Llamalndex Workflows framework. A central planning agent orchestrates five domain-specific research agents, each aligned with a major oncology reasoning function. These include: (1) guidelines and protocols interpretation, (2) drug information synthesis, (3) evidence-based medicine review, (4) clinical trial matching, and (5) research and literature enrichment. Each research agent operates as a tool-using sub-agent with scoped access to specific corpus segments via structured API tools.
[0315] Upon receiving a patient-specific query, the primary agent formulates a structured clinical plan, parsing molecular diagnostics, disease stage, prior treatment history, and comorbid conditions. The agent identifies any missing context and recommends additional testing if necessary, such as for example, but not limited to, germline sequencing or molecular subtyping. The primary agent then decomposes the query into sub-tasks and delegates execution to the research agents.
[0316] Each agent conducts retrieval via tool-assisted vector search, using embeddings appropriate to the query type. Agents return structured outputs with citations, which are synthesized by the primary agent into a final markdown response. The final response includes: (1) a stepwise reasoning trace; (2) a reference list for all retrieved documents; and (3) a summary of key clinical insights. Intermediate outputs are accessible to the user and facilitate auditability, transparency, and extension.
[0317] Response times average between 40 and 70 seconds, with first-step planning initiated within seconds of submission. This staged reasoning architecture simulates a tumor board-like process and enables iterative refinement across multiple rounds. Tool re-engagements are performed when information gaps remain or when user queries require layered inference across agents. This approach supports the generation of evidence-backed, verifiable, and context-rich oncology recommendations.Working Example Methodology for Deep Clinical and Genomic Patient Dossier Construction and Model Evaluation
[0318] To enable clinically grounded benchmarking of gSage, a structured evaluation framework was developed using Deep Clinical and Genomic Patient Dossiers (DCGPDs). These dossiers represent high-fidelity reconstructions of real-world breast cancer patient journeys, integrating longitudinal clinical histories, diagnostic findings, treatment courses, and next-generation sequencing (NGS) results. This framework enables comprehensive assessment of Al-driven clinical decision support systems in a precision oncology context.
[0319] The source cohort comprised 42 breast cancer patients with diverse molecular subtypes who underwent tumor profiling using the GeneSilico Deep Precision™ - Breast Cancer ST NGS Panel. Sequencing was performed at CAP-accredited, NABL-certified laboratories in accordance with ethical approvals from relevant institutional review boards. Tumor tissue samples were obtained under standardized collection protocols and processed using validated laboratory workflows, including DNA extraction from formalin-fixed paraffin-embedded (FFPE) tissue, hybrid-capture target enrichment, and high-depth sequencing via Illumina NovaSeq X Plus. Rigorous quality control metrics ensured sequencing integrity, with median unique coverage exceeding 1500X and stringent thresholds applied for DNA yield, coverage uniformity, and contamination.
[0320] Somatic variant detection and annotation were performed using a hybrid computational and expert-validated pipeline. Reads were aligned to GRCh38 using Qiagen CLC Workbench, and variants were called via the DRAGEN somatic variant pipeline. Copy number alterations and structural variants were validated through Integrative Genomics Viewer (IGV) and annotated using ClinVar, COSMIC, OncoKB, and gnomAD. All clinically relevant variants were reviewed by a multidisciplinary tumor board and categorized using AMP / ASCO / CAP standards to ensure reproducibility and translational accuracy.
[0321] Of the 42 sequenced patients, 15 were selected for inclusion in the evaluation cohortbased on completeness of clinical and molecular records. These cases underwent rigorous transformation into structured DCGPDs, which preserved the full therapeutic and diagnostic timeline. Dossiers included patient demographics, cancer subtype and stage, immunohistochemistry (IHC) results, radiologic findings, pathology reports, systemic therapy regimens, response assessments, and NGS-based molecular findings. All information was anonymized and curated into a unified case representation format to simulate real-world oncology workflows.
[0322] A panel of board-certified oncologists developed a structured set of clinical queries for each DCGPD, typically three to four per case. These queries spanned diagnostic reasoning, treatment decision-making, molecular interpretation, and longitudinal care planning. Each question was constructed to require context-sensitive, multi-step reasoning anchored in real patient data. Ground truth responses were developed via expert consensus and underwent iterative refinement to ensure accuracy, clarity, and alignment with current clinical standards.
[0323] A multi-stage evaluation framework was established to rigorously assess model performance. First, multiple LLMs — including gSage and comparator models — were tasked with responding to the 46 curated case-based queries. Next, oncologist-defined criteria were applied to evaluate each response across three dimensions: evidence-based accuracy, clinical reasoning depth, and patient-centricity. These metrics were chosen to reflect real clinical priorities rather than surface-level language quality.
[0324] To increase scalability and reduce subjectivity, LLM-based evaluators were deployed using custom prompts authored by oncologists. These automated judges assessed each response against the predefined clinical metrics. Finally, to ensure reliability and mitigate automation bias, all LLM-based assessments were independently reviewed by practicing oncologists in a blinded setting. Discrepancies were resolved through adjudication, and final scores were used to benchmark each model.
[0325] This hybrid evaluation process — combining structured automation with blinded expert review — ensures rigorous, reproducible, and clinically meaningful model evaluation. The resulting benchmark dataset and framework represent a valuable tool for assessing Al copilots in oncology, providing a template for evaluating the capacity of LLM-based systems to reason through complex patient cases and deliver evidence-aligned clinical recommendations.BRCA Gene Panel
[0326] The customized BRCA gene panel, in embodiments of the method and system disclosed herein, can be designed for a high degree of specificity and sensitivity in the detection of genetic aberrations associated with breast cancer. The custom BRCE gene panel can include a curated set of genes known to be involved in the onset and progression of breast cancer, including but not limited to BRCA1 and BRCA2, which have been robustly associated with an increased risk of breast and ovarian cancers.
[0327] In certain embodiments, the selection of genes for inclusion in the custom panel can be informed by extensive analysis of genetic and clinical data derived from global cancer databases, current literature, and ongoing clinical trials. Thereby, in such embodiments, the genes included in the custom BRCA gene panel can be chosen for their clinically actionable mutations, where therapeutic interventions are available or are currently being tested in clinical trials. In some embodiments, the custom BRCA gene panel can encompass not just the genes but also their variants that have been shown to respond to specific therapies, supporting precision medicine initiatives.
[0328] The construction of the panel can be, in some embodiments, further refined through computational algorithms that evaluate the frequency and distribution of specific genetic mutations across large patient populations. In such embodiments, these algorithms can take into account the biological significance of mutations, their prevalence in various subtypes of breast cancer, and their implications for treatment outcomes.
[0329] In synthesizing the custom gene panel, the method can leverage a high-throughput sequencing approach that ensures comprehensive coverage of each gene of interest. In some embodiments, the coverage of each gene of interest can allow for the identification of the common mutations as well as rare variants that may be of particular significance to individual patients.
[0330] In some embodiments, the depth of sequencing can be set to a level that balances the need for a comprehensive mutational analysis with the practicalities of clinical application.
[0331] Quality control measures can be applied at each stage of custom BRCA gene panel development to ensure the fidelity of the gene sequences included. In some embodiments, the quality control measure can involve, but are not limited to, cross-referencing with established genomic databases to verify the accuracy of the genetic data and utilizing in predictive tools to assess the potential impact of detected mutations on protein function.
[0332] To maintain the relevance of the custom BRCA gene panel in an evolving therapeutic landscape, in some embodiments, the system can include a mechanism for the periodic updating of the panel to incorporate new genetic discoveries and therapeutic targets. The dynamic nature of the panel can thus be used to ensure that the genomic analysis provided to patients reflects the most current understanding of breast cancer genomics and the latest advancements in treatment options.
[0333] In some embodiments, the custom BRCA gene panel can be utilized to provide a comprehensive genomic profiling of breast cancer patients, enabling the identification of individualized treatment strategies. This level of personalization in treatment planning has the potential to significantly improve patient outcomes by matching each patient with the most effective therapies based on their unique genetic makeup.
[0334] Furthermore, in some embodiments, the system and method can allow for the stratification of patients into various risk categories based on their genetic profiles. Thisstratification enables the anticipation of treatment responses and the potential for resistance to certain therapies, thereby guiding the selection of second-line treatments and the management of the disease over time.
[0335] FIG. 29 depicts a block diagram of a process for utilizing a custom BRCA gene panel for the treatment of breast cancer, in accordance with certain embodiments of the present disclosure.
[0336] FIG. 29 illustrates a block diagram of a process 2900 for employing a custom gene panel in the treatment of breast cancer, consistent with certain disclosed embodiments. This diagram presents a hierarchical workflow that begins with the aggregation of data from multiple knowledge sources 2901.
[0337] Knowledge sources 2901 may include various databases and repositories, including but not limited to, OncoKB, DrugBank, and ClinGen. In some embodiments, these knowledge sources 2901 provide information regarding drugs, therapeutic targets, and genetic data pertinent to breast cancer. These knowledge sources 2901 can serve as a bedrock for the subsequent steps in the process by offering a comprehensive suite of data points that include drug efficacy profiles, genetic mutation libraries, and clinical trial outcomes.
[0338] Building on these foundational knowledge sources 2901, the process 2900 involves the creation of therapy and clinical trial targets 2902. In some embodiments, this portion of the process 2900 allows for determining frequently altered genes and mapping genes, which are vital in understanding the genetic landscape of breast cancer. These mapped genes and identified targets can then be used to refine the information for a custom gene panel.
[0339] The process 2900, as depicted in FIG. 29, continues with the identification and validation of genomic coordinates 2903. In some embodiments, this is a technical phase where genomic data is scrutinized to confirm the precise locations of genes within the genome. Validation at this stage ensures that the subsequent analysis using the custom gene panel isaccurate and that the genomic data corresponds correctly to the targeted genetic markers.
[0340] The process 2900 continues with expert validation 2904 to discern feasibility, which may encompass considerations of cost and coverage. This phase in process 2900 can act as a quality assurance step where specialists review and verify the genomic coordinates and therapy targets, ensuring that the custom gene panel is both clinically and economically viable for use in breast cancer treatment.
[0341] At the apex of the process 2900 is the creation of a custom gene panel database 2905, which synthesizes all the validated data into an actionable form. This database 2905 can, in certain embodiments, serve as the centralized repository from which the digital clinical decision support system can draw information to generate personalized treatment reports. The database 2905 is structured to be dynamic, allowing for updates and refinements as new data becomes available, thus ensuring that the treatment recommendations derived from the custom gene panel remain current and are based on the most comprehensive and relevant data.
[0342] FIG. 30 depicts a schematic diagram of panel design process for preparing a custom BRCA gene panel, in accordance with certain embodiments of the present disclosure.
[0343] FIG. 30 provides a visual representation of the process 3000 for the development of a custom BRCA gene panel as set forth in various embodiments of the current disclosure. This schematic delineates the systematic approach taken to design a gene panel that is both precise in its scope and personalized for breast cancer treatment.
[0344] The initial stage of the process 3000 is represented by block 3001, where consolidation of knowledge sources and trial data takes place. This consolidation in block 3001 involves the gathering and synthesis of information from several key databases and repositories, including, for example, but not limited to, OncoKB, CKB boost, QIAGEN, TTD, TCGA, HSMD, cBioPortal, Sophia Genetics, FoundationOne, MSK-Impact, and combinations thereof. These knowledge sources provide critical data on genetic markers, therapy efficacy, and ongoingclinical trials that are integral to the design of the gene panel.
[0345] In some embodiments, during the initial stage in block 3001, the genes associated with breast cancer are sourced from three primary knowledge source databases, including (1) HMSD, a somatic database by QIAGEN, (2) TCGA and cBioPortal, publicly accessible databases, and (3) OncoKB, a publicly accessible database of genomic alterations in cancer. In such embodiments, the knowledge source consolidation step in block 3001 may utilize the information retrieved from OncoKB for validation of selected cancer genes.
[0346] In some embodiments, using block 3002, the genes from commercially available breast cancer and solid tumor panels can be identified. In such embodiments, the commercially available breast cancer and solid tumor panels are identified from Foundation Medicine, Sophia Genetics, MSK-IMPACT, and similar databases and resources, or combinations thereof.
[0347] In certain embodiments, the information collected from the knowledge sources regarding the genes includes alteration frequency in population database, presence in commercial panels, gene length, availability of targeted FDA-approved therapies, participation in targeted clinical trials, and combinations thereof.
[0348] Subsequently, in block 3002, a computational environment is utilized to process and analyze the consolidated information. In certain embodiments, the computational environment may be Jupyter Notebook (Python). In some embodiments, block 3002 involves scripting and data manipulation within a Python environment, enabling the detailed analysis and processing of the genetic data and trial results collected from the knowledge sources.
[0349] The information flow continues into block 3003, where the analyzed data is inputted into a BRCA database. This database, as previously described above in respect to FIG. 29, integrates the data comprehensively and facilitates the determination of relevant biomarkers for inclusion in the custom gene panel.
[0350] At block 3004, a selection process begins, wherein the genes are ranked based onparticular criteria. In certain embodiments, the gene ranking criteria includes the information collected from the knowledge sources regarding the genes, including the alteration frequency in population database, presence of genes across commercial panels, gene length, availability of targeted FDA-approved therapies, participation in targeted clinical trials, HRR genes, PGx genes, and combinations thereof.
[0351] Continuing with the selection process, at block 3005, process 3000 includes genomic coordination optimization. Within this block 3005, in certain embodiments, process 3000 includes pre-processing of the features includes imputation of missing values by median and standardization.
[0352] Further, in some embodiments, in block 3005 the process 3000 includes the utilization of a Robust Rank Aggregation algorithm to extract approximately top 50 genes of high significance towards breast cancer. In such embodiments, the standardized features can include, but are not limited to, overall length of exonic regions of a gene, log-normalize frequency for each gene in HSMD, TCGA, and cBioPortal databases based on captured exonic lengths for statistical analysis, frequency of occurrence of each gene in commercial panels, frequency of FDA-approved targeted therapies for each gene, frequency of Phase 4 Clinical Trials for each gene, frequency of Phase 3 Clinical Trials for each gene, or combinations thereof. In certain embodiments, the use of the Robust Rank Aggregation algorithm can be utilized to reveal the genes with high mutation rates and known targeted therapies.
[0353] Continuing the selection process in block 3005, process 3000 can include accounting for DNA repair mechanism and Pharmacodynamics through an analysis of homologous recombination repair (HRR) genes, such as for example FANCC or RAD 15, and pharmacogenomic (PGx) genes, such as for example, DPYD or CYP2B6.
[0354] In some embodiments, block 3006 utilizes manual team testing, diagnostic SMEs, and oncologists to verify and validate top genes for the panel. In some embodiments, the feedbackis used continuously to update the top genes for a gene panel.
[0355] Following, at block 3007, the gene panel is constructed. In some embodiments, the gene panel includes exonic regions with alterations to streamline the panel size. Further, in some embodiments, process 3000 can include categorizing the final genes into various groups based on therapy options and importance. For example, in some embodiments, Group 1 may contain genes with FDA-approved drugs, Phase 4 and Phase 3 clinical trials. Additionally, for example, in some embodiments, Group 2 may contain genes included in Group 1 along with those involved in clinical trials for other phases, Homologous Recombination Repair (HRR) genes, Pharmacogenomic (PGx) genes, and combinations thereof. Additionally, for example, in some embodiments, Group 3 may contain the genes included in Group 1 as well as highly researched genes with high frequent alterations but no targeted therapies.
[0356] Process 3000, as shown in FIG. 30, culminates in block 3007, which illustrates the development of the custom BRCA gene panel design. This panel is intentionally designed to include approximately 50 genes, ensuring it is significantly more targeted than general cancer panels. In some embodiments, the gene panel design includes at least 50 genes. In other embodiments, the gene panel design includes exactly 50 genes. Further, in other embodiments, the gene panel design includes between 40 and 50 genes. In certain embodiments, the panel encompasses 2.5 MB of genetic data, striking a balance between comprehensiveness and specificity.
[0357] In certain embodiments, the key panel considerations include choosing the right target regions, sequencing coverage versus depth, reducing experimental noise, and validating panel design.
[0358] FIG. 31 depicts a schematic of the report generation process based on the use of the custom BRCA gene panel, in accordance with certain embodiments of the present disclosure.
[0359] FIG. 31 illustrates the report generation process 3100 utilizing the custom BRCA genepanel in accordance with certain embodiments of the present disclosure. This schematic details the procedural steps taken from initial patient screening to the final report generation, using the gene panel for accurate breast cancer assessment.
[0360] The report generation process 3100 begins with block 3101, where an individual undergoing breast cancer screening provides a sample from which FASTA formatted genomic data is obtained through the cloud services of a sequencing partner. In certain embodiments, the process begins with the utilization of the gene panel as designed from process 3000, detailed above in respect to FIG. 30.
[0361] In some embodiments, in block 3101, the report generation process 3100 begins with the tumor specimen collection from an individual that is known or suspected to have a disease, such as, for example, but not limited to breast cancer. In such an embodiment, the tumor speiciment collection may include collecting Formalin-Fixed Paraffin-Embedded (FFPE) tissue or circulating tumor DNA (ctDNA).
[0362] In certain embodiments, block 3101 can include the establishment of MoU. In such embodiments, the MoU may be established with companies or Contract Research Organizations (CRO) for DNA sequencing using custom BRCA Panel, as generated by the process 3000 outlined above in respect to FIG. 30.
[0363] Further, in block 3101, the report generation process 3100 may include utilization of cloud-based services for the storage of sequencing data, such as for example in a FASTA or FASTQ format. In some embodiments, the cloud-based services may include the BaseSpace Sequence Hub of Illumina.
[0364] Subsequent to data acquisition in block 3101 of the report generation process 3100 continues with an analysis and processing of data to ensure that the report includes the desired outcomes. Accordingly, the report generation process 3100 can proceed to the quality control block 3102. At this juncture, the data undergoes a rigorous quality control process to removeany unreliable sequences and ensure the integrity of the data for precise and reliable analysis. In some embodiments, this block 3102 allows for maintaining the quality of the genomic data that will inform the report generation.
[0365] In some embodiments, the quality control block 3102 can be used to conduct read preprocessing for quality control using tools such as Trimmomatic for high quality reads. In some embodiments, in block 3102 the quality-control check includes ERBB2 (HER2) amplification.
[0366] Following the quality check performed in block 3102, the report generation process 3100 advances to block 3103, which in some embodiments encompasses several analytical procedures including read alignment, variant calling, and variant annotation. In some embodiments, block 3103 can include reference genome indexing and read alignment utilizing tools such as BWA or Bowtie2. In some embodiments, block 3103 can include variant calling for identification of single nucleotide polymorphisms and insertions / deletions using software such as for example GATK or FreeBayses. In some embodiments, block 3103 can include annotation of variants by ANNOVAR tool for final storage within a proprietary database.
[0367] In certain embodiments, the analytical procedures in block 3103 further include, but are not limited to, single nucleotide polymorphism (SNP) calling, microsatellite instability (MSI) assessment, tumor mutational burden (TMB) calculation, and transcription factor (TF) activity evaluation. Each of these processes contributes to a comprehensive understanding of the patient's genomic landscape and potential cancer-related mutations.
[0368] Continuing in report generation process 3100, as shown in FIG. 31, in block 3104, the refined genomic data is inputted into the BRCA database where it is augmented with relevant literature and interpreted in the context of the Quality Control Interpretation (QCI) guidelines. In some embodiments, this integration ensures that the genetic findings are viewed in the light of current medical knowledge and practice standards, resulting in informed therapeuticdecision-making.
[0369] In some embodiments, in block 3104, the integration of variant information can include adding variant call format (VCF) files into metadata curation layer. In some embodiments, in block 3104, the integration of variant information can include the incorporation with curated literature from knowledge sources, such as those analyzed in process 3000, as described above in respect to FIG. 30.
[0370] The final stage of the report generation process 3100 is represented in block 3105, which depicts the generation of a detailed report. In some embodiments, in block 3105, the generation of personalized therapy reports can include the utilization of QCI Interpret and VCF files for report generation.
[0371] As shown in FIG. 31, the report generated in block 3105 encapsulates a variety of treatment-related information including, but not limited to, genes, gene alterations, FDA- approved drugs, drugs under clinical trials, indications for other cancers, familial cancer markers such as BRCA1 and BRCA2, experimental indications, references, and combinations thereof.
[0372] In certain embodiments, each section of the report is tailored to the patient's specific genomic findings from the custom BRCA gene panel utilized in block 3101 as developed using process 3000, providing a personalized treatment roadmap that is both comprehensive and precise.
[0373] In some embodiments, in block 3105, the personalized therapy reports can include fields including, but not limited to, disease, gene mutation, description of genes and gene alterations, FDA-approved drugs and clinical trials, personalized therapies, indications to other cancers, or combinations thereof.
[0374] Accordingly, in such an embodiment, the report generated through this process 3100 provides a tool in precision oncology, offering actionable insights to clinicians and patients forinformed therapeutic decision-making.
[0375] The foregoing description, for purposes of explanation, used specific nomenclature to provide a thorough understanding of the described embodiments. However, it should be apparent to one skilled in the art that the specific details are not required in order to practice the described embodiments. Thus, the foregoing descriptions of specific embodiments are presented for purposes of illustration and description. They are not intended to be exhaustive or to limit the described embodiments to the precise forms disclosed. It should be apparent to one of ordinary skill in the art that many modifications and variations are possible in view of the above teachings.
[0376] While embodiments of the disclosure have been shown and described, modifications thereof can be made by one skilled in the art without departing from the spirit and teachings of the disclosure. The embodiments described and the examples provided herein are exemplary only, and are not intended to be limiting. Many variations and modifications of the disclosure disclosed herein are possible and are within the scope of the disclosure. The scope of protection is not limited by the description set out above, but is only limited by the claims which follow, that scope including all equivalents of the subject matter of the claims.
[0377] At least some of the methods (e.g., microservices) can be implemented as computer readable instructions that can be executed by one or more computational devices, such as the status identification engine of the systems disclosed in FIGS. 30-31. For example, an implementation of one or more embodiments of the methods and systems as described above may include microservices included in a digital and laboratory health care platform that can generate a patient's treatment based upon the patient's next generation sequencing results.
[0378] Further microservices may include implementation of a DNA / RNA analysis, a Bioinformatics analysis, and a reporting analysis where each respective analysis may be implemented via a series of intertwined microservices managed by an order managementserver.
[0379] In embodiments of the present disclosure, the machine learning techniques utilized may include, but are not limited to, one or more of the following: Ordinary Least Squares Regression (OLSR), Linear Regression, Logistic Regression, Stepwise Regression, Multivariate Adaptive Regression Splines (MARS), Locally Estimated Scatterplot Smoothing (LOESS), Instancebased Algorithms, k-Nearest Neighbor (KNN), Learning Vector Quantization (LVQ), SelfOrganizing Map (SOM), Locally Weighted Learning (LWL), Regularization Algorithms, Ridge Regression, Least Absolute Shrinkage and Selection Operator (LASSO), Elastic Net, Least-Angle Regression (LARS), Decision Tree Algorithms, Classification and Regression Tree (CART), Iterative Dichotomizer 3 (ID3), C4.5 and C5.0 (different versions of a powerful approach), Chi-squared Automatic Interaction Detection (CHAID), Decision Stump, M5, Conditional Decision Trees, Naive Bayes, Gaussian Naive Bayes, Causality Networks (CN), Multinomial Naive Bayes, Averaged One-Dependence Estimators (AODE), Bayesian Belief Network (BBN), Bayesian Network (BN), k-Means, k-Medians, K-cluster, Expectation Maximization (EM), Hierarchical Clustering, Association Rule Learning Algorithms, A-priori algorithm, Eclat algorithm, Artificial Neural Network Algorithms, Perceptron, Back- Propagation, Hopfield Network, Radial Basis Function Network (RBFN), Deep Learning Algorithms, Deep Boltzmann Machine (DBM), Deep Belief Networks (DBN), Convolutional Neural Network (CNN), Deep Metric Learning, Stacked Auto-Encoders, Dimensionality Reduction Algorithms, Principal Component Analysis (PCA), Principal Component Regression (PCR), Partial Least Squares Regression (PLSR), Collaborative Filtering (CF), Latent Affinity Matching (LAM), Cerebri Value Computation (CVC), Multidimensional Scaling (MDS), Projection Pursuit, Linear Discriminant Analysis (LDA), Mixture Discriminant Analysis (MDA), Quadratic Discriminant Analysis (QDA), Flexible Discriminant Analysis (FDA), Ensemble Algorithms, Boosting, Bootstrapped Aggregation(Bagging), AdaBoost, Stacked Generalization (blending), Gradient Boosting Machines (GBM), Gradient Boosted Regression Trees (GBRT), Random Forest, Computational intelligence (evolutionary algorithms, etc.), Computer Vision (CV), Natural Language Processing (NLP), Recommender Systems, Reinforcement Learning, Graphical Models, or combinations thereof.
[0380] The foregoing description, for purposes of explanation, used specific nomenclature to provide a thorough understanding of the described embodiments. However, it should be apparent to one skilled in the art that the specific details are not required in order to practice the described embodiments. Thus, the foregoing descriptions of specific embodiments are presented for purposes of illustration and description. They are not intended to be exhaustive or to limit the described embodiments to the precise forms disclosed. It should be apparent to one of ordinary skill in the art that many modifications and variations are possible in view of the above teachings.
[0381] Amounts and other numerical data may be presented herein in a range format. It is to be understood that such range format is used merely for convenience and brevity and should be interpreted flexibly to include not only the numerical values explicitly recited as the limits of the range, but also to include all the individual numerical values or sub-ranges encompassed within that range as if each numerical value and sub-range is explicitly recited. For example, a numerical range of approximately 1 to approximately 4.5 should be interpreted to include not only the explicitly recited limits of 1 to approximately 4.5, but also to include individual numerals such as 2, 3, 4, and sub-ranges such as 1 to 3, 2 to 4, etc. The same principle applies to ranges reciting only one numerical value, such as “less than approximately 4.5,” which should be interpreted to include all of the above-recited values and ranges. Further, such an interpretation should apply regardless of the breadth of the range or the characteristic being described. The symbolis the same as “approximately”.
[0382] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which the presently disclosed subject matter belongs. Although any methods, devices, and materials similar or equivalent to those described herein can be used in the practice or testing of the presently disclosed subject matter, representative methods, devices, and materials are now described.
[0383] The above discussion is meant to be illustrative of the principles and various embodiments of the present disclosure. Numerous variations and modifications will become apparent to those skilled in the art once the above disclosure is fully appreciated. It is intended that the following claims be interpreted to embrace all such variations and modifications.
[0384] Those skilled in the art will appreciate that although the previous paragraphs relate to embodiments where steps may be described as occurring in a certain order, no ordering is required unless otherwise stated. In fact, steps described in the previous paragraphs may occur in any order. Furthermore, although one step may be described in one figure and another step may be described in another figure, embodiments of the present disclosure are not limited to such combinations, as any of the steps described above may be combined in particular embodiments.
[0385] Those skilled in the art will further appreciate that although the examples described above relate to embodiments where an artificial intelligence infrastructure supports the execution of machine learning models, the artificial intelligence infrastructure may support the execution of a broader class of Artificial Intelligence algorithms, including production algorithms. In fact, the steps described above may similarly apply to such a broader class of Al algorithms.
[0386] Those skilled in the art will further appreciate that although the embodiments described above relate to embodiments where the artificial intelligence infrastructure includes one or more storage systems and one or more GPU servers, in other embodiments, other technologiesmay be used. For example, in some embodiments the GPU servers may be replaced by a collection of GPUs that are embodied in a non-server form factor. Likewise, in some embodiments, the GPU servers may be replaced by some other form of computer hardware that can execute computer program instructions, where the computer hardware that can execute computer program instructions may be embodied in a server form factor or in a non-server form factor.
[0387] Example embodiments are described largely in the context of a fully functional computer system. Those having skill in the art will recognize, nonetheless, that the present disclosure also may be embodied in a computer program product disposed upon computer readable storage media for use with any suitable data processing system. Such computer readable storage media may be any storage medium for machine-readable information, including magnetic media, optical media, or other suitable media. Examples of such media include magnetic disks in hard drives or diskettes, compact disks for optical drives, magnetic tape, and others as will occur to those of skill in the art. Persons skilled in the art will immediately recognize that any computer system having suitable programming means will be capable of executing the steps of the method as embodied in a computer program product. Persons skilled in the art will recognize also that, although some of the example embodiments described in this specification are oriented to software installed and executing on computer hardware, nevertheless, alternative embodiments implemented as firmware or as hardware are well within the scope of the present disclosure.
[0388] Embodiments can include be a system, a method, and / or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0389] The computer readable storage medium can be a tangible device that can retain andstore instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electro-magnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electro-magnetic waves, electro-magnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0390] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0391] Computer readable program instructions for carrying out operations of the presentdisclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’s computer, partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0392] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, systems, systems-of-systems, and computer program products according to some embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0393] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processingapparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein includes an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0394] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0395] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the blockdiagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
[0396] Those skilled in the art will appreciate that the steps described herein may be carried out in a variety ways and that no particular ordering is required. It will be further understood from the foregoing description that modifications and changes may be made in various embodiments of the present disclosure without departing from its true spirit. The descriptions in this specification are for purposes of illustration only and are not to be construed in a limiting sense.
[0397] Consistent with the above disclosure, the examples of systems and methods enumerated in the following clauses are specifically contemplated and are intended as a non-limiting set of examples.
[0398] Clause 1. A system for personalized breast cancer treatment decision support includes: a central processing unit (CPU); a computer-readable memory; a computer-readable storage media; a set of program instructions including: first program instructions configured to analyze patient-specific genomic data and compare it with a database of genomic profiles to provide personalized treatment recommendations, where the analysis includes constructing genomic data profiles for the patient and comparing these profiles to the database to identify matching treatment options; second program instructions configured to integrate electronic health record (EHR) data with the patient-specific genomic data, where integration includes analyzing patient health history and current medical conditions; third program instructions configured to analyze imaging data, where the imaging data includes a repository comprising known treatment outcomes; fourth program instructions configured to retrieve and utilize public data sources; fifth program instructions to generate a treatment plan based on compiled data in the set program instructions; and where the set of program instructions are stored on the computer-readable storage media for execution by the CPU via the computer-readable memory.
[0399] Clause 2. The system of any foregoing clause, where the treatment plan is based on the analysis of integrated patient data includes one or more selected from the group consisting of on-label prescriptions, clinical trial allocations, and investigational drug opportunities.
[0400] Clause 3. The system of any foregoing clause, where the first program instructions further include analyzing genomic alterations for FDA-approved therapy matching.
[0401] Clause 4. The system of any foregoing clause, where the first program instructions further include providing targeted therapy recommendations based on genomic data analysis.
[0402] Clause 5. The system of any foregoing clause, where the second program instructions further include processing and analyzing patient history data, including previous treatments and outcomes, to personalize the treatment recommendations.
[0403] Clause 6. The system of any foregoing clause, where the set of program instructions further includes sixth program instructions configured to provide genetic counseling referrals as part of the treatment plan based on genetic markers identified in the genomic data analysis.
[0404] Clause 7. The system of any foregoing clause, where the sixth program instructions further include analyzing genetic counseling needs based on the comprehensive integration of genomics data, EHR, and patient history data.
[0405] Clause 8. The system of any foregoing clause, where the set of program instructions further includes seventh program instructions to coordinate multidisciplinary team meetings and tumor board discussions for case preparation and treatment plan validation.
[0406] Clause 9. The system of any foregoing clause, where the seventh program instructions further include documenting the decisions made and recommendations provided during tumor board discussions for future reference and follow-up.
[0407] Clause 10. The system of any foregoing clause, where the set of program instructions further includes eighth program instructions to summarize the treatment plan in aconversational format that mimics the decision-making process of a human-led tumor board.
[0408] Clause 11. The system of any foregoing clause, where the set of program instructions further includes ninth program instructions to continuously update the system’s database with new clinical oncology knowledge, treatment guidelines, and emerging therapeutic agents and strategies.
[0409] Clause 12. The system of any foregoing clause, where the ninth program instructions further include instructions for the system to learn and adapt to new information about drugdrug interactions and drug response predictions using artificial intelligence.
[0410] Clause 13. The system of any foregoing clause, where the set of program instructions further includes tenth program instructions to facilitate discussion and case preparation by a multidisciplinary team using an integrated web-based platform.
[0411] Clause 14. The system of any foregoing clause, where the tenth program instructions further include providing an interface for clinicians to input, review, and discuss the treatment plan, facilitating real-time adjustments based on clinician expertise and patient feedback.
[0412] Clause 15. A method for personalized breast cancer treatment decision support implemented in a computer infrastructure having computer executable code tangibly embodied on a computer-readable storage medium including programming instructions to provide a personalized treatment plan, including the steps of: analyzing patient-specific genomic data to identify targeted therapy options; integrating electronic health record (EHR) data with genomic data to contextualize the treatment recommendations; conducting imaging data analysis to refine the treatment recommendations; retrieving public data sources; and generating a treatment plan based on steps (a)-(d).
[0413] Clause 16. The method of any foregoing clause, where the step of retrieving public data sources includes identifying potential clinical trial participations and investigational drug opportunities.
[0414] Clause 17. The method of any foregoing clause, further including coordinating multidisciplinary team meetings and tumor board discussions for case preparation and treatment plan validation.
[0415] Clause 18. The method of any foregoing clause, further including summarizing the treatment plan in a conversational format for clinician review.
[0416] Clause 19. The method of any foregoing clause, further including continuously updating the system’s database with clinical oncology knowledge and treatment guidelines.
[0417] Clause 20. The method of any foregoing clause, further including retrieving feedback from clinicians, wherein the feedback is based on the treatment plan.
[0418] Clause 21. A method for generating a personalized cancer treatment report, including collecting a sample from a patient, obtaining genomic data using a custom BRCA gene panel, performing quality control on the genomic data, resultant from performing quality control on the genomic data, removing unreliable sequences from the genomic data to produce quality- controlled genomic data, analyzing the quality-controlled genomic data using a custom BRCA gene panel to perform variant calling to obtain called variants, annotating the called variants to obtain annotated variants, integrating the annotated variants with the patient’s electronic health records (EHR), imaging data, and public clinical trials information, and generating a personalized cancer treatment report based on the integrating, where the report comprises treatment-related information, gene alterations, FDA-approved drugs, investigational drugs, and clinical trial opportunities.
[0419] Clause 22. The method of any foregoing clause where the custom BRCA gene panel includes at least 50 genes relevant to breast cancer diagnosis, prognosis, and treatment.
[0420] Clause 23. The method of any foregoing clause further including providing descriptions of terminologies used in the personalized cancer treatment report in a pop-up window interface.
[0421] Clause 24. The method of any foregoing clause further including utilizing bothformalin-fixed paraffin-embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA) for genomic data analysis.
[0422] Clause 25. The method of any foregoing clause further including identifying approved drugs and investigational drugs suitable for the patient as indicated by the personalized cancer treatment report.
[0423] Clause 26. The method of any foregoing clause further including allocating clinical trial opportunities based on the integrated data analysis.
[0424] Clause 27. The method of any foregoing clause further including providing genetic counseling referrals based on the personalized cancer treatment report.
[0425] Clause 28. The method of any foregoing clause where the personalized cancer treatment report includes drug-drug interaction information and adverse reaction information for prescribed medications.
[0426] Clause 29. The method of any foregoing clause further including, prior to generating the personalized cancer treatment report, removing unreliable data in the quality control on the genomic data.
[0427] Clause 30. The method of any foregoing clause where the integrated data analysis further includes analysis of one or more of patient’s history data, genomics data, and PET scan data.
[0428] Clause 31. A digital clinical decision support system (DCDSS) for precision oncology, including a processor; and a memory storing instructions that, when executed by the processor, cause the system to analyze genomic data using a custom BRCA gene panel, integrate analyzed genomic data with patient's electronic health records (EHR), imaging data, and public clinical trials information, and generate a personalized cancer treatment report based on the integrated data analysis.
[0429] Clause 32. The system of any foregoing clause where the custom BRCA gene panelincludes at least 50 genes relevant to breast cancer diagnosis, prognosis, and treatment.
[0430] Clause 33. The system of any foregoing clause where the system is further configured to provide descriptions of terminologies used in the personalized cancer treatment report in a pop-up window interface.
[0431] Clause 34. The system of any foregoing clause where the system is further configured to utilize both formalin-fixed paraffin-embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA) for genomic data analysis.
[0432] Clause 35. The system of any foregoing clause where the system is further configured to identify approved drugs and investigational drugs suitable for the patient as indicated by the personalized cancer treatment report.
[0433] Clause 36. The system of any foregoing clause where the system is further configured to allocate clinical trial opportunities based on the integrated data analysis.
[0434] Clause 37. The system of any foregoing clause where the system is further configured to provide genetic counseling referrals based on the personalized cancer treatment report.
[0435] Clause 38. The system of any foregoing clause where the personalized cancer treatment report includes drug-drug interaction information and adverse reaction information for prescribed medications.
[0436] Clause 39. The system of any foregoing clause where the system is further configured to, prior to generating the personalized cancer treatment report, remove unreliable data in the quality control on the genomic data.
[0437] Clause 40. The system of any foregoing clause where the integrated data analysis further includes analysis of one or more of patient’s history data, genomics data, and PET scan data.
[0438] Clause 41. A method for breast cancer genomic analysis, including consolidating genetic information from a plurality of knowledge sources to obtain consolidated geneticinformation, where the plurality of knowledge sources include databases and trial data on breast cancer treatment; utilizing a computational environment to process the consolidated genetic information to facilitate the ranking of genes, where the ranking of genes is based on criteria comprising alteration frequency, gene length, and the presence of genes across commercial panels; employing a robust rank aggregation algorithm to prioritize genes for inclusion based on their significance in breast cancer to obtain ranked genes; creating a custom BRCA gene panel from the ranked genes comprising genes most relevant to breast cancer treatment; storing the custom BRCA gene panel in a genomic database designed to facilitate ongoing updates and integrations.
[0439] Clause 42. The method of any foregoing claim, where the biological sample is selected from the group consisting of formalin-fixed paraffin-embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA).
[0440] Clause 43. The method of any foregoing claim, where the genomic database is configured to continuously update with new genetic information from the knowledge sources and new patient genomic data.
[0441] Clause 44. The method of any foregoing claim, further including using the custom BRCA gene panel to analyze a patient’s biological sample for genetic markers associated with breast cancer, where the analysis includes assessing the presence of specific genetic alterations linked to breast cancer prognosis and treatment responsiveness.
[0442] Clause 45. The method of any foregoing claim, further including testing the patient’s biological sample using the custom gene panel to identify mutations indicative of breast cancer.
[0443] Clause 46. The method of any foregoing claim, where testing for genetic markers includes utilizing the custom gene panel to identify and interpret mutations in the patient’s biological sample, including the assessment of homologous recombination repair (HRR) genes and pharmacogenomic (PGx) genes.
[0444] Clause 47. A system for breast cancer genomic analysis, including a processor; a memory, where the processor and the memory are configured to receive and consolidate genetic information from multiple knowledge sources; a computational module programmed to use a robust rank aggregation algorithm to analyze and rank genes based on their relevance to breast cancer treatment; a genomic database configured to store the ranked genes and facilitate the creation of a custom gene panel based on the ranked genes.
[0445] Clause 48. The system of any foregoing claim, where the biological sample is obtained from FFPE tissue blocks or ctDNA of the patient.
[0446] Clause 49. The system of any foregoing claim, where the custom gene panel includes approximately 50 genes, which are categorized into groups based on their clinical relevance and therapy options, facilitating targeted and precise breast cancer treatment strategies.
[0447] Clause 50. The system of any foregoing claim, where the genomic database continuously updates to include new genetic information from ongoing research and trial data, enhancing the accuracy and relevance of the custom gene panel.REFERENCES
[0448] Abdin, M. et al. Phi-4 Technical Report. Preprint at https: / / arxiv.org / abs / 2412.08905 (2024).
[0449] Ankit Pal, M. S. OpenBioLLMs: Advancing Open-Source Large Language Models for Healthcare and Life Sciences. Hugging Face repository Preprint at (2024).
[0450] Arnold, M. et al. Current and future burden of breast cancer: Global statistics for 2020 and 2040. The Breast 66, 15-23 (2022).
[0451] AWS, Guidance for Conversational Chatbots Using Retrieval Augmented Generation on AWS, https: / / aws.amazon.com / solutions / guidance / conversational-chatbots-using-retrieval- augmented-generation-on-aws / .
[0452] AWS, Retrieval Augmented Generation (RAG) in Amazon SageMaker,https: / / docs.aws.amazon.com / sagemaker / latest / dg / jumpstart-foundation-models-customize- rag.html.
[0453] Bamford, S. et al. The COSMIC (Catalogue of Somatic Mutations in Cancer) database and website. British Journal of Cancer 200491:2 91, 355-358 (2004).
[0454] Chang, T.-G., Park, S, Shaffer, A. A., Jiang, P. & Ruppin, E., Hallmarks of artificial intelligence contributions to precision oncology. Nat Cancer (2025).
[0455] Chen, X., Ji, Z. L. & Chen, Y. Z. TTD: Therapeutic Target Database. Nucleic Acids Res 30, 412-415 (2002).
[0456] Ferber, D. et al. GPT-4 for Information Retrieval and Comparison of Medical Oncology Guidelines. NEJM Al 1, (2024).
[0457] Formal, T., Piwowarski, B. & Clinchant, S. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. SIGIR 2021 - Proceedings of the 44 th International ACM SIGIR Conference on Research and Development in Information Retrieval 2288-2292 (2021).
[0458] Genomoncology, Molecular Tumor Board, https: / / www.genomoncology.com / molecular-tumor-board.
[0459] Gilbert, S., Harvey, H., Melvin, T., Vollebregt, E. & Wicks, P. Large language model Al chatbots require approval as medical devices. Nature Medicine 2023 29:1029, 2396-2398 (2023).
[0460] Haltaufderheide, J. & Ranisch, R. The ethics of ChatGPT in medicine and healthcare: a systematic review on Large Language Models (LLMs). npj Digital Medicine 2024 7:1 7, 1- 11 (2024).
[0461] Haslam, A., Kim, M. S. & Prasad, V. Updated estimates of eligibility for and response to genome-targeted oncology drugs among US cancer patients, 2006-2020. Annals of Oncology 32, 926-932 (2021).
[0462] Hou, Y. et al. Fine-tuning a local LLaMA-3 large language model for automatedprivacy-preserving physician letter generation in radiation oncology. Front Artif Intell 7, (2025).
[0463] Hsu Lin, L., et al., Comparison of solid tissue sequencing and liquid biopsy accuracy in identification of clinically relevant gene mutations and rearrangements in lung adenocarcinomas, 34 Modern Pathology 2168-2174 (Dec. 2021), https: / / www.sciencedirect.com / science / article / pii / S0893395222003775.
[0464] Huang, X., Chadha, R., Singh, H., Khetan, A., Dadarkar, M., Ulrich, K., Question answering using Retrieval Augmented Generation with foundation models in Amazon SageMaker JumpStart (May 2, 2023), https: / / aws.amazon.com / blogs / machine- learning / question-answering-using-retrieval-augmented-generation-with-foundation-models- in-amazon-sagemaker-jumpstart / .
[0465] Jin, D. et al. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. Applied Sciences 2021, Vol. 11, Page 6421 11, 6421 (2021).
[0466] Jin, H. et al. Comparative study of Claude 3.5-Sonnet and human physicians in generating discharge summaries for patients with renal insufficiency: assessment of efficiency, accuracy, and quality. Front Digit Health 6, (2024).
[0467] Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W. & Lu, X. PubMedQA: A Dataset for Biomedical Research Question Answering. EMNLP-IJCNLP 2019 - 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, Proceedings of the Conference 2567-2577 (2019).
[0468] Joshi, P., What is Retrieval-Augmented Generation?, https: / / prateekjoshi.substack.com / p / what-is-retrieval-augmented- generation?utm_medium=reader2.
[0469] Kazmi, M., Enrich LEMS with Retrieval-Augmented Generation (RAG), GoPenAI (July31, 2023), https: / / blog.gopenai.com / enrich-llms-with-retrieval-augmented-generation-rag- 17b82a96b6f0.
[0470] Labrak, Y. et al. BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains. Preprint at https: / / arxiv.org / abs / 2402.10373 (2024).
[0471] Landrum, M. J. etal. ClinVar: public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res 42, D980-D985 (2014).
[0472] Lazris, D., Schenker, Y. & Thomas, T. Exploring ALgenerated content and professional guidelines in cancer symptom management: A comparative analysis between ChatGPT and NCCN guidelines. Journal of Clinical Oncology 42, el3610-el3610 (2024).
[0473] Lee, J. et al. BioBERT : a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36, 1234-1240 (2020).
[0474] Lee, P., Bubeck, S. & Petro, J. Benefits, Limits, and Risks of GPT-4 as an Al Chatbot for Medicine. N. Engl. J. Med. 388, 1233-1239 (2023).
[0475] Lewis, P., Perez, E. Piktus, A., Petroni, F., et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Cornell Comput. Sci. (May 22, 2020), https: / / arxiv.org / abs / 2005.11401.
[0476] Li, Y. etal. RefAL a GPT-powered retrieval-augmented generative tool for biomedical literature recommendation and summarization. Journal of the American Medical Informatics Association (2024).
[0477] Luo, R. et al. BioGPT: generative pre-trained transformer for biomedical text generation and mining. Brief Bioinform 23, 1-11 (2022).
[0478] Meta Al Blog, Retrieval-Augmented Generation: Streamlining the Creation of Intelligent Natural Language Processing Models, https: / / ai.meta.com / blog / retrieval- augmented-generation-streamlining-the-creation-of-intelligent-natural -language-processing- models / .
[0479] Namazi, M. J. et al. RadOnc-GPT (gpt-4o) versus human data extraction for prostate cancer clinical research. Journal of Clinical Oncology 43, 425-425 (2025).
[0480] Omiye, J. A., Gui, H., Rezaei, S. J., Zou, J. & Daneshjou, R. Large Language Models in Medicine: The Potentials and Pitfalls. Ann Intern Med 177, 210-220 (2024).
[0481] Pal, A., Umapathi, L. K. & Sankarasubbu, M. MedMCQA: A Large-scale MultiSubject Multi-Choice Dataset for Medical domain Question Answering. Proceedings of Machine Learning Research vol. 174 248-260 Preprint at https: / / proceedings.mlr.press / vl74 / pal22a.html (2022).
[0482] Prompting Guide, Techniques: RAG, https: / / www.promptingguide.ai / techniques / rag.
[0483] Rydzewski, N. R. etal. Comparative Evaluation of LLMs in Clinical Oncology. NEJM Al 1, (2024).
[0484] Simon, J., Retrieval-Augmented Generation chatbot, part 1: LangChain, Hugging Face, FAISS, AWS, https: / / www.youtube.com / watch?v=7kDaMz3Xnkw.
[0485] Singhal, K., Azizi, S., Tu, T, Large language models encode clinical knowledge. Nature 2023 620:7972 620, 172-180 (2023).
[0486] Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat Med 31, 943-950 (2025).
[0487] Sorin, V., Klang, E., Sklair-Levy, M. et al., Large language model (ChatGPT) as a support tool for breast tumor board, npj Breast Cancer 9:1 9, 1-4 (2023).
[0488] Steen, EL, Wahlin, D., Retrieval Augmented Generation (RAG) in Azure Al Search (Nov. 20, 2023), https: / / learn.microsoft.com / en-us / azure / search / retrieval-augmented- generation-overview.
[0489] Tamborero, D., et al., Support systems to guide clinical decision-making in precision oncology: The Cancer Core Europe Molecular Tumor Board Portal. Nature Medicine. 2020.26. 1-3.
[0490] Tan, R. et al. Retrieval-augmented large language models for clinical trial screening. https: / / doi.org / 10.1200 / JCO.2024.42.16_suppl.el3611 42, el361 l-el3611 (2024).
[0491] Tempus, Tempus One, https: / / www.tempus.com / oncology / tempus-one / .
[0492] Thirunavukarasu, A. J. et al. Large language models in medicine. Nature Medicine 2023 29:8 29, 1930-1940 (2023).
[0493] Wishart, D. S. et al. DrugBank 5.0: a major update to the DrugBank database for 2018. Nucleic Acids Res 46, D1074-D1082 (2018).
[0494] Yalamanchili, A. et al. Quality of Large Language Model Responses to Radiation Oncology Patient Care Questions. JAMA Netw Open 7, e244630-e244630 (2024).
[0495] Yang, H., Yue, S. & He, Y. Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions. (2023).
[0496] Yang, X. et al. A large language model for electronic health records, npj Digital Medicine 2022 5:1 5, 1-9 (2022).
[0497] Yao, S. et al. ReAct: Synergizing Reasoning and Acting in Language Models. (2022).
[0498] Zhang, K. et al. UltraMedical: Building Specialized Generalists in Biomedicine. Preprint at https: / / arxiv.org / abs / 2406.03949 (2024).
[0499] Zhao, H. et al. Explainability for Large Language Models: A Survey. ACM Trans Intell Syst Technol 15, 38 (2024).
[0500] Zhou, S., Wang, N., Wang, L., Liu, H. & Zhang, R. CancerBERT: a cancer domainspecific language model for extracting breast cancer phenotypes from electronic health records. Journal of the American Medical Informatics Association 29, 1208-1216 (2022).
[0501] Zhu, N., Zhang, N., Shao, Q., Cheng, K. & Wu, H. OpenATs GPT-4o in surgical oncology: Revolutionary advances in generative artificial intelligence. Eur J Cancer 206, 114132 (2024).
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A system for personalized breast cancer treatment decision support comprising:(a) a central processing unit (CPU);(b) a computer-readable memory;(c) a computer-readable storage media;(d) a set of program instructions comprising:(i) first program instructions configured to analyze patient-specific genomic data and compare it with a database of genomic profiles to provide personalized treatment recommendations, wherein the analysis includes constructing genomic data profiles for the patient and comparing these profiles to the database to identify matching treatment options;(ii) second program instructions configured to integrate electronic health record (EHR) data with the patient-specific genomic data, wherein integration comprises analyzing patient health history and current medical conditions;(iii) third program instructions configured to analyze imaging data, wherein the imaging data comprises a repository comprising known treatment outcomes;(iv) fourth program instructions configured to retrieve and utilize public data sources;(v) fifth program instructions to generate a treatment plan based on compiled data in the set program instructions; and(e) wherein the set of program instructions are stored on the computer-readable storage media for execution by the CPU via the computer-readable memory.
2. The system of Claim 1, wherein the treatment plan based on the analysis of integrated patient data comprises one or more selected from the group consisting of on-label prescriptions, clinical trial allocations, and investigational drug opportunities3. The system of Claim 1, wherein the first program instructions further comprise analyzing genomic alterations for FDA-approved therapy matching.
4. The system of Claim 1, wherein the first program instructions further comprise providing targeted therapy recommendations based on genomic data analysis.
5. The system of Claim 1, wherein the second program instructions further comprise processing and analyzing patient history data, including previous treatments and outcomes, to personalize the treatment recommendations.
6. The system of Claim 1, wherein the set of program instructions further comprises sixth program instructions configured to provide genetic counseling referrals as part of the treatment plan based on genetic markers identified in the genomic data analysis.
7. The system of Claim 6, wherein the sixth program instructions further comprise analyzing genetic counseling needs based on the comprehensive integration of genomics data, EHR, and patient history data.
8. The system of Claim 1, wherein the set of program instructions further comprises seventh program instructions to coordinate multidisciplinary team meetings and tumor board discussions for case preparation and treatment plan validation;9. The system of Claim 8, wherein the seventh program instructions further comprise documenting the decisions made and recommendations provided during tumor board discussions for future reference and follow-up.
10. The system of Claim 1 , wherein the set of program instructions further comprises eighth program instructions to summarize the treatment plan in a conversational format that mimics the decision-making process of a human-led tumor board;11. The system of Claim 1, wherein the set of program instructions further comprises ninth program instructions to continuously update the system’s database with new clinical oncology knowledge, treatment guidelines, and emerging therapeutic agents and strategies.
12. The system of Claim 11, wherein the ninth program instructions further comprise instructions for the system to learn and adapt to new information about drug-drug interactions and drug response predictions using artificial intelligence.
13. The system of Claim 1, wherein the set of program instructions further comprises tenth program instructions to facilitate discussion and case preparation by a multidisciplinary team using an integrated web-based platform.
14. The system of Claim 13, wherein the tenth program instructions further compriseproviding an interface for clinicians to input, review, and discuss the treatment plan, facilitating real-time adjustments based on clinician expertise and patient feedback.
15. A method for personalized breast cancer treatment decision support implemented in a computer infrastructure having computer executable code tangibly embodied on a computer- readable storage medium including programming instructions to provide a personalized treatment plan, comprising the steps of:(a) analyzing patient-specific genomic data to identify targeted therapy options;(b) integrating electronic health record (EHR) data with genomic data to contextualize the treatment recommendations;(c) conducting imaging data analysis to refine the treatment recommendations;(d) retrieving public data sources; and(e) generating a treatment plan based on steps (a)-(d).
16. The method of Claim 15, wherein the step of retrieving public data sources comprises identifying potential clinical trial participations and investigational drug opportunities.
16. The method of Claim 15, wherein the step of generating a treatment plan comprises incorporating genetic counseling referrals.
17. The method of Claim 15, further comprising coordinating multidisciplinary team meetings and tumor board discussions for case preparation and treatment plan validation.
18. The method of Claim 15, further comprising summarizing the treatment plan in a conversational format for clinician review.
19. The method of Claim 15, further comprising continuously updating the system’s database with clinical oncology knowledge and treatment guidelines.
20. The method of Claim 15, further comprising retrieving feedback from clinicians, wherein the feedback is based on the treatment plan.
21. A method for generating a personalized cancer treatment report, comprising:(a) collecting a sample from a patient;(b) obtaining genomic data using a custom BRCA gene panel;(c) performing quality control on the genomic data;(d) resultant from performing quality control on the genomic data, removing unreliable sequences from the genomic data to produce quality-controlled genomic data;(e) analyzing the quality-controlled genomic data using a custom BRCA gene panel to perform variant calling to obtain called variants;(f) annotating the called variants to obtain annotated variants;(g) integrating the annotated variants with the patient’s electronic health records (EHR), imaging data, and public clinical trials information; and(h) generating a personalized cancer treatment report based on the integrating, wherein the report comprises treatment-related information, gene alterations, FDA-approved drugs, investigational drugs, and clinical trial opportunities.
22. The method of Claim 21, wherein the custom BRCA gene panel includes at least 50 genes relevant to breast cancer diagnosis, prognosis, and treatment.
23. The method of Claim 21 further comprising providing descriptions of terminologies used in the personalized cancer treatment report in a pop-up window interface.
24. The method of Claim 21 further comprising utilizing both formalin-fixed paraffin- embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA) for genomic data analysis.
25. The method of Claim 21 further comprising identifying approved drugs and investigational drugs suitable for the patient as indicated by the personalized cancer treatment report.
26. The method of Claim 21 further comprising allocating clinical trial opportunities based on the integrated data analysis.
27. The method of Claim 21 further comprising providing genetic counseling referrals based on the personalized cancer treatment report.
28. The method of Claim 21, wherein the personalized cancer treatment report comprises drug-drug interaction information and adverse reaction information for prescribed medications.
29. The method of Claim 21 further comprising, prior to generating the personalized cancer treatment report, removing unreliable data in the quality control on the genomic data.
30. The method of Claim 21, wherein the integrated data analysis further includes analysisof one or more of patient’s history data, genomics data, PET scan data.
31. A digital clinical decision support system (DCDSS) for precision oncology, comprising:(a) a processor; and(b) a memory storing instructions that, when executed by the processor, cause the system to(i) analyze genomic data using a custom BRCA gene panel,(ii) integrate analyzed genomic data with patient's electronic health records (EHR), imaging data, and public clinical trials information, and(iii) generate a personalized cancer treatment report based on the integrated data analysis.
32. The system of Claim 31, wherein the custom BRCA gene panel includes at least 50 genes relevant to breast cancer diagnosis, prognosis, and treatment.
33. The system of Claim 31, wherein the system is further configured to provide descriptions of terminologies used in the personalized cancer treatment report in a pop-up window interface.
34. The system of Claim 31, wherein the system is further configured to utilize both formalin-fixed paraffin-embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA) for genomic data analysis.
35. The system of Claim 31, wherein the system is further configured to identify approveddrugs and investigational drugs suitable for the patient as indicated by the personalized cancer treatment report.
36. The system of Claim 31, wherein the system is further configured to allocate clinical trial opportunities based on the integrated data analysis.
37. The system of Claim 31, wherein the system is further configured to provide genetic counseling referrals based on the personalized cancer treatment report.
38. The system of Claim 31, wherein the personalized cancer treatment report comprises drug-drug interaction information and adverse reaction information for prescribed medications.
39. The system of Claim 31, wherein the system is further configured to, prior to generating the personalized cancer treatment report, remove unreliable data in the quality control on the genomic data.
40. The system of Claim 31, wherein the integrated data analysis further includes analysis of one or more of patient’s history data, genomics data, and PET scan data.
41. A method for breast cancer genomic analysis, comprising:(a) consolidating genetic information from a plurality of knowledge sources to obtain consolidated genetic information, wherein the plurality of knowledge sources comprise databases and trial data on breast cancer treatment;(b) utilizing a computational environment to process the consolidated genetic information to facilitate the ranking of genes, wherein the ranking of genes isbased on criteria comprising alteration frequency, gene length, and the presence of genes across commercial panels;(c) employing a robust rank aggregation algorithm to prioritize genes for inclusion based on their significance in breast cancer to obtain ranked genes;(d) creating a custom BRCA gene panel from the ranked genes comprising genes most relevant to breast cancer treatment;(e) storing the custom BRCA gene panel in a genomic database designed to facilitate ongoing updates and integrations.
42. The method of Claim 41, wherein the biological sample is selected from the group consisting of formalin-fixed paraffin-embedded (FFPE) tissue blocks and circulating tumor DNA (ctDNA).
43. The method of Claim 41, wherein the genomic database is configured to continuously update with new genetic information from the knowledge sources and new patient genomic data.
44. The method of Claim 41, further comprising using the custom BRCA gene panel to analyze a patient’s biological sample for genetic markers associated with breast cancer, wherein the analysis includes assessing the presence of specific genetic alterations linked to breast cancer prognosis and treatment responsiveness.
45. The method of Claim 44, further comprising testing the patient’s biological sample using the custom gene panel to identify mutations indicative of breast cancer.
46. The method of Claim 45, wherein testing for genetic markers includes utilizing the custom gene panel to identify and interpret mutations in the patient’s biological sample, including the assessment of homologous recombination repair (HRR) genes and pharmacogenomic (PGx) genes.
47. A system for breast cancer genomic analysis, comprising:(a) a processor;(b) a memory, wherein the processor and the memory are configured to receive and consolidate genetic information from multiple knowledge sources;(c) a computational module programmed to use a robust rank aggregation algorithm to analyze and rank genes based on their relevance to breast cancer treatment;(d) a genomic database configured to store the ranked genes and facilitate the creation of a custom gene panel based on the ranked genes.
48. The system of Claim 47, wherein the biological sample is obtained from FFPE tissue blocks or ctDNA of the patient.
49. The system of Claim 47, wherein the custom gene panel includes approximately 50 genes, which are categorized into groups based on their clinical relevance and therapy options, facilitating targeted and precise breast cancer treatment strategies.
50. The system of Claim 47, wherein the genomic database continuously updates to include new genetic information from ongoing research and trial data, enhancing the accuracy and relevance of the custom gene panel.
Citation Information
Patent Citations
telegenetics
US20170098053A1
Data based cancer research and treatment systems and methods
US20210090694A1
System and method for predicting diseases in its early phase using artificial intelligence
US20230248998A1