Omics automation using large language models

An AI system using LLMs automates lipidomics data analysis, addressing manual processing challenges and enhancing reproducibility and accuracy in lipidomics data interpretation.

WO2025175288A1PCT designated stage Publication Date: 2025-08-21PURDUE RES FOUND +4
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/016254
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2025-02-17
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Current lipidomics data analysis is largely manual, prone to human error, and lacks standardized processing and statistical analysis workflows, making it difficult to interpret large-scale lipidomics datasets and maintain reproducibility.

Method used

An AI system utilizing multi-modal large language models (LLMs) to automate data acquisition and analysis, including automated worklist generation, raw data parsing, and robust bioinformatic analysis, with a user interface for interactive data interpretation.

Benefits of technology

Facilitates efficient, accurate, and reproducible lipidomics data processing, enabling detailed lipid structural identification and pathway analysis, reducing human error and enhancing throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025016254_21082025_PF_FP_ABST
    Figure US2025016254_21082025_PF_FP_ABST
Patent Text Reader

Abstract

.An artificial intelligence (Al) system includes a processor executing software residing on a non- transient memory, the processor is configured to establish an Al manager configured to communicate with an external entity, establish at least one Al agent each coupled to the Al manager and each configured to communicate with one or more multi-modal large language models (LLMs), whereby communication between the external entity and the Al agent is analyzed and generated by the one or more LLMs thus establishing conversational communication between the external entity and the Al manager, analyze queries from the external entity, determine if a tool from a predetermined toolkit exist to be used to address the external entity query.
Need to check novelty before this filing date? Find Prior Art

Description

PRF-70566-02 OMICS AUTOMATION USING LARGE LANGUAGE MODELS CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present non-provisional patent application is related to and claims the priority benefit of U.S. Provisional Patent Application Serial 63 / 554,874, filed February 16, 2024, the contents of which are hereby incorporated by reference in its entirety into the present disclosure. STATEMENT REGARDING GOVERNMENT FUNDING

[0002] This invention was made with government support under TR004146 awarded by the National Institutes of Health. The government has certain rights in the invention. TECHNICAL FIELD

[0003] The present disclosure generally relates to an artificial intelligence (AI) based tool, and in particular to an AI system that can communicate with a client, whether a human or a computer- based system, via large language models (LLMs). BACKGROUND

[0004] This section introduces aspects that may help facilitate a better understanding of the disclosure. Accordingly, these statements are to be read in this light and are not to be understood as admissions about what is or is not prior art.

[0005] Lipids are universal biomolecules that are integral to an array of cellular processes, such as cell signaling, energy conservation, metabolic regulation, and the maintenance of cellular structure. In the past decade, there has been a significant increase in lipidomic research due to advancements in mass spectrometry (MS) technologies and recognition in the community thatPRF-70566-02 changes in lipid metabolism plays a major role in disease pathologies. For example, chronic diseases, such as Alzheimer's disease (AD), cancer, type 2 diabetes, cardiovascular disease, among others, have shown altered lipid profiles related to disease onset and progression. Therefore, detailed lipidome investigations can identify how specific metabolic pathways influences cellular functions and how specific lipids affect health and disease pathogenesis.

[0006] Despite the importance of detailed lipidome profiling, the precise characterization, identification, and quantitation of lipids presents significant challenges. MS has emerged as a popular tool in lipidomic analysis, owing to its exceptional sensitivity, specificity, and versatility. There exists a set of hierarchical structural information for detailed lipid identification. Lipids can be separated based on class and sum composition that includes the sum of carbon atoms and degree of unsaturation at the lowest hierarchy. The lipid classification system proposed by the LIPID MAPS®consortium classifies lipid structures into eight main categories: fatty acyls (FA), glycerolipids (GL), glycerophospholipids (GP), sphingolipids (SP), sterol lipids (ST), prenol lipids (PR), saccharolipids (SL), and polyketides (PK). Within each lipid category, there exists further structural hierarchical levels including an array of lipid classes and subclasses. Such structural assignments can be readily obtained from conventional MS and tandem-MS (MS / MS) experiments. Independent of ionization mode, lipid precursor ions can be identified at a sum compositional level via accurate mass measurements (i.e., observed mass-to-charge (m / z) ratios). Additional structural information can be obtained by utilizing collision-induced dissociation (CID). For example, individual GP classes like glycerophosphoethanolamines (PEs), glycerophosphoglycerols (PGs), glycerophosphoserines (PSs), glycerophosphoinositols (PIs), glycerophosphocholines (PCs) and sphingomyelin (SM) can be detected and identified based on class-specific fragmentation that relate to GP headgroup composition. Additionally, in negative ion mode, MS / MS of deprotonated acidic GP anions results in the cleavage of ester bonds at the sn-1 and sn-2 positions of anionic GP. In turn, fatty acyl chains are liberated from the anionic GP precursor ion, yielding abundant carboxylate anions that permit the assignment of fatty acyl sum composition and, in some cases, GP subclass assignment.

[0007] Recently, multiple reaction monitoring (MRM)-profiling has demonstrated success for large-scale lipid profiling. Briefly, MRM-profiling is a shotgun MS / MS technique for thePRF-70566-02 exploratory analysis of lipids. MRM-profiling facilitates the concurrent analysis of up to thousands of individual precursor / product ion pair transitions in a single sample injection, allowing for the efficient generation of large amounts of lipidomic data that should be analyzed using statistically robust methods to identify up- and down-regulated lipids. MRM transitions can be predicted and subsequently generated by exploiting lipid class-specific fragmentation patterns in conjunction with lipid database information. In general, the precursor ion is the expected ionized m / z of the lipid molecule at its class or species level, while the product ion is the expected m / z for a headgroup or fatty acyl neutral loss from the lipid precursor ion. The combination of the precursor and product ion pair constitutes the MRM transition. Although MRM-profiling is limited to exploratory analysis, the identified lipid targets and are in good agreement with LC-MS / MS lipidomics profiling results.

[0008] While conventional tandem-MS experiments including MRM-profiling are effective, they have inherent limitations specifically related to isomeric lipids resolution varying in carbon- carbon double bond position and geometry. We note that while basic lipid identification using lipid class and sum composition identifiers can provide substantial biological insights, deep lipid structural identification is required to fully understand lipid roles, behavior, and identify new targets and biomarkers in metabolic and disease state perturbations. For example, lipid identification at the carbon-carbon double bond (C=C) level has proven useful in differentiating between cancerous and healthy cells. In addition, it was demonstrated that only quantitative analysis of lipid C=C location was sufficient to discriminate gefitinib-resistant cells from a population of gefitinib-sensitive cells.

[0009] To aid in isomeric resolution, MS-based methods are often coupled with chromatography, novel-ion activation techniques, and chemical conjugation strategies. Ozone electrospray ionization (OzESI) and ozone induced dissociation (OzID) have proven highly successful for the detailed identification of unsaturated lipid structures. Briefly, OzESI-MS permits the elucidation of C=Cs within unsaturated lipid structures using ozone-induced fragmentation with the source of a conventional ESI mass spectrometer. In-source ozone can be generated either via (1) a corona discharge at the ESI capillary in an oxygen atmosphere or via (2) an external ozone generator. Upon exposure to ozone, unsaturated lipid ions will undergoPRF-70566-02 facile cleavage at each C=C position. Thus, OzESI-MS results in chemically induced fragment ions that reveal C=C position(s). While effective and easy to implement on commercial platforms, OzESI-MS relies on ozone exposure prior to mass selection. In turn, lipidomics datasets can quickly become difficult to interpret. Therefore, interfacing with liquid chromatography (LC) can reduce complexity in these datasets, but manual interpretation can still be overwhelming and laborious, especially when developing high throughput workflows. In turn, large-scale lipidomics dataset processing, such as those generated by MRM-profiling, are inherently prone to human error and bias.

[0010] Currently, comprehensive solutions and standards for MRM-lipidomics data analysis have yet to be established. Various public tools exist that cater to specific parts of the lipidomics data collection and analysis workflow, but overall, the process remains largely manual. Moreover, there is no standardized data processing and statistical analysis workflow for large- scale lipidomics data analysis which encompasses parsing raw data files, peak annotation, lipid identification, robust statistical and pathway analysis. Consequently, variations in data processing pipelines and choice of statistical analysis leads to vastly different interpretations and presentations of lipidomics data. In addition, one of the largest challenges faced by modern lipidomics is to keep informatic pipelines sustainable, adaptable, and reproducible, given the rapidly changing landscape of MS-based lipidomics technologies. Therefore, it is essential that future lipidomics analysis pipelines are reliable and robust, and incorporate updated lipidomics standardization, such as suggested nomenclature and data presentation guidelines.

[0011] Therefore, there is an unmet need for a novel system and method that can automate data acquisition and analysis via a computer-based tool that can work with clients such as humans and other computers using large langue models. SUMMARY

[0012] An artificial intelligence (AI) system is disclosed. The AI system include a processor executing software residing a non-transient memory. The processor is configured to establish an AI manager configured to communicate with an external entity, establish at least one AI agentPRF-70566-02 each coupled to the AI manager and each configured to communicate with one or more multi- modal large language models (LLMs), whereby communication between the external entity and the AI agent is analyzed and generated by the one or more LLMs thus establishing conversational communication between the external entity and the AI manager, analyze queries from the external entity, determine if a tool from a predetermined toolkit exist to be used to address the external entity query. If a tool is identified to exist, the AI agents generates input for the identified tool and executes the identified tool with the generated input. If no tool is identified to exist, the AI agent generates code and new-tool inputs for a new tool and executes the new tool with the generated code and the new-tool inputs and adds the new tool to the toolkit.

[0013] In the above AI system, the external entity is a human.

[0014] In the above AI system, the external entity is a computing device.

[0015] In the above AI system, the if the during execution of the identified tool, the identified tool generates an error, the AI agent analyzes the error, corrects the input, and re-executes the identified tool with the corrected input.

[0016] An artificial intelligence (AI)-based method is also disclosed for providing a toolkit. The method includes establishing an AI manager configured to communicate with an external entity. The method further includes establishing at least one AI agent each coupled to the AI manager and each configured to communicate with one or more multi-modal large language models (LLMs), whereby communication between the external entity and the AI agent is analyzed and generated by the one or more LLMs thus establishing conversational communication between the external entity and the AI manager. Additionally, the method includes analyzing queries from the external entity, and determining if a tool from a predetermined toolkit exist to be used to address the external entity query. If a tool is identified to exist, the method further includes the AI agents generating input for the identified tool and executing the identified tool with the generated input. If no tool is identified to exist, the method further includes the AI agent generating code and new-tool inputs for a new tool and executing the new tool with the generated code and the new- tool inputs and adding the new tool to the toolkit.

[0017] In the above method, the external entity is a human.

[0018] In the above method, the external entity is a computing device.PRF-70566-02

[0019] In the above method, if the during execution of the identified tool, the identified tool generates an error, the AI agent analyzes the error, corrects the input, and re-executes the identified tool with the corrected input. BRIEF DESCRIPTION OF FIGURES

[0020] FIG.1a is a simplified schematic showing an automated data acquisition system according to the present disclosure.

[0021] FIG.1b is a simplified schematic showing raw data processing to statistically correct differential lipid analysis, according to the present disclosure.

[0022] FIG.1c is a simple schematic of a lipid informatic analysis system, according to the present disclosure.

[0023] FIG.1d is a simple schematic of an artificial intelligence (AI) system according to the present disclosure which includes a computer interface for a human or another computer system (not shown), the computer interface interacts with an AI agent which uses one or more multi- modal large language models (LLMs) along with a tool kit in order to generate the workflow.

[0024] FIGs.2a and 2b are schematics of Language User Interface (LUI) with chatbot and interactive artificial intelligent (AI) agents for lipidomics data analysis.

[0025] FIGs.3a, 3b, 3c, 3d, 3e, 3f, 3g, and 3h provide comparison of brain region resolved lipid droplet (LD) composition for Alzheimer’s disease mice (5×FAD).

[0026] FIGs.3i and 3j provide an ion count distribution of all samples in an experimental design which illustrates where the ion count follows a negative binomial distribution for each sample.

[0027] FIG.3k provides plots which display Principal component analyses (PCAs) for 5×FAD vs. wild type (WT) lipids specifically within the cerebellum, cortex, diencephalon, and hippocampus.

[0028] FIG.4a is a schematic which shows how a user selects n-7 and n-9 double bond positions to evaluate double bond positions with the CLAW OzESI method, using a Jupyter Notebook.PRF-70566-02

[0029] FIG.4b, is a schematic in which possible m / z values are automatically generated based on the user selected double bond positions as provided in FIG.4a and then matched and annotated with the experimental data.

[0030] FIG.4c provides a comparison of ozone off / on chromatogram for bleached and degummed (RBD) canola oil with triacylglycerol (TG) retention time identified.

[0031] FIGs.4d and 4e reflect the Extracted ion chromatograms (EICs) of the above multiple reaction monitoring (MRM) transitions for TG 54:4 with ozone off and on, respectively.

[0032] FIG.4f is a plot which presents lipid ratio between the n-9 and n-7 isomers of a variety of TG molecular species in canola oil containing FA 18:1.

[0033] FIG.4g provides results of CLAW max intensity ratios and manually calculated area ratios for the OzESI crude canola oil data.

[0034] FIG.4h provides results of CLAW max intensity ratios and manually calculated area ratios for the OzESI degummed canola oil data.

[0035] FIG.4i provides results of CLAW max intensity ratios and manually calculated area ratios for the OzESI RBD canola oil data.

[0036] FIG.5 is a block diagram of the processes of the present disclosure including an AI agent.

[0037] FIG.6 is a similar flow as shown in FIG.5 but now different AI agents are chosen to accomplish different tasks.

[0038] FIG.7 is an example of a computer system that can interface with the above-discussed system. DETAILED DESCRIPTION

[0039] For the purposes of promoting an understanding of the principles of the present disclosure, reference will now be made to the embodiments illustrated in the drawings, and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of this disclosure is thereby intended.PRF-70566-02

[0040] In the present disclosure, the term “about” can allow for a degree of variability in a value or range, for example, within 10%, within 5%, or within 1% of a stated value or of a stated limit of a range.

[0041] In the present disclosure, the term “substantially” can allow for a degree of variability in a value or range, for example, within 90%, within 95%, or within 99% of a stated value or of a stated limit of a range.

[0042] A novel system and method are described herein that can automate data acquisition and analysis via a computer-based tool that can work with clients such as humans and other computers using one or more multi-modal large langue models (LLMs). Towards this end, an end-to-end integrated and automated multiple reaction monitoring (MRM) lipidomics platform referred to herein as Comprehensive Lipidomic Automation Workflow (CLAW) are disclosed herein. The CLAW platform encompasses automated worklist generation for data acquisition shown in FIG.1a, which is a simplified schematic showing an automated data acquisition system according to the present disclosure; raw data parsing and annotation, shown in FIG.1b which is a simplified schematic showing raw data processing to statistically correct differential lipid analysis, according to the present disclosure; and identification of related genes in lipid pathways shown in FIG.1c, which is a simple schematic of a lipid informatic analysis system, according to the present disclosure. Additionally, FIG.1d shows a simple schematic of an artificial intelligence (AI) system according to the present disclosure which includes a computer interface for a human or another computer system (not shown), the computer interface interacts with an AI agent which uses one or more LLMs along with a tool kit in order to generate the workflow. In particular, FIG.1a, according to one embodiment, shows how an Agilent 6495C QQQ instrument labeled samples are formatted by the worklist generator and pasted into the MassHunter worklist. After the MRM experiment, proprietary binary (.d) files are exported. In FIG.1b, the proprietary binary files are converted to mzML using MSConvert. A script utilizing the pymzml package reads the mzML files, which are then parsed and annotated with a custom lipid database, and then stored in a pandas dataframe. In FIG.1c a Language User Interface (LUI) and Graphical User Interface (GUI) are shown that provide the user with the ability to compare data from different samples. These interfaces help the user perform robust statisticalPRF-70566-02 analysis to visualize and identify significantly expressed lipids. Finally, tools for identifying lipid-based pathways are provided for target gene identification for biological samples. According to one embodiment, the present workflow bridges current gaps in mass spectrometry (MS)-based lipidomics to develop automated and modular solutions ranging from data collection to processing that are easily extendable to different MS-based methodologies and instruments. The platform of the present disclosure automates worklist generation, lipid annotation with structural details (i.e., C=C localization) and statistically robust bioinformatic analysis. In addition, the present disclosure introduces a Language User Interface (LUI) and custom AI agents to assist users in proper data collection, annotation, and interpretation based on LLMs. To demonstrate the versatility of the CLAW platform of the present disclosure, we report results by using traditional MRM-profiling and online LC-OzESI-MRM profiling to differentiate lipid profiles from different sample types, including lipid droplet extracts from brains of mice and canola oil samples.

[0043] A specific nomenclature for LIPID MAPS®is adopted. Briefly, GP classes include phosphatidylcholine (PC), phosphatidylethanolamine (PE), phosphatidylglycerol (PG), phosphatidylinositol (PI), and phosphatidylserine (PS). Fatty acyl substituents are described with the total number of carbon atoms and double bonds before and after the colon, respectively. If known, identified double bond position(s) are indicated with parentheses following the shorthand notation for FA sum composition. For example, 18:1(9) represents an 18-carbon chain with 1 double bond between carbon-9 and carbon-10 as numbered from the carboxylate moiety. In some cases, isomeric FAs are referred to using the “n-x” nomenclature where the unsaturation site occurs “x” carbons away from the terminal, i.e. methyl end of the aliphatic chain. To describe triacylglycerol (TG) structure, we first describe TG sum composition, reflecting the sum of carbon atoms and double bonds within the three FA chain substituents. For example, TG 54:2 indicates a TG containing 3 acyl chains whose sum composition adds up to 54 total carbons and 2 carbon-carbon double bonds. To further describe TG structures obtained using traditional- MRM profiling, the following nomenclature is used to express TG sum compositions carrying specific FA chain substituents. For example, TG 54:2_FA 18:1 indicates a TG molecular species with sum composition of 54:2 that carries at least one 18:1 acyl chain. In the case of OzESI-PRF-70566-02 MRM profiling experiments, confirmed double bond location in a given FA substituent is indicated as TG 54:2_FA 18:1n-9.

[0044] MRM experiments were conducted on an Agilent 6495C triple quadrupole mass spectrometer equipped with a jet stream technology ion source (AJS) that has been modified to perform OzESI. Samples were introduced via flow injection (i.e., no chromatography) using an Agilent 1290 II Infinity LC system. For traditional MRM-profiling experiments, dried lipid extracts were dissolved in 200 uL of 1:1 methanol:chloroform with 10 mM ammonium formate. Prior to analysis, reconstituted lipid extracts were diluted 2-fold in 70:30 methanol:acetonitrile with 10 mM ammonium formate. Injection solvent without lipids was used as the ‘blank’ sample, while injection solvent containing EquiSPLASH™ LIPIDOMIX® (Avaniti® Polar Lipids, Alabaster, AL, USA) at a concentration of 0.02 µg / mL was used as a quality control sample to monitor instrument status throughout the run. Briefly, 8 µL of diluted lipid extract was delivered to the AJS source of the mass spectrometer using an Agilent 67167B autosampler. MRM methods were established for 11 lipid classes, covering approximately 1500 individual species tabulated in Table 1, provided below.PRF-70566-02 Table 1 – Number of MRM transitions analyzed for each lipid class with their abbreviation and total transition count Lipid Class Abbreviation Transition Count

[0045] An online LC-OzESI-MRM method was established to enable unsaturated lipid identification at the C=C level. To facilitate OzESI, high concentration ozone generated via an external generator was delivered to the nebulizer on the Agilent 6495C mass spectrometer. Several canola oil samples were examined for triacylglycerol (TG) content using the developed OzESI-MRM pipeline. Crude, degummed and refined, bleached and degummed (RBD) canola oil samples were diluted 10,000-fold in 1:1 methanol: isopropyl alcohol (IPA). TGs were first separated chromatographically. Briefly, 4 µL of the diluted oil samples was injected onto anPRF-70566-02 EclipsePlus C18 RRHD (1.8 µm, 2.1 x 100 mm) column held at 50°C. To separate TG species based on equivalent carbon number (ECN), we used a gradient consisting of A = IPA with 10 mM ammonium formate and B = acetonitrile (ACN) with 10 mM ammonium formate with a flow rate of 0.6 mL / min. The gradient progresses linearly from 20% A to 60% A over the course of 24 min. The mobile phase composition is held at 60% A for 6 min, before returning to initial conditions (20% A) for another 5 min before terminating the run.

[0046] The mass spectrometer was operated in MRM mode. Custom MRMs and OzESI-MRMs were established for unsaturated TG species. Briefly, traditional MRMs monitored the neutral loss of designated acyl chains from ammoniated TG adduct ions. For example, to screen for TG species containing an 18:1 fatty acyl chain, traditional MRMs monitor the neutral loss of 299.2 Da from the [TG + NH4]+precursor ion. OzESI-MRMs were established using previously tabulated neutral losses that are indicative of C=C position formed via the reaction of the unsaturated lipid ion with ozone. In general, in-source ozonolysis results in the efficient formation of product ions termed OzESI-aldehydes, reflecting decomposition of the unsaturated lipid at the C=C location and resulting in formation of the diagnostic aldehyde product ion. Ultimately, the investigation of “ozone on” and “ozone off” mass spectra were analyzed by the CLAW automated informatic pipeline for detailed lipid molecular identification, including acyl chain composition and the localization of unsaturation sites. In the “ozone on” experiments, we note that ozonolysis is occurring prior to mass selection. To ensure accurate lipid identification, retention times were matched to correlate precursor and product ion pairs between the “ozone on” and “ozone off” experiments.

[0047] Wild-type mice (C57BL / 6J) and Alzheimer’s disease mice (5×FAD) were obtained from the Jackson Laboratory and maintained in a pathogen-free facility. All experiments, including breeding, involving mice were performed in accordance with Purdue University’s Institutional Animal Care and Use Committee (IACUC) guidelines. Specifically, 18-24-month-old 5×FAD (n=5 mice) and wild-type mice (n = 3 mice) were perfused, and the brain tissues were isolated and snap-frozen in dry ice chilled iso-pentane and stored at -80 °C. Prior to processing, the brain tissue was divided into four regions (hippocampus, cortex, cerebellum, and diencephalon). Next, the lipid droplets (LDs) from different brain regions were isolated using a lipid droplet isolationPRF-70566-02 kit from Cell Biolabs (San Diego, CA) according to the manufacturer’s instructions and then stored at −80 °C. All LD samples were processed together for lipid extraction using the Bligh and Dyer protocol. Briefly, the frozen LDs were thawed in 4 °C and 450 µL of cold methanol and 250 µL of chloroform was added. The samples were mixed and vortexed for 10 seconds, resulting in a one-phase solution, which was incubated at 4 °C for 15 mins. Next, 250 µL ultrapure water and 250 µL chloroform was added, followed by centrifugation at 16,000 x g for 10 mins, resulting in three phases in the tubes. The bottom organic phase that contains the lipids was transferred to new tubes and evaporated in a speed-vac, leaving behind the dried lipid mixture.

[0048] CLAW is designed to automate data acquisition as shown in FIG.1a, data processing as shown in FIG.1b and statistical and bioinformatic analysis as shown in FIG.1c. For more efficient worklist creation, CLAW incorporates an automated worklist generator for MRM lipidomics experiments. By automating the worklist construction process, CLAW eliminates the slow and error-prone process of manual worklist creation. Next, the user can initiate data acquisition for the MRM experiments. Although the program has been tested on the Agilent MassHunter software, the generated worklist can be used with other MS acquisition software that support copy and paste functions. CLAW’s automatic worklist generator reduces worklist preparation time and prevent user errors.

[0049] Currently, a significant portion of data analysis for different lipidomics methods are processed manually, which is tedious, time-consuming, and error prone. This becomes especially relevant for large data sets that are often obtained from MRM-profiling experiments in a high- throughput manner. To reduce analysis time and help provide more consistent, accurate results, CLAW automates the majority of the MRM-based lipidomics data processing pipeline. Following data acquisition, raw data files are exported and prepared for the CLAW data processing step: parsing and annotation. To prepare data files for parsing in CLAW, raw data files converted to mzML format using MSConvert. Converting the proprietary binary (.d) file format into a more universally accessible mzML format facilitates several downstream applications.PRF-70566-02

[0050] The entire parsing and annotation script is implemented in a Jupyter notebook. This process begins by parsing the data using a python script with the pymzml (version 2.5.2) package to extract MRM transitions and MS ion count based on the user selected MRM transition database. During the parsing step, lipids are identified based on given precursor and product ion pair transitions as defined by a user-constructed MRM database. The selected database, uploaded as a CSV or excel file, must contain known lipids, classes, and corresponding transitions. An example database with the number of MRM transitions for each lipid class with respective abbreviation is shown in Table 1. Potential LIPID MAPS® ID’s (LM_IDs) are also assigned if naming convention is compatible with LIPID MAPS® or PubChem through their respective RESTful APIs.

[0051] The parsing script operates by matching the precursor and product ion data from the MRM database with the information present in the mzML files. However, it is possible that the m / z for each MRM transition value may not be an exact match when rounding to 1 decimal place. To make CLAW capture slight m / z differences across low to high-resolution instrumentation, a resolution tolerance parameter is selected that takes the absolute difference between two values allowing non-identical transitions to be matched. Since mass spectrometry has a high selectivity, the default tolerance is set to m / z 0.1 but may be adjusted per user specifications.

[0052] Each matching transition is annotated with the corresponding lipid from the MRM database and stored in a pandas (version 1.5.2) dataframe. Pandas dataframes offer flexible and intuitive data structures with comprehensive built-in functions, making data manipulation and analysis streamlined. Their compatibility with various data formats, readable syntax, and integration with other key Python libraries further enhance their ease of use.

[0053] CLAW’s parsing and annotation is designed for both direct infusion and LC raw data files. To identify lipid structures with C=C specificity, an LC-OzESI-MRM approach was developed. Briefly, in the LC-OzESI-MRM method, data processing follows the same parsing and annotation procedure as the standard MRM method, with addition of several key steps.

[0054] To establish lipid retention times, “ozone off” experiments were first employed. Using a Jupyter notebook interface, a detailed table is compiled within a pandas data frame to capture allPRF-70566-02 identified lipids along with their corresponding retention times. Importantly, established retention times serve as the ground truths for “ozone on” experiments. Next, OzESI-MRM transitions for identified lipids are automatically generated and tabulated, providing a targeted data-driven subset of OzESI-MRM predictions for each LC-OzESI-MS / MS experiment. Additionally, the annotated dataframe can be subsequently refined to incorporate one or more labels derived from each sample name, adhering to the naming convention established by the user. These labels are utilized to categorize the lipids into groups, facilitating the distinction of desired variables within the pandas dataframe.

[0055] After running the LC-OzESI-MRM method with the “ozone on”, experimental data is first parsed and the subsequently matched with the previously tabulated data-driven subset of OzESI-MRM predictions corresponding to each lipid C=C position within the subset. Successful matches made within a user-specified tolerance are then annotated in a pandas dataframe with the associated lipid identity, including C=C position.

[0056] After categorizing the lipids in the enhanced dataframe, the ensuing step employs a Gaussian mixture model (GGM) to define peak profiles. This model is essential for clustering the data and defining the distinct peak shapes for each group. The clustering is centered around a retention time window, which is user-defined and based on the ground truth retention time derived from the ozone off experiment. The cluster that most closely aligns with this ground truth retention time is retained, while other clusters corresponding to the same group are excluded from the annotated pandas dataframe. For each lipid species defined by a unique MRM transition and retention time, the maximum MRM intensity value from the associated cluster is then selected and stored.

[0057] An option to manually visualize chromatograms from specific lipids is available to validate peak clustering. If a peak was not clustered correctly the user has an option to add back in the previous clusters and then manually adjust the retention time for that specific lipid. Every maximum intensity value for each lipid is systematically annotated, and following this, a relative quantitation analysis of unsaturated C=C isomeric ratios is performed on the double-bond locations specified by the user.PRF-70566-02

[0058] One option to interactively select lipid MRM data for comparison in CLAW is a graphical user interface (GUI). The GUI is built using Jupyter notebook (IPython) widgets and is freely available in CLAW’s GitHub repository. The GUI filters are applied to the pandas dataframe prior to performing statistical analysis. The filters are stored as JavaScript Object Notation (JSON) files. In general, JSON files are an organized, hierarchical file format that permit filter information saved as key-value pairs to be easily accessed. One of the useful features of CLAW’s GUI is the ability to create multiple filters sequentially for datasets with several parameters simultaneously. Specifically, Table 2, provided below, shows an example where a user can simultaneously filter data based on genotype, sex, brain region, or any combination, to create two JSON files allowing for multiple statistical comparisons to identify differential lipids. The GUI displays parameters from the labels.csv file, previously created at the worklist generation step. The user can continuously generate comparisons by selecting the 'Add More JSON Pairs' button and choosing new parameters. Selections are finalized by clicking the ‘Finish’ button. Finally, the generated JSON files are used in the Jupyter notebook for data analysis. Table 2 – An example of CLAW’s interactive GUI Genotype Wild Type (WT)

[0059] Table 2 provides an example of CLAW’s interactive GUI. CLAW’s GUI allows quick comparison of multiple groups. The groups available are based on the labels.csv file. ForPRF-70566-02 instance, the left column showcases Genotype: 5×FAD, Brain Region: cerebellum, Sex: male, while the right presents Genotype: WT, Brain Region: cerebellum, Sex: male. The analysis will compare these groups, highlighting lipids that are upregulated in the 5×FAD mouse from the left column. Users can incorporate additional filters for the comparison by selecting the “Add more JSON pairs” button. Once all selections are made, the “Finish” button finalizes the process.

[0060] The second option in CLAW for interacting with data is a language user interface (LUI), which interfaces LLMs with custom action agents to interact with various components of the workflow. The LUI uses a chatbot style interface which can provide text-based assistance to a user. Specifically, the LUI interacts with tailor-made AI agents to perform functions. Each AI agent uses the LLM to dynamically process user text based on OpenAI's GPT-4 model, interprets relevant instructions, and select from a predefined set of tools to fulfill tasks. The AI agents repeatedly cycle through a thought-action-observation process based on user input and selected tool output until it is determined that the final answer is obtained based on instructional prompt. Examples are provided in FIGs.2a and 2b which are schematics of Language User Interface (LUI) with chatbot and interactive artificial intelligent (AI) agents for lipidomics data analysis. In FIG.2a, a human uses CLAW’s LUI to interact via a chatbot with customized AI agents for data manipulation. These AI agents use OpenAI’s GPT-4 LLM to process user requests on dataset filtration and then select from a list of tools developed for lipidomics analyses to take action to fulfill those requests. The custom prompt, used by AI agents for direction, instructs the agent to use tools in a guided manner, seeking user clarification and confirmation on requests to ensure robustness. FIG.2b is a simplified example of a human interacting with the AI agent via a chatbot interface. The human seeks information on available brain regions and then requests a comparison on the cortex vs cerebellum data. The AI agent analyzes the appropriate tools to fulfill the requests (bold text) and responds in a conversational manner (underlined text).

[0061] An example of the LUI interaction is provided in script 1 below: Script 1 - Example of LUI interaction > Entering new AgentExecutor chain... Thought: I should use the json_spec_list_keys tool to see what keys exist in the JSON Action: json_spec_list_keysPRF-70566-02 Action Input: data Observation: ['Cage', 'Sex', 'Genotype', 'Brain Region'] Thought: I should ask the human what key and values they would like to filter on Action: Human Action Input: what key and values would you like to filter on? What key and values would you like to filter on? What are the brain option and Gene options? Observation: What are the brain option and Gene options? Thought: I should use the json_spec_get_value tool to see what values exist for the Brain Region and Genotype keys Action: json_spec_get_value Action Input: data["Brain Region"] Observation: ['hippocampus', 'cortex', 'cerebellum', 'diencephalon'] Thought: I should use the json_spec_get_value tool to see what values exist for the Brain Region and Genotype keys Action: json_spec_get_value Action Input: data["Genotype"] Observation: ['WT', '5×FAD'] Thought: I should use the json_filter_json_from_lists tool to filter the 350N based on the keys and values provided by the huma Action: json_filter_json_from_lists Action Input: ["Brain Region", "Genotype"]; ["hippocampus", "WT"] Observation: 'Genotype': ['WT'], 'Brain Region': ['hippocampus"]} Thought: I should use the Human tool to ask the human to confirm that this is the JSON they wanted to filter Action: Human Action Input: Is this the JSON you wanted to filter? Is this the JSON you wanted to filter? No I want 5×FAD and cortex Observation: No I want 5×FAD and cortex Thought: I should use the json_filter_json_from_lists tool to filter the JSON based on the keys and values provided by the huma Action: json_filter_json_from_listsPRF-70566-02 Action Input: ["Brain Region", "Genotype"]; ["cortex","5×FAD"] Observation: Genotype': ['5×FAD' ], 'Brain Region': ['cortex']} Thought: I should use the Human tool to ask the human to confirm that this is the JSON they wanted to filter Action: Human Action Input: Is this the JSON you wanted to filter? Is this the JSON you wanted to filter? yes Observation: yes Thought: I should use the write file tool to save the JSON Action: write_file Action Input: { "path": ". / demo_data / JSON / filter_keys_1.json", "text": {'Genotype': ['5×FAD'], 'Brain Region': ['cortex']} } observation: File written successfully to . / demo_data / JSON / filter_keys_1.json.

[0062] Script 1 provides an example of LUI interaction. The user asks, “What are the brain option and gene options?” The LUI then infers that the user would like to know the potential labels for both genotype and brain region. The LUI then creates a filter for WT and hippocampus based on user input and will then ask for clarification. The user specifies the following parameters: 5×FAD and cortex. The LUI then creates the correct JSON filter but will seek user confirmation before proceeding. After confirmation the LUI will then save the JSON.

[0063] Script 2 is a continuation of the LUI interaction. Script 2 – Continuation of the LUI interaction Observation: File written successfully to . / demo_data / JSON / filter_keys_1.json. Thought: I should ask the human what key and values they would like to filter on Action: Human Action Input: What key and values would you like to filter on? What key and values would you like to filter on?PRF-70566-02 wild type from the cortex Observation: wild type from the cortex Thought: I should use the json_filter_json_from_lists tool to filter the JSON based on the keys and values provided by the huma n Action: json_filter_json_from_lists Action Input: ["Brain Region", "Genotype"]; ["cortex","WT"] Observation: (Genotype': ['WT'], 'Brain Region': ['cortex']} Thought: I should use the Human tool to ask the human to confirm that this is the JSON they wanted to filter Action: Human Action Input: Is this the JSON you wanted to filter? Is this the JSON you wanted to filter? yes Observation: yes Thought: I should use the write file tool to save the JSON Action: write_file Action Input: { } "path": ". / demo_data / JSON / filter_keys_2.json", "text": {"Genotype': ['WT'], 'Brain Region': ['cortex']} Observation: File written successfully to . / demo_data / JSON / filter_keys_2.json. Thought: I now know the final answer Final Answer: Meow! I have successfully saved two filtered JSONS based on the original JSON. You can find them in the demo_dat a / JSON folder. > Finished chain. "Meow! I have successfully saved two filtered JSONS based on the original JSON. You can find them in the demo_data / JSON folder."

[0064] Script 2 is an example of continuation of LUI interaction. After creation of the first JSON file the LUI will then ask for the parameters of the second JSON file. The user responds withPRF-70566-02 “wild type from the cortex.” Although the label for genotype is “WT” the LUI can infer the meaning of wild type and can recognize the cortex is a brain region. The LUI will once again confirm before creating the second JSON file.

[0065] The LangChain Python package was used and modified to create a custom AI agents that parses and saves JSON files specific to dataset labeling and filtration. Specifically, LangChain’s toolkit was extended, and a custom prompt was created to enable the agent to interact with the relevant JSON data in a highly structured manner. The toolkit included LangChain’s built in JSON interaction tools, LangChain’s tools for human input, a custom tool for filtering JSON files based on user input, and a modified version of LangChain’s file writing tool to save JSON files that serve as CLAW statistical filters shown in Scripts 3, 4, and 5 provided below. The agent’s instructional prompt was customized to ensure robust operation by clearly defining goals, strictly outlining tool usage, and ensuring clarification is requested, if needed as provided in script 6 provided below. The main goal of the agent is to write and save JSON filters for the current MRM dataset, with a subtask of answering user questions about the dataset. Each tool that the AI agent accesses contains its own section in the prompt that outlines conditions on when and how the tool can be used. Additionally, wording is added to instruct the agent to ask the user for clarification, whenever needed. This is done to make the AI agents behave in a robust and consistent manner. Specific instructions in the prompt allows the agent to generate two JSON files consecutively for comparison, facilitating potential enhancements for interactive data analysis. The LUI generates two JSON files to compare lipid MRM data from samples for 5×FAD vs wild-type genotypes for the cortex brain region of mice. Even with deliberate user spelling errors and ambiguous queries, the LUI accurately identifies user’s intent and generates the desired JSON files for analysis.

[0066] Script 3 – LUI tool setup - Custom LangChain agent tool for filtration of JSON files based on key and value information provided by the user. These filtered JSON files will be used as filters for data for statistical analysis. The tools private run function (called upon agent use) is assigned to return the output of the custom filtration function (filter_json_nested) provided with a list of keys and values. The custom filtration function recursively searches through the JSON file and only pulls entries which match the key and values provided. If provided keys and / orPRF-70566-02 values are not present, the function will return an error string which will be caught by the agent and the agent will attempt to find keys and values which are valid and similar to user input or ask the user for clarification. Once a suitable filtration of the original JSON has been obtained, the filter will be provided to the agent for future use. spec: JsonSpec def _run( self, tool_input: str, run_manager: Optional[CallbackManagerForToolRun] = None, ) -> str: return self.filter_json_nested( self.spec.dict_, tool_input.split(“;”)[0], tool_input.split(“;”)[1] ) async def _arun( self, tool_input: str, run_manager: Optional[AsyncCallbackManagerForToolRun] = None, ) -> str: return self._run(tool_input)

[0067] Script 4 - LUI tool setup continued - Modified LangChain agent file writing tool for writing a JSON filter to disk. LangChain base file writing tool was modified to be specifically compatible with the JSON filters which we use. This tool attempts to write a file to disk at the provided path with the provided text. In CLAW usage, the text will be JSON filter information that was the successful output of the custom JSON filtration tool. class CustomWriteFileTool(BaseFileToolMixin, BaseTool): name: str = "write_file" # args_schema: Type[BaseModel] = WriteFileInputPRF-70566-02 description: str = "Write file to disk" def _run( self, info: str, append: bool = False, run_manager: Optional[CallbackManagerForToolRun] = None, ) -> str: file_path = str(ast.literal_eval(info)["path"]) text = str(ast.literal_eval(info)["text"]) try: write_path = self.get_relative_path(file_path) except FileValidationError: return INVALID_PATH_TEMPLATE.format(arg_name="file_path", value=file_path) try: write_path.parent.mkdir(exist_ok=True, parents=False) mode = "a" if append else "w" with write_path.open(mode, encoding="utf-8") as f: f.write(text) return f"File written successfully to {file_path}." except Exception as e: return "Error: " + str(e) async def _arun( self, file_path: str, text: str, append: bool = False, run_manager: Optional[AsyncCallbackManagerForToolRun] = None, ) -> str:PRF-70566-02 # TODO: Add aiofiles method raise NotImplementedError

[0068] Script 5 - LUI custom toolkit setup - Custom toolkit provided to the AI agent which is used in CLAW’s language user interface (LUI). This toolkit contains LangChain’s built-in tools for JSON file interaction (JSONListKeysTool and JSONGetValueTool), LangChain’s builtin tool for human interaction (HumanInputRun), the custom tool for JSON file filtration, and the modified version of LangChain’s write file tool. Class CustomToolkit(BaseToolkit): spec: JsonSpec def get_tools(self) -> List[BaseTool]: return [ JsonListKeysTool(spec=self.spec), JsonGetValueTool(spec=self.spec), JsonFilterFromKeysTool(spec=self.spec), HumanInputRun(), CustomWriteFileTool() ]

[0069] Script 6 - Example LUI prompt designed to instruct the LLM model - Custom prompt designed to provide LUI instruction to the AI agent. The agent is provided with this prompt at object definition and references the prompt as it cycles through its thought-action-observation process. The beginning of the prompt defines the overall goal of the agent which is to filter up to two JSON based on user specifications with a subtasks of providing the user with information which they request. The middle of prompt outlines correct usage of the tools and how to handle any errors that may result. The end of the prompt tells the agent how to start its interaction with the user and tells the agent to ask for guidance if needed. This strict structure and detailed wording of the prompt promotes robust agent interaction. prefix = """You are an agent designed to interact with a JSON and a human. Your overall goal is to interact with the user to create and save TWO (2) new filtered JSONs based on the original JSON.PRF-70566-02 During interaction with the human prior to your final answer, they may ask for information about the keys or values of the JSON which you should provide to them. You should talk like a cat when responding, make sure to use lots of meows. ONLY provide your final answer after you have successfully used the `write_file` tool twice without an error. You MUST fulfill this prior to providing your final answer. You have access to the following tools which help you learn more about the JSON you are interacting with and provide the human with information on the JSON you are interacting with. Only use the information returned by the below tools to construct your final answer. You MUST use the `json_filter_json_from_lists` tool in the following manner: 1. Your input to the `json_filter_json_from_lists` tool must be exactly two Python formatted lists in the form '["key1","key2"];["value1","value2"]''. 16 the first list is a list of dictionary keys and the second list is a list of dictionary values. The input to the tool must have its list be separated by a semicolon and nothing else. You can infer any keys from previous observations of the `json_spec_list_keys` tool and you can infer any values from previous observations of the `json_spec_get_value` tool. 2. If the observation from the `json_filter_json_from_lists` tool returns an error, you must use the `json_spec_list_keys` and `json_spec_get_value` tools to find valid keys and values for this tool and then infer based on the humans request. 3. After using the `json_filter_json_from_lists` tool without an error, you MUST use the `Human` tool with an input of a string of the filtered JSON and ask them to confirm that it is the JSON they wanted to filter. You MUST use the `write_file` tool in the following manner: 1. Before using the `write_file` tool, if you do not have a previous observation from the `json_filter_json_from_lists` tool, use the `Human` tool with an input asking they would like to filter on. 2. Your input for the `write_file` tool should be a text representation in JSON format containing the keys "path" and "text".PRF-70566-02 The path should be ". / demo_data / JSON / filter_keys_<n>.json" where <n> is a string of the number 1 or 2.1 if it is the first JSON that is being save, 2 if it is the second JSON being saved The text should be a JSON formatted string based on a previous observation from the `json_filter_json_from_lists` tool. 3. Check to see how many times you have used the `write_file` tool without error, if it is 2 or more times, you should return your final answer. 4. After using the write_file tool without error, use the `Human` tool with an input asking what key and values they would like to filter on. You MUST use the `json_spec_list_keys` and `json_spec_get_value` tools in the following manner: 1. Your input to tools named `json_spec_list_keys` and `json_spec_get_value` should be in the form of `data["key"]` where `data` is the JSON blob you are interacting with, and the syntax used is Python. You should only use keys that you know for a fact exist. You must validate that a key exists by seeing it previously when calling `json_spec_list_keys`. If you have not seen a key in one of those responses, you cannot use it. 17 You should only add one key at a time to the path. You cannot add multiple keys at once. If you encounter a "KeyError", go back to the previous key, look at the available keys, and try again. 2. If the human asked you to provide them with keys or values of the JSON and you used either the `json_spec_list_keys` or `json_spec_get_value` tool, you MUST use the `Human` tool with an input of what you observed. You MUST use the `Human` tool in the following manner: 1. Do not give the `Human` tool any input on keys or value that you did not directly observe as a result of using the `json_spec_list_keys` or `json_spec_get_value` tools 1. If you are providing the human with keys or values from the JSON, you must first use the `json_spec_list_keys` or `json_spec_get_value` tools to confirm those keys and values exist in the JSON.PRF-70566-02 You can use the `json_spec_list_keys` or `json_spec_get_value` tools any number of times prior to using the `Human` tool. 2. If you just used another tool besides 'Human' in the previous action. The input to the `Human` tool should include your observation of that action. 3. If you are asking the human for clarification or confirmation, you input the `Human` tool should be explicit on what you need clarification or confirmation on. If the question does not seem to be related to the JSON, use the `Human` tool and ask for guidance Always begin your interaction with the `json_spec_list_keys` tool with input "data" to see what keys exist in the JSON. Then for each key that you observe use the `json_spec_get_value` tool for that key to see what values exists in the JSON. Then start your interaction with the human. Note that sometimes the value at a given path is large. In this case, you will get an error "Value is a large dictionary, should explore its keys directly". In this case, you should ALWAYS follow up by using the `json_spec_list_keys` tool to see what keys exist at that path. Do not simply refer the user to the JSON or a section of the JSON, as this is not a valid answer. Keep digging until you find the answer and explicitly return it. """ suffix = """Begin!" 18 Remember to format your answers talking as a cat would. {chat_history} Question: {input} {agent_scratchpad}""" memory = ConversationBufferMemory(memory_key="chat_history")

[0070] Generally, scientific data analysis tools require some level of programming experience to properly process relevant data. The LUI addresses this problem by assisting the user with CLAWPRF-70566-02 data analysis. The LUI functions like a copilot or co-scientist, performing user-defined queries and clarifying specifics regarding data analysis in an interactive manner. This enables users to issue commands via chat-based input, guiding the LUI agent to filter data for comparative analysis. To ensure robust performance, the LUI will request clarification when provided with invalid or unclear instructions (see Script 1 and Script 2). We achieved this robust response by integrating native LLM with a specialized agent to utilize a select set of tools. To our knowledge, this represents the first use of an LLM to aid in lipidomics profiling, establishing a novel, robust and comprehensive approach for data analysis.

[0071] Ion count measurements from MRM-profiling experiments follow a negative binomial distribution in a similar manner to data from RNA-seq experiments. FIGs.3a-3h provide comparison of brain region resolved LD composition for 5×FAD. Specifically, FIG.3a provides an experimental design which illustrates the processing of a typical negative binomial distribution of lipid MRM counts, employing edgeR's GLM for statistical analysis of significant lipids. In particular, FIG.3a illustrates a representative distribution, while FIGs.3i and 3j provide an ion count distribution of all samples where the ion count follows a negative binomial distribution for each sample. Statistical models used are tailored to this distribution. In particular FIGs.3i and 3j showcase the true ion count distributions across all samples. Considering these distributions, tools such as edgeR’s generalized linear model (GLM) developed for RNA-seq count data can be used for analysis of lipid MRM ion counts. By accounting for the inherent variability and structure in ion counts data, the EdgeR GLM approach ensures statistically robust calculation of changes in lipid abundance while incorporating experiment-specific variations in the blank (injection medium) for several injections. The EdgeR package fits a GLM to the log– linear relationship for the mean variance as follows: log μ ^^^ = ^^ ^^ + log^^where μ^^is the expected ion counta sample ^ and ^^represents the ion intensity for sample ^. ^^represents the regression coefficients associated with each lipid ^ that captures the experimental conditions on lipid expression. The coefficient of variance (CV) is then calculated for lipid ion count for a given sample ^^^using the following formula:PRF-70566-02 CV^(^^^) = 1 / μ^^ + Φ^

[0072] Ion count dispersion wasbiological replicates are present, it is calculated using the common dispersion method. When no biological replicates are present the dispersion term is set to 0.1 which is recommended for genetically identical model organisms. Fold change is calculated between the groups of interest and p-values are obtained through the likelihood ratio test. The p-values are corrected using the Benjamini-Hochberg procedure to calculate false discovery rate (FDR). We considered FDR value of less than 0.1 (10%) as significant to identify differential lipids between comparisons.

[0073] Predicting the relationship between lipid molecules and the genes that encode proteins related to their biosynthetic pathways can help identify new targets for novel biological insights. Bioinformatics Methodology for Pathway Analysis (BioPAN) is a web-based tool that allows is users to upload lipidomic data for pathway analysis. CLAW exports lipid MRM results as CSV files in a BioPAN-compatible format for each comparison selected. Instead of manually formatting lipidomics data, the CSV files generated by CLAW can be directly uploaded to the BioPAN website for downstream analysis.

[0074] CLAW was used to evaluate the differential expression of lipids within lipid droplets (LDs) isolated from specific brain regions (hippocampus, cortex, cerebellum, and diencephalon) of 18–24 month-old Alzheimer’s disease model (5×FAD) and age-matched wild-type (WT) male mice. LDs are dynamic cellular organelles that not only serve as lipid reservoirs but play significant roles in cellular signaling, detoxification and inflammation. Accordingly, their accumulation is correlated to pathophysiology of lipid imbalance linked disorders like Alzheimer’s disease and even aging. Despite being pointed out as “adipose saccules” in postmortem brain of AD patients by Alois Alzheimer’s in 1907, functional significance of LDs has been largely overlooked. Since dyshomeostatic environment in disease conditions manifest as aberrations in lipid droplet composition, we are elucidating the lipid signatures to find druggable targets for restoring brain homeostasis in AD and aging.

[0075] CLAW’s comparative analysis GUI (see Table 2) was used to rapidly select multiple comparisons from various brain regions from the 5×FAD and wild-type mice. This streamlinedPRF-70566-02 the analysis of MRM experimental data which was obtained using an Agilent 6495C mass spectrometer. We used tailored MRMs that were categorized into 10 main classes, including glycerophospholipids, glycerolipids, sphingolipids, fatty acyls (FA), and sterol lipids. Such broad coverage and depth of profiling enabled us to first identify a detailed LD composition. In brief, LDs were found to be rich in both TGs and CEs as traditionally characterized, but also contained a variety of lipid species spanning the acyl carnitine (CAR), sphingomyelin (SM), phosphatidylethanolamine (PE), and ceramide (CER) subclasses. For example, CAR 14:2, CAR 18:4, SM(d16:1 / 24:1), PE 38:0, and Cer(d18:1 / 16:0) were found across all LDs. Importantly, our bioinformatic analysis also facilitated the development of lipid-profile signature for LDs from aged-WT and 5×FAD brains, revealing that FA 24:6 was found to be downregulated in all 5×FAD samples. The combination of all brain regions resulted in 10 differentially expressed lipids between 5×FAD and WT samples. Specifically, FA 24:6, FA 24:5, and PG 32:5 down regulated, while 22:3 campestral ester, 22:2 campestral ester, 22:0 cholesteryl ester, and 22:1 cholesteryl ester were significantly upregulated lipids in 5×FAD mice.

[0076] Next, we constructed region-specific lipidomic profiles for LDs isolated from the brains of aged-WT and 5×FAD mice. The presence of cholesteryl ester-rich LDs in hippocampus and cortex, which are the hotspots for amyloid plaques, underscores the impact of environmental changes on composition in different brain regions. Specifically, PG 32:5, PS 32:5, and FA 24:6 were down regulated in the hippocampus LDs of 5×FAD mice. In the diencephalon LDs from 5×FAD mice, FA 24:6 was also downregulated, while PS 32:5, 22:3 campestral ester, 22:1 campestral ester, 22:2 campestral ester, 22:0 cholesteryl ester, and 22:1 cholesteryl ester was found to be upregulated in LDs. In the cerebellum, there was only a single differentially expressed lipid, fatty acid 24:6 when comparing LDs from 5×FAD to WT mice. FIGs.3d-3f illustrate the log-fold change (logFC) of lipids in 5×FAD mice relative to wild type across various lipid classes, highlighting the differential expression that varies depending on the brain region. In the hippocampus, acyl carnitines, cholesterol esters, phosphatidylserine (PS), and triacylglycerols are upregulated in the 5×FAD model. While, phosphatidylcholine (PC), phosphatidylinositol (PI), phosphatidylglycerol (PG), and sphingomyelin (SM) classesPRF-70566-02 are downregulated in this region for the 5×FAD mice. In the diencephalon, there is an upregulation of acyl carnitine, cholesterol esters, phosphatidylcholine, and sphingomyelin lipid classes. However, certain fatty acids exhibit downregulation in the diencephalon of the 5×FAD model. For the cerebellum, the lipid classes of acyl carnitines, ceramides, and phosphatidylserine are downregulated, while phosphatidylinositol, sphingomyelin, and phosphatidylcholine classes exhibit upregulation. In the cortex, specific cholesterol esters and phosphatidylglycerol classes are downregulated, with phosphatidylinositol showing an upregulation. Notably, acyl carnitines, phosphatidylcholine, and sphingomyelin display bimodal distributions, with lipids from these classes being both up and downregulated. Intriguingly, fatty acids displayed the most expansive distribution across all brain regions, with instances of both upregulation and downregulation. The differential lipidomic profiles across various brain regions underscore the importance of understanding the regional differences in lipid metabolism resulting from neurodegenerative disorders. The consistent downregulation of fatty acid 24:6, regardless of the brain region, suggests it may play a pivotal role in the pathophysiology of 5×FAD mice, warranting further investigation into its potential role. Overall, these results demonstrate CLAW’s ability to manage complex lipidomic datasets, allowing efficient and meaningful comparisons that could shed light on the role of lipid droplets in the pathogenesis of neurodegenerative diseases. Significant lipids and their corresponding LogFC are shown in Tables 3-7 provided below. All MRM transitions utilized in this study, along with their corresponding lipid classes and abbreviations, are detailed in Table 1. A visual representation of the MRM transition count distribution is shown in FIG.3b in which distribution of MRM transitions selected for screening lipids are shown. A total of 1497 transitions (used to ID lipid species) were organized into 10 MRM-based mass spectrometry methods for lipid classes. Principal component analysis (PCA) was used to visualize variance among the samples. In FIG.3c, a PCA was constructed by summing the intensities across all brain regions for each sample, demonstrating the variation in the LD lipidome of four brain regions. The analysis revealed marginally greater variance among the wildtype samples compared to the 5×FAD. This observationPRF-70566-02 aligns with the heatmap generated using the same summation approach in FIG.3d, in which a heat map is provided that displays z-scores of normalized lipid intensities across samples. Referring to FIG.3k, plots display PCAs for 5×FAD vs. WT lipids specifically within the cerebellum, cortex, diencephalon, and hippocampus. High variance is shown in both 5×FAD and wildtype mice regardless of brain region. However, these PCAs do not present any discernible patterns, making it challenging to draw definitive conclusions. This lack of clear patterns aligns with the observation of a limited number of differentially expressed lipids. Additionally, the analysis was based on only a few wild- type samples, which limits the generalizability of the findings. CLAW automatically exports a variety of plots, including PCAs, to assist users in interpreting their data. Referring to FIG.3e, ridge plot are provided displaying distribution of logFC values for all lipid species within each class in Hippocampus. Similarly, referring to FIG.3f, 3g, and 3h ridge plots of Diencephalon (FIG.3f), Cerebellum (FIG.3g), and Cortex region (FIG.3h) of 5×FAD vs. WT mice brain are provided. Data represented for LDs obtained from hippocampus, diencephalon, cerebellum, and cortex region isolated from 5×FAD (n=5 mice) and WT mice (n=3) brain. Table 3 - Significant lipids in across brain regions cortex, cerebellum, hippocampus, and diencephalon combined of 5×FAD vs WT male mice. Lipid Class logFC PValue FDR 8 5 5 4 1 1 1PRF-70566-02 22:2 Cholesteryl ester, 20:1 Stigmasteryl ester, 20:2 Sitosteryl ester 1 8 2Lipid logFC PValue FDR ClassTable 5 – Significant lipids in brain region cortex of 5×FAD vs WT male mice Lipid logFC PValue FDR ClassTable 6 – Significant lipids in brain region diencephalon of 5×FAD vs WT male mice Lipid Class logFC PValue FDRPRF-70566-02 22:3 Campesteryl ester CE 2.41 1.94E-05 0.01Table 7 – Significant lipids in brain region Hippocampus of 5×FAD vs WT male mice Lipid logFC PValue FDR Class

[0077] CLAW was used to investigate the lipid profiles of canola oil across three stages in the refinement process. The LC-OzESI-MRM experiments were performed on an Agilent 6495C QQQ modified for OzESI. Triacylglycerols (TGs) containing an unsaturated fatty acyl chain FAPRF-70566-02 18:1 (i.e., 18 carbons, 1 C=C) were targeted due to their known high abundance in canola oil. Due to their known biological prevalence, we sought to determine the relative ratio of TGs containing FA 18:1 with an unsaturation site either ∆9 or ∆7 carbons away from the terminal, methyl end of the aliphatic chain. Note that we denote the first isomer as TG_FA 18:1(n-9) and the second as TG_FA 18:1(n-7). To evaluate these double bond positions with the CLAW OzESI method, the user selected these positions shown in FIG.4a using a Jupyter Notebook. FIG.4a is a schematic which shows a user selects n-7 and n-9 double bond positions to evaluate. Referring to FIG.4b, a schematic is shown in which possible m / z values are automatically generated based on the user selected double bond positions and then matched and annotated with the experimental data. Briefly, LC was used to first separate TG molecular species based on their equivalent carbon number (ECN), as highlighted with the total ion chromatogram (TIC) shown in FIG.4c which provides a comparison of ozone off / on chromatogram for RBD canola oil with TG retention time identified. In general, the larger the ECN, the greater the retention time. For example, the most abundant TG molecular species, TG 54:3 and TG 52:2, both characterized by an ENC of 48, eluted at 18.1 min. TG 54:2 (ECN = 50) had a retention time of 20.0 min, while TG 52:3 and TG 54:4 (ECNs = 46) coeluted at 16.1 min. FIG.4c depicts the results of “ozone on” and “ozone off” experiments. Briefly, the black and red traces depict the TIC for “ozone off” and “ozone on” experiments, respectively. When the ozone gas is supplied to the nebulizer, the unsaturated TGs readily react with gaseous ozone, resulting in a decrease in the observed TIC signal.

[0078] Utilizing tabulated ozonolysis neutral loss values that have been extensively reported in the literature, we first tabulated a list of predicted TG precursor ion values ([TG + NH4]+) and the corresponding ozonolysis product ions. This tabulation occurs automatically by a python script in FIG.4b allowing CLAW to rapidly generate possible m / z values for each user selected double bond location for TGs in canola oil. For example, the ammonium cation adduct of TG 54:4 is observed at m / z 900.8. To monitor TG 54:4 molecular species containing the FA 18:1, we established an MRM precursor / product ion pair of m / z 900.8 à 601.6 by exploiting the NL of 299.2 Da indicative of FA 18:1. Based on predicted NL values following ozonolysis for monounsaturated lipids, TG 54:4 that carries an n-9 and n-7 double bond generated thePRF-70566-02 diagnostic product ions at m / z 790.6 and m / z 818.7, respectively. Next, exploiting the NL of FA 18:1, OzESI-MRMs can be established to monitor unsaturated lipid profiles with C=C specificity. Thus, to monitor the presence of TG 54:4_FA 18:1n-9 the precursor / product ion pair of m / z 790.6 à 690.9 was used, while the precursor / product ion pair of m / z 818.7 à 690.9 was used to profile TG 54:4_FA 18:1n-7. To demonstrate the LC-OzESI-MRM approach, FIGs.4d and 4e reflect the Extracted ion chromatograms (EIC)s of the above MRM transitions for TG 54:4 with ozone off and on, respectively. Briefly, when the ozone is off, only signal for the TG 54:4_FA 18:1 MRM transition (i.e., m / z 900.8 à 601.6) is observed. However, when ozone is admitted to the nebulizer, additional MRM signal is observed for the n-9 and n-7 C=C isomer channels for TG 54:4_FA 18:1. Referring to FIG.4f, a plot presents lipid ratio between the n-9 and n-7 isomers of a variety of TG molecular species in canola oil containing FA 18:1. In general, the isomeric ratio of TGs containing FAs 18:1-n-9 and n-7 remained relatively consistent in canola oil through its refinement stages. Thus, we can conclude that the refinement process does not significantly impact isomer ratios. While we were not surprised by consistent isomer ratios displayed throughout the refinement process, it is important to note that lipid precursor ion populations previously grouped as a single entity, in fact are composed of at least two distinct isomeric populations, previously unresolved with conventional LC-MS / MS approaches. To our knowledge, this is the first critical investigation of TG profiles at the C=C level in canola oil throughout the refinement stages.

[0079] To validate CLAW isomer ratio calculation results, CLAW outcomes were cross- referenced with the manual LC peak area data analysis. These results are shown in Tables 8 and 9 provided below and FIGs.4g, 4h, and 4i. Specifically, FIG.4g provides results of CLAW max intensity ratios and manually calculated area ratios for the OzESI crude canola oil data. The ratio of the n-9 / n-7 double bond location was compared at seven different TGs. FIG.4h provides results of CLAW max intensity ratios and manually calculated area ratios for the OzESI degummed canola oil data. The ratio of the n-9 / n-7 double bond location was compared at seven different TGs. FIG.4i provides results of CLAW max intensity ratios and manually calculated area ratios for the OzESI RBD canola oil data. The ratio of the n-9 / n-7 double bond location was compared at seven different TGs. Traditionally, relative abundance calculations such as thePRF-70566-02 isomer ratios portrayed herein, can be achieved via exploiting integrated peak area values generated within commercial data-processing software. However, peak area calculations are most reliable when analyte baseline separation can be achieved, requiring highly optimized and precise LC methods. Thus, we chose to employ maximum MRM intensity values as a basis for CLAW’s isomer ratio calculations as an alternate strategy. Notably, both CLAW and manual analysis showed consistent agreement. For example, an overall average standard deviation of 0.75 was observed between these two methods across all oil samples. Highlighting the precision and reliability of the CLAW's data processing capabilities, CLAW TG isomer ratio values calculated using maximum MRM intensity values revealed an average standard deviation of 0.29 for seven lipids across the three canola oil purities. In contrast, a higher average standard deviation of 0.37 was observed for TG isomer ratio values obtained via manual LC peak area calculations. Moreso, standards deviations of 0.84, 0.92, and 0.50 for crude, degummed, and RBD canola oils, respectively, were observed across automated CLAW and manual isomer ratio calculations. The observed discrepancies can be ascribed to CLAW selecting the maximum MRM intensity value from the raw data, while manual interpretation is subject to variations in peak area selection. Future work aims to further investigate the discrepancies arising between isomer ratio calculations that rely on LC peak area versus maximum MRM intensity values.

[0080] In the present disclosure, CLAW has been described as a comprehensive MRM -based pipeline for the detailed identification of lipids in complex biological samples. By automating processes traditionally done manually, CLAW significantly reduces analysis time and boosts throughput. In large-scale lipidomics, experimental worklist generation is tedious and susceptible to human-related errors which CLAW provides a solution by facilitating the generation of acquisition worklists. Following MRM-based experiments, raw data is converted to an mzML file format, which is then parsed by a CLAW python script using the pymzml package. During parsing, the data is annotated by matching the experimental results to a user selected MRM database. The annotated results are stored in a flexible pandas dataframe to easily extract the necessary data for statistical and bioinformatic analysis. An important statistical analysis feature in CLAW is the integration of EdgeR GLMs, marking a significant advancement in lipidomics data analysis by addressing the inherent variability of lipid MRM ion count distributions withPRF-70566-02 enhanced reliability in differential expression analysis. While CLAW represents a comprehensive solution, we advocate in general for the broader adoption of statistical methodologies, like edgeR’s GLM, to effectively handle over dispersed ion counts in MS-based lipidomics. To appeal to a broad range of users, CLAW’s GUI / LUI simplifies complex dataset analysis by enabling simultaneous comparison of multiple parameters. To improve user experience, the LUI provides a chatbot-style interaction that aids in analysis and data processing.

[0081] To demonstrate the capabilities of CLAW, we initially used the developed pipeline on biological samples from four distinct mouse brain regions. In the first example, traditional MRM-profiling of roughly 1500 individual lipid species was conducted on LDs isolated from specific brain regions of 5×FAD and WT mice. The data shows clear lipidome distinctions amongst LDs obtained from aged and AD-diseased brains, indicating the LDs related to aging and AD are not the same. Furthermore, distinct lipid signatures for the LDs isolated from the hippocampus, cortex, cerebellum, and diencephalon regions were observed.

[0082] In the second example, an online LC-OzESI-MRM method previously developed in- house was used to examine TGs with C=C specificity from several samples of canola oil taken at various stages of refinement. TGs profiled using OzESI-MRMs revealed little to no effect of the refinement process of TG isomer composition. While not surprising, the utilization of OzESI- MRMs in conjugation with CLAW successfully resolved, identified, and relatively quantified isomeric populations of TG molecular species that would otherwise remain unidentified using conventional or traditional workflows.

[0083] Referring to FIG. 5, a block diagram of the present disclosure is provided. A processor further discussed with reference to FIG. 7, executing software held on a non-transitory memory manages the system of the present disclosure. In operation, a human interfaces with the AI engine of the present disclosure via an AI manager. They AI manager through an AI agent communicates with LLMs, previously discussed, to generate questions and responses for the human user. The LLMs may be held locally or remotely on a cloud. The LLMs understand conversational queries from the human and provide conversational responses therefor. Once the human queries have been understood and analyzed, the AI agent along with the LLM decide which tools are needed to accomplish the task presented by the human. The system then reviews the available tools andPRF-70566-02 proceeds with choosing a tool that would provide a high chance of success. Once the tool is chosen, the AI agent and the LLM develop code to run the tool. If the tool returns errors during execution of the code, the system tries to correct the code and retries. If, however, no tool is identified, the system develops code for a new tool and after a successful execution of the code, adds the new tool to its toolkit (shown as Tool n+1).

[0084] Referring to FIG.6, a similar flow is shown as in FIG.5 but now different AI agents are chosen to accomplish different tasks. Each of these AI agents then proceeds with choosing a tool available in the toolkit to accomplish a corresponding task.

[0085] Referring to FIG.7, an example of a computer system is provided that can interface with the above-discussed system. Referring to FIG.7, a high-level diagram showing the components of an exemplary data-processing system 1000 for analyzing data and performing other analyses described herein, and related components. The system includes a processor 1086, a peripheral system 1020, a user interface system 1030, and a data storage system 1040. The peripheral system 1020, the user interface system 1030 and the data storage system 1040 are communicatively connected to the processor 1086. Processor 1086 can be communicatively connected to network 1050 (shown in phantom), e.g., the Internet or a leased line, as discussed below. The imaging described in the present disclosure may be obtained using imaging sensors 1021 and / or displayed using display units (included in user interface system 1030) which can each include one or more of systems 1086, 1020, 1030, 1040, and can each connect to one or more network(s) 1050. Processor 1086, and other processing devices described herein, can each include one or more microprocessors, microcontrollers, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), programmable logic devices (PLDs), programmable logic arrays (PLAs), programmable array logic devices (PALs), or digital signal processors (DSPs).

[0086] Processor 1086 can implement processes of various aspects described herein. Processor 1086 can be or include one or more device(s) for automatically operating on data, e.g., a central processing unit (CPU), microcontroller (MCU), desktop computer, laptop computer, mainframe computer, personal digital assistant, digital camera, cellular phone, smartphone, or any other device for processing data, managing data, or handling data, whether implemented with electrical, magnetic, optical, biological components, or otherwise. Processor 1086 can include Harvard-PRF-70566-02 architecture components, modified-Harvard-architecture components, or Von-Neumann- architecture components.

[0087] The phrase “communicatively connected” includes any type of connection, wired or wireless, for communicating data between devices or processors. These devices or processors can be located in physical proximity or not. For example, subsystems such as peripheral system 1020, user interface system 1030, and data storage system 1040 are shown separately from the data processing system 1086 but can be stored completely or partially within the data processing system 1086.

[0088] The peripheral system 1020 can include one or more devices configured to provide digital content records to the processor 1086. For example, the peripheral system 1020 can include digital still cameras, digital video cameras, cellular phones, or other data processors. The processor 1086, upon receipt of digital content records from a device in the peripheral system 1020, can store such digital content records in the data storage system 1040.

[0089] The user interface system 1030 can include a mouse, a keyboard, another computer (connected, e.g., via a network or a null-modem cable), or any device or combination of devices from which data is input to the processor 1086. The user interface system 1030 also can include a display device, a processor-accessible memory, or any device or combination of devices to which data is output by the processor 1086. The user interface system 1030 and the data storage system 1040 can share a processor-accessible memory.

[0090] In various aspects, processor 1086 includes or is connected to communication interface 1015 that is coupled via network link 1016 (shown in phantom) to network 1050. For example, communication interface 1015 can include an integrated services digital network (ISDN) terminal adapter or a modem to communicate data via a telephone line; a network interface to communicate data via a local-area network (LAN), e.g., an Ethernet LAN, or wide-area network (WAN); or a radio to communicate data via a wireless link, e.g., WiFi or GSM. Communication interface 1015 sends and receives electrical, electromagnetic or optical signals that carry digital or analog data streams representing various types of information across network link 1016 to network 1050. Network link 1016 can be connected to network 1050 via a switch, gateway, hub, router, or other networking device.PRF-70566-02

[0091] Processor 1086 can send messages and receive data, including program code, through network 1050, network link 1016 and communication interface 1015. For example, a server can store requested code for an application program (e.g., a JAVA applet) on a tangible non-volatile computer-readable storage medium to which it is connected. The server can retrieve the code from the medium and transmit it through network 1050 to communication interface 1015. The received code can be executed by processor 1086 as it is received, or stored in data storage system 1040 for later execution.

[0092] Data storage system 1040 can include or be communicatively connected with one or more processor-accessible memories configured to store information. The memories can be, e.g., within a chassis or as parts of a distributed system. The phrase “processor-accessible memory” is intended to include any data storage device to or from which processor 1086 can transfer data (using appropriate components of peripheral system 1020), whether volatile or nonvolatile; removable or fixed; electronic, magnetic, optical, chemical, mechanical, or otherwise. Exemplary processor- accessible memories include but are not limited to: registers, floppy disks, hard disks, tapes, bar codes, Compact Discs, DVDs, read-only memories (ROM), erasable programmable read-only memories (EPROM, EEPROM, or Flash), and random-access memories (RAMs). One of the processor-accessible memories in the data storage system 1040 can be a tangible non-transitory computer-readable storage medium, i.e., a non-transitory device or article of manufacture that participates in storing instructions that can be provided to processor 1086 for execution.

[0093] In an example, data storage system 1040 includes code memory 1041, e.g., a RAM, and disk 1043, e.g., a tangible computer-readable rotational storage device such as a hard drive. Computer program instructions are read into code memory 1041 from disk 1043. Processor 1086 then executes one or more sequences of the computer program instructions loaded into code memory 1041, as a result performing process steps described herein. In this way, processor 1086 carries out a computer implemented process. For example, steps of methods described herein, blocks of the flowchart illustrations or block diagrams herein, and combinations of those, can be implemented by computer program instructions. Code memory 1041 can also store data, or can store only code.PRF-70566-02

[0094] Various aspects described herein may be embodied as systems or methods. Accordingly, various aspects herein may take the form of an entirely hardware aspect, an entirely software aspect (including firmware, resident software, micro-code, etc.), or an aspect combining software and hardware aspects. These aspects can all generally be referred to herein as a “service,” “circuit,” “circuitry,” “module,” or “system.”

[0095] Furthermore, various aspects herein may be embodied as computer program products including computer readable program code stored on a tangible non-transitory computer readable medium. Such a medium can be manufactured as is conventional for such articles, e.g., by pressing a CD-ROM. The program code includes computer program instructions that can be loaded into processor 1086 (and possibly also other processors), to cause functions, acts, or operational steps of various aspects herein to be performed by the processor 1086 (or other processors). Computer program code for carrying out operations for various aspects described herein may be written in any combination of one or more programming language(s), and can be loaded from disk 1043 into code memory 1041 for execution. The program code may execute, e.g., entirely on processor 1086, partly on processor 1086 and partly on a remote computer connected to network 1050, or entirely on the remote computer. Referring to code below, software according to the present disclosure is provided to generate code by the AI agent where a tool is not present in the predetermined toolkit. --------------------------------------------------------------------------------------------------------------------- import os from langchain.agents import AgentType, initialize_agent from langchain.chat_models import ChatOpenAI from langchain.tools import E2BDataAnalysisTool os.environ["E2B_API_KEY"] = "e2b_d9ff840679da88a932be4bb6e93d524df295ee4b"PRF-70566-02 os.environ["OPENAI_API_KEY"] = "sk- m7QXcgpD5DGZ0pzdQ2McT3BlbkFJkG86xgzmlpe9ZQTrmLmR" # Artifacts are charts created by matplotlib when `plt.show()` is called def save_artifact(artifact): print("New matplotlib chart generated:", artifact.name) # Download the artifact as `bytes` and leave it up to the user to display them (on frontend, for example) file = artifact.download() basename = os.path.basename(artifact.name) # Save the chart to the `charts` directory with open(f". / charts / {basename}", "wb") as f: f.write(file) e2b_data_analysis_tool = E2BDataAnalysisTool( # Pass environment variables to the sandbox env_vars={"MY_SECRET": "secret_value"}, on_stdout=lambda stdout: print("stdout:", stdout), on_stderr=lambda stderr: print("stderr:", stderr), on_artifact=save_artifact, )PRF-70566-02 path_to_load_csv_result # with open(". / Genotype_ 5xFAD__Region_ Brain vs Genotype_ WT__Region_ BrainDropped2nd_full_reordered.csv") as f: with open(path_to_load_csv_result) as f: remote_path = e2b_data_analysis_tool.upload_file( file=f, description=""" Data about 5xFAD vs WT mouse models in the brain showing PValue, FDR, lipid, logFC commonly called log fold change and type which refers to lipid class, etc. """, ) print(remote_path) tools = [e2b_data_analysis_tool.as_tool()] llm = ChatOpenAI(model="gpt-4", temperature=0) agent = initialize_agent( tools, llm, agent=AgentType.OPENAI_FUNCTIONS,PRF-70566-02 verbose=True, handle_parsing_errors=True, ) from IPython.display import Image, display import io import re def display_plot_from_sandbox(final_answer: str) -> None: try: # This pattern matches any markdown image syntax with a sandbox path pattern = r'!\[.*?\]\(sandbox:(.*?)\)' match = re.search(pattern, final_answer) if match: extracted_info = match.group(1) plot_as_bytes = e2b_data_analysis_tool.session.download_file(extracted_info) # Assuming the method to download image = Image(data=plot_as_bytes) display(image) except Exception as e:PRF-70566-02 print(e) # It's helpful to print out the exception for debugging return final_answer ---------------------------------------------------------------------------------------------------------------------

[0096] Those having ordinary skill in the art will recognize that numerous modifications can be made to the specific implementations described above. The implementations should not be limited to the particular limitations described. Other implementations may be possible.

Claims

PRF-70566-02 Claims:

1. An artificial intelligence (AI) system, comprising: a processor executing software residing on a non-transient memory, the processor configured to: establish an AI manager configured to communicate with an external entity; establish at least one AI agent each coupled to the AI manager and each configured to communicate with one or more multi-modal large language models (LLMs), whereby communication between the external entity and the AI agent is analyzed and generated by the one or more LLMs thus establishing conversational communication between the external entity and the AI manager; analyze queries from the external entity; determine if a tool from a predetermined toolkit exist to be used to address the external entity query, if a tool is identified to exist, the AI agents generates input for the identified tool and executes the identified tool with the generated input; if no tool is identified to exist, the AI agent generates code and new-tool inputs for a new tool and executes the new tool with the generated code and the new-tool inputs and adds the new tool to the toolkit.

2. The AI system of claim 1, wherein the external entity is a human.

3. The AI system of claim 1, wherein the external entity is a computing device.

4. The AI system of claim 1, wherein if the during execution of the identified tool, the identified tool generates an error, the AI agent analyzes the error, corrects the input, and re- executes the identified tool with the corrected input.PRF-70566-02 5. An artificial intelligence (AI)-based method for providing a toolkit, comprising: establishing an AI manager configured to communicate with an external entity; establishing at least one AI agent each coupled to the AI manager and each configured to communicate with one or more multi-modal large language models (LLMs), whereby communication between the external entity and the AI agent is analyzed and generated by the one or more LLMs thus establishing conversational communication between the external entity and the AI manager; analyzing queries from the external entity, determining if a tool from a predetermined toolkit exist to be used to address the external entity query, if a tool is identified to exist, the AI agents generating input for the identified tool and executing the identified tool with the generated input; if no tool is identified to exist, the AI agent generating code and new-tool inputs for a new tool and executing the new tool with the generated code and the new-tool inputs and adding the new tool to the toolkit.

6. The AI-based method of claim 5, wherein the external entity is a human.

7. The AI-based method of claim 5, wherein the external entity is a computing device.

8. The AI-based method of claim 5, wherein if the during execution of the identified tool, the identified tool generates an error, the AI agent analyzes the error, corrects the input, and re- executes the identified tool with the corrected input.

Citation Information

Patent Citations

  • User dialog-based automated system design for programmable integrated circuits

    US10922463B1

  • Multi-lingual virtual personal assistant

    US20190332680A1

  • Utilizing artificial intelligence to improve productivity of software development and information technology operations (devops)

    US20210064361A1