Secure agentic data collaboration with ai-driven governance and polymorphic identity management

WO2026178153A1PCT designated stage Publication Date: 2026-08-27LIVERAMP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/015707
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-20
Filing Date
2026-02-18
Publication Date
2026-08-27

Smart Images

  • Figure US2026015707_27082026_PF_FP_ABST
    Figure US2026015707_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A system for secure multi-party data collaboration with Al-driven governance and polymorphic identity management comprises a secure data cleanroom environment, a polymorphic data identity management framework that encodes each party's data within unique identity spaces to prevent direct correlation of PH or confidential business data across parties, an agentic data connection component for defining data-source-level Al collaboration rules, an agentic dataset component for per-column / per-field access control policies subordinate to source-level rules, an Al-driven governance engine that dynamically enforces and adapts access policies, and a logging component tracking Al agent activities. The hierarchical control structure enables organizations to leverage Al for collaborative insights, RAG systems, and LLM development while maintaining strict data privacy and regulatory compliance.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. RAMP-00325-WOSECURE AGENTIC DATA COLLABORATION WITH AI-DRIVEN GOVERNANCE AND POLYMORPHIC IDENTITY MANAGEMENTCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of US Provisional Patent Application No. 63 / 761,005, filed on February 20, 2025. Such application is incorporated by reference herein.BACKGROUND OF THE INVENTION

[0002] Existing data cleanroom solutions typically provide basic security features like encryption and access controls. However, they often lack the granular control and dynamic governance capabilities to manage Al agent access. Current approaches to access control, such as role-based access control (RBAC) or attribute-based access control (ABAC), are often too static and inflexible to effectively manage the complex interactions between artificial intelligence (Al) agents and sensitive data.

[0003] Furthermore, existing solutions often fail to address the dynamic tension between enduser enthusiasm for Al and the more cautious approach of legal, data policy, and data ethics teams. They do not provide a mechanism for balancing these competing interests within a unified governance framework. Critically, they often fail to address the risks associated with linking or correlating personally identifiable information (PH) across different partner datasets within the cleanroom, even if the data is anonymized or pseudonymized.

[0004] References mentioned in this background section are not admitted to be prior art with respect to the invention as defined by the claims set forth herein.SUMMARY OF THE INVENTION

[0005] The invention is directed to a system and method that allows multiple organizations to collaborate on data analysis, data activation, and Al model development within a secure data cleanroom. It employs a polymorphic data identity management framework to encode each party's data with a unique identifier. This encoding prevents direct access by humans, Al agents, or linkage of Pll or confidential business data, even within the cleanroom. On top of this, an Al- driven governance engine enforces fine-grained access control policies for Al agents, ensuring that they can only access data in ways specifically permitted.Attorney Docket No. RAMP-00325-WO

[0006] The invention allows parties to share data and work together on Al projects without compromising data privacy or security. This is particularly valuable for industries dealing with sensitive information, such as advertising, healthcare or finance, where collaboration can lead to significant advancements but carries inherent risks. The invention aims to balance the desire for data-driven insights, data activation, and advanced Al activities— such as retrieval-augmented generation (RAG) and large language model (LLM) creation— with the need for strict data protection.

[0007] The invention is directed to a system and method to secure agentic data collaboration, particularly within data cleanrooms. It addresses the critical need for granular control over Al agent access to sensitive data while enabling organizations to leverage the power of Al for collaborative insights, activation, RAG, and LLM training.

[0008] The invention in various embodiments is designed to solve the problem of securely and effectively collaborating on data analysis, data activation, and Al development when dealing with sensitive information from multiple parties. In today's data-driven world, organizations recognize the immense value of combining datasets for insights, but doing so safely presents significant challenges. Existing data cleanroom solutions often lack the fine-grained Al control and dynamic Al governance necessary to prevent unauthorized access to Personally Identifiable Information (PH) or confidential business data. This invention directly addresses these issues.

[0009] The invention in various embodiments tackles the tension between the desire for collaboration and the need for data privacy . It specifically focuses on the risks of linking PH or confidential / business data across different datasets within a cleanroom, even if the data is anonymized or pseudonymized. By implementing a polymorphic data identity management system, such as but not limited to LiveRamp's RampID®, the invention ensures that each party's data remains within its unique identity space. This means data from different parties cannot be directly correlated or linked, even while being analyzed together either by humans or Al agents.

[0010] The system according to an embodiment addresses limitations of certain prior art systems by introducing several innovations. One innovation is granular Al agent access control. The system employs a policy-based access control mechanism that allows organizations to define highly specific rules governing Al agent access to data. These rules can be based onAttorney Docket No. RAMP-00325-WOvarious factors, including the type of Al agent, the specific data being accessed (represented within a specific identity space), the purpose of the access, and the context of the collaboration.

[0011] Another innovation is the dynamic governance engine. The Al-driven governance engine dynamically adapts access control policies based on real-time conditions, user roles, and data sensitivity. This allows organizations to respond quickly to changing regulatory requirements, internal policies, and ethical considerations.

[0012] Another innovation is polymorphic data identity management. In an embodiment, the system's polymorphic data identity management framework, using a service like RampID®, ensures that each partner's data remains within its own unique identity space within the cleanroom. This prevents partners from directly correlating or linking PH, or confidential business data across datasets, even if the data is pseudonymized, anonymized or aggregated within the cleanroom. Only the identity provider can perform the necessary transcoding for authorized data sharing or analysis.

[0013] Another innovation is integration with existing workflows. The system in an embodiment is designed to integrate seamlessly with existing Al workflows and data science tools, minimizing disruption to end-users while ensuring compliance with organizational policies. This can be achieved through application programming interfaces (APIs), software development kits (SDKs), and other integration mechanisms.

[0014] The system according to an embodiment provides a number of advantages. One such advantage is improved data security. The system's granular access control, polymorphic data identity management framework, and Al-driven governance significantly reduce the risk of data breaches, unauthorized access, and unintended Pll linkage or confidential business data within the cleanroom.

[0015] Another such advantage is enhanced control over Al usage. The Al-driven governance engine empowers organizations to define and enforce clear policies regarding Al agent access to data, ensuring compliance with legal, ethical, and policy requirements.

[0016] Another such advantage is increased collaboration opportunities. By providing a secure and controlled environment for data collaboration, the system enables organizations to unlock the full potential of their data while mitigating the risks associated with data sharing withinAttorney Docket No. RAMP-00325-WOagentic workflows (insights generation, activation recommendation, audience expansion, RAG, and LLM training).

[0017] Another such advantage is streamlined compliance. The system simplifies compliance with data privacy regulations such as GDPR, CCPA, and others by providing a centralized platform for managing data access and enforcing data governance policies.

[0018] In certain embodiments, the system encompasses a unique combination of granular Al agent access control, a dynamic Al-driven governance engine, and polymorphic data identity management within a data cleanroom context. Specifically, the novel integration of these elements, particularly the use of distinct data identity spaces within the cleanroom to prevent Pll linkage across partners and agentic workflows, to address the dynamic tension between enduser Al enthusiasm and organizational policy, represents a significant advancement.

[0019] An advantage of the system in certain embodiments is that it significantly reduces the hardware footprint in multi-party data collaboration by employing granular control mechanisms for data access, coupled with a centralized data cleanroom and polymorphic identity management. In traditional models, each bilateral collaboration within a data cleanroom requires separate compute and storage instances, leading to a rapid increase in hardware needs as the number of partners grows. The proposed system consolidates these resources by providing a unified, shared data cleanroom environment. This consolidation is advantageous for managing the intensive computational demands of Al agents, RAG systems that fetch and process information, and LLMs that require substantial memory and processing power.

[0020] By utilizing a single shared environment, the system according to certain embodiments avoids the redundancy inherent in multiple isolated setups. Al agents can access and process data without needing separate instances for each partner interaction. Collaborative RAG benefits from a centralized knowledge base, reducing the storage and retrieval overhead that comes with distributed data silos. Similarly, collaborative LLMs can be trained and deployed within this unified space, optimizing resource allocation and preventing the fragmentation of training data and compute power.

[0021] In one aspect, the invention is directed to a system for secure multi-party data collaboration with Al-driven governance, comprising a secure data cleanroom environment configured to receive and store data from a plurality of collaborating parties, a polymorphic dataAttorney Docket No. RAMP-00325-WOidentity management framework in communication with the secure data cleanroom environment, wherein the polymorphic data identity management framework is configured to encode data from each collaborating party within a respective unique identity space, such that data from different collaborating parties cannot be directly correlated or linked, and such that direct access to personally identifiable information (PH) or confidential business data between parties is prevented, an agentic data connection and sources configuration component configured to manage connections to data sources feeding the secure data cleanroom environment and to define, for each data source, a first level of Al collaboration rules specifying permitted Al agent interactions with data originating from that data source, an agentic dataset cleanroom configuration component configured to define, at a per-column (or per-field) granularity within datasets in the secure data cleanroom environment, a second level of access control policies governing Al agent access to individual data elements, wherein the second level of access control policies is subordinate to the first level of Al collaboration rules, an Al-driven governance engine configured to enforce the first level of Al collaboration rules and the second level of access control policies for Al agents accessing data within the secure data cleanroom environment, and to dynamically adapt at least one of the first level of Al collaboration rules or the second level of access control policies based on at least one of real-time conditions, user roles, or data sensitivity, and a cleanroom agentic logging and monitoring component configured to record Al agent activities within the secure data cleanroom environment, including at least identifiers of data sources, datasets, and columns accessed by each Al agent activity.

[0022] In another aspect, the invention is directed to a computer-implemented method for governing Al agent access to data in a secure multi-party data collaboration environment, comprising receiving data from a plurality of collaborating parties into a secure data cleanroom environment, encoding the data from each collaborating party within a respective unique identity space using a polymorphic data identity management framework, such that data from different collaborating parties cannot be directly correlated or linked and such that direct access to personally identifiable information ( PI I) or confidential business data between parties is prevented, configuring, for each data source connected to the secure data cleanroom environment, a first level of Al collaboration rules defining permitted Al agent interactions with data originating from that data source, including at least one of permitted Al workflow types, allowed Al models, or query tracking levels, configuring, at a per-column or per-field granularityAttorney Docket No. RAMP-00325-WOwithin datasets derived from the data sources, a second level of access control policies governing Al agent access to individual data elements, the second level being hierarchically subordinate to the first level such that a column or field is available for Al agent access only if the corresponding data source has a permitting first-level rule, enforcing the first level of Al collaboration rules and the second level of access control policies via an Al-driven governance engine when Al agents access or process data within the secure data cleanroom environment, and logging Al agent activities within the secure data cleanroom environment, including recording identifiers of accessed data sources, datasets, and columns for each Al agent activity.

[0023] In another aspect, the invention is directed to a computer system for managing secure Al-driven multi-party data collaboration, comprising one or more processors and memory storing instructions that, when executed by the one or more processors, cause the computer system to provide a first user interface configured to enable configuration of data source connections to a secure data cleanroom environment and definition of Al collaboration rules for each data source, the Al collaboration rules specifying at least permitted large language model (LLM) models, permitted agentic workflow types, and query tracking levels, a second user interface configured to enable configuration of per-column or per-field access control policies for datasets within the secure data cleanroom environment, including selection of an Al agent access level for each enabled column or field from among column name only, column name and row data, and column name and synthetic data for rows, a third user interface configured to enable management of Al agents and agentic workflows operating within the secure data cleanroom environment, and a fourth user interface configured to display logs of Al agent activities within the secure data cleanroom environment, the logs including at least identifiers of originating cleanrooms, accessed datasets, accessed columns, and Al workflow types for each logged activity.

[0024] In another aspect, the invention is directed to a computer-implemented method for collaborative Al model development within a secure data cleanroom, comprising onboarding data from a plurality of collaborating parties into a secure data cleanroom environment, encoding the onboarded data using a polymorphic data identity management framework to create a respective unique identity space for each collaborating party's data, thereby preventing direct correlation or linkage of personally identifiable information (PH) or confidential business data across parties, configuring, via a hierarchical policy structure comprising a data-source-levelAttorney Docket No. RAMP-00325-WOset of Al collaboration rules and a column-level or field-level set of access control policies subordinate thereto, permitted interactions between Al agents and the onboarded data, collaboratively training or building an Al model, selected from one or more of a Retrieval- Augmented Generation (RAG) system and a Large Language Model (LLM), on combined datasets within the secure data cleanroom environment in accordance with the hierarchical policy structure, and logging activities of Al agents involved in the collaborative training or building, including identifiers of accessed data sources, datasets, and columns.

[0025] These and other features, objects and advantages of the present invention will become better understood from a consideration of the following detailed description of the preferred embodiments and appended claims in conjunction with the drawings as described following:BRIEF DESCRIPTION OF DRAWINGS

[0026] Fig. 1 depicts an overall user interface (Ul) for data connections in a collaborative cleanroom according to an embodiment of the present invention.

[0027] Fig. 2 depicts data source selection at the Ul according to an embodiment of the present invention.

[0028] Fig. 3 depicts the entry of configuration information at the Ul according to an embodiment of the present invention.

[0029] Fig. 4 depicts naming and storage of configuration rules at the Ul according to an embodiment of the present invention.

[0030] Fig. 5 depicts definition of specific rules for Al agents at the Ul according to an embodiment of the present invention.

[0031] Fig. 6 depicts control over Al agent data access at the Ul according to an embodiment of the present invention.

[0032] Fig. 7 depicts control over access to specific columns or fields within a dataset at the Ul according to an embodiment of the present invention.

[0033] Fig. 8 depicts Al workflow activity control at the Ul according to an embodiment of the present invention.Attorney Docket No. RAMP-00325-WO

[0034] Fig. 9 depicts Al insights generated about a brand buyer according to an embodiment of the present invention.

[0035] Fig. 10 depicts platform activation according to an embodiment of the present invention.

[0036] Fig. 11 depicts a computer system used for implementing a cloud computing system and network according to an embodiment of the present invention.DETAILED DESCRIPTION

[0037] Before the present invention is described in further detail, it should be understood that the invention is not limited to the particular embodiments described, and that the terms used in describing the particular embodiments are for the purpose of describing those particular embodiments only, and are not intended to be limiting, since the scope of the present invention will be limited only by the claims.

[0038] To enable the operation and management of the secure multi-party data collaboration agentic system with Al-driven governance and polymorphic identity management, a suite of intuitive user interfaces (Uls) may be employed. The following four core systems, each with dedicated Uls, are identified for implementation. These systems, and their corresponding Uls, work in concert to provide the functionality outlined herein. Each system addresses a distinct aspect of the overall architecture and contributes to the system's ability to provide secure, governed, and privacy-preserving Al-driven collaboration. The Uls are designed to be accessible to both technical and non-technical users, facilitating seamless interaction with the system's powerful capabilities. These Uls include:

[0039] (i) Agentic Data Connection / Sources Configuration and Access

[0040] (ii) Agentic Dataset Cleanroom Configuration and Access

[0041] (iii) Al Agent, Agentic Workflow & Al Model Management

[0042] (iv) Cleanroom Agentic Logging and Monitoring

[0043] Implementation begins with the Agentic Data Connection / Sources Configuration and Access system 100, as shown in Fig. 1. This system's Ul allows users to define and manage the data sources that will feed into the collaboration environment. This includes connecting to various data repositories, setting up data pipelines, and defining access permissions for different data sources. The Ul provides controls for authentication, authorization, and data validation,Attorney Docket No. RAMP-00325-WOensuring that only authorized and clean data enters the system and is then exposed to Al agents and humans.

[0044] Next is the Agentic Dataset Cleanroom Configuration and Access system 200, as shown in Figs. 3-6. Its Ul enables users to establish rules regarding dataset columns and rows, defining what data is included in which analysis and by whom. This involves setting up granular access controls at the dataset level, specifying which parties can access which data (column and row level), and for what purposes. This Ul also provides tools for data exploration and transformation, allowing users to prepare their data for analysis within the data cleanroom environment.

[0045] The Al Agent, Agentic Workflow & Al Model Management system Ul 300, shown in Fig.7, is used for configuring the data cleanroom itself. This includes defining system-wide settings, managing Al agents, and creating and deploying agentic workflows. Users can leverage the Ul to define the behavior of Al agents, setting up their access permissions, defining the tasks they are allowed to perform, and monitoring their activity. Furthermore, this interface also handles model management and versioning.

[0046] Finally, the Cleanroom Agentic Logging and Monitoring system Ul 400, shown in Fig. 8, provides for rich Al event logging and Al monitoring. This interface presents detailed logs of all activities within the cleanroom, including data access, Al agent actions, and policy enforcement events. This allows administrators to track system usage, identify potential security issues, and ensure compliance with data governance policies. It also allows for error handling and troubleshooting with specific logging per agent and action taken.

[0047] In essence, the system is implemented through these four interconnected systems, each with user-friendly Uls that abstract away the complexities of the underlying technology. This Ul- centric approach allows for widespread adoption and ease of use, making powerful data collaboration and Al development accessible to a broader range of users. Each of the four Uls and their operation are discussed in more detail below.

[0048] Agentic Data Connection / Sources Configuration and Access Ul 100 allows data operators and administrators to configure and connect data sources, or data connections, to a data collaboration cleanroom. Agentic, in this context, refers to the use of Al agents to automate and enhance data processing and collaboration. Data connections can include databases, cloud storage, APIs, and other data sources.Attorney Docket No. RAMP-00325-WO

[0049] Data source 102 represents an example data source and stores the metadata and configuration required to access and connect the data for future usage in one or many data cleanrooms. Data operators and administrators will recognize this pattern present in multiple ETL (Extract, Transform, Load) products, which are tools used to move and transform data between systems. Data source name, status, and type are all expected metadata properties.

[0050] Data operator interface 104 allows, for each selected data source, data operators and administrators to view additional metadata and configuration aspects (details, linked assets / dataset / tables). Al Collaboration Rules details the Al agentic workflows that are enabled for any data coming from this data source to one or many cleanrooms. This includes which LLM models are allowed, if Al Agents can analyze and query the data coming from this data source, how the Al activities are tracked and logged, and finally, if the Al agent can assess data activation status for any or a selected list of destinations (advertising platforms). Al collaboration rules might specify which Al models can access the data, what types of queries Al agents can perform, and how Al activity is logged and audited.

[0051] It is important to note that this configuration is the first level of configuration and control where data operators and administrators can enforce the downstream usage of Al agents and agentic workflows. This Al collaboration rules section allows for the first level of fine-grained control over how Al agents interact with data within the data cleanroom environment. This granular level of control is crucial for maintaining data security and privacy within the multi-party collaboration environment. The next level of fine-grained control will be introduced at the cleanroom dataset level, as described in the following sections.

[0052] The Ul flows for creating a new data source 200, shown in Fig. 2, enables data collaboration within the system with step-by-step details for creating a new data source that can be utilized within one or many data cleanrooms.

[0053] At data source creation 202 in Fig. 2, the data source creation flow begins by selecting the method of data access, which can include various types of data connections such as cloud storage (e.g., AWS, Azure, GCP), SFTP, manual uploads, APIs, or connectors for specific platforms like Salesforce or Linkedln. In this particular case, the data source will connect to an AWS cloud bucket.

[0054] At configure source 204 in Fig. 3, the configuration information that is required to connect to the AWS data is gathered. Each data source type will have a different set of configuration requirements at this step.Attorney Docket No. RAMP-00325-WO

[0055] At configuration name block 206 in Fig. 4, the Ul provides a simple affordance to name and store the Al collaboration rules so they can be reused across multiple data sources. This increases efficiency for data operators and administrators who manage multiple data sources across multiple projects and need to apply similar rules to many data sources.

[0056] At allowed large language model (LLM) models 208, the data operators and administrators can pick and allow different LLM models that will be enabled against these data sources and Al agentic workflows in the cleanroom. This list can be curated, and specific configuration screens will also be added to the organization setting pages for Al keys and location. This provides control to administrators for specifying LLM models and Al providers that their data ethics and legal team have approved for their organization.

[0057] At Al collaborative agents and workflows 210, the data operators and administrators can pick and allow which agentic workflows are enabled for a given data source. This provides a high-level control on downstream data activities that will be enabled in all cleanrooms that are pulling data from this data source. A potential use case would be that a Zendesk data source could be used for only collaborative RAG development and not be available for other Al agents or collaborative LLMs.

[0058] At allow Al agents to query 212 in Fig. 5, this Ul selection allows data operators and administrators to define specific rules for Al agents to perform analytics on the data originating from the current data source. This enables fine-grained control over how Al is utilized for analytical purposes, such as generating insights, identifying trends, and creating reports. For instance, in a retail media network (RMN) context, these rules can empower Al agents to analyze customer purchase history from a specific data source, providing brand partners with valuable information about customer segments, product preferences, and campaign performance. This level of control is useful for scenarios like enhancing audience targeting and measurement, where Al agents need to derive insights from combined datasets from different partners (Brand, Retailer, Third-Party Data Seller) securely and in compliance with data policies.

[0059] By configuring these analytic rules, data operators can specify which Al agents are authorized to perform particular types of analysis and what data points they can access. This ensures that Al-driven analytics are conducted in a governed and privacy-preserving manner. In the following sections, it will be explained how this Ul selection and column / row controls can let an RMN allow Al agents to analyze aggregated sales data to identify popular product categories but restrict access to individual customer purchase records to maintain privacy.Attorney Docket No. RAMP-00325-WO

[0060] At query tracking 212, the Ul allows for the management of the level of Al agent query tracking, specifically for human oversight and understanding of Al agent activities within the data cleanroom. Given that these queries are auto-generated by the Al agents, this tracking is crucial for transparency and accountability. Three levels of monitoring can be configured. One level is "no tracking." This level disables all logging of Al agent queries. This provides no visibility into the Al's actions, hindering auditability and troubleshooting. This option would rarely be used in a data cleanroom environment where auditing and transparency are paramount.

[0061] A second level is query statement-level tracking. This logs the query statements made by Al agents. This allows data operators and administrators to see the questions the Al is asking of the data, providing valuable insights into the Al's reasoning and analysis. This level is suitable for understanding the general flow of Al queries and ensuring they align with intended use cases.

[0062] A third level is query statements and result tracking. This level logs both the query statements and their corresponding results. This provides the most comprehensive level of tracking, enabling detailed auditing, debugging, and performance analysis of Al agents. By seeing both the queries and their results, human operators can verify the accuracy of the Al's analysis, identify potential biases or errors, and fine-tune the Al's behavior. This level of tracking is essential for ensuring compliance, maintaining data quality, and building trust in the Al-driven processes within the cleanroom.

[0063] For example, in a retail media network (RMN), a retailer might choose to enable query statement-level tracking to monitor the types of analytical queries that Al agents are making on customer purchase data. This would help them understand how Al is being used to generate insights without exposing the specific details of those insights. This is very useful for understanding what the Al is doing without looking into the specific result.

[0064] On the other hand, for a critical data source used in financial analysis, the retailer might enable query statements and result tracking to ensure full accountability and auditability of all Al-driven operations. This level would allow human operators to investigate the data that the Al used to come to specific results or debug a new Al agent; having the query and result will significantly improve the overall speed of debugging. A human review of the results can ensure that the Al performs within the required guardrails.Attorney Docket No. RAMP-00325-WO

[0065] At activation rules 216 in Fig. 6, this Ul element provides the initial control for data operators and administrators to enable or disable the Al agent's ability to assess data activation status and related metrics for various ad tech platforms (e.g., Meta, X, Trade Desk) for the data originating from the current data source. By toggling this setting, data operators can manage whether Al agents can interact with and retrieve information related to data activation from these external platforms. This allows for granular control over the Al's role in monitoring and optimizing data activation campaigns.

[0066] In a particular use case example, a retail media network (RMN) is working on a targeted advertising campaign with a brand partner . The brand wants to understand how well its data is activating across platforms like Meta and Trade Desk. The RMN's data operator enables the activation assessment Al agent for the data source containing the brand's customer segments. Now, Al agents can analyze activation metrics, such as match rates, audience reach, and campaign performance on these platforms. If the RMN wants to pause or restrict this data flow, it can simply disable the activation assessment Al agent forthat specific data source, ensuring data privacy and control. For instance, if there's a temporary issue with a platform's API, the RMN can quickly disable the agent to prevent errors or data inaccuracies.

[0067] At destination selection 218, this Ul element allows data operators and administrators to further refine the behavior of the activation assessment Al agent. It provides controls to specify the destinations (e.g., ad tech platforms) for the agent to assess data activation status. Three options are available. One option is "Only Selected Destinations (Whitelist)." Using this option, the Al agent can only assess the activation status for the platforms explicitly listed in a provided list. This provides the most restrictive control and ensures that the agent only interacts with approved destinations.

[0068] A second option is "All Except Selected Destinations (Blacklist)." Using this option, the Al agent can assess the activation status for all platforms except those specifically listed in a provided list. This allows for excluding platforms where activation assessment is not desired or permitted.

[0069] A third option is "All Destinations." Using this option, the Al agent can assess activation status and related metrics for all available ad tech platforms (e.g., Meta, X, Trade Desk).

[0070] In a use case example, an RMN has partnered with several ad tech platforms. However, due to data-sharing agreements, only a subset of these platforms are approved for automated activation assessment. The data operator configures destination selection 218 to "Only SelectedAttorney Docket No. RAMP-00325-WODestinations" and adds Meta and Trade Desk to the list. This ensures that the activation assessment Al agent only interacts with and retrieves data from these two platforms, maintaining compliance with data policies. Suppose a new platform is added to the approved list. In that case, the operator can easily update the selection, granting the Al agent access to assess the activation status for that platform as well.

[0071] At destination choice 220, this section presents the list of destinations (e.g., ad tech platforms) that are affected by the selection made in destination selection 218. If "Only Selected Destinations" is chosen, this list shows the platforms where the activation assessment Al agent can operate. If "All Except Selected Destinations" is chosen, this list shows the platforms where the agent cannot operate.

[0072] The description now turns to agentic dataset cleanroom configuration and access 300, as shown in Fig. 7. This shows how cleanroom collaborators and data scientists configure datasets or tables for a data collaboration cleanroom. This step usually follows once a data source has been configured. A data source can be used and re-used in the configuration of datasets and tables and their Al workflow configuration / rules / permission will influence what is available to the collaborators and data scientists.

[0073] The first steps (Select Intended Use, Configure Table, Configure Identity Resolution...) related to adding a dataset / tables to a collaborative cleanroom will be skipped as they follow the standard ETL, Data processing patterns from the industry. The "Map Fields" steps can be seen in Fig. 7. At Al workflows 302 in Fig. 7, this input governs whether a specific column or field within a dataset or table can be included in Al agent workflows, also including collaborative RAG and collaborative LLM creation. This control is exercised at a granular, column-by-column level, allowing for precise management of data accessibility for Al.

[0074] Al workflows 302 enables global enablement dependency. A column / field can only be enabled for Al workflows if Al collaboration rules have been previously configured and activated at the data source level 210. This ensures a hierarchical access control structure, where broader data source permissions must first be in place.

[0075] Al workflows 302 also provides for column-specific toggle by providing a toggle or checkbox for each individual column / field. When this toggle is activated (checked), it signifies that the corresponding column / field is permitted to participate in Al workflows within the data cleanroom.Attorney Docket No. RAMP-00325-WO

[0076] Al workflows 302 also provides detailed side panel controls. Activating the toggle for a column / field in 302 unlocks additional, more detailed Al agentic workflow controls in a side panel. These controls are further defined in 308, 310, 312, 314, and 316, as described below. They allow precise management of how Al agents can interact with the specific data within that column / field.

[0077] Al workflows 302 also provides for selective data exposure. Its primary function is to enable selective data exposure to Al agents. This allows data operators and administrators to carefully choose which data elements are relevant and safe for Al processing, and which should be excluded to protect sensitive information or maintain operational control.

[0078] It may be seen then that Al workflows 302 is the "gatekeeper" for each column / field. It determines whether that specific piece of data is allowed to "talk" to the Al agents or not. Before any column can be opened up to the Al, the higher-level "Al Collaboration Rules" data source level 210 must already be set. One may think of data source level 210 as the overall permission slip, and Al workflows 302 as the specific pass for each individual data element. If the column / field is enabled here, then the detailed side panel controls come into play for even more specific instructions on what the Al can do with that data. This level of control is essential in a secure multi-party data collaboration environment to balance the power of Al with the need for data privacy, security, and regulatory compliance. It ensures that only authorized data is used by Al agents and that the data is used consistently with organizational policies.

[0079] At first example row 304, unchecking the box for 'distributor_key' and ' transaction d' means these columns are effectively hidden from Al agent workflows. The Al agents will not be able to access or process this data, providing an additional layer of security and focusing the Al on the data intended for analysis. For example, a commerce media movie operator can enable Al agent access for "movie_title", "genre", and "chain_name" to allow Al to analyze trends and optimize ad targeting; disable Al agent access for "Maid" to prevent the Al from directly linking viewing habits to specific individuals, mobile devices; or disable Al agent access for 'distributor key' and 'transaction id' because they are operational data irrelevant to what the Al agents need to perform. This also mitigates against any leakage of data to external parties.

[0080] Input at Al agent description 308 provides a mechanism to add detailed descriptions and specific instructions for Al agents interacting with those columns / fields. This enhances the Al's understanding of the data and optimizes its performance within the data cleanroom environment.Attorney Docket No. RAMP-00325-WO

[0081] The column / field description is a free-text field is provided for each enabled column / field, allowing users to input a detailed description of the column's contents, purpose, and context. This description can include information about the data type, format, units, and any special considerations or potential biases. This enhances the Al's ability to interpret and utilize the data effectively, especially for complex or nuanced datasets.

[0082] Al agent instructions is a separate free-text field is available for each enabled column / field, where users can provide specific instructions to the Al agents. These instructions can guide the Al on best using the data for particular tasks or workflows. For example, they could provide context for potential outliers or anomalies; guides to the Al on how to handle missing or null values; instructions for the Al to focus on specific data ranges or categories; requests to the Al to generate specific types of outputs or reports based on the data in this column; or explicit disallowal of certain operations or interpretations to align with policy or privacy needs.

[0083] Workflow context descriptions and instructions are linked to the specific Al workflows enabled for the column / field. Therefore, the guidance provided directly applies to how the Al agents interact with the data in those workflows. In an example, one may consider a column named "Customer Purchase Amount" in a retail media network dataset. The column / field description could be, "This column represents the total purchase amount for each transaction in USD. Values may range from $0.01 to $10,000. Includes taxes and discounts. Note that unusually high values may indicate bulk purchases or potential data entry errors."

[0084] An example of Al Agent Instructions would be, "When analyzing this column for sales trends, use a box plot to identify potential outliers. For predictive modeling, consider segmenting customers based on purchase amount quartiles. If a value is zero, it might indicate a promotional item or an error; flag these for review."

[0085] Al agent description 308 improves Al Accuracy by providing detailed descriptions and instructions helps Al agents better understand the data, leading to more accurate analysis and insights. It enhances collaboration because clear instructions enable effective collaboration between data operators, scientists, and Al agents. It allows for customized Al behavior because specific instructions allow for tailoring Al behavior to meet specific analytical or operational goals. It mitigates risk by explicit disallowances or restrictions that can help prevent misuse or misinterpretation of sensitive data by Al agents. And finally, it provides documentation becauseAttorney Docket No. RAMP-00325-WOthe descriptions and instructions serve as valuable documentation for the dataset and its intended use with Al.

[0086] By incorporating agent description 308, the system transforms from a basic data access control system to a more intelligent and nuanced platform for Al-driven data collaboration. It empowers users to guide Al behavior, improve data quality, and ensure that Al agents operate within the desired parameters.

[0087] Insight and analytic Al agents 310 becomes available when "Allow Al Agents" is enabled at the data source level 210. It provides granular control over the enablement and capabilities of analytic Al agents for specific fields / columns within a dataset. This allows for tailored Al- driven analysis while maintaining data privacy and security.

[0088] Insight and analytic Al agents 310 provides for conditional availability. It is only accessible if "Allow Al Agents" has been activated for the corresponding data source. This ensures a hierarchical control structure, where data source-level permissions precede field / column-level permissions.

[0089] It further allows for field / column-specific enablement: For each field / column where Al agent interaction is enabled, insight and analytic Al agents 310 provides a dropdown menu to control the level of access and capabilities for analytic Al agents. The dropdown menu offers the following options:

[0090] (i) Column Name Only: The Al agent can only access the name or identifier of the column. This allows for metadata analysis, schema understanding, and basic data discovery without exposing the actual data values.

[0091] (ii) Column Name and Row: The Al agent can access both the column name and the actual data values within each row of that column. This enables comprehensive data analysis, trend identification, and insight generation.

[0092] (iii) Column Name and Synthetic Data for Rows: The Al agent can access the column name and a set of synthetically generated data points that mimic the statistical properties of the actual data. This is particularly useful for standard fields like ages, gender, or geographic locations, where synthetic data can provide meaningful insights without revealing sensitive individual information.

[0093] In an example, consider a retail media network (RMN) dataset with the following columns: "customer_id", "age", "gender", "purchase_amount", and "product_category".

[0094] 1. "customer d":Attorney Docket No. RAMP-00325-WO

[0095] Setting: Column Name Only

[0096] Use Case: The RMN wants to allow Al agents to understand the existence and number of customer identifiers for data management purposes but prevent any analysis that could link other attributes to specific individuals. This setting allows tracking the number of unique customers without revealing who they are.

[0097] Why it's useful: Maintaining customer privacy is paramount. This prevents the Al from revealing the association of customer IDs with their purchase behavior, demographics, etc.

[0098] 2. "age":

[0099] Setting: "Column Name and Synthetic Data for Rows"

[0100] Use Case: The RMN wants Al agents to analyze customer age distribution and its relationship to purchase behavior. Generating synthetic age data allows for cohort analysis (e.g., "customers aged 25-34") without disclosing the actual age of any specific individual.

[0101] Why it's useful: Protects individual privacy while enabling valuable demographic insights that can inform marketing strategies and product recommendations. Synthetic data preserves the overall statistical patterns of the real data.

[0102] 3. "purchase_amount":

[0103] Setting: "Column Name and Row"

[0104] Use Case: The RMN wants Al agents to perform detailed financial analysis, such as average purchase value, spending patterns over time, and correlation with other factors like product category or marketing campaigns.

[0105] Why it's useful: Provides full analytical capability for the Al to derive insights. Also, having the real 'purchase_amount' helps in accurately measuring the effectiveness of marketing campaigns and identifying high-value customer segments.

[0106] 4. "product_category":

[0107] Setting: 'Column Name and Row'

[0108] Use Case: The RMN wants Al agents to analyze product preferences, category trends, and cross-selling opportunities.

[0109] Why it's useful: Enables the Al to understand product popularity, which product categories are often purchased together, and how product choices correlate with other customer attributes.

[0110] By using insight and analytic Al agents 310, the RMN can fine-tune how Al agents interact with different data elements, balancing the need for analytical insights with theAttorney Docket No. RAMP-00325-WOimperative to protect sensitive information and maintain privacy. It also allows for various levels of data protection depending on the field, and thus the RMN has much more fine-grained control over their Al agents and what they do.

[0111] Activation Al agents 310, accessible when "Allow Al Agents" is enabled, governs how activation Al agents interact with specific fields / columns for audience segmentation and activation purposes. It follows the same pattern as just described.

[0112] It features conditional availability, i.e., it is available only if "Allow Al Agents" is enabled at the Data Source level. It further features field / column-specific enablement, and thus for Al- enabled fields / columns, the dropdown controls activation Al agent access. It further features activation Al agent access levels of column-name only (Al agent accesses only the column name for metadata and schema understanding); column name and row (Al agent accesses column name and data values for detailed segment creation and activation); and column name and synthetic data for rows (Al agent accesses column name and synthetic data for privacy-sensitive fields like age or gender). Examples may include RMN activation and segment creation use cases.

[0113] Collaborative RAG 314 becomes available if "Allow Collaborative RAG" is enabled if data source level 210 is enabled at the data source level. It provides granular control over how Al agents utilize specific fields / columns for collaborative Retrieval-Augmented Generation (RAG) development. This allows for tailored RAG systems while maintaining data privacy and security within the cleanroom.

[0114] It features conditional availability, i.e., it is only accessible when "Allow Collaborative RAG" has been activated for the corresponding data source. It further features field / column- specific enablement, by which, for each field / column where Al agent interaction is enabled, it provides a dropdown menu to control the level of access and capabilities for collaborative RAG Al agents. It further features collaborative RAG Al agent access levels, by which the dropdown menu offers three options. The first option is column name only. The Al agent can only access the name or identifier of the column. This allows for metadata analysis, schema understanding, and basic data discovery for RAG knowledge base construction without exposing the actual data values. The second option is column name and row. The Al agent can access both the column name and the actual data values within each row of that column. This enables the use of the full data for building the RAG knowledge base, enabling rich and contextually relevant retrieval. The third option is column name and synthetic data for rows. The Al agent can access the columnAttorney Docket No. RAMP-00325-WOname and a set of synthetically generated data points that mimic the statistical properties of the actual data. This is particularly useful for standard or sensitive fields like ages, gender, or locations, where synthetic data can provide useful contextual information for the RAG system without revealing individual details.

[0115] Examples include Commerce Media Network (CMN) or Retail Media Network (RMN) Use Cases. Consider a CMN / RMN dataset with the following columns: "product_id", "product_description", "price", "category", "customer_age".

[0116] 1. "product d":

[0117] Setting: Column Name Only

[0118] Use Case: To allow the RAG system to identify and track products within the knowledge base, but prevent it from linking product IDs to specific customer data during retrieval.

[0119] Why it's useful: Maintains product tracking while preventing direct association of products with specific individuals or transactions.

[0120] 2. "product_description":

[0121] Setting: Column Name and Row

[0122] Use Case: To enable the RAG system to use full product descriptions to retrieve relevant information based on customer queries. This allows the RAG to provide detailed and contextrich answers.

[0123] Why it's useful: The core of a product-related RAG system relies on detailed product descriptions for effective retrieval and response generation.

[0124] 3. "price":

[0125] Setting: Column Name and Row

[0126] Use Case: To allow the RAG system to provide pricing information in response to customer queries. This enhances the RAG's ability to provide complete and practical answers.

[0127] Why it's useful: Pricing is essential information for customers, and its inclusion in the RAG system improves the overall customer experience.

[0128] 4. "category":

[0129] Setting: Column Name and Row

[0130] Use Case: To enable the RAG system to categorize and filter products based on category, improving the relevance of search results.

[0131] Why it's useful: Category information is crucial for organizing the RAG's knowledge base and providing relevant results based on category-specific queries.Attorney Docket No. RAMP-00325-WO

[0132] 5. "customer_age":

[0133] Setting: Column Name and Synthetic Data for Rows

[0134] Use Case: To allow the RAG system to understand general age-related preferences and provide age-appropriate recommendations without exposing individual customer ages.

[0135] Why it's useful: Balances personalization with privacy, allowing the RAG to provide relevant recommendations based on age demographics without revealing personal information.

[0136] By utilizing collaborative RAG 314, the CMN / RMN can construct a collaborative RAG system that is both powerful and privacy-preserving. The different data access levels ensure that sensitive data remains protected while still allowing the RAG to leverage the necessary information for effective product discovery, recommendations, and customer support. This granular control is essential for building trust and ensuring compliance within a multi-party data collaboration environment.

[0137] Collaborative LLM 316 becomes available if "Allow Collaborative LLM" is enabled at the data source level. It provides granular control over how Al agents utilize specific fields / columns for collaborative Large Language Model (LLM) development and training. This allows for tailored LLM systems while maintaining data privacy and security within the cleanroom.

[0138] For each field / column where Al agent interaction is enabled , Collaborative LLM 316 provides a dropdown menu to control the level of access and capabilities for collaborative LLM Al agents. The dropdown menu offers three options. The first option is column name only. The Al agent can only access the name or identifier of the column. This allows for metadata analysis, schema understanding, and basic data discovery for LLM development without exposing the actual data values. This can help an LLM understand what different columns of data exist, without providing any access to the content within those columns. The second option is column name and row. The Al agent can access both the column name and the actual data values within each row of that column. This enables the use of the full data for building and training the LLM, enabling rich and contextually relevant knowledge infusion. The third option is column name and synthetic data for rows. The Al agent can access the column name and a set of synthetically generated data points that mimic the statistical properties of the actual data. This is particularly useful for standard or sensitive fields like age and gender, where synthetic data can provide useful contextual information for the LLM without revealing individual details.Attorney Docket No. RAMP-00325-WO

[0139] As an example, one may consider a commerce media network (CMN) or retail media network (RMN) use case. The data set may have the following columns: "product_id", "product_description", "price", "category", "customer_age", "review_text".

[0140] 1. "productjd":

[0141] Setting: Column Name Only

[0142] Use Case: To allow the LLM to identify and track products within its knowledge base but prevent it from linking product IDs to specific customer data during training. This setting helps to protect privacy and avoid memorization of particular IDs within the LLM weights.

[0143] Why it's useful: Maintains product tracking while preventing direct association of products with specific individuals or transactions in the LLM model.

[0144] 2. "product_description":

[0145] Setting: Column Name and Row

[0146] Use Case: To enable the LLM to use full product descriptions for training, allowing it to understand and generate detailed and context-rich responses. This allows the LLM to become very knowledgeable about the products in the catalog.

[0147] Why it's useful: Detailed product descriptions are essential for the LLM to learn about product features, benefits, and use cases, which is crucial for effective response generation.

[0148] 3. "price":

[0149] Setting: Column Name and Row

[0150] Use Case: To allow the LLM to understand and generate responses related to pricing, such as price comparisons, discounts, and price trends. The LLM can learn general concepts of pricing and the relationship between products and their prices.

[0151] Why it's useful: Pricing is a key factor in customer decision-making, and its inclusion in LLM training enhances the model's ability to provide relevant and practical answers.

[0152] 4. "category":

[0153] Setting: Column Name and Row

[0154] Use Case: To enable the LLM to understand product categorization and generate responses related to category-specific queries and recommendations. The LLM can become an expert in each product category.

[0155] Why it's useful: Category information is crucial for organizing the LLM's knowledge base and providing relevant results based on category-specific queries and user intents.

[0156] 5. "customer_age":Attorney Docket No. RAMP-00325-WO

[0157] Setting: Column Name and Synthetic Data for Rows

[0158] Use Case: To allow the LLM to understand general age-related preferences and tailor responses accordingly, without exposing individual customer ages. The LLM can learn about general trends across different age groups without learning anything about particular individuals.

[0159] Why it's useful: Balances personalization with privacy, allowing the LLM to provide relevant recommendations and responses based on age demographics without revealing personal information.

[0160] 6. "review_text":

[0161] Setting: Column Name and Row

[0162] Use Case: To enable the LLM to learn from customer reviews, understanding sentiment, identifying popular features, and generating helpful responses to customer inquiries. This can help the LLM become very good at customer service tasks.

[0163] Why it's useful: Customer reviews provide valuable insights into product strengths and weaknesses, and training the LLM on this data enhances its ability to provide insightful and relevant information.

[0164] By utilizing collaborative LLM 316, the CMN / RMN can construct a powerful and specialized LLM in a privacy-preserving manner. The different data access levels ensure that sensitive data remains protected while still allowing the LLM to leverage the necessary information for effective product discovery, recommendations, and customer support. This granular control is essential for building trust and ensuring compliance within a multi-party data collaboration environment, especially where the risk of data leakage from LLM memorization is a significant concern.

[0165] Fig. 8 illustrates the Cleanroom Al Agentic Workflow Logging and Monitoring Ul 400.This Ul ensures transparency and accountability of Al agent activities within the data cleanroom. Given the automated nature of Al agent queries, robust tracking mechanisms are essential for maintaining transparency and enabling human oversight. The following details the levels of monitoring and logging, which provide various degrees of visibility into Al agent activities, ensuring compliance and auditability within the data collaboration framework.

[0166] Data operators and administrators will have access to a general "Al Workflow Activity" section in which all the Al workflow activities will be displayed. The granularity of the activity tracking will be controlled for each data source, and this type of control will also be available forAttorney Docket No. RAMP-00325-WOAl activation agent, and collaborative RAG, LLM activities (not currently integrated in previous Ul screens).

[0167] At collaborative cleanroom column 402 in Fig. 8, for each Al workflow activity, the lineage back to the originating data collaborative cleanroom is tracked and displayed. This feature allows data operators and administrators to quickly identify which cleanroom generated specific Al activity, providing crucial context for auditing, troubleshooting, and understanding workflow dependencies. This ensures clear accountability and facilitates the management of complex multi-party collaborations.

[0168] At dataset(s) column 402, for each Al workflow activity, the specific datasets and tables accessed and utilized during the Al workflow are logged and linked to the activity. This detailed tracking allows data operators and administrators to identify precisely which data resources were involved in any given Al operation. This facilitates in-depth audits, impact analysis of data changes, and the ability to trace the flow of data throughout the system. For instance, if an Al agent generates a report based on specific sales data and customer demographics, this feature will log the exact sales table and customer demographics table used. Administrators can use this information to verify data usage, troubleshoot discrepancies, and ensure that Al agents are adhering to data access policies.

[0169] At query column(s) 406, for each Al workflow activity, the specific columns or fields accessed and utilized within the datasets are logged and linked to the activity. This granular tracking allows data operators and administrators to identify precisely which data elements were involved in any given Al operation. This detailed logging is crucial for upholding the fieldlevel permissions and Al collaboration rules defined during the dataset mapping phase.

[0170] For example, one may consider a scenario where a data operator has configured a dataset to allow Al agents to access only specific columns like "product_name" and "category," while restricting access to sensitive columns like "customer_id." If an Al agent generates a report, query column(s) 406 will log which columns were actually accessed during that activity. If the log shows that the Al agent attempted to access "customer id," it would immediately flag a potential policy violation. This allows administrators to quickly identify and address any unauthorized data access, ensuring that the meticulously defined dataset mapping field rules are strictly preserved and enforced throughout all Al workflows.

[0171] At activity detail 406, when a specific Al workflow activity is selected, a detailed side panel is displayed. This panel provides additional options and information related to thatAttorney Docket No. RAMP-00325-WQactivity. Depending on the type of Al activity, the panel may include "view results," which is an option to view the output or results generated by the Al agent's activity. This could include reports, analyses, prompt exchanges or any data transformations performed. The panel may also include a "view SQL" tab (top section next to "detail"). For Al activities involving dataset queries, the side panel provides an option to view the SQL query that was automatically generated by the Al agent. This allows data operators and administrators to understand the exact queries executed against the data, ensuring transparency and enabling troubleshooting.

[0172] This detailed side panel enhances transparency, audibility, and control over Al agent activities, enabling users to inspect the Al's work and verify its compliance with data policies. For example, if an Al agent generated a report on sales trends, the "View Results" option would display the report. Additionally, the "View Generated SQL" option would show the specific query used if the agent queried a database to gather sales data.

[0173] At Al workflow 412, the type of originating Al workflow is tracked and displayed for each Al workflow activity. This includes identifying whether an analytic Al agent, an activation Al agent, a collaborative RAG system, or a collaborative LLM initiated the activity. This categorization provides context for the type of Al processing involved and allows for targeted analysis and auditing of specific Al workflows. For instance, administrators can filter or sort activities based on the workflow type to analyze the effectiveness of activation campaigns versus analytical tasks or to review the queries made by the collaborative LLM. This level of detail enhances understanding of Al agent behavior and ensures that each type of Al activity performs as expected and within policy guidelines.

[0174] At view query results 410, the detailed side panel provides a contextual "view result" button. This button's label dynamically adapts based on the specific type of Al workflow activity. "View query result," for activities involving database queries, this button displays the results of the SQL query executed by the Al agent. "View prompt result," for activities involving collaborative LLMs or RAG systems, this button shows the response generated by the model based on the given prompt. "View recommendation result" presents the generated recommendations or activation strategies for activities involving activation Al agents or recommendation engines.

[0175] This contextual button provides quick access to the specific output of the Al activity, enhancing usability and facilitating efficient review and verification of the Al's work. For example, the "View Query Result" button displays the sales data summary after an Al agentAttorney Docket No. RAMP-00325-WQanalyzes sales data and generates a trend report. If a collaborative LLM generates personalized product recommendations, the "View Prompt Result" button will show the list of recommended products.

[0176] This detailed logging, combined with the dynamic "View Result" buttons and contextual side panels, empowers users to verify Al actions, troubleshoot issues, and enforce data policies effectively. The different levels of tracking (No Tracking, Query Statement-Level Tracking, and Query Statement and Result Tracking) offer flexibility, allowing organizations to tailor their monitoring practices to the sensitivity and criticality of the data and Al workflows. This robust monitoring framework is essential for fostering trust, ensuring compliance, and unlocking the full potential of Al-driven data collaboration within the secure confines of the data cleanroom.

[0177] Several scenarios may be used to describe certain embodiments of the invention in greater detail. One category is general ad tech experiences, extending current existing use cases. One scenario in this category is enhanced audience targeting and measurement. As an example in this scenario, a large retail chain wants to enable its brand partners to target specific customer segments more effectively on their RMN. The retailer has rich first-party data (purchase history, browsing behavior, loyalty program data), while the brands have their own customer data. The system allows the retailer and brand partners onboard their data into the secure environment, data cleanroom. Using an identifier system such as RampID®, customer data is encoded within each party's identity space, ensuring privacy. Al agents within the system can analyze the combined datasets to identify granular audience segments (e.g., "loyal customers who frequently buy organic groceries" or "new customers who purchased a specific brand of running shoes"). The brands can then target these segments with personalized ads on the RMN, without either party directly seeing the other's raw customer data. The system can also measure the effectiveness of the ad campaigns by analyzing conversion rates and other metrics, all within the secure environment.

[0178] Another scenario in this category is personalized product recommendations. For example, perhaps an e-commerce platform wants to improve its product recommendation engine by leveraging data from multiple sources, including its own customer data, partner data (e.g., from suppliers), and contextual data (e.g., trending products, social media buzz). The platform and its partners contribute their data to the system. Al agents within the system can analyze the combined datasets to generate highly personalized product recommendations for individual users. For example, an agent might identify that a user who recently purchased aAttorney Docket No. RAMP-00325-WOspecific coffee maker also frequently views related accessories like coffee grinders and filters. The platform can then recommend these accessories to the user, increasing the likelihood of a cross-sell. Because of the polymorphic identity management, a customer's Pll is never directly exposed, ensuring privacy and trust.

[0179] Another scenario within this category is optimizing advertising campaigns across multiple channels. One may suppose that a brand wants to maximize its advertising campaigns across multiple channels (e.g., the retailer's website, mobile app, in-store displays, and social media) based on real-time data. Data from all channels is ingested into the system. Al agents can analyze this data to identify which channels are performing best for different customer segments and adjust the campaigns accordingly. For instance, the agents might discover that younger customers are more responsive to social media ads, while older customers are more likely to engage with in-store displays. The brand can then allocate its advertising budget more effectively across these channels. The dynamic governance engine ensures that all data usage is compliant with privacy regulations and brand guidelines across the various channels.

[0180] Another scenario within this category is collaborative development of new advertising products. For example, a technology company may wish to collaborate with a retailer to develop new advertising products and features. The two companies can use the system to share data and insights securely, enabling them to identify opportunities for innovation. For example, they might use the system to analyze customer behavior and identify pain points in the current advertising experience. They can then use this information to design and test new products that address those pain points. The secure environment fosters trust and allows for more open collaboration, accelerating development.

[0181] In all these examples, the system's key features (secure multi-party collaboration, Al- driven governance, and polymorphic identity management) are crucial for enabling effective and privacy-preserving advertising and AdTech innovation within RMNs (retail Media Networks) and CMNs (Commerce Media Networks). It allows retailers, brands, and technology providers to work together to create better experiences for consumers, while still protecting sensitive data.

[0182] A second overall category is new Al-enabled use cases around RAG / LLM development with this system. One scenario then would be collaborative RAG development for enhanced product discovery and customer support. One may suppose that a large RMN wants to enhance its customer experience by building a sophisticated RAG system. This system will power several key functions. It can provide improved product discovery by allowing customers to ask complex,Attorney Docket No. RAMP-00325-WOnatural language questions about products, compare features, and find items that precisely match their needs. The RAG can provide contextually relevant personalized recommendations based on individual customer history and preferences. The system can provide enhanced customer support by quickly answering customer inquiries, drawing on a vast knowledge base of product information, FAQs, and troubleshooting guides.

[0183] To build this special RAG as just described, the RMN owner collaborates with several partners. Brand partners provide detailed product specifications, marketing materials, and other information. Data providers / aggregators offer aggregated customer behavior data, trend analysis, and market insights. Technology Providers / AI Research Firms contribute Al expertise, natural language processing (NLP) models, and the infrastructure for the RAG system. This systems allows for secure knowledge base creation. First, each collaborator onboards their respective data into the system. Then brand partners upload product catalogs, manuals, and marketing copy. Data providers contribute to customer segmentation and trend data. All this information becomes part of the RAG system's knowledge base.

[0184] In this scenario as just described, polymorphic data identity provides context. An identity resolution system such as RampID® ensures that sensitive customer data used for personalization remains within its identity space. The RAG system can access and use this data to provide personalized responses without exposing the underlying PH to other collaborators or the system itself. The technology partners use the system's secure environment to build and train the RAG models collaboratively. They can test different approaches, evaluate performance, and iterate on the system in a way that respects each collaborator's privacy and data policies. The RMN uses the Al-driven governance engine to define fine-grained rules about how the RAG system can access and use the data. For instance, they can allow the RAG to use aggregated trend data to generate reports but prevent it from revealing individual customer purchase histories. The system enables ongoing updates to the RAG's knowledge base and model. As brand partners introduce new products or data providers offer fresh insights, the system can be dynamically updated without requiring a complete overhaul. The Al-driven governance ensures that all these updates adhere to existing data policies. The resulting RAG system provides a vastly improved customer experience, with more accurate search results, highly personalized recommendations, and faster, more effective customer support. This, in turn, drives sales, increases customer satisfaction, and strengthens the RMN's competitive position.Attorney Docket No. RAMP-00325-WO

[0185] In this example for the scenario just described, the system facilitates constructing a powerful RAG system that would otherwise be difficult or impossible to build due to data privacy concerns and the complexities of multi-party collaboration. The system allows the RMN and its partners to combine their unique expertise and data assets securely and controlled, resulting in a superior customer experience.

[0186] A second scenario in this category is collaborative LLM development for next-generation retail experiences. For example, one may suppose that a leading RMN wants to develop a highly specialized LLM finely tuned for the retail domain. This LLM will power advanced conversational commerce by enabling natural, intuitive conversations between customers and the RMN's platform for product search, recommendations, and support. It can generate compelling product descriptions, ad copy, and marketing content. It can also tailor product suggestions and interactions based on a deep understanding of individual customer preferences and behavior.

[0187] To build this specialized LLM, the RMN owner collaborates with several partners. Brand partners provide product information, customer insights, and marketing materials. Data Providers / Aggregators offer aggregated customer trend data and market analysis. Technology Providers / AI Research Firms contribute expertise in LLM development, training, and optimization.

[0188] The system in this scenario allows for secure data ingestion for LLM training in that the RMN and its partners onboard vast amounts of retail-specific data into the system. This includes product catalogs, customer reviews, purchase histories, website content, marketing materials, and customer support logs. All this data becomes training data for the LLM.

[0189] The system in this scenario allows for polymorphic data identity for privacy-preserving training, since the identity resolution system ensures that sensitive customer data used in LLM training remains within its own identity space. The LLM can learn patterns and relationships within this data without exposing the underlying Pll to other collaborators or even the Al research firms.

[0190] The system in this scenario allows for collaborative LLM development and training. The Al research firms use the system's secure environment to develop and train the LLM collaboratively. They can access and process the combined training data while respecting the privacy boundaries enforced by polymorphic identity management. They can experiment with different model architectures and training techniques, evaluate performance, and iterate on the LLM, all within a safe environment.Attorney Docket No. RAMP-00325-WO

[0191] The system in this scenario allows for Al-Driven governance of data usage in LLM training. The RMN uses the Al-driven governance engine to set detailed rules for how the training data can be used. For example, they might allow the Al researchers to use aggregated sales data to improve the LLM's understanding of product trends but strictly forbid using individual customer purchase records for this purpose.

[0192] The system in this scenario allows for dynamic updates and fine-tuning. The system enables continuous updating and fine-tuning of the LLM. As new data becomes available (e.g., from new product launches or evolving customer trends), it can be securely added to the training data, and the LLM can be updated without disrupting the RMN's operations. The Al- driven governance ensures that all updates adhere to data privacy policies.

[0193] This system in this scenario allows for deployment and real-world application. Once trained and fine-tuned, the LLM can be deployed within the RMN's platform to power the advanced conversational commerce, automated content creation, and personalized shopping experiences as intended.

[0194] In this scenario, the system is useful for building a highly specialized LLM in a privacypreserving and collaborative manner. It allows the RMN to leverage the expertise and data of its partners while maintaining strict control over sensitive customer information. This results in a powerful Al asset that enhances the RMN's capabilities and provides a superior customer experience.

[0195] Finally, a particular application can be described as an embodiment of the invention that facilitates enhanced audience targeting and measurement in a retail media network (RMN). Agentic systems can effectively enhance audience targeting and measurement within an RMN. In this scenario, a large retail chain seeks to empower its brand partners to target specific customer segments more precisely. The retailer possesses rich first-party data, including purchase history, browsing behavior, and loyalty program information, while the brand partners have their own valuable customer datasets. This application focuses on how the system facilitates secure and privacy-preserving collaboration between the retailer and its partners to achieve superior audience targeting and measurement outcomes. The brand end-user can collaborate with an Al agent that has both access to the retailer and brand data. One can imagine a query such as, "Give me segment suggestions about my buyers." An Al Agent can generate general insights about brand buyers by accessing point of sale or transaction data from the retailer, such as shown in Fig. 9.Attorney Docket No. RAMP-00325-WO

[0196] The brand end-user then continues collaborating with the Al agent to bring additional data points from a data seller to this multi-party cleanroom. Because the Al agent has access to the retailer, brand data, and further data seller (i.e., a data marketplace seller), brands are capable of expanding their reach across multiple ad tech platforms. The user could issue a query such as, "I like these! Can you extend my reach with the data marketplace." The brand end-user can now directly activate ad tech platforms (i.e., X, Meta, TradeDesk), with unique Audience / Segmentation (High-Spenders Expanded, Lapsed Purchaser Expended) by clicking the "Send for Activation button," with a display as shown in Fig. 10.

[0197] Fig. 11 is a block diagram illustrating an example computer hardware system, according to various embodiments, that can be used as part of a multi-component system to implement the invention as has been described in a cloud-computing environment. Computer system 440 may implement a hardware portion of a cloud computing system as forming parts of the various implementations of the present invention.

[0198] Computer system 440 may be any of various types of hardware devices, including, but not limited to, a commodity server, personal computer system, desktop computer, laptop or notebook computer, mainframe computer system, handheld computer, workstation, network computer, a consumer device, application server, physical storage device, telephone, mobile telephone, or in general any type of computing node, compute node, compute device, and / or hardware computing device.

[0199] Computer system 440 includes one or more hardware processors 441a, 441b...441n (any of which may include multiple processing cores, which may be single or multi-threaded) coupled to a physical system memory 442 via an input / output (I / O) interface 444. Computer system 440 further may include a network interface 446 coupled to I / O interface 444.

[0200] In various embodiments, computer system 440 may be a single processor system including one hardware processor 441a, or a multiprocessor system including multiple hardware processors 441a, 441b...441n. Processors 441a, etc. may be any suitable processors capable of executing computing instructions. For example, in various embodiments, processors 441a, etc. may be general-purpose or embedded processors implementing any of a variety of instruction set architectures.

[0201] In multiprocessor systems, each of processors 441a, etc. may commonly, but not necessarily, implement the same instruction set. The computer system 440 also includes one or more hardware network communication devices (e.g., network interface 446) forAttorney Docket No. RAMP-00325-WOcommunicating with other systems and / or components over a communications network, such as a local area network, wide area network, or the Internet. For example, a client application executing on system 440 may use network interface 446 to communicate with a server application executing on a single hardware server or on a cluster of hardware servers that implement one or more of the components of the systems described herein in a cloud computing environment as implemented in various sub-systems. In another example, an instance of a server application executing on computer system 440 may use network interface 446 to communicate with other instances of an application that may be implemented on other computer systems.

[0202] In the illustrated embodiment, computer system 440 also includes one or more physical persistent storage devices 448 and / or one or more I / O devices 450. In various embodiments, persistent storage devices 448 may correspond to disk drives, tape drives, solid-state memory or drives, other mass storage devices, or any other persistent storage devices. Computer system 440 (or a distributed application or operating system operating thereon) may store instructions and / or data in persistent storage devices 448, as desired, and may retrieve the stored instructions and / or data as needed. For example, in some embodiments, computer system 440 may implement one or more nodes of a control plane or control system, and persistent storage 448 may include the solid-state drives (SSDs) attached to that server node. Multiple computer systems 440 may share the same persistent storage devices 448 or may share a pool of persistent storage devices, with the devices in the pool representing the same or different storage technologies, including such technologies as described above.

[0203] Computer system 440 includes one or more physical system memories 442 that may store code / instructions 443 and data 445 accessible by processor(s) 441a, etc. The system memories 442 may include multiple levels of memory and memory caches in a system designed to swap information in memories based on access speed, for example.

[0204] The interleaving and swapping may extend to persistent storage devices 448 in a virtual memory implementation, where memory space is mapped onto the persistent storage devices 448. The technologies used to implement the system memories 442 may include, by way of example, static random-access memory (RAM), dynamic RAM, read-only memory (ROM), nonvolatile memory, solid-state memory, or flash-type memory.

[0205] As with persistent storage devices 448, multiple computer systems 440 may share the same system memory systems 442 or may share a pool of system memories 442. SystemAttorney Docket No. RAMP-00325-WOmemory or memory systems 442 may contain program instructions 443 that are executable by processor(s) 441a, etc. to implement the routines described herein.

[0206] In various embodiments, program instructions 443 may be encoded in binary, Assembly language, any interpreted language such as Java, compiled languages such as C / C++, or in any combination thereof; the particular languages given here are only examples. In some embodiments, program instructions 443 may implement multiple separate clients, server nodes, and / or other components.

[0207] In some implementations, program instructions 443 may include instructions executable to implement an operating system (not shown), which may be any of various operating systems, such as UNIX, LINUX, Solaris™, MacOS™, or Microsoft Windows™. Any or all of program instructions 443 may be provided as a computer program product, or software, that may include a non-transitory computer-readable storage medium having stored thereon instructions, which may be used to program a computer system (or other electronic devices) to perform a process according to various implementations.

[0208] A non-transitory computer-readable storage medium may include any mechanism for storing information in a form (e.g., software or processing application) readable by a machine (e.g., a physical computer). Generally speaking, a non-transitory computer-accessible medium may include computer-readable storage media or memory media such as magnetic or optical media, e.g., disk or DVD / CD-ROM, coupled to or in communication with computer system 440 via I / O interface 444.

[0209] A non-transitory computer-readable storage medium may also include any volatile or non-volatile media such as RAM or ROM that may be included in some embodiments of computer system 440 as system memory 442 or another type of memory. In other implementations, program instructions may be communicated using optical, acoustical or other form of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.) conveyed via a communication medium such as a network and / or a wired or wireless link, such as may be implemented via network interface 446. Network interface 446 may be used to interface with other devices 442, which may include other computer systems or any type of external electronic device.

[0210] In some embodiments, system memory 442 may include data store 445, as described herein. In general, system memory 442 and persistent storage 448 may be accessible on other devices 452 through a network and may store data blocks, replicas of data blocks, metadataAttorney Docket No. RAMP-00325-WOassociated with data blocks, and / or their state, database configuration information, and / or any other information usable in implementing the routines described herein.

[0211] In one embodiment, I / O interface 444 may coordinate I / O traffic between processors 441a, etc., system memory 442, and any peripheral devices in the system, including through network interface 446 or other peripheral interfaces. In some embodiments, I / O interface 444 may perform any necessary protocol, timing or other data transformations to convert data signals from one component (e.g., system memory 442) into a format suitable for use by another component (e.g., processors 441a, etc.).

[0212] In some embodiments, I / O interface 444 may include support for devices attached through various types of peripheral buses, such as a variant of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard, as examples. Also, in some embodiments, some or all of the functionality of I / O interface 444, such as an interface to system memory 442, may be incorporated directly into processor(s) 441a, etc.

[0213] Network interface 446 may allow data to be exchanged between computer system 440 and other devices attached to a network, such as other computer systems (which may implement one or more storage system server nodes, primary nodes, read-only node nodes, and / or clients of the database systems described herein), for example. In addition, I / O interface 444 may allow communication between computer system 440 and various I / O devices 450 and / or remote storage 448.

[0214] Input / output devices 450 may, in some embodiments, include one or more display terminals, keyboards, keypads, touchpads, scanning devices, voice or optical recognition devices, or any other devices suitable for entering or retrieving data by one or more computer systems 440. These may connect directly to a particular computer system 440 or generally connect to multiple computer systems 440 in a cloud computing environment, grid computing environment, or other system involving multiple computer systems 440.

[0215] Multiple input / output devices 450 may be present in communication with computer system 440 or may be distributed on various nodes of a distributed system that includes computer system 440. In some embodiments, similar input / output devices may be separate from computer system 440 and may interact with one or more nodes of a distributed system that includes computer system 440 through a wired or wireless connection, such as over network interface 446.Attorney Docket No. RAMP-00325-WO

[0216] Network interface 446 may commonly support one or more wireless networking protocols (e.g., Wi-Fi / I EEE 802.11, or another wireless networking standard). Network interface 446 may support communication via any suitable wired or wireless general data networks, such as other types of Ethernet networks, for example. Additionally, network interface 446 may support communication via telecommunications / telephony networks such as analog voice networks or digital fiber communications networks, via storage area networks or via any other suitable type of network and / or protocol. In various embodiments, computer system 440 may include more, fewer, or different components (e.g., displays, video cards, audio cards, peripheral devices, or an Ethernet interface).

[0217] Any of the distributed system embodiments described herein, or any of their components, may be implemented as one or more network-based services in the cloud computing environment. For example, a read-write node and / or read-only nodes within the database tier of a hardware database system may present database services and / or other types of physical data storage services that employ the distributed storage systems described herein to clients as network-based services.

[0218] In some embodiments, a network-based service may be implemented by a software and / or hardware system designed to support interoperable machine-to-machine interaction over a network. A web service may have an interface described in a machine-processable format. Other systems may interact with the network-based service in a manner prescribed by the description of the network-based service's interface. For example, the network-based service may define various operations that other systems may invoke, and may define a particular application programming interface (API) to which other systems may be expected to conform when requesting the various operations.

[0219] In various embodiments, a network-based service may be requested or invoked through the use of a message that includes parameters and / or data associated with the network-based services request. Such a message may be formatted according to a particular markup language such as Extensible Markup Language (XML), and / or may be encapsulated using a protocol. To perform a network-based services request, a network-based services client may assemble a message including the request and convey the message to an addressable endpoint (e.g., a Uniform Resource Locator (URL)) corresponding to the web service, using an Internet-based application layer transfer protocol such as Hypertext Transfer Protocol (HTTP).Attorney Docket No. RAMP-00325-WO

[0220] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0221] Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, a limited number of the exemplary methods and materials are described herein. It will be apparent to those skilled in the art that many more modifications are possible without departing from the inventive concepts herein.

[0222] All terms used herein should be interpreted in the broadest possible manner consistent with the context.

[0223] When a grouping is used herein, all individual members of the group and all combinations and sub-combinations possible of the group are intended to be individually included.

[0224] When a range is stated herein, the range is intended to include all sub-ranges within the range, as well as all individual points within the range.

[0225] When "about," "approximately," or like terms are used herein, they are intended to include amounts, measurements, or the like that do not depart significantly from the expressly stated amount, measurement, or the like, such that the stated purpose of the apparatus or process is not lost.

[0226] All references cited herein are hereby incorporated by reference to the extent that there is no inconsistency with the disclosure of this specification.

[0227] The present invention has been described with reference to certain preferred and alternative embodiments that are intended to be exemplary only and not limiting to the full scope of the present invention, as set forth in the appended claims.

Claims

Attorney Docket No. RAMP-00325-WOClaims1. A system for secure multi-party data collaboration with Al-driven governance, comprising:a secure data cleanroom environment configured to receive and store data from a plurality of collaborating parties;a polymorphic data identity management framework in communication with the secure data cleanroom environment, wherein the polymorphic data identity management framework is configured to encode data from each collaborating party within a respective unique identity space, such that data from different collaborating parties cannot be directly correlated or linked, and such that direct access to personally identifiable information (PH) or confidential business data between parties is prevented; an agentic data connection and sources configuration component configured to manage connections to data sources feeding the secure data cleanroom environment and to define, for each data source, a first level of Al collaboration rules specifying permitted Al agent interactions with data originating from that data source;an agentic dataset cleanroom configuration component configured to define, at a per-column or perfield granularity within datasets in the secure data cleanroom environment, a second level of access control policies governing Al agent access to individual data elements, wherein the second level of access control policies is subordinate to the first level of Al collaboration rules;an Al-driven governance engine configured to enforce the first level of Al collaboration rules and the second level of access control policies for Al agents accessing data within the secure data cleanroom environment, and to dynamically adapt at least one of the first level of Al collaboration rules or the second level of access control policies based on at least one of real-time conditions, user roles, or data sensitivity; anda cleanroom agentic logging and monitoring component configured to record Al agent activities within the secure data cleanroom environment, including at least identifiers of data sources, datasets, and columns accessed by each Al agent activity.

2. The system of claim 1, wherein the agentic data connection and sources configuration component is further configured to allow a data operator to specify, for each data source, one or more of: allowed large language model (LLM) models, permitted agentic workflow types, Al agent query permissions, query tracking levels, and activation assessment rules.Attorney Docket No. RAMP-00325-WO3. The system of claim 1, wherein the agentic dataset cleanroom configuration component is configured to provide, for each column or field enabled for Al agent access, a selectable access level comprising one of: column name only, column name and row data, or column name and synthetic data for rows, wherein synthetic data mimics statistical properties of actual data without revealing individual data values.

4. The system of claim 3, wherein the selectable access level is independently configurable for each of a plurality of Al workflow types including insight and analytic Al agents, activation Al agents, collaborative Retrieval-Augmented Generation (RAG) agents, and collaborative Large Language Model (LLM) training agents.

5. The system of claim 1, wherein the agentic dataset cleanroom configuration component is further configured to accept, for each column or field enabled for Al agent access, a column description and Al agent instructions that guide Al agent interpretation and processing of the data within that column or field.

6. The system of claim 1, wherein the cleanroom agentic logging and monitoring component is configured to support a plurality of tracking levels comprising: no tracking, query statement-level tracking that logs query statements generated by Al agents, and query statement and result tracking that logs both query statements and corresponding results generated by Al agents.

7. The system of claim 1, wherein the cleanroom agentic logging and monitoring component is further configured to provide a contextual detail view for each logged Al agent activity, the contextual detail view dynamically adapted based on a type of Al workflow that initiated the activity, the contextual detail view including at least one of a view query result option, a view prompt result option, or a view recommendation result option.

8. The system of claim 1, wherein the polymorphic data identity management framework utilizes an external identifier service to perform identity resolution, and wherein only the external identifier service is capable of performing transcoding between the unique identity spaces of different collaborating parties for authorized data sharing or analysis.

9. A computer-implemented method for governing Al agent access to data in a secure multi-party data collaboration environment, comprising:receiving data from a plurality of collaborating parties into a secure data cleanroom environment;Attorney Docket No. RAMP-00325-WOencoding the data from each collaborating party within a respective unique identity space using a polymorphic data identity management framework, such that data from different collaborating parties cannot be directly correlated or linked and such that direct access to personally identifiable information (PH) or confidential business data between parties is prevented;configuring, for each data source connected to the secure data cleanroom environment, a first level of Al collaboration rules defining permitted Al agent interactions with data originating from that data source, including at least one of permitted Al workflow types, allowed Al models, or query tracking levels;configuring, at a per-column or per-field granularity within datasets derived from the data sources, a second level of access control policies governing Al agent access to individual data elements, the second level being hierarchically subordinate to the first level such that a column or field is available for Al agent access only if the corresponding data source has a permitting first-level rule;enforcing the first level of Al collaboration rules and the second level of access control policies via an Al-driven governance engine when Al agents access or process data within the secure data cleanroom environment; andlogging Al agent activities within the secure data cleanroom environment, including recording identifiers of accessed data sources, datasets, and columns for each Al agent activity.

10. The method of claim 9, wherein configuring the second level of access control policies comprises, for each column or field enabled for Al agent access, selecting an access level from a group consisting of: column name only, column name and row data, and column name and synthetic data for rows.

11. The method of claim 10, wherein selecting the access level is performed independently for each of a plurality of Al workflow types associated with the column or field, such that a first Al workflow type is granted a different access level than a second Al workflow type for the same column or field.

12. The method of claim 9, further comprising dynamically adapting at least one of the first level of Al collaboration rules or the second level of access control policies based on at least one of a change in real-time conditions, a change in user roles, or a change in data sensitivity classification.

13. The method of claim 9, further comprising configuring, for each data source, activation assessment rules that control whether Al agents can assess data activation status for external advertising technologyAttorney Docket No. RAMP-00325-WOplatforms, including specifying permitted destinations via a whitelist, a blacklist, or an all-destinations selection.

14. The method of claim 9, wherein the logging comprises supporting configurable tracking levels on a per-data-source basis, the tracking levels including at least: a query statement-level tracking mode that captures Al-agent-generated query statements, and a query statement and result tracking mode that captures both query statements and their corresponding results.

15. The method of claim 9, further comprising providing, for each column or field enabled for Al agent access, descriptive metadata and Al agent instructions that constrain how Al agents interpret and process data in that column or field.

16. A computer system for managing secure Al-driven multi-party data collaboration, comprising one or more processors and memory storing instructions that, when executed by the one or more processors, cause the computer system to provide:a first user interface configured to enable configuration of data source connections to a secure data cleanroom environment and definition of Al collaboration rules for each data source, the Al collaboration rules specifying at least permitted large language model (LLM) models, permitted agentic workflow types, and query tracking levels;a second user interface configured to enable configuration of per-column or per-field access control policies for datasets within the secure data cleanroom environment, including selection of an Al agent access level for each enabled column or field from among column name only, column name and row data, and column name and synthetic data for rows;a third user interface configured to enable management of Al agents and agentic workflows operating within the secure data cleanroom environment; anda fourth user interface configured to display logs of Al agent activities within the secure data cleanroom environment, the logs including at least identifiers of originating cleanrooms, accessed datasets, accessed columns, and Al workflow types for each logged activity.

17. The computer system of claim 16, wherein the second user interface is further configured to present, for each enabled column or field, independently configurable access level selections for each of a plurality of Al workflow categories including insight and analytic Al agents, activation Al agents, collaborative RAG agents, and collaborative LLM training agents.Attorney Docket No. RAMP-00325-WO18. The computer system of claim 16, wherein the fourth user interface is further configured to provide a contextual detail panel for a selected Al agent activity, the detail panel dynamically presenting a result view option labeled according to a type of Al workflow that initiated the selected activity.

19. A computer-implemented method for collaborative Al model development within a secure data cleanroom, comprising:onboarding data from a plurality of collaborating parties into a secure data cleanroom environment; encoding the onboarded data using a polymorphic data identity management framework to create a respective unique identity space for each collaborating party's data, thereby preventing direct correlation or linkage of personally identifiable information (PI I) or confidential business data across parties;configuring, via a hierarchical policy structure comprising a data-source-level set of Al collaboration rules and a column-level or field-level set of access control policies subordinate thereto, permitted interactions between Al agents and the onboarded data;collaboratively training or building an Al model, selected from one or more of a Retrieval-Augmented Generation (RAG) system and a Large Language Model (LLM), on combined datasets within the secure data cleanroom environment in accordance with the hierarchical policy structure; andlogging activities of Al agents involved in the collaborative training or building, including identifiers of accessed data sources, datasets, and columns.

20. The method of claim 19, wherein the column-level or field-level set of access control policies specifies, for at least one column or field, an access level of column name and synthetic data for rows, wherein the synthetic data preserves statistical properties of actual data while preventing exposure of individual data values during Al model training.