Edge Cloud Intermediary for User-Specific LLM Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional large language models (LLMs) face challenges in providing user-specific context due to privacy, security, and resource constraints, limiting their effectiveness in user-specific applications, especially in embedded devices with limited resources.

Innovation Solution

The proposed solution involves a system architecture that captures instantaneous user context using edge devices and aggregates it through an intermediary cloud service, enabling user-specific embedding vectors and access control to provide personalized data to LLMs without the need for continuous session persistence, thereby optimizing resource allocation and privacy management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If user-specific data is provided to cloud-based LLMs, then personalized responses are improved, but privacy and security risks increase

Engineering Contradiction:
Improveuser-specific customizationVSAvoidprivacy and security risks
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces an intermediary cloud service layer between edge devices and LLMs. This intermediary aggregates user context from multiple edge devices, manages access control, and provides user-specific embedding vectors to LLMs without exposing raw user data. The intermediary acts as a mediator that enables personalization while maintaining privacy through centralized management and selective data sharing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts user-specific information from raw edge device data by generating embedding vectors that capture user context without storing actual user data. The system extracts only the necessary contextual features (user preferences, behavior patterns) and transforms them into compressed embedding representations that can be shared with LLMs while minimizing privacy exposure.

Inventive Principle:
Principle #2Taking out (Extraction)

2Adaptability or versatility

If continuous session persistence is maintained for user context, then user-specific responses are improved, but resource consumption increases

Engineering Contradiction:
Improveuser context availabilityVSAvoidresource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary action by aggregating and processing user context from edge devices before it is needed for LLM inference. The cloud service proactively collects user data, generates embedding vectors, and stores them in advance. This eliminates the need for continuous real-time processing during user interactions, reducing computational resource consumption while maintaining user context availability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a copy of user context in the form of embedding vectors that can be efficiently stored and retrieved. Instead of maintaining continuous live sessions with full user data, the system creates compressed representations (embeddings) that capture essential user characteristics. These embeddings can be quickly loaded into LLM contexts without requiring continuous data transmission or processing.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If user context is aggregated from multiple edge devices, then user-specific intelligence is improved, but system complexity increases

Engineering Contradiction:
Improveuser-specific intelligenceVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent designs the cloud service as a universal platform that handles multiple functions: aggregating user context from diverse edge devices (smartphones, tablets, wearables), processing different data types (text, audio, visual), generating embedding vectors, managing access control, and providing data to LLMs. This multi-functional design consolidates complexity into a single service layer rather than requiring complex integration at each edge device.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The cloud service acts as an intermediary that simplifies the interaction between edge devices and LLMs. Instead of edge devices needing to directly communicate with LLMs and manage complex authentication and data sharing protocols, the intermediary handles these complexities centrally, providing a simplified API for data aggregation and embedding generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Object-affected harmful factors

If access control mechanisms are implemented for user data, then security is improved, but ease of operation decreases

Engineering Contradiction:
Improvedata securityVSAvoiduser interaction simplicity
Core Design Contradiction:
Object-affected harmful factorsVSEase of operation

Solution Approach 1:

The patent implements self-service access control where the system automatically manages user data access based on pre-configured policies and user profiles. The cloud service autonomously determines which users can access which data, manages authentication tokens, and enforces access rights without requiring manual intervention. This maintains strong security while simplifying user interaction to basic authentication.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240420491A1Network infrastructure for user-specific generative intelligence
Publication Date: 2024.12.19 SOFTEYE INC
  • US20240420491A1 patent drawing
  • US20240420491A1 patent drawing
  • US20240420491A1 patent drawing

AI summary

Network infrastructure for user-specific generative intelligence. Providing user-specific context to a generically trained LLM introduces a variety of complications (privacy, resource utilization, training costs, etc.). Various aspects of the present disclosure provide novel user-specific data structures, privacy and access control, layers of data, and session management, within a network infrastructure for generative intelligence. For example, user-specific embedding vectors may be used to provide user context to a generically trained foundation model. In some variants, edge devices capture multiple modalities of user context (images, audio; not just text). Privacy and access control mechanisms also allow a user to control information that is captured and sent to the foundation model. Session management further decouples a user's conversational state from the foundation model's session state. These concepts and others may be used to emulate e.g., a chatbot based virtual assistant that responds based on user context.