Device for the secure
hybrid execution of context-aware Large
Language Model (LLM) deployments, comprising: an edge
execution unit (101) configured to execute at least a first part of an LLM
inference process on a local device; a cloud
inference processing unit (102) configured to execute at least a second part of the LLM
inference process on a
remote computing infrastructure; a
context sensitivity analyzer (103) configured to analyze an input query and associated
context data and generate a sensitivity classification for said
context data;a
hybrid execution
orchestration engine (104) that functions with the
context sensitivity analyzer (103), the edge
execution unit (101), and the cloud inference
processing unit (102), wherein the
hybrid execution
orchestration engine (104) is configured to dynamically split the LLM inference process into an edge-executed part and a cloud-executed part, based on at least the sensitivity classification; a
secure communication interface (105) configured to transmit intermediate inference outputs between the edge
execution unit (101) and the cloud inference
processing unit (102) over an encrypted
communication channel; a TEE module (106) integrated into the edge execution unit (101) to perform at least one privacy-relevant inference operation in an isolated, hardware-protected environment;a
key management and
encryption controller (107) configured to generate, store, distribute, and rotate cryptographic keys and to encrypt and decrypt intermediate inference outputs transmitted between the edge execution unit (101) and the cloud inference processing unit (102); a
policy enforcement and
access control unit (108) configured to validate user
authorization and enforce execution policies that define whether at least some of the
context data is restricted to local execution;and an audit
logging and
threat monitoring module (109) configured to
record inference execution events, security decisions,
policy enforcement events and communication events and detect
anomalous behavior related to the LLM inference process, wherein the device securely generates an output response to the input query by performing context-aware LLM inference while preventing sensitive context data from being disclosed to unauthorized external environments.