Methods for enabling virtual personal assistant mixed reality service using integrated communication and computing
The distributed AI service platform optimizes computational workload distribution between user devices and network nodes by splitting AI/ML tasks, addressing inefficiencies in existing systems and enhancing performance and user experience.
Patent Information
- Application Number
- PCT/CN2025/129013
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-21
- Filing Date
- 2025-10-21
- Publication Date
- 2026-04-30
AI Technical Summary
Existing wireless communication systems, particularly 5G NR, lack efficient methods for integrating artificial intelligence services that optimize computational workload distribution between user devices and network nodes, leading to suboptimal performance and resource utilization.
A distributed AI service platform that splits AI/ML tasks between user devices and edge servers, enabling local computation and offloading intensive tasks to network devices, reducing unnecessary data transmission and optimizing resource sharing.
Enhances system performance and user experience by reducing uplink traffic, minimizing bandwidth usage, and improving processing times while maintaining personalized assistance across various domains.
Smart Images

Figure CN2025129013_30042026_PF_FP_ABST
Abstract
Description
METHODS FOR ENABLING VIRTUAL PERSONAL ASSISTANT MIXED REALITY SERVICE USING INTEGRATED COMMUNICATION AND COMPUTINGCROSS-REFERENCE TO RELATED APPLICATION (S)
[0001] This application claims the benefits of U.S. Provisional Application Serial No. 63 / 709, 637, entitled “Methods for Enabling a Virtual Personal Assistant Mixed Reality Service Using Integrated Communication and Computing” and filed on October 21, 2024, which is expressly incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates generally to wireless communications, and more particularly, to techniques for enabling a virtual personal assistant mixed reality service using integrated communication and computing.BACKGROUND
[0003] The statements in this section merely provide background information related to the present disclosure and may not constitute prior art.
[0004] Wireless communication systems are widely deployed to provide various telecommunication services such as telephony, video, data, messaging, and broadcasts. Typical wireless communication systems may employ multiple-access technologies capable of supporting communication with multiple users by sharing available system resources. Examples of such multiple-access technologies include code division multiple access (CDMA) systems, time division multiple access (TDMA) systems, frequency division multiple access (FDMA) systems, orthogonal frequency division multiple access (OFDMA) systems, single-carrier frequency division multiple access (SC-FDMA) systems, and time division synchronous code division multiple access (TD-SCDMA) systems.
[0005] These multiple access technologies have been adopted in various telecommunication standards to provide a common protocol that enables different wireless devices to communicate on a municipal, national, regional, and even global level. An example telecommunication standard is 5G New Radio (NR) . 5G NR is part of a continuous mobile broadband evolution promulgated by Third Generation Partnership Project (3GPP) to meet new requirements associated with latency, reliability, security, scalability (e.g., with Internet of Things (IoT) ) , and other requirements. Some aspects of 5G NR may be based on the 4G Long Term Evolution (LTE) standard. There exists a need for further improvements in 5G NR technology. These improvements may also be applicable to other multi-access technologies and the telecommunication standards that employ these technologies.SUMMARY
[0006] The following presents a simplified summary of one or more aspects in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements of all aspects nor delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.
[0007] In an aspect of the disclosure, a method, a computer-readable medium, and an apparatus are provided. The method is performed by a user device for providing an Artificial Intelligence (AI) service. The user device receives a user query for the AI service. The user device assesses one or more conditions. The one or more conditions may include at least one of: a local resource of the user device, a network connection resource between the user device and a server, or an availability of information on the user device to respond to the user query. The user device determines, based on the assessment, an operational mode from a plurality of operational modes for executing the AI service. The plurality of operational modes may include a local-only mode and at least one split-computation mode. The user device executes the AI service to provide a response to the user query according to the determined operational mode. In the local-only mode, the AI service may be executed entirely on the user device. In the at least one split-computation mode, at least a portion of the AI service may be offloaded to the server.
[0008] To the accomplishment of the foregoing and related ends, the one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the annexed drawings set forth in detail certain illustrative features of the one or more aspects. These features are indicative, however, of but a few of the various ways in which the principles of various aspects may be employed, and this description is intended to include all such aspects and their equivalents.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] FIG. 1 is a diagram illustrating an example of a wireless communications system and an access network.
[0010] FIG. 2 is a diagram illustrating a base station in communication with a UE in an access network.
[0011] FIG. 3 illustrates an example logical architecture of a distributed access network.
[0012] FIG. 4 illustrates an example physical architecture of a distributed access network.
[0013] FIG. 5 is a diagram illustrating a distributed AI service platform design overview.
[0014] FIG. 6 is a diagram illustrating a virtual personal assistant mixed reality service deployment architecture for retail industry.
[0015] FIG. 7 is a diagram illustrating a message flow for a centralized service proxy with CPN connection.
[0016] FIG. 8 is a diagram illustrating a message flow for a centralized service proxy without CPN connection.
[0017] FIG. 9 is a diagram illustrating a message flow without using centralized service proxy and without CPN connection.
[0018] FIG. 10 is a diagram illustrating a virtual personal assistant mixed reality service design.
[0019] FIG. 11 is a diagram illustrating a scenario where the network is capable, the user device resources are insufficient and / or the information is not available locally.
[0020] FIG. 12 is a diagram illustrating a scenario where the network is poor, the user device resources are sufficient, and the information is not available locally.
[0021] FIG. 13 is a diagram illustrating a scenario where the user device resources are sufficient, and the information is available locally.
[0022] FIG. 14 illustrates a flow chart of a process for enabling a virtual personal assistant mixed reality service using integrated communication and computing.
[0023] FIG. 15 illustrates a flow chart of another process for enabling a virtual personal assistant mixed reality service using integrated communication and computing.
[0024] FIG. 16 illustrates a flow chart of a process for preparing a service model for a virtual assistant service.DETAILED DESCRIPTION
[0025] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well known structures and components are shown in block diagram form in order to avoid obscuring such concepts.
[0026] Several aspects of telecommunications systems will now be presented with reference to various apparatus and methods. These apparatus and methods will be described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as “elements” ) . These elements may be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.
[0027] By way of example, an element, or any portion of an element, or any combination of elements may be implemented as a “processing system” that includes one or more processors. Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs) , central processing units (CPUs) , application processors, digital signal processors (DSPs) , reduced instruction set computing (RISC) processors, systems on a chip (SoC) , baseband processors, field programmable gate arrays (FPGAs) , programmable logic devices (PLDs) , state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functionality described throughout this disclosure. One or more processors in the processing system may execute software. Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software components, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0028] Accordingly, in one or more example aspects, the functions described may be implemented in hardware, software, or any combination thereof. If implemented in software, the functions may be stored on or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media. Storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise a random-access memory (RAM) , a read-only memory (ROM) , an electrically erasable programmable ROM (EEPROM) , optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of the aforementioned types of computer-readable media, or any other medium that can be used to store computer executable code in the form of instructions or data structures that can be accessed by a computer.
[0029] FIG. 1 is a diagram illustrating an example of a wireless communications system and an access network 100. The wireless communications system (also referred to as a wireless wide area network (WWAN) ) includes base stations 102, UEs 104, an Evolved Packet Core (EPC) 160, and another core network 190 (e.g., a 5G Core (5GC) ) . The base stations 102 may include macrocells (high power cellular base station) and / or small cells (low power cellular base station) . The macrocells include base stations. The small cells include femtocells, picocells, and microcells.
[0030] The base stations 102 configured for 4G LTE (collectively referred to as Evolved Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access Network (E-UTRAN) ) may interface with the EPC 160 through backhaul links 132 (e.g., SI interface) . The base stations 102 configured for 5G NR (collectively referred to as Next Generation RAN (NG-RAN) ) may interface with core network 190 through backhaul links 184. In addition to other functions, the base stations 102 may perform one or more of the following functions: transfer of user data, radio channel ciphering and deciphering, integrity protection, header compression, mobility control functions (e.g., handover, dual connectivity) , inter cell interference coordination, connection setup and release, load balancing, distribution for non-access stratum (NAS) messages, NAS node selection, synchronization, radio access network (RAN) sharing, multimedia broadcast multicast service (MBMS) , subscriber and equipment trace, RAN information management (RIM) , paging, positioning, and delivery of warning messages. The base stations 102 may communicate directly or indirectly (e.g., through the EPC 160 or core network 190) with each other over backhaul links 134 (e.g., X2 interface) . The backhaul links 134 may be wired or wireless.
[0031] The base stations 102 may wirelessly communicate with the UEs 104. Each of the base stations 102 may provide communication coverage for a respective geographic coverage area 110. There may be overlapping geographic coverage areas 110. For example, the small cell 102’ may have a coverage area 110’ that overlaps the coverage area 110 of one or more macro base stations 102. A network that includes both small cell and macrocells may be known as a heterogeneous network. A heterogeneous network may also include Home Evolved Node Bs (eNBs) (HeNBs) , which may provide service to a restricted group known as a closed subscriber group (CSG) . The communication links 120 between the base stations 102 and the UEs 104 may include uplink (UL) (also referred to as reverse link) transmissions from a UE 104 to a base station 102 and / or downlink (DL) (also referred to as forward link) transmissions from a base station 102 to a UE 104. The communication links 120 may use multiple-input and multiple-output (MIMO) antenna technology, including spatial multiplexing, beamforming, and / or transmit diversity. The communication links may be through one or more carriers. The base stations 102 / UEs 104 may use spectrum up to 7 MHz (e.g., 5, 10, 15, 20, 100, 400, etc. MHz) bandwidth per carrier allocated in a carrier aggregation of up to a total of Yx MHz (x component carriers) used for transmission in each direction. The carriers may or may not be adjacent to each other. Allocation of carriers may be asymmetric with respect to DL and UL (e.g., more or fewer carriers may be allocated for DL than for UL) . The component carriers may include a primary component carrier and one or more secondary component carriers. A primary component carrier may be referred to as a primary cell (PCell) and a secondary component carrier may be referred to as a secondary cell (SCell) .
[0032] Certain UEs 104 may communicate with each other using device-to-device (D2D) communication link 158. The D2D communication link 158 may use the DL / UL WWAN spectrum. The D2D communication link 158 may use one or more sidelink channels, such as a physical sidelink broadcast channel (PSBCH) , a physical sidelink discovery channel (PSDCH) , a physical sidelink shared channel (PSSCH) , and a physical sidelink control channel (PSCCH) . D2D communication may be through a variety of wireless D2D communications systems, such as for example, FlashLinQ, WiMedia, Bluetooth, ZigBee, Wi-Fi based on the IEEE 802.11 standard, LTE, or NR.
[0033] The wireless communications system may further include a Wi-Fi access point (AP) 150 in communication with Wi-Fi stations (STAs) 152 via communication links 154 in a 5 GHz unlicensed frequency spectrum. When communicating in an unlicensed frequency spectrum, the STAs 152 / AP 150 may perform a clear channel assessment (CCA) prior to communicating in order to determine whether the channel is available.
[0034] The small cell 102’ may operate in a licensed and / or an unlicensed frequency spectrum. When operating in an unlicensed frequency spectrum, the small cell 102’ may employ NR and use the same 5 GHz unlicensed frequency spectrum as used by the Wi-Fi AP 150. The small cell 102’ , employing NR in an unlicensed frequency spectrum, may boost coverage to and / or increase capacity of the access network.
[0035] A base station 102, whether a small cell 102’ or a large cell (e.g., macro base station) , may include an eNB, gNodeB (gNB) , or another type of base station. Some base stations, such as gNB 180 may operate in a traditional sub 6 GHz spectrum, in millimeter wave (mmW) frequencies, and / or near mmW frequencies in communication with the UE 104. When the gNB 180 operates in mmW or near mmW frequencies, the gNB 180 may be referred to as an mmW base station. Extremely high frequency (EHF) is part of the RF in the electromagnetic spectrum. EHF has a range of 30 GHz to 300 GHz and a wavelength between 1 millimeter and 10 millimeters. Radio waves in the band may be referred to as a millimeter wave. Near mmW may extend down to a frequency of 3 GHz with a wavelength of 100 millimeters. The super high frequency (SHF) band extends between 3 GHz and 30 GHz, also referred to as centimeter wave. Communications using the mmW / near mmW radio frequency band (e.g., 3 GHz -300 GHz) has extremely high path loss and a short range. The mmW base station 180 may utilize beamforming 182 with the UE 104 to compensate for the extremely high path loss and short range.
[0036] The base station 180 may transmit a beamformed signal to the UE 104 in one or more transmit directions 108a. The UE 104 may receive the beamformed signal from the base station 180 in one or more receive directions 108b. The UE 104 may also transmit a beamformed signal to the base station 180 in one or more transmit directions. The base station 180 may receive the beamformed signal from the UE 104 in one or more receive directions. The base station 180 / UE 104 may perform beam training to determine the best receive and transmit directions for each of the base station 180 / UE 104. The transmit and receive directions for the base station 180 may or may not be the same. The transmit and receive directions for the UE 104 may or may not be the same.
[0037] The EPC 160 may include a Mobility Management Entity (MME) 162, other MMEs 164, a Serving Gateway 166, a Multimedia Broadcast Multicast Service (MBMS) Gateway 168, a Broadcast Multicast Service Center (BM-SC) 170, and a Packet Data Network (PDN) Gateway 172. The MME 162 may be in communication with a Home Subscriber Server (HSS) 174. The MME 162 is the control node that processes the signaling between the UEs 104 and the EPC 160. Generally, the MME 162 provides bearer and connection management. All user Internet protocol (IP) packets are transferred through the Serving Gateway 166, which itself is connected to the PDN Gateway 172. The PDN Gateway 172 provides UE IP address allocation as well as other functions. The PDN Gateway 172 and the BM-SC 170 are connected to the IP Services 176. The IP Services 176 may include the Internet, an intranet, an IP Multimedia Subsystem (IMS) , a PS Streaming Service, and / or other IP services. The BM-SC 170 may provide functions for MBMS user service provisioning and delivery. The BM-SC 170 may serve as an entry point for content provider MBMS transmission, may be used to authorize and initiate MBMS Bearer Services within a public land mobile network (PLMN) , and may be used to schedule MBMS transmissions. The MBMS Gateway 168 may be used to distribute MBMS traffic to the base stations 102 belonging to a Multicast Broadcast Single Frequency Network (MBSFN) area broadcasting a particular service, and may be responsible for session management (start / stop) and for collecting eMBMS related charging information.
[0038] The core network 190 may include a Access and Mobility Management Function (AMF) 192, other AMFs 193, a location management function (LMF) 198, a Session Management Function (SMF) 194, and a User Plane Function (UPF) 195. The AMF 192 may be in communication with a Unified Data Management (UDM) 196. The AMF 192 is the control node that processes the signaling between the UEs 104 and the core network 190. Generally, the SMF 194 provides QoS flow and session management. All user Internet protocol (IP) packets are transferred through the UPF 195. The UPF 195 provides UE IP address allocation as well as other functions. The UPF 195 is connected to the IP Services 197. The IP Services 197 may include the Internet, an intranet, an IP Multimedia Subsystem (IMS) , a PS Streaming Service, and / or other IP services.
[0039] The base station may also be referred to as a gNB, Node B, evolved Node B (eNB) , an access point, a base transceiver station, a radio base station, a radio transceiver, a transceiver function, a basic service set (BSS) , an extended service set (ESS) , a transmit reception point (TRP) , or some other suitable terminology. The base station 102 provides an access point to the EPC 160 or core network 190 for a UE 104. Examples of UEs 104 include a cellular phone, a smart phone, a session initiation protocol (SIP) phone, a laptop, a personal digital assistant (PDA) , a satellite radio, a global positioning system, a multimedia device, a video device, a digital audio player (e.g., MP3 player) , a camera, a game console, a tablet, a smart device, a wearable device, a vehicle, an electric meter, a gas pump, a large or small kitchen appliance, a healthcare device, an implant, a sensor / actuator, a display, or any other similar functioning device. Some of the UEs 104 may be referred to as IoT devices (e.g., parking meter, gas pump, toaster, vehicles, heart monitor, etc. ) . The UE 104 may also be referred to as a station, a mobile station, a subscriber station, a mobile unit, a subscriber unit, a wireless unit, a remote unit, a mobile device, a wireless device, a wireless communications device, a remote device, a mobile subscriber station, an access terminal, a mobile terminal, a wireless terminal, a remote terminal, a handset, a user agent, a mobile client, a client, or some other suitable terminology.
[0040] Although the present disclosure may reference 5G New Radio (NR) , the present disclosure may be applicable to other similar areas, such as LTE, LTE-Advanced (LTE-A) , Code Division Multiple Access (CDMA) , Global System for Mobile communications (GSM) , or other wireless / radio access technologies.
[0041] FIG. 2 is a block diagram of a base station 210 in communication with a UE 250 in an access network. In the DL, IP packets from the EPC 160 may be provided to a controller / processor 275. The controller / processor 275 implements layer 3 and layer 2 functionality. Layer 3 includes a radio resource control (RRC) layer, and layer 2 includes a packet data convergence protocol (PDCP) layer, a radio link control (RLC) layer, and a medium access control (MAC) layer. The controller / processor 275 provides RRC layer functionality associated with broadcasting of system information (e.g., MIB, SIBs) , RRC connection control (e.g., RRC connection paging, RRC connection establishment, RRC connection modification, and RRC connection release) , inter radio access technology (RAT) mobility, and measurement configuration for UE measurement reporting; PDCP layer functionality associated with header compression / decompression, security (ciphering, deciphering, integrity protection, integrity verification) , and handover support functions; RLC layer functionality associated with the transfer of upper layer packet data units (PDUs) , error correction through ARQ, concatenation, segmentation, and reassembly of RLC service data units (SDUs) , re-segmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and transport channels, multiplexing of MAC SDUs onto transport blocks (TBs) , demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction through HARQ, priority handling, and logical channel prioritization.
[0042] The transmit (TX) processor 216 and the receive (RX) processor 270 implement layer 1 functionality associated with various signal processing functions. Layer 1, which includes a physical (PHY) layer, may include error detection on the transport channels, forward error correction (FEC) coding / decoding of the transport channels, interleaving, rate matching, mapping onto physical channels, modulation / demodulation of physical channels, and MIMO antenna processing. The TX processor 216 handles mapping to signal constellations based on various modulation schemes (e.g., binary phase-shift keying (BPSK) , quadrature phase-shift keying (QPSK) , M-phase-shift keying (M-PSK) , M-quadrature amplitude modulation (M-QAM) ) . The coded and modulated symbols may then be split into parallel streams. Each stream may then be mapped to an OFDM subcarrier, multiplexed with a reference signal (e.g., pilot) in the time and / or frequency domain, and then combined together using an Inverse Fast Fourier Transform (IFFT) to produce a physical channel carrying a time domain OFDM symbol stream. The OFDM stream is spatially precoded to produce multiple spatial streams. Channel estimates from a channel estimator 274 may be used to determine the coding and modulation scheme, as well as for spatial processing. The channel estimate may be derived from a reference signal and / or channel condition feedback transmitted by the UE 250. Each spatial stream may then be provided to a different antenna 220 via a separate transmitter 218TX. Each transmitter 218TX may modulate an RF carrier with a respective spatial stream for transmission.
[0043] At the UE 250, each receiver 254RX receives a signal through its respective antenna 252. Each receiver 254RX recovers information modulated onto an RF carrier and provides the information to the receive (RX) processor 256. The TX processor 268 and the RX processor 256 implement layer 1 functionality associated with various signal processing functions. The RX processor 256 may perform spatial processing on the information to recover any spatial streams destined for the UE 250. If multiple spatial streams are destined for the UE 250, they may be combined by the RX processor 256 into a single OFDM symbol stream. The RX processor 256 then converts the OFDM symbol stream from the time-domain to the frequency domain using a Fast Fourier Transform (FFT) . The frequency domain signal comprises a separate OFDM symbol stream for each subcarrier of the OFDM signal. The symbols on each subcarrier, and the reference signal, are recovered and demodulated by determining the most likely signal constellation points transmitted by the base station 210. These soft decisions may be based on channel estimates computed by the channel estimator 258. The soft decisions are then decoded and deinterleaved to recover the data and control signals that were originally transmitted by the base station 210 on the physical channel. The data and control signals are then provided to the controller / processor 259, which implements layer 3 and layer 2 functionality.
[0044] The controller / processor 259 can be associated with a memory 260 that stores program codes and data. The memory 260 may be referred to as a computer-readable medium. In the UL, the controller / processor 259 provides demultiplexing between transport and logical channels, packet reassembly, deciphering, header decompression, and control signal processing to recover IP packets from the EPC 160. The controller / processor 259 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.
[0045] Similar to the functionality described in connection with the DL transmission by the base station 210, the controller / processor 259 provides RRC layer functionality associated with system information (e.g., MIB, SIBs) acquisition, RRC connections, and measurement reporting; PDCP layer functionality associated with header compression / decompression, and security (ciphering, deciphering, integrity protection, integrity verification) ; RLC layer functionality associated with the transfer of upper layer PDUs, error correction through ARQ, concatenation, segmentation, and reassembly of RLC SDUs, re-segmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and transport channels, multiplexing of MAC SDUs onto TBs, demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction through HARQ, priority handling, and logical channel prioritization.
[0046] Channel estimates derived by a channel estimator 258 from a reference signal or feedback transmitted by the base station 210 may be used by the TX processor 268 to select the appropriate coding and modulation schemes, and to facilitate spatial processing. The spatial streams generated by the TX processor 268 may be provided to different antenna 252 via separate transmitters 254TX. Each transmitter 254TX may modulate an RF carrier with a respective spatial stream for transmission. The UL transmission is processed at the base station 210 in a manner similar to that described in connection with the receiver function at the UE 250. Each receiver 218RX receives a signal through its respective antenna 220. Each receiver 218RX recovers information modulated onto an RF carrier and provides the information to a RX processor 270.
[0047] The controller / processor 275 can be associated with a memory 276 that stores program codes and data. The memory 276 may be referred to as a computer-readable medium. In the UL, the controller / processor 275 provides demultiplexing between transport and logical channels, packet reassembly, deciphering, header decompression, control signal processing to recover IP packets from the UE 250. IP packets from the controller / processor 275 may be provided to the EPC 160. The controller / processor 275 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.
[0048] New radio (NR) may refer to radios configured to operate according to a new air interface (e.g., other than Orthogonal Frequency Divisional Multiple Access (OFDMA) -based air interfaces) or fixed transport layer (e.g., other than Internet Protocol (IP) ) . NR may utilize OFDM with a cyclic prefix (CP) on the uplink and downlink and may include support for half-duplex operation using time division duplexing (TDD) . NR may include Enhanced Mobile Broadband (eMBB) service targeting wide bandwidth (e.g. 80 MHz beyond) , millimeter wave (mmW) targeting high carrier frequency (e.g. 60 GHz) , massive MTC (mMTC) targeting non-backward compatible MTC techniques, and / or mission critical targeting ultra-reliable low latency communications (URLLC) service.
[0049] A single component carrier bandwidth of 100 MHz may be supported. In one example, NR resource blocks (RBs) may span 12 sub-carriers with a sub-carrier bandwidth of 60 kHz over a 0.25 ms duration or a bandwidth of 30 kHz over a 0.5 ms duration (similarly, 50MHz BW for 15kHz SCS over a 1 ms duration) . Each radio frame may consist of 10 subframes (10, 20, 40 or 80 NR slots) with a length of 10 ms. Each slot may indicate a link direction (i.e., DL or UL) for data transmission and the link direction for each slot may be dynamically switched. Each slot may include DL / UL data as well as DL / UL control data. UL and DL slots for NR may be as described in more detail below with respect to FIGs. 5 and 6.
[0050] The NR RAN may include a central unit (CU) and distributed units (DUs) . A NR BS (e.g., gNB, 5G Node B, Node B, transmission reception point (TRP) , access point (AP) ) may correspond to one or multiple BSs. NR cells can be configured as access cells (ACells) or data only cells (DCells) . For example, the RAN (e.g., a central unit or distributed unit) can configure the cells. DCells may be cells used for carrier aggregation or dual connectivity and may not be used for initial access, cell selection / reselection, or handover. In some cases DCells may not transmit synchronization signals (SS) in some cases DCells may transmit SS. NR BSs may transmit downlink signals to UEs indicating the cell type. Based on the cell type indication, the UE may communicate with the NR BS. For example, the UE may determine NR BSs to consider for cell selection, access, handover, and / or measurement based on the indicated cell type.
[0051] FIG. 3 illustrates an example logical architecture of a distributed RAN 300, according to aspects of the present disclosure. A 5G access node 306 may include an access node controller (ANC) 302. The ANC may be a central unit (CU) of the distributed RAN. The backhaul interface to the next generation core network (NG-CN) 304 may terminate at the ANC. The backhaul interface to neighboring next generation access nodes (NG-ANs) 310 may terminate at the ANC. The ANC may include one or more TRPs 308 (which may also be referred to as BSs, NR BSs, Node Bs, 5G NBs, APs, or some other term) . As described above, a TRP may be used interchangeably with “cell. ”
[0052] The TRPs 308 may be a distributed unit (DU) . The TRPs may be connected to one ANC (ANC 302) or more than one ANC (not illustrated) . For example, for RAN sharing, radio as a service (RaaS) , and service specific ANC deployments, the TRP may be connected to more than one ANC. A TRP may include one or more antenna ports. The TRPs may be configured to individually (e.g., dynamic selection) or jointly (e.g., joint transmission) serve traffic to a UE.
[0053] The local architecture of the distributed RAN 300 may be used to illustrate fronthaul definition. The architecture may be defined that support fronthauling solutions across different deployment types. For example, the architecture may be based on transmit network capabilities (e.g., bandwidth, latency, and / or jitter) . The architecture may share features and / or components with LTE. According to aspects, the next generation AN (NG-AN) 310 may support dual connectivity with NR. The NG-AN may share a common fronthaul for LTE and NR.
[0054] The architecture may enable cooperation between and among TRPs 308. For example, cooperation may be preset within a TRP and / or across TRPs via the ANC 302. According to aspects, no inter-TRP interface may be needed / present.
[0055] According to aspects, a dynamic configuration of split logical functions may be present within the architecture of the distributed RAN 300. The PDCP, RLC, MAC protocol may be adaptably placed at the ANC or TRP.
[0056] FIG. 4 illustrates an example physical architecture of a distributed RAN 400, according to aspects of the present disclosure. A centralized core network unit (C-CU) 402 may host core network functions. The C-CU may be centrally deployed. C-CU functionality may be offloaded (e.g., to advanced wireless services (AWS) ) , in an effort to handle peak capacity. A centralized RAN unit (C-RU) 404 may host one or more ANC functions. Optionally, the C-RU may host core network functions locally. The C-RU may have distributed deployment. The C-RU may be closer to the network edge. A distributed unit (DU) 406 may host one or more TRPs. The DU may be located at edges of the network with radio frequency (RF) functionality.
[0057] The traditional and common role of user devices has been to provide connectivity to purpose-built servers within a network, as seen in server-client based services, or to facilitate peer-to-peer communication through such servers, exemplified by phone and video calls. However, next-generation networks, including mobile networks, are anticipated to support a wide variety of user devices, artificial intelligence (AI) applications, and distributed compute and communication sharing.
[0058] The distributed AI service platform addressed in this disclosure represents a strategic advancement in utilizing distributed computing principles within AI and machine learning (ML) frameworks for enabling virtual personal assistant mixed reality services. The platform is designed to optimize computational workload distribution between user devices and network nodes by splitting AI / ML tasks based on real-time conditions. When a customer visits a shopping mall or store, a virtual personal assistant is assigned and helps the customer throughout their visit until they leave the store. This is achieved through intelligent distribution of computational tasks between smartphones that handle initial and local AI / ML computation, and edge servers that perform substantial computation offloaded from user devices.
[0059] The platform addresses several key business elements that distinguish it from conventional approaches. First, it positions AI as an assistant with omnichannel capabilities that can address various service challenges. Second, it provides a customer service virtual assistant that accompanies the customer throughout their entire store visit, offering continuous support. Third, the architecture splits the roles of AI / ML processing between smartphones and edge servers, with smartphones processing the initial phases of AI / ML computation before offloading more intensive tasks to network devices. Fourth, edge servers handle significant AI / ML computation that has been offloaded from user or end devices. Fifth, the system reduces network traffic by avoiding the transmission of unnecessary information from user devices to network devices, thereby reducing the amount of uplink traffic.
[0060] This distributed approach provides multiple advantages for system performance and user experience. By distributing compute tasks across available resources in both the network and local devices, the system achieves efficient resource sharing that enhances overall system performance and scalability. The service significantly reduces uplink traffic by transmitting only necessary data to network devices, which minimizes bandwidth usage and associated costs. Users benefit from faster processing times and improved battery life on their devices, as heavy computations are offloaded to edge servers when appropriate. Additionally, this disclosure establishes a foundation for future value creation opportunities for device vendors, service providers, and network providers.
[0061] The distributed AI service platform provides an Integrated Communication and Compute (ICC) service architecture design and mechanism that specifically enables virtual personal assistant mixed reality services using virtual assistants. The platform supports domain-specific virtual personal assistant services which can be integrated with mixed reality capabilities. When the service initiates, a dedicated virtual assistant is assigned to assist the user until task completion, with the capability to continue personalized service based on the user’s preferences and previous interactions. The deployment of this distributed AI service can be flexible, with dedicated servers or compute resources either installed in the Customer Premises Network (CPN) within the enterprise network, or leased from a network operator such as a Mobile Network Operator (MNO) or cloud operator. This flexibility allows service providers to choose the most appropriate deployment model based on their specific requirements and resources.
[0062] FIG. 5 is a diagram 500 illustrating a distributed AI service platform design overview. The platform demonstrates that the same distributed AI service infrastructure for virtual personal assistant services can be applied across various domains, including healthcare, retail, manufacturing, agriculture, food services, libraries, museums, transportation, travel services, real-time language services, and services for disabled people. This versatility allows for flexible and personalized assistance while maintaining personal information on the user’s device and service provider information in the provider’s managed facilities, whether owned or leased. The distributed AI service can be deployed with a dedicated server or computing resource installed in the Customer Premises Network (CPN) within an enterprise network, or it can be leased from a network operator such as a Mobile Network Operator (MNO) or cloud operator. This distributed AI service platform creates new value creation opportunities for device vendors, service providers, and network providers by enabling them to participate in different aspects of the distributed service delivery ecosystem.
[0063] The distributed AI service platform comprises various functional modules that can be flexibly deployed across user devices and service engines. These modules are selectively activated based on real-time conditions including device capabilities and status such as low battery or low resource availability, and network conditions such as low link bandwidth availability. Some modules may reside on user devices like smartphones, while others may be deployed in nearby systems such as Customer Premises Equipment (CPE) or within mobile operator networks. Mobile operators such as AT&T or Verizon may host certain computing machines that participate in the service delivery. By grouping these distributed user devices and network resources together, they collectively provide the service through coordinated operation.
[0064] The user device incorporates several key functional modules to support the distributed service. A service model proxy serves as a simplified service model located on the device for initial processing, containing only essential components while critical service resources remain in the service engine managed by the service provider. A personal assistant module provides a virtual assistant that interacts with the user, local AI / ML modules, and the service model in the service engine. The device also includes a multi-lingual service module containing various small-sized ML models such as LLMs and Llava, which provides different language services based on user preferences when services are delivered locally. A device analysis module continuously monitors or predicts the user device’s operational status, including CPU and memory utilization, battery level, and battery charging status. A network analysis module monitors or predicts wireless channel conditions, including link bandwidth availability and achievable end-to-end throughput between the user device and the service engine. Additionally, the device includes a Speech to Text (STT) module that converts voice audio prompts or commands into text, and a Text to Speech (TTS) module that converts text responses into audio streams.
[0065] The service engine contains complementary modules that work with the user device to deliver the complete service. A service model serves as a comprehensive knowledge base containing all domain-specific information required to support clients and is managed by the service provider. The service engine’s multi-lingual service module houses various large-sized ML models such as LLMs and Llava, which are utilized when services are provided via high-performance servers and deliver different language services based on user preferences. A network analysis module in the service engine interacts with the corresponding module in the user device, enabling the device to accurately monitor network conditions. The service engine also includes a TTS module that converts text responses into audio streams when this processing is performed server-side rather than on the device.
[0066] The virtual personal assistant mixed reality service finds particular application in the retail industry, where it addresses challenges faced by both traditional stores and the emerging trend of unmanned stores. Unmanned stores, also known as autonomous or cashier-less stores, have gained significant traction in the retail industry as part of the broader shift toward automation and smart technology integration. These stores rely heavily on technologies such as RFID, computer vision, IoT, and AI to track purchases and manage inventory. Their primary consumer benefit is streamlining the shopping experience by eliminating checkout lines and making the shopping process more efficient. These stores collect vast amounts of consumer behavior data that can be used to optimize inventory levels, store layouts, and marketing strategies. While they potentially reduce operational costs by minimizing staffing requirements, the initial technology investment can be substantial.
[0067] Major technology companies are driving this unmanned store trend forward. Amazon with its Amazon Go stores and other tech giants are expanding their presence in this sector, demonstrating growing confidence in the unmanned store model’s viability. Global adoption varies significantly by country, with some nations implementing these technologies more rapidly than others. While retail remains the primary application, the concept extends to food services, libraries, and transportation hubs. As technology continues advancing, more sophisticated and diverse applications of unmanned systems are expected to emerge in retail and beyond. However, the trend faces significant challenges, particularly in addressing privacy concerns related to data collection and maintaining personalized, interactive customer services that extend beyond basic purchase tracking and inventory management.
[0068] The virtual personal assistant service directly addresses these challenges by providing the personalized interaction that unmanned stores typically lack. Consider a retail store such as a supermarket, where customers often struggle to locate products. At stores like Costco, frequent reorganization of item locations creates customer frustration when trying to find desired products. Customers entering the store often need assistance from customer service representatives, but human staff are frequently unavailable, occupied in back offices or other areas of the store.
[0069] When a store implements the distributed AI customer service system, virtual assistants replace the need for physical customer service staff. As a customer enters a shopping mall or store, the system can detect their presence and activate the virtual personal assistant on their device to provide assistance throughout their visit. Alternatively, customers who wish to use the virtual assistant can voluntarily activate it by scanning a Near Field Communication (NFC) tag upon entering the store or mall. The NFC tag scanning prompts users to install the application if not already present, or activates the virtual assistant if the application is already installed.
[0070] For first-time visitors, a default 3D avatar appears on the user’s device, whether smart glasses, smartphone, or other display-capable device, and provides a greeting. The avatar may be rendered over the real-world view as mixed reality (MR) or displayed separately from the real view as augmented reality (AR) . The interaction between the user and virtual assistant can occur through audio, text, or both, depending on device capabilities and user preferences.
[0071] Following an initial greeting such as "I am your assistant" or "You may select a virtual avatar or build one for yourself, " users can choose different virtual avatars including male, female, child, or pet representations, or customize the avatar’s appearance. This customization option is particularly relevant for first-time visitors or users without existing customer information stored in the system.
[0072] Returning customers benefit from continuity as the same avatar can continue serving them across visits. This personalization feature enhances customer engagement and may increase the likelihood of repeat store visits.
[0073] The system requests informed consent to capture the customer’s photograph for registration and future identification purposes. Alternative identification methods such as phone numbers are also available, and customers retain the option to decline and skip this registration step.
[0074] The virtual assistant may present an additional privacy option through a voice prompt stating "Would you allow me to remember you and your visit? I can serve you better next time when you visit again. " Communications during the visit can be stored locally on the user’s device for use during future store visits.
[0075] When customers have questions while shopping in the store, the virtual store assistant provides comprehensive support services. The assistant delivers detailed product information including manufacturer details, expiration dates, pricing, discounts, warranties, ingredients, nutritional content, and caloric values per serving. It also provides operational guidance such as assembly instructions, installation procedures, and usage directions. The system facilitates purchase-related services including order placement and order status tracking. Additionally, it provides product availability information, identifying the number of items in the current store, locating nearby stores with the desired product in stock, and providing backorder details. When users grant permission, the virtual store assistant delivers personalized services based on individual preferences, shopping history, and other relevant factors.
[0076] The system enables customers to ask natural language questions such as "Where can I find this item? " and receive immediate guidance to the appropriate aisle. When customers scan items, the virtual assistant provides comprehensive information about pricing, applicable discounts, sales status, backorder availability, and any other product details customers may request. This service cannot operate solely on customer devices such as phones or smart glasses because the requested information is specific to each store’s operations and inventory.
[0077] The system infrastructure resides on store servers, implemented through proxies or CPE owned by the store. Alternatively, stores like Costco may lease server capacity from network operators, with services delivered from that external infrastructure. Even as customers interact with virtual assistants on their personal devices, the underlying services originate from either nearby in-store servers, more distant mobile operator servers, or in certain cases, run directly on customer devices when sufficient local resources and cached information are available.
[0078] The distributed system makes dynamic decisions about function allocation based on multiple real-time factors. Network conditions and device capabilities determine where each function executes. High-end smartphones with substantial processing power can handle more functions locally, while smart glasses with limited battery life or hardware capacity rely more heavily on external servers, whether local proxies or mobile network infrastructure. This dynamic allocation optimizes performance while managing device resources and network bandwidth efficiently.
[0079] FIG. 6 is a diagram 600 illustrating a virtual personal assistant mixed reality service deployment architecture for retail industry. The virtual assistant service network comprises user devices, a product registration server, and network services connecting the user device, service proxy, and servers containing the service model and service engine. User devices include smartphones, smart glasses, and smartwatches that interact with the distributed service infrastructure. While a Retrieval Augmented Generation (RAG) framework is used for service model implementation in this disclosure, the application scope extends beyond RAG to include other suitable frameworks.
[0080] The service engine incorporates multiple functional modules including Speech to Text (STT) , Text to Speech (TTS) , a collection of Large Language Models (LLMs) , and submodules for 3D virtual assistant avatar generation and rendering. These submodules may be distributed across the service network rather than being collocated in a single machine, allowing for flexible deployment based on resource availability and performance requirements. The server infrastructure shown in FIG. 6 may be deployed in various configurations: within a dedicated server in the Customer Premises Network (CPN) as indicated by index ’ 1’ , hosted on shared computing and storage resources managed by Mobile Network Operators (MNO) as indicated by index ’ 2’ , operated by fixed network operators including DSL, fiber optic, cable, Wi-Fi, or satellite network operators as indicated by index ’ 3’ , or deployed on cloud network operator infrastructure as indicated by index ’ 4’ .
[0081] To optimize performance and reliability, the server may be deployed across multiple sites or networks for load balancing, improved fault tolerance, and enhanced end-to-end performance. The service model and service engine modules within the server may be deployed in different physical locations, with the possibility of maintaining a single service model while operating multiple service engine instances to support the virtual assistant service. The service proxy functions as the initial communication endpoint where user devices establish contact before connecting to the servers.
[0082] The user device, whether a smartphone, smart glasses, or smartwatch, operates as the initial processing hub for AI / ML tasks. It processes preliminary computational phases before offloading resource-intensive operations to the server infrastructure. This approach accelerates initial response times while reducing network data transmission. The device handles lightweight AI / ML computations including initial data gathering, preprocessing steps, and basic analytical operations that do not require extensive computational resources.
[0083] The service proxy serves as the primary service landing point for user device communications. When a virtual assistant application launches on a user device, it establishes connection with a known proxy site. This proxy may be located within Customer Premises Equipment (CPE) in the CPN or accessible through any network connection available to the client application. Based on configured service operation policies, the proxy performs various functions including redirecting the virtual assistant application to specific service servers, providing message forwarding capabilities, or implementing load balancing across multiple server instances.
[0084] The Evolved Residential Gateway (eRG) provides gateway functionality between public 5G network operators and Customer Premises Networks within residences, offices, or retail locations. The 5G system architecture supports network access through gateway components such as Personal IoT Network (PIN) elements with gateway capabilities or eRG devices for authorized User Equipment (UE) and authorized non-3GPP devices within a PIN or CPN. Premises Radio Access Stations (PRAS) , which are base stations installed within CPNs, function similarly to standard base stations and can operate using licensed frequency bands, unlicensed frequency bands, or a combination of both. The connectivity between the eRG and connected devices, including UEs, non-3GPP devices, or PRAS units, can utilize any appropriate non-3GPP technology such as Ethernet, optical connections, or WLAN.
[0085] The server functions as a network device containing both the service model and service engine components, executing significant AI / ML computations offloaded from user devices. The server infrastructure is equipped with substantial computational power and storage capacity to handle complex processing tasks. The service model component builds and maintains embedding databases for products during the registration process, while the service engine processes user queries using information retrieved from the service model. The server’s computational capabilities include executing deep learning processes, performing complex analytics, and managing storage-intensive operations that exceed the capacity of user devices.
[0086] Multiple technologies are available for detecting user devices and enabling interaction within the service area, with selection based on required accuracy and user convenience considerations. Table 1 provides detailed specifications of these detection technologies. Bluetooth and Wi-Fi Direct technologies are widely available in consumer devices, with their effective range determined by the radio propagation characteristics of the deployment environment.Table 1 Technologies for exchanging data over short distances
[0087] For applications requiring activation only when customers are physically inside the store, without false triggering from passersby, precise location detection becomes necessary. Ultra-wideband (UWB) technology provides high-precision localization capabilities but requires deployment of UWB anchor devices throughout the service area, resulting in substantially higher implementation costs. Near Field Communication (NFC) and Quick Response (QR) code technologies are particularly well-suited for virtual assistant activation, as they require deliberate user action to scan an NFC tag or QR code upon store entry, providing clear user intent with minimal effort.
[0088] When a customer enters a retail location, the system initiates a series of procedures to activate the virtual assistant service. Customers who wish to engage with the virtual assistant scan an NFC tag or QR code positioned at the store entrance. This scanning action triggers the system to either prompt for application installation if not already present on the device, or activate the virtual assistant if the application is already installed. The system then requests informed consent to capture a photograph of the customer for registration and identification during future visits. Alternative identification methods such as phone numbers are also supported, and customers retain the option to decline this registration step. The virtual assistant may present this request through natural speech, for example stating "Would you allow me to remember you and your visit? I can serve you better next time when you visit again. "
[0089] The virtual assistant also requests permission regarding interaction recording during the store visit. Customers can choose to have interactions stored in the store’s system for improved future service, saved only on their local device for personal reference and training purposes, or not recorded at all to address privacy concerns. This granular control over data handling allows customers to balance service personalization with privacy preferences.
[0090] The system accommodates various network connectivity scenarios that may occur during service operation. Similar to the variable usage of public Wi-Fi in locations like Starbucks, customers may or may not establish direct connections to the store’s local network infrastructure. The message flow adapts dynamically based on whether the user device connects to the store’s CPN. When no connection is established, the service operates through the user’s subscribed mobile network operator. Some retail locations may operate without any onsite service infrastructure, relying entirely on services hosted by mobile network operators or cloud providers.
[0091] This architecture supports comprehensive end-to-end communications through various message flow patterns adapted to different deployment and operational configurations. When users launch the virtual assistant client application, their devices typically are not initially connected to the CPN access point such as the store’s Wi-Fi. Instead, the client application contacts the service proxy through either the CPN or internet connectivity. The initial message from the user device reaches the service proxy through the user’s subscribed network operator and internet infrastructure. The proxy then coordinates subsequent communication steps according to its operational configuration. When proxy functionality is integrated directly into the client application, updates become necessary whenever the service network configuration changes, such as when adding, relocating, or removing service servers.
[0092] The user-subscribed network refers to the mobile network operator with which the customer maintains service, such as AT&T or Verizon. When customers enter a store, they typically do not have immediate access to the store’s local network credentials. If the store’s web service URL is known to the application, the initial service request routes through the customer’s subscribed operator network to reach the store’s proxy server. The proxy can then provide network access credentials including authentication IDs and passwords to enable local network connection. When the user device accepts these credentials and establishes a local connection, subsequent communications bypass the mobile operator network, reducing latency through direct interaction with the store’s service infrastructure. This mechanism optimizes data exchange efficiency once the local network connection is established.
[0093] FIG. 7 is a diagram 700 illustrating a message flow for a centralized service proxy with Customer Premises Network (CPN) connection. This message flow represents one of three primary deployment scenarios for the virtual personal assistant service, where the service proxy acts as an intermediary between user devices and distributed service servers. The CPN represents the local network infrastructure within a retail location or enterprise facility, which may host service servers directly or provide optimized connectivity to remote service resources.
[0094] The message flow begins when a user device launches the virtual assistant client application. At this initial stage, the user device typically connects through its subscribed network operator, such as AT&T or Verizon, rather than directly to the store’s local network infrastructure. The client application is configured with the service proxy’s address, which serves as the primary landing point for all initial service connections.
[0095] Step 1 involves the user device sending a service discovery request message to the service proxy via its subscribed network operator and the internet. This message contains comprehensive device information that enables the system to make optimal service delivery decisions. The device capability information includes hardware specifications such as the number of CPU cores, which indicates whether the device is a high-end smartphone capable of local AI processing or a resource-constrained device like smart glasses. The message also includes the device’s current location within the store, which helps determine proximity to local service resources. Battery status information, including whether the device is charging and the current charge level, such as 20 percent or 80 percent, provides critical input for workload distribution decisions. When battery levels are low, the system minimizes local processing requirements and relies more heavily on network-based resources to preserve device operation.
[0096] The service discovery request also includes Radio Access Technology (RAT) availability, specifying which wireless technologies the device supports, such as Wi-Fi, 4G, 5G, or Bluetooth. Each RAT option offers different bandwidth capabilities and power consumption characteristics that influence service delivery decisions. Radio quality metrics, including signal strength and connection stability, determine the maximum achievable data throughput between the user device and the network. This information directly impacts decisions about data-intensive operations, such as whether to stream high-quality video avatars or use lower-bandwidth text-based interactions.
[0097] Step 2 occurs when the service proxy responds with CPN access credentials and connection information. The proxy provides the network name, authentication credentials, and connection parameters that enable the user device to establish a direct connection to the store’s local Wi-Fi access point or other CPN infrastructure. This local connection opportunity offers reduced latency and higher bandwidth compared to routing through the public internet. However, the user device retains autonomy over this connection decision and may choose to opt-in or opt-out based on user preferences, privacy settings, or security considerations.
[0098] Step 3 represents the scenario where the user device accepts the local connection opportunity and establishes a connection to the CPN access point. This connection provides the most efficient communication path for subsequent service interactions, as data travels directly between the user device and local service infrastructure without traversing the public internet.
[0099] Step 4 confirms the successful CPN connection when the user device sends an acknowledgment message to the service proxy via the newly established CPN connection. This acknowledgment indicates that subsequent communications can utilize the optimized local network path rather than the longer route through the subscribed network operator.
[0100] Steps 5 and 6 demonstrate the standard service interaction pattern when the service server resides within the CPN. The user sends a service request to the service proxy, which may contain natural language questions about products, requests for assistance, or commands to the virtual assistant. The service proxy forwards this request to the appropriate server within the CPN network, which processes the request using its AI models and local database of product information. The server generates a response and sends it back through the service proxy to the user device. This local processing path provides the lowest possible latency since all communication occurs within the store’s network infrastructure.
[0101] Steps 7 and 8 illustrate an alternative service path when the required service server is deployed in a mobile network edge cloud rather than locally within the CPN. This configuration might be used for computationally intensive operations that require more powerful servers than those available locally, or for accessing centralized services shared across multiple store locations. The user sends a service request to the service proxy, which determines that the mobile network edge server should handle this particular request. The proxy forwards the message to the server in the mobile network. When the mobile network edge server is not directly accessible from the public internet, the traffic from the proxy passes through an evolved Residential Gateway (eRG) . The eRG serves as a gateway between the CPN and the 5G network infrastructure, providing authorized access to 5G network services for both 3GPP-compliant User Equipment (UEs) and non-3GPP devices within the CPN. The response from the mobile network edge server follows the reverse path through the eRG and service proxy back to the user device.
[0102] The message flow in FIG. 7 represents the optimal latency scenario where the user device utilizes the local CPN connection for communication with service infrastructure. This configuration minimizes network delays by eliminating unnecessary routing through public networks. However, the system design accommodates situations where devices may choose not to utilize available local connections. After receiving CPN access credentials in Step 2, a user device might decline the local connection opportunity due to privacy preferences, security policies, or user configuration settings. In such cases, the device continues routing all traffic through its subscribed mobile network operator, resulting in longer network paths and higher latency. While this represents a suboptimal performance configuration, it remains a valid operational mode that the system fully supports. This flexibility reflects the platform’s commitment to user autonomy and privacy preferences, allowing customers to balance performance optimization against their personal data handling preferences.
[0103] FIG. 8 is a diagram 800 illustrating a message flow for a centralized service proxy without CPN connection. This configuration represents an alternative operational mode where the user device maintains connectivity through its subscribed network operator rather than establishing a direct connection to the store’s local network infrastructure. While this approach results in higher latency compared to direct CPN connectivity, it remains a fully supported configuration that accommodates user preferences for privacy, security policies, or situations where automatic connection to local networks is disabled on the device.
[0104] The message flow begins with Step 1, where the user device sends a service discovery request message to the service proxy via its subscribed network operator and the internet. This initial communication follows the same pattern as in the CPN-connected scenario, with the message containing comprehensive device information including device capabilities, current location within the store, battery life status, device type, Radio Access Technology (RAT) availability, radio quality metrics, and user preference configurations. These parameters enable the service proxy to make informed decisions about service redirection and resource allocation, accounting for the fact that all subsequent communications will traverse the public internet rather than a local network path.
[0105] In Step 2, the service proxy responds with CPN access information and credentials, offering the user device an opportunity to connect to the CPN Access Point. This would enable reduced network service latency and achieve higher throughput connectivity by establishing a direct path to local service resources. The system presents this option to all devices regardless of their eventual connection decision, maintaining consistency in the service discovery protocol.
[0106] Step 3 represents the distinguishing characteristic of this scenario, where the user device actively chooses not to connect to the CPN Access Point despite having received valid credentials. The device sends a negative acknowledgement (NACK) message to the proxy via its subscribed network and the internet, confirming its decision to maintain routing through the public network. This choice may reflect user privacy preferences, corporate security policies that prohibit automatic connection to third-party networks, or device configurations that prioritize cellular connectivity. While this decision results in increased network delays due to the longer routing path, the system fully accommodates this choice as a valid operational mode.
[0107] Steps 4 and 5 demonstrate service interaction when the target server resides within the CPN network. The user sends a service request, such as a natural language question about products, to the service proxy through the internet. The service proxy forwards this request to the appropriate server in the CPN network, which processes the query and generates a response. The response from the server travels back to the user device through the service proxy, following the reverse path through the internet and the user’s subscribed network. Despite the server being located within the CPN, the communication path traverses the public internet because the user device has opted not to establish a local connection.
[0108] Steps 6 and 7 illustrate an alternative service fulfillment path when the required processing occurs on a server deployed in a Mobile Network Edge cloud rather than within the CPN. The user sends a service request to the service proxy via the internet, and based on factors such as computational requirements, data locality, or load balancing considerations, the proxy determines that the mobile network edge server should handle this particular request. The proxy forwards the message to the server in the mobile network. When the mobile network edge server is not directly accessible from the public internet, which is common for operator-managed infrastructure, the traffic from the proxy passes through an evolved Residential Gateway (eRG) . The eRG provides the necessary gateway functionality, supporting access to the 5G network and its services for both authorized User Equipment (UEs) and authorized non-3GPP devices. The response from the mobile network server follows the reverse path through the eRG and service proxy back to the user device.
[0109] The message flow illustrated in FIG. 8 demonstrates the platform’s flexibility in accommodating different user preferences and network configurations. While the absence of a direct CPN connection results in suboptimal latency characteristics compared to the locally-connected scenario shown in FIG. 7, this configuration provides important benefits for users who prioritize network isolation or have specific security requirements. The system’s ability to maintain full functionality regardless of the chosen network path demonstrates the robustness of the distributed architecture, where service delivery adapts to available connectivity options while maintaining consistent user experience.
[0110] FIG. 9 is a diagram 900 illustrating a message flow without using centralized service proxy and without CPN connection. This configuration represents the scenario with the highest latency among the three deployment options, as it combines the absence of both proxy-mediated communication optimization and local CPN connectivity. The message flow demonstrates how the system operates when the user device maintains independence from both the store’s local infrastructure and the centralized proxy’s ongoing mediation.
[0111] Step 1 initiates when the user device sends a service discovery request message to the service proxy via its subscribed network operator and the internet. The message contains comprehensive device information including device capabilities, location, battery life, device type, Radio Access Technology (RAT) availability, radio quality metrics, and user preference configuration. These parameters inform the proxy’s decision-making process for providing appropriate server information to the user device.
[0112] In Step 2, the service proxy responds with CPN access information and credentials, offering the user device an opportunity to establish a connection to the CPN Access Point. This connection would provide reduced network service latency and higher throughput connectivity. The proxy presents this option regardless of the user device’s eventual connection decision.
[0113] Step 3 reflects the user device’s decision not to connect to the CPN Access Point. The device sends a negative acknowledgement (NACK) message to the proxy via its subscribed network and the internet, confirming its preference to maintain routing through public networks. This decision may result from privacy preferences, security policies, or device configuration settings that prioritize maintaining network isolation from local infrastructure.
[0114] Step 4 represents a critical distinction in this scenario. The proxy provides suitable server information to the user device based on various operational factors including load distribution, geographic location, and server availability. When multiple servers exist across the distributed infrastructure, the proxy selects an appropriate server for direct communication. These servers may be deployed in diverse locations as illustrated in FIG. 6, including within Customer Premises Equipment near the store, in mobile operator networks, on fixed network operator infrastructure, or in cloud platforms such as ChatGPT, Amazon Services, or other cloud providers. Each server maintains identical functional capabilities, allowing the system to select based on current operational conditions and server availability rather than functional differences.
[0115] In Step 5, the user device sends an acknowledgement (ACK) message confirming receipt of the server information. This acknowledgement may be omitted if the previous message was delivered via TCP, which provides its own delivery confirmation mechanism.
[0116] Steps 6 and 7 demonstrate direct communication between the user device and the selected server. The user device sends service requests, such as natural language questions about products, directly to the designated server through the internet without routing through the service proxy. This direct communication path distinguishes this scenario from the proxy-mediated configurations. The server processes the request and sends its response directly back to the user device, following the reverse network path through the internet and the user’s subscribed network operator.
[0117] Steps 8 and 9 illustrate an alternative service fulfillment path when the service server operates within a fixed network infrastructure. Similar to the previous steps, the user device communicates directly with the server without proxy intermediation, with requests and responses traversing the internet and the user’s subscribed network.
[0118] This configuration results in the highest latency among the three deployment scenarios due to the combination of two factors. First, the absence of proxy-mediated forwarding after initial setup means that each request must establish its own routing path through the public internet without the benefit of proxy optimization. Second, the lack of CPN connection forces all traffic to traverse the longer path through the user’s subscribed network operator and the public internet, rather than utilizing potentially faster local network infrastructure. While this approach introduces additional latency, it provides maximum network isolation for users who prioritize privacy and security over performance optimization. The services in this configuration originate from servers deployed in mobile networks, fixed networks, or cloud infrastructure, but always reach the user device through the public internet path.
[0119] FIG. 10 is a diagram 1000 illustrating a virtual personal assistant mixed reality service design. The diagram depicts the system architecture between a user device, such as a smartphone, and a service server infrastructure, abstracting away intermediate network elements including Customer Premises Equipment (CPE) , evolved Residential Gateway (eRG) , Customer Premises Network (CPN) , and Mobile Network Operator (MNO) components. The design encompasses both the preparatory procedures for configuring the service server, which comprises the service engine and service model components, and the complete message flow sequences from initial user requests through final response delivery.
[0120] For the retail industry implementation, the service model preparation follows a structured three-step process. In Step 1, product list and information registration establishes the foundational knowledge base. The virtual assistant service requires comprehensive product information access, necessitating a complete service model build-out with all item details. The system stores all product items with their associated metadata in an inventory database, capturing unique item names, categories, detailed descriptions, regular pricing, discount rates, warranty terms, expiration dates, best-by dates, display inventory counts, and storage inventory counts. This data may be stored as plain text or converted to context vectors for efficient retrieval. The top vector database (DB) 1002 shown in FIG. 10 represents this product list and information registration database, serving as the primary repository for textual product information.
[0121] Step 2 implements product item image registration through a smartphone application interface. The process involves scanning each product item with the device camera while capturing the unique product item name identifier. The system stores multiple image captures of each item from different angles along with their corresponding item name embeddings in a dedicated vector database. This embedding storage process repeats for the complete product catalog, with the bottom vector DB 1004 in FIG. 10 representing the product item image registration database. This dual-database architecture enables efficient cross-referencing between visual product representations and their textual descriptions.
[0122] The two databases operate as complementary components of the service model. The product list and information registration database provides detailed textual product data, while the product item image registration database enables visual product identification through image embeddings. These databases can be referenced using either product names or product indices as keys, creating flexible query pathways. While the databases may be combined into a single unified database for simplified management, maintaining them separately optimizes query performance for different types of searches. The system employs pre-trained deep learning models for constructing the embedding database from product item images and their corresponding names. OpenAI’s Contrastive Language-Image Pre-Training (CLIP) model serves as the primary implementation choice, generating embeddings suitable for both image and text processing. CLIP’s dual-modal capability makes it particularly effective for matching visual product representations to their textual item names. The model operates under the Massachusetts Institute of Technology (MIT) License, providing flexibility for commercial deployment.
[0123] Step 3 involves fine-tuning the service model through validation and iterative improvement. After completing initial product item registration, the system validates accuracy by scanning sample items from the registered inventory. The validation process confirms whether the system correctly identifies items and retrieves accurate information. When misidentifications occur, the system updates existing embeddings or registers additional image embeddings until achieving reliable target item identification. Once the service model completes preparation and establishes connection with the service engine, the virtual personal assistant service becomes operational.
[0124] The system implements dynamic computational distribution based on real-time assessment of multiple operational factors. Network connectivity status determines available bandwidth and latency characteristics for data transmission. User device capability encompasses processing power, available memory, and current resource utilization levels. Resource availability includes battery status, with the system considering both current charge levels and whether the device is actively charging. Information availability on the user device affects whether queries can be resolved locally or require network access. Based on these factors, user inquiries may be processed at different locations within the distributed architecture, including the user device itself or various network servers. Response delivery similarly adapts to available resources, with processing occurring at the optimal location for current conditions.
[0125] The platform implements three primary operational scenarios that accommodate different combinations of network and device conditions. Scenario 1 addresses situations where the network provides capable connectivity while user device resources are insufficient or required information is unavailable locally. In this configuration, the user device minimizes local processing by streaming voice input directly to network servers. The server infrastructure handles computationally intensive tasks including speech-to-text conversion, query processing through Large Language Models, response generation, and text-to-speech conversion. The server transmits the final audio response back to the device, which simply plays the received audio without additional processing. This approach preserves device battery life and enables operation on resource-constrained devices such as smart glasses.
[0126] Scenario 2 accommodates conditions where network connectivity is poor but the user device possesses sufficient computational resources. Although the device can execute processing functions locally, it lacks the required information to answer queries, necessitating network communication. To minimize data transmission over the degraded network connection, the device performs speech-to-text conversion locally, transmitting only the compact text representation rather than bandwidth-intensive audio streams. The network server processes the text query through its AI models and returns a textual response, similar to typical ChatGPT interactions. The device then performs local text-to-speech conversion, generating audio output for the user. This approach maintains voice-based interaction while adapting to network limitations through efficient data transmission strategies.
[0127] Scenario 3 represents complete local processing when the user device possesses both sufficient computational resources and locally cached information. This situation typically occurs when customers make repeated queries that can be satisfied from previously cached responses. The device confirms two prerequisites before engaging local processing: first, that cached data adequately addresses the current request, and second, that sufficient battery capacity exists for local computation. The processing workflow executes entirely on-device, beginning with speech-to-text conversion of the user’s vocal query. The local Large Language Model generates a textual response using cached information, followed by text-to-speech transformation for audio output. The system completes the interaction by playing the synthesized audio response to the user. This configuration eliminates network dependency, providing rapid response times while preserving user privacy through complete on-device processing.
[0128] FIG. 11 is a diagram 1100 illustrating Scenario 1 of the distributed AI service platform operation, where the network provides capable connectivity while the user device has insufficient computational resources or lacks the required information locally. This scenario demonstrates the platform’s ability to dynamically offload computational tasks from resource-constrained devices to network servers, optimizing the distribution of AI / ML processing across the distributed architecture.
[0129] In this configuration, the distributed AI service platform allocates processing responsibilities based on real-time assessment of network capabilities and device limitations. The network’s robust connectivity enables efficient data transmission between the user device and service infrastructure, while the device’s resource constraints or information gaps necessitate reliance on server-side processing. This dynamic allocation represents a capability of the platform’s distributed computing architecture, where computational tasks migrate to the most appropriate processing location based on current operational conditions.
[0130] Step 1 initiates when the user requires information about a specific item. The user scans the item using the smartphone camera and generates an inquiry through voice or text input. In the exemplary diagram shown in FIG. 11, the camera image on the left side of the smartphone displays the target item being captured, either through the smartphone’s integrated camera or via smart glasses worn by the customer. This scenario assumes dual-mode communication between the customer and the virtual assistant, supporting both audio and text interactions to accommodate different user preferences and device capabilities.
[0131] Step 2 performs speech-to-text conversion on the smartphone, transforming the voice query into text format and subsequently into embeddings for processing. The converted text appears on the smartphone screen for user confirmation. The system stores this query locally and privately in the device’s database, creating a historical record that can enhance future service interactions through personalized response optimization.
[0132] Step 3 involves the smartphone converting both the input prompt text and captured images into embedding representations. This conversion process utilizes the CLIP model or similar embedding generation techniques to create compact vector representations of the multimodal input data.
[0133] Step 4 transmits the generated embeddings from the smartphone to the service engine through the selected network connection. The embedding-based transmission significantly reduces network traffic compared to sending raw images. While an uncompressed image at 1920x1080 resolution would require approximately 6 MB of data transmission, and even JPEG compression would result in 300-600 KB transfers, the embedding representation requires only about 2 KB. This represents a reduction of over 99%compared to uncompressed images and 85-93%compared to JPEG images, demonstrating the platform’s efficiency in minimizing network bandwidth consumption.
[0134] Step 5 occurs when the service engine receives the embeddings and retrieves the most relevant information for the user query from the service model. The service model may be collocated with the service engine for optimal performance or deployed remotely based on infrastructure considerations. The retrieval process utilizes the embedding vectors to identify matching products and associated information from the vector databases maintained within the service model.
[0135] Step 6 involves the service engine preparing a comprehensive response using its Large Language Model (LLM) capabilities combined with the relevant information retrieved from the service model. The service engine processes this information to generate a natural language response tailored to the user’s query, then delivers this response to the user device. This server-side processing leverages the computational resources available in the network infrastructure, compensating for the limited processing capabilities of the user device.
[0136] Step 7 extends the response delivery by converting the text-based response into speech format. The service engine utilizes a Text-to-Speech (TTS) Application Programming Interface (API) that supports Speech Synthesis Markup Language (SSML) , providing precise control over speech synthesis parameters. The TTS system generates phoneme or viseme data alongside the audio output, creating the synchronization information necessary for avatar lip-sync animation. This server-side audio generation reduces the computational burden on the user device while maintaining high-quality speech output.
[0137] Step 8 addresses the rendering of the virtual assistant’s 3D avatar, which can be performed either on the server or the client device depending on available resources and network conditions. The distributed AI platform supports both configurations, dynamically selecting the optimal approach based on current operational parameters.
[0138] For Option 1, when the virtual assistant avatar renders on the server, the system must address several technical considerations inherent in server-side rendering. While this approach offloads computational requirements from the client device, it introduces challenges related to network infrastructure and real-time streaming. Server-side rendering requires robust network connectivity to handle streaming data, with sufficient bandwidth to transmit high-quality video streams and low enough latency to maintain interactive responsiveness. The implementation utilizes streaming protocols optimized for real-time content delivery, while the server infrastructure must possess adequate computational power to handle simultaneous rendering and streaming operations.
[0139] FIG. 11 specifically illustrates the server-side avatar rendering configuration. The implementation creates a virtual assistant avatar that executes on the network server and streams to the user device’s display. WebSocket connections facilitate real-time bidirectional communication between the client and server, enabling synchronized control and feedback mechanisms. For real-time 3D avatars that perform text reading with synchronized lip movements, the system employs dynamic rendering rather than pre-rendered content. The server generates animations in real-time based on the phoneme data from the TTS process, creating lip shapes or visemes that correspond to each phoneme in the speech output.
[0140] The server-side rendering process involves multiple coordinated steps. Following TTS conversion from Step 7, the system performs avatar animation by mapping phonemes to corresponding facial animations. Each phoneme associates with specific lip shapes or visemes through predefined animation mappings. A 3D rendering engine processes these mappings in real-time, generating animated sequences synchronized with the audio timeline. The rendering infrastructure utilizes professional 3D engines such as Unity or Unreal Engine, which provide server-side rendering capabilities controllable through scripting interfaces.
[0141] The rendered avatar content streams to client devices using appropriate streaming technologies. WebRTC provides peer-to-peer real-time communication capabilities suitable for low-latency interactions. Alternatively, video streaming protocols such as Real-Time Messaging Protocol (RTMP) or HTTP Live Streaming (HLS) handle the content distribution. RTMP optimizes for real-time communication with minimal delay, making it suitable for interactive avatar streaming. HLS utilizes standard HTTP for data transfer, providing broad compatibility with web infrastructure, though it typically introduces higher latency than RTMP, which may impact real-time interaction quality.
[0142] The client-side configuration for server-rendered avatars involves embedding a video player capable of receiving and displaying live streaming content. The client application synchronizes the received audio and video streams, potentially applying timing adjustments to maintain precise alignment between the avatar’s lip movements and the speech audio. This synchronization compensates for any differential delays introduced during separate audio and video transmission paths.
[0143] For Option 2, when the virtual assistant avatar renders on the client device, the computational distribution shifts to favor client-side processing. This configuration performs rendering within the user’s web browser or native application environment, creating and displaying the animated 3D avatar locally. The distributed AI platform divides responsibilities strategically: the server handles computationally intensive TTS processing and phoneme generation, while the client device focuses on rendering and synchronization tasks that benefit from local execution.
[0144] The client-side rendering approach offers several advantages within the distributed architecture. Local rendering eliminates video streaming bandwidth requirements, reducing network traffic to only audio and synchronization data transmission. This configuration also minimizes latency in avatar responses since rendering occurs directly on the display device without network round-trip delays. The server generates the audio response and phoneme synchronization data using the TTS API infrastructure described in Step 7. Real-time communication between server and client occurs through WebSocket connections, transmitting audio data and phoneme timing information rather than video streams.
[0145] The client device receives the audio and phoneme data from the server and performs local avatar rendering. The browser or application loads the 3D avatar model and animation system, then processes the received phoneme data to generate corresponding facial animations. The rendering engine on the client device creates smooth transitions between visemes, synchronized with the audio playback timeline. The client simultaneously plays the received audio while animating the avatar’s facial features, maintaining precise synchronization through the phoneme timing data. This distributed approach balances computational load across the network, utilizing server resources for complex AI / ML processing while leveraging client devices for display-proximate rendering tasks.
[0146] Step 9 completes the interaction cycle by storing the response locally and privately in the device’s database. This local storage creates a persistent record of the interaction that can enhance future service encounters through improved personalization and response caching. The storage operation can occur at any point after the device receives the text-based response from Step 6, operating asynchronously with the audio and visual presentation steps to optimize the user experience. This local caching mechanism contributes to the platform’s ability to transition between network-dependent and local processing modes, as cached responses may enable future queries to be satisfied without network access.
[0147] FIG. 12 is a diagram 1200 illustrating Scenario 2 of the distributed AI service platform operation, which addresses conditions where network connectivity is poor while the user device possesses sufficient computational resources. This scenario specifically optimizes for degraded network conditions by strategically distributing processing tasks to minimize data transmission requirements while maintaining full service functionality. The configuration applies when the user device contains adequate computing resources including CPU, GPU, and memory capacity, along with sufficient battery power to support local processing operations. Despite these local resources, the device lacks the product information required to answer user queries, necessitating communication with network servers.
[0148] The distinction between Scenario 1 and Scenario 2 lies in the strategic relocation of bandwidth-intensive processing tasks from the network to the user device. In Scenario 2, both 3D avatar rendering and text-to-speech conversion execute locally on the device, as depicted in FIG. 12. This architectural decision directly addresses the network’s inability to support high-bandwidth streaming of video or audio content with acceptable latency. By performing these operations locally, the system reduces network traffic to only essential text-based communications, transmitting queries and responses as compact text strings rather than bandwidth-intensive multimedia streams.
[0149] Steps 1 through 6 follow the same operational sequence as described in Scenario 1, but with a critical difference in data transmission strategy. The user initiates the interaction by scanning a product item and providing a voice inquiry, which the device converts to text locally. The device then generates embeddings from both the text query and captured images. These embeddings, requiring only approximately 2 KB of data transmission, are sent to the service engine rather than raw images or audio. The service engine retrieves relevant product information from the service model and generates a textual response using its Large Language Model capabilities. Notably, the server transmits only the text-based response back to the user device, avoiding the bandwidth requirements of audio or video streaming that would be problematic over the degraded network connection.
[0150] Step 7 represents a key adaptation to network limitations, where the user device assumes responsibility for converting the text-based response received from the server into speech format. The device employs a Text-to-Speech API supporting Speech Synthesis Markup Language (SSML) , which provides precise control over speech synthesis parameters. The TTS system generates phoneme or viseme data alongside the audio output, creating the synchronization information necessary for avatar lip-sync animation. This local processing eliminates the need to stream audio content from the server, significantly reducing network bandwidth consumption while maintaining high-quality speech output for the user.
[0151] Step 8 consolidates all presentation-layer processing on the user device, including TTS operations, phoneme generation, 3D avatar rendering, and audio-visual synchronization. The device performs the complete set of operations that would typically be distributed between client and server in optimal network conditions. This comprehensive local processing parallels the client-side rendering approach described in Scenario 1’s Option 2, but extends beyond that configuration by also incorporating the TTS and phoneme generation tasks that would normally execute on the server. The device creates and animates the 3D avatar based on locally generated phoneme data, synchronizing lip movements with the audio output in real-time. This configuration maintains a smooth and engaging user experience despite the poor network conditions, demonstrating the platform’s ability to adapt its computational distribution based on available resources.
[0152] Step 9 completes the interaction by storing the response locally and privately in the device’s database, creating a persistent record for future service enhancement. This storage operation can occur at any point after the device receives the text-based response from Step 6, operating independently of the presentation processing. The local caching of responses contributes to the platform’s resilience, as frequently accessed information becomes available for future queries without requiring network access, potentially allowing the system to transition to fully local operation for repeated queries as described in Scenario 3.
[0153] This configuration exemplifies the distributed AI platform’s adaptive capabilities, dynamically adjusting the computational workload distribution to accommodate network limitations while maintaining service quality. By transmitting only compact text data over the constrained network connection while performing resource-intensive multimedia processing locally, the system achieves efficient operation even under challenging network conditions.
[0154] FIG. 13 is a diagram 1300 illustrating Scenario 3 of the distributed AI service platform operation, representing the configuration where the user device possesses both sufficient computational resources and locally cached information to fulfill user queries without network dependency. This scenario demonstrates the platform’s capability to operate completely autonomously on the user device, providing the fastest response times and maximum privacy protection by eliminating all network communications during query processing.
[0155] Scenario 3 represents a departure from both Scenario 1 and Scenario 2 in that the user device operates independently without any interaction with network servers, as depicted in FIG. 13. This configuration becomes viable when the user device satisfies two critical prerequisites. First, the device must contain adequate computational resources including processing power from CPU and GPU cores, sufficient available memory for model execution, and adequate battery capacity to support local AI / ML operations. Second, the information required to answer the user’s query must be available in the device’s local cache, either from previous interactions during the current session or from historical data preserved from earlier visits to the store.
[0156] The availability of local information typically occurs through two mechanisms. When customers make repeated queries about the same products during a shopping session, the responses from initial network-based queries are cached locally, enabling subsequent identical or similar queries to be satisfied from the cache. Additionally, information from previous store visits may persist in the device’s local storage, creating a personalized knowledge base that grows over time. This accumulated local information reduces dependency on network communications and enables progressively more queries to be handled entirely on-device.
[0157] The processing sequence begins with Steps 1 through 3, which follow the same pattern established in Scenario 1. The user initiates the interaction by scanning a product item with the device camera while providing a voice inquiry about the item. The smartphone performs local speech-to-text conversion, transforming the audio input into text format and displaying the converted text on the screen for user confirmation. The device then generates embeddings from both the text query and the captured product images using the locally deployed CLIP model or similar embedding generation framework.
[0158] Step 4 marks the divergence from network-dependent scenarios, as the user device utilizes the generated embeddings to retrieve relevant information directly from its local vector database. This local database contains previously cached product information, including embeddings and associated metadata from earlier queries. The retrieval process operates entirely within the device’s memory, eliminating network latency and providing near-instantaneous access to stored information.
[0159] Step 5 continues the local processing as the user device employs its locally deployed Large Language Model to prepare a natural language response. The local LLM, which may be a smaller, optimized version such as Llama3 8B or similar compact models suitable for mobile deployment, processes the retrieved information to generate a contextually appropriate response to the user’s query. This local inference eliminates the need for cloud-based AI processing while maintaining response quality appropriate for the retail assistant use case.
[0160] Step 6 involves the device converting the generated text response into speech format using a locally executed Text-to-Speech API. The TTS system, supporting Speech Synthesis Markup Language for precise speech control, generates both the audio output and accompanying phoneme or viseme data required for avatar animation synchronization. This local TTS processing maintains consistency with the fully autonomous operation mode, preventing any network dependencies in the response delivery pipeline.
[0161] Step 7 encompasses the complete avatar rendering and synchronization process, executing the same comprehensive set of operations described as Step 8 in Scenario 2. The device performs all avatar animation tasks locally, including loading and rendering the 3D avatar model, processing phoneme data to generate corresponding facial animations, and synchronizing lip movements with the audio playback timeline. This local rendering provides smooth, responsive avatar interactions without the latency or bandwidth constraints associated with network-based rendering.
[0162] Step 8 completes the interaction cycle by storing both the query and response in the device’s local database, paralleling Step 9 from Scenario 2. This persistent storage expands the local knowledge base, potentially enabling future related queries to be satisfied locally without network access. The continuous accumulation of interaction history creates an increasingly capable local assistant that can handle a growing proportion of user queries autonomously.
[0163] The distributed AI service platform’s architecture enables significant network traffic reduction through its use of embedding-based representations, a benefit that becomes particularly apparent when comparing the data transmission requirements across different scenarios. The platform employs pre-trained image embedding models to convert visual information into compact vector representations, dramatically reducing the amount of data that must be transmitted when network communication is required.
[0164] The platform utilizes various embedding models for this purpose, with common implementations including models from the Vision Transformer family, Convolutional Neural Networks such as ResNet and EfficientNet, and OpenAI’s Contrastive Language-Image Pre-Training model. The CLIP model serves as the primary implementation example in this disclosure, providing a concrete basis for quantifying the network traffic reduction achieved through embedding-based representations.
[0165] The CLIP model architecture comprises two primary components that work together to create unified embeddings for both visual and textual information. The image encoder, implemented as either a Vision Transformer or a ResNet, processes visual inputs to generate fixed-dimensional embeddings. When using the Vision Transformer variant, the model treats images as sequences of patches, while the ResNet variant employs convolutional layers for hierarchical feature extraction. The text encoder, implemented as a Transformer model, processes textual descriptions to generate embeddings in the same dimensional space as the image embeddings, enabling direct comparison and matching between visual and textual representations.
[0166] The specific CLIP variant employed in this implementation, designated as ViT-B / 32, utilizes a Vision Transformer backbone with specific architectural parameters. The vision encoder operates with a hidden size of 768 dimensions, representing the dimensionality of the hidden states within the transformer layers, and comprises 12 transformer layers for progressive feature refinement. The text encoder operates with a hidden size of 512 dimensions and similarly comprises 12 transformer layers. These dimensional specifications directly impact both the model’s representational capacity and the size of the generated embeddings.
[0167] The embedding generation process produces remarkably compact representations compared to the original image data. For the ViT-B / 32 model, the projection dimension is 512, meaning each embedding contains 512 numerical values. With each value stored as a 32-bit floating-point number occupying 4 bytes, the total embedding size equals 512 dimensions multiplied by 4 bytes per dimension, resulting in 2048 bytes or approximately 2 kilobytes per embedding.
[0168] This compact embedding representation provides dramatic data reduction compared to transmitting raw images. Consider a typical smartphone camera capturing images at 1080p resolution, comprising 1920 by 1080 pixels. With standard 24-bit color depth allocating 8 bits per color channel in the RGB model, an uncompressed image requires 1920 times 1080 times 3 bytes, totaling 6, 220, 800 bytes or approximately 6 megabytes. Even with JPEG compression, which typically achieves compression ratios between 10: 1 and 20: 1, the same image would require between 300 and 600 kilobytes for transmission.
[0169] The contrast between embedding size and image size demonstrates the platform’s efficiency in minimizing network traffic. Transmitting a 2-kilobyte embedding instead of a 6-megabyte uncompressed image represents a 99.97 percent reduction in data transmission. Even compared to JPEG-compressed images, the embedding approach reduces transmission requirements by 93 to 99 percent. This dramatic reduction in data transmission requirements enables the platform to maintain responsive operation even over bandwidth-constrained network connections, while also reducing the power consumption associated with wireless data transmission on battery-powered devices.
[0170] FIG. 14 illustrates a flow chart 1400 of a process for enabling a virtual personal assistant mixed reality service using integrated communication and computing. The process may be performed by a user device for providing an Artificial Intelligence (AI) service.
[0171] At block 1402, the user device receives a user query for the AI service.
[0172] At block 1404, the user device assesses one or more conditions. The one or more conditions may include at least one of: a local resource of the user device, a network connection resource between the user device and a server, or an availability of information on the user device to respond to the user query.
[0173] At block 1406, the user device determines, based on the assessment, an operational mode from a plurality of operational modes for executing the AI service. The plurality of operational modes may include a local-only mode and at least one split-computation mode.
[0174] At block 1408, the user device executes the AI service to provide a response to the user query according to the determined operational mode. In the local-only mode, the AI service may be executed entirely on the user device. In the at least one split-computation mode, at least a portion of the AI service may be offloaded to the server.
[0175] In certain configurations, the AI service may include a virtual personal assistant mixed reality service applicable to one or more domains.
[0176] In certain configurations, the user device may further activate the AI service via a detection mechanism, the detection mechanism may include one of: scanning a Near Field Communication (NFC) tag, scanning a Quick Response (QR) code, Bluetooth detection, Wi-Fi Direct, ultra-wideband (UWB) , or infrared.
[0177] In certain configurations, the local resource of the user device may include one or more of:a CPU utilization, a memory utilization, a battery level, or a battery charging status. The network connection resource may include one or more of: a link bandwidth availability, an achievable end-to-end throughput, a Radio Access Technology (RAT) availability, or a radio quality between the user device and the server.
[0178] In certain configurations, the server may include one or more of: a dedicated server in a Customer Premises Network (CPN) , a server managed by a mobile network operator, a server managed by a fixed network operator, a satellite network operator, or a cloud-based server. The server may be selected in real-time based on service requirements and resource availability.
[0179] In certain configurations, assessing the availability of information may include checking a local database on the user device for cached information sufficient to respond to the user query. The user device may further store the response in the local database for future use.
[0180] In certain configurations, the at least one split-computation mode may be selected when the assessment determines that a network connection is capable, and at least one of the local resource is insufficient or the information is not available on the user device. Executing the AI service may include: transmitting data derived from the user query to the server, the data including at least one of an audio stream of the user query or an embedding of the user query; and receiving a response from the server. The response may include at least one of: (a) an audio stream and a video stream of a rendered virtual assistant; or (b) an audio stream and phoneme data for synchronizing lip movements of a virtual assistant rendered on the user device.
[0181] In certain configurations, the at least one split-computation mode may be selected when the assessment determines that a network connection is poor, the local resource is sufficient, and the information is not available on the user device. Executing the AI service may include: generating, by the user device, an embedding based on the user query; transmitting the embedding to the server; receiving a text response from the server; generating, by the user device, an audio stream from the text response; and rendering, by the user device, a virtual assistant with lip movements synchronized to the generated audio stream.
[0182] In certain configurations, the local-only mode may be selected when the assessment determines that the local resource is sufficient and the information is available in a local database on the user device. Executing the AI service may include: retrieving, from the local database, data to respond to the user query; generating, by the user device, a text response based on the retrieved data using a local language model; and generating, by the user device, an audio stream and a corresponding synchronized virtual assistant rendering based on the text response.
[0183] In certain configurations, the user device may further, prior to assessing the one or more conditions: send a service discovery request to a service proxy, the service discovery request including information related to the one or more conditions; and receive, from the service proxy, information for communicating with the server, including load balancing or redirection to one of a plurality of servers.
[0184] In certain configurations, the user device may further receive, from the service proxy, credentials for a Customer Premises Network (CPN) ; and based on a user preference, either connect to the CPN to communicate with the server or refuse to connect to the CPN and communicate with the server via a subscribed mobile network.
[0185] In certain configurations, the user device may further obtain user consent for one or more of:capturing a photograph for identification, storing interaction data locally, or recording interactions for future personalization.
[0186] In certain configurations, the user device may further provide multi-lingual support for the AI service using one or more machine learning models on the user device or the server, and the multi-lingual support may be selected based on user preference.
[0187] FIG. 15 illustrates a flow chart 1500 of another process for enabling a virtual personal assistant mixed reality service using integrated communication and computing. The process may be performed by a user device for providing a virtual assistant service.
[0188] At block 1502, the user device receives a user query including at least one of an image of an object or a voice input related to the object.
[0189] At block 1504, the user device generates an embedding vector based on the user query using a pre-trained model.
[0190] At block 1506, the user device transmits the embedding vector to a server over a network. A data size of the embedding vector may be substantially smaller than a data size of the at least one of the image or the voice input to reduce network traffic.
[0191] At block 1508, the user device receives a response from the server, the response being generated by the server based on the embedding vector and providing information related to the object.
[0192] In certain configurations, the pre-trained model may include a Contrastive Language-Image Pre-Training (CLIP) model. The embedding vector may have a data size of approximately 2 kilobytes.
[0193] In certain configurations, the response received from the server may include an audio stream and a video stream of a rendered virtual assistant. The user device may further play the audio stream and display the video stream of the rendered virtual assistant.
[0194] In certain configurations, the response received from the server may include an audio stream and phoneme data corresponding to the audio stream. The user device may further render an avatar of a virtual assistant with lip movements synchronized to the audio stream based on the phoneme data.
[0195] In certain configurations, the response received from the server may include a text response. The user device may further generate an audio stream from the text response using a local Text-to-Speech (TTS) module; and render an avatar of a virtual assistant with lip movements synchronized to the generated audio stream.
[0196] In certain configurations, the user device may further select or customize an avatar for the virtual assistant service based on user input. The avatar may include one of a default avatar, a male avatar, a female avatar, a child avatar, or a pet avatar.
[0197] FIG. 16 illustrates a flow chart 1600 of a process for preparing a service model for a virtual assistant service.
[0198] At block 1602, textual information for a plurality of items in one or more domains is stored in a first database on a server, the textual information including at least a unique item name for each item.
[0199] At block 1604, a camera of a user device is used to capture a plurality of images for each of the plurality of items.
[0200] At block 1606, a pre-trained model is used to generate an image embedding for each of the captured images.
[0201] At block 1608, the image embeddings and their corresponding unique item names are stored in a second database on the server.
[0202] At block 1610, the service model is fine-tuned by validating retrieved information for sample items and updating image embeddings as needed. The first and second databases may be configured to be queried by a service engine to respond to user queries about the plurality of items.
[0203] In certain configurations, the pre-trained model may include a Contrastive Language-Image Pre-Training (CLIP) model. The one or more domains may include retail, healthcare, manufacturing, or agriculture.
[0204] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed is an illustration of exemplary approaches. Based upon design preferences, it is understood that the specific order or hierarchy of blocks in the processes / flowcharts may be rearranged. Further, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in a sample order, and are not meant to be limited to the specific order or hierarchy presented.
[0205] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language claims, wherein reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more. ” The word “exemplary” is used herein to mean “serving as an example, instance, or illustration. ” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects. Unless specifically stated otherwise, the term “some” refers to one or more. Combinations such as “at least one of A, B, or C, ” “one or more of A, B, or C, ” “at least one of A, B, and C, ” “one or more of A, B, and C, ” and “A, B, C, or any combination thereof” include any combination of A, B, and / or C, and may include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as “at least one of A, B, or C, ” “one or more of A, B, or C, ” “at least one of A, B, and C, ” “one or more of A, B, and C, ” and “A, B, C, or any combination thereof” may be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combinations may contain one or more member or members of A, B, or C. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims. The words “module, ” “mechanism, ” “element, ” “device, ” and the like may not be a substitute for the word “means. ” As such, no claim element is to be construed as a means plus function unless the element is expressly recited using the phrase “means for.”
Claims
1.A method performed by a user device for providing an Artificial Intelligence (AI) service, comprising:receiving a user query for the AI service;assessing one or more conditions, wherein the one or more conditions comprise at least one of: a local resource of the user device, a network connection resource between the user device and a server, or an availability of information on the user device to respond to the user query;determining, based on the assessment, an operational mode from a plurality of operational modes for executing the AI service, the plurality of operational modes comprising a local-only mode and at least one split-computation mode; andexecuting the AI service to provide a response to the user query according to the determined operational mode, wherein in the local-only mode, the AI service is executed entirely on the user device, and wherein in the at least one split-computation mode, at least a portion of the AI service is offloaded to the server.2.The method of claim 1, wherein the AI service comprises a virtual personal assistant mixed reality service applicable to one or more domains.3.The method of claim 1, further comprising activating the AI service on the user device via a detection mechanism, the detection mechanism comprising one of: scanning a Near Field Communication (NFC) tag, scanning a Quick Response (QR) code, Bluetooth detection, Wi-Fi Direct, ultra-wideband (UWB) , or infrared.4.The method of claim 1, wherein the local resource of the user device comprises one or more of: a CPU utilization, a memory utilization, a battery level, or a battery charging status, and wherein the network connection resource comprises one or more of: a link bandwidth availability, an achievable end-to-end throughput, a Radio Access Technology (RAT) availability, or a radio quality between the user device and the server.5.The method of claim 1, wherein the server comprises one or more of: a dedicated server in a Customer Premises Network (CPN) , a server managed by a mobile network operator, a server managed by a fixed network operator, a satellite network operator, or a cloud-based server, and wherein the server is selected in real-time based on service requirements and resource availability.6.The method of claim 1, wherein assessing the availability of information comprises checking a local database on the user device for cached information sufficient to respond to the user query, and wherein the method further comprises storing the response in the local database for future use.7.The method of claim 1, wherein the at least one split-computation mode is selected when the assessment determines that a network connection is capable, and at least one of the local resource is insufficient or the information is not available on the user device, and wherein executing the AI service comprises:transmitting data derived from the user query to the server, the data comprising at least one of an audio stream of the user query or an embedding of the user query; andreceiving a response from the server, wherein the response comprises at least one of:(a) an audio stream and a video stream of a rendered virtual assistant; or(b) an audio stream and phoneme data for synchronizing lip movements of a virtual assistant rendered on the user device.8.The method of claim 1, wherein the at least one split-computation mode is selected when the assessment determines that a network connection is poor, the local resource is sufficient, and the information is not available on the user device, and wherein executing the AI service comprises:generating, by the user device, an embedding based on the user query;transmitting the embedding to the server;receiving a text response from the server;generating, by the user device, an audio stream from the text response; andrendering, by the user device, a virtual assistant with lip movements synchronized to the generated audio stream.9.The method of claim 1, wherein the local-only mode is selected when the assessment determines that the local resource is sufficient and the information is available in a local database on the user device, and wherein executing the AI service comprises:retrieving, from the local database, data to respond to the user query;generating, by the user device, a text response based on the retrieved data using a local language model; andgenerating, by the user device, an audio stream and a corresponding synchronized virtual assistant rendering based on the text response.10.The method of claim 1, further comprising, prior to assessing the one or more conditions:sending a service discovery request to a service proxy, the service discovery request comprising information related to the one or more conditions; andreceiving, from the service proxy, information for communicating with the server, including load balancing or redirection to one of a plurality of servers.11.The method of claim 10, further comprising:receiving, from the service proxy, credentials for a Customer Premises Network (CPN) ; andbased on a user preference, either connecting to the CPN to communicate with the server or refusing to connect to the CPN and communicating with the server via a subscribed mobile network.12.The method of claim 1, further comprising:obtaining user consent for one or more of: capturing a photograph for identification, storing interaction data locally on the user device, or recording interactions for future personalization.13.The method of claim 1, further comprising providing multi-lingual support for the AI service using one or more machine learning models on the user device or the server, the multi-lingual support selected based on user preference.14.A method performed by a user device for providing a virtual assistant service, comprising:receiving a user query comprising at least one of an image of an object or a voice input related to the object;generating an embedding vector based on the user query using a pre-trained model;transmitting the embedding vector to a server over a network, wherein a data size of the embedding vector is substantially smaller than a data size of the at least one of the image or the voice input to reduce network traffic; andreceiving a response from the server, the response being generated by the server based on the embedding vector and providing information related to the object.15.The method of claim 14, wherein the pre-trained model comprises a Contrastive Language-Image Pre-Training (CLIP) model, and wherein the embedding vector has a data size of approximately 2 kilobytes.16.The method of claim 14, wherein the response received from the server comprises an audio stream and a video stream of a rendered virtual assistant, and wherein the method further comprises:playing the audio stream and displaying the video stream of the rendered virtual assistant.17.The method of claim 14, wherein the response received from the server comprises an audio stream and phoneme data corresponding to the audio stream, and wherein the method further comprises:rendering an avatar of a virtual assistant with lip movements synchronized to the audio stream based on the phoneme data.18.The method of claim 14, wherein the response received from the server comprises a text response, and wherein the method further comprises:generating an audio stream from the text response using a local Text-to-Speech (TTS) module; andrendering an avatar of a virtual assistant with lip movements synchronized to the generated audio stream.19.The method of claim 14, further comprising:selecting or customizing an avatar for the virtual assistant service based on user input, wherein the avatar comprises one of a default avatar, a male avatar, a female avatar, a child avatar, or a pet avatar.20.A method for preparing a service model for a virtual assistant service, comprising:storing, in a first database on a server, textual information for a plurality of items in one or more domains, the textual information comprising at least a unique item name for each item;capturing, using a camera of a user device, a plurality of images for each of the plurality of items;generating an image embedding for each of the captured images using a pre-trained model;storing the image embeddings and their corresponding unique item names in a second database on the server; andfine-tuning the service model by validating retrieved information for sample items and updating image embeddings as needed, wherein the first and second databases are configured to be queried by a service engine to respond to user queries about the plurality of items.21.The method of claim 20, wherein the pre-trained model comprises a Contrastive Language-Image Pre-Training (CLIP) model, and wherein the one or more domains comprise retail, healthcare, manufacturing, or agriculture.22.A user device for providing an Artificial Intelligence (AI) service, comprising:a memory; andat least one processor coupled to the memory and configured to:receive a user query for the AI service;assess one or more conditions, wherein the one or more conditions comprise at least one of: a local resource of the user device, a network connection resource between the user device and a server, or an availability of information on the user device to respond to the user query;determine, based on the assessment, an operational mode from a plurality of operational modes for executing the AI service, the plurality of operational modes comprising a local-only mode and at least one split-computation mode; andexecute the AI service to provide a response to the user query according to the determined operational mode, wherein in the local-only mode, the AI service is executed entirely on the user device, and wherein in the at least one split-computation mode, at least a portion of the AI service is offloaded to the server.23.A computer-readable medium storing computer executable code for providing an Artificial Intelligence (AI) service performed by a user device, comprising code to:receive a user query for the AI service;assess one or more conditions, wherein the one or more conditions comprise at least one of: a local resource of the user device, a network connection resource between the user device and a server, or an availability of information on the user device to respond to the user query;determine, based on the assessment, an operational mode from a plurality of operational modes for executing the AI service, the plurality of operational modes comprising a local-only mode and at least one split-computation mode; andexecute the AI service to provide a response to the user query according to the determined operational mode, wherein in the local-only mode, the AI service is executed entirely on the user device, and wherein in the at least one split-computation mode, at least a portion of the AI service is offloaded to the server.24.A user device for providing a virtual assistant service, comprising:a memory; andat least one processor coupled to the memory and configured to:receive a user query comprising at least one of an image of an object or a voice input related to the object;generate an embedding vector based on the user query using a pre-trained model;transmit the embedding vector to a server over a network, wherein a data size of the embedding vector is substantially smaller than a data size of the at least one of the image or the voice input to reduce network traffic; andreceive a response from the server, the response being generated by the server based on the embedding vector and providing information related to the object.25.A computer-readable medium storing computer executable code for providing a virtual assistant service performed by a user device, comprising code to:receive a user query comprising at least one of an image of an object or a voice input related to the object;generate an embedding vector based on the user query using a pre-trained model;transmit the embedding vector to a server over a network, wherein a data size of the embedding vector is substantially smaller than a data size of the at least one of the image or the voice input to reduce network traffic; andreceive a response from the server, the response being generated by the server based on the embedding vector and providing information related to the object.26.An apparatus for preparing a service model for a virtual assistant service, comprising:a memory; andat least one processor coupled to the memory and configured to:store, in a first database on a server, textual information for a plurality of items in one or more domains, the textual information comprising at least a unique item name for each item;capture, using a camera of a user device, a plurality of images for each of the plurality of items;generate an image embedding for each of the captured images using a pre-trained model;store the image embeddings and their corresponding unique item names in a second database on the server; andfine-tune the service model by validating retrieved information for sample items and updating image embeddings as needed, wherein the first and second databases are configured to be queried by a service engine to respond to user queries about the plurality of items.27.A computer-readable medium storing computer executable code for preparing a service model for a virtual assistant service, comprising code to:store, in a first database on a server, textual information for a plurality of items in one or more domains, the textual information comprising at least a unique item name for each item;capture, using a camera of a user device, a plurality of images for each of the plurality of items;generate an image embedding for each of the captured images using a pre-trained model;store the image embeddings and their corresponding unique item names in a second database on the server; andfine-tune the service model by validating retrieved information for sample items and updating image embeddings as needed, wherein the first and second databases are configured to be queried by a service engine to respond to user queries about the plurality of items.
Citation Information
Patent Citations
Techniques for invoking services based on patterns in context determined using context mining
US20040225654A1
Intelligent layer to power cross platform, edge-cloud hybrid artificial intelligence services
US20210297494A1
Systems and methods for optimized execution of program operations on cloud-based services
US20210349764A1
Group action fulfillment across multiple user devices
US20220067767A1
METHODS, APPARATUS AND SYSTEMS FOR DECOMPOSING MOBILE APPLICATIONS INTO MICRO-SERVICES (MSs) AT RUNTIME FOR DISTRIBUTED EXECUTION
US20240086250A1