Edge computing interactive system capable of receiving data from external server via network

The edge computing interactive system addresses the lack of natural interaction and privacy concerns by hosting a local large language model to process user inputs and provide context-aware responses without transmitting user data, enhancing user experience and compliance.

WO2026075898A1PCT designated stage Publication Date: 2026-04-09AVANTI R&D INC
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing systems lack the ability for passengers in autonomous vehicles and users of edge devices to interact naturally, provide context-specific assistance, and ensure privacy in data transmission, especially in environments with limited internet connectivity, raising compliance and privacy concerns.

Method used

An edge computing interactive system that hosts a local large language model, processes user inputs locally without transmitting user-specific data, and receives real-time environmental data to provide context-aware responses, using input devices like cameras and microphones, and integrates with vehicle systems for seamless interaction.

Benefits of technology

Enhances user experience with intuitive and secure interactions, simplifies content creation and distribution, and maintains privacy by processing data locally, ensuring compliance with regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025048279_09042026_PF_FP_ABST
    Figure US2025048279_09042026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure describes an edge computing interactive system receiving data from an external server via a network. This system incorporates a compute module hosting a local large language model (LLM), executing speech-to-text and text-to-speech functions, and tuning the LLM with custom domain datasets. It features input devices for user capture, output devices for response delivery, and artificial intelligence agents for real-time data processing and response coordination. The system enables natural, contextual interaction by processing user queries locally, avoiding user data upload to the external server. Activation occurs via gaze detection or wake word. The LLM analyzes speech tone, accesses vehicle data, and clears conversation memory on external triggers. This approach offers enhanced privacy and responsive user interaction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AVAN.004PCT PATENT

[0002] EDGE COMPUTING INTERACTIVE SYSTEM CAPABLE OE RECEIVING DATA FROM EXTERNAL SERVER VIA NETWORK

[0003] CROSS-REFERENCE TO RELATED APPLICATION

[0004] This application claims priority to U.S. Provisional Patent Application No. 63 / 701,768, filed October 1, 2024, the entire contents of which are hereby incorporated by reference herein in their entirety and for all purposes.

[0005] BACKGROUND

[0006] The present disclosure generally relates to interactive systems, particularly to edge computing systems enabling contextual interactions in diverse environments. More particularly, to an edge computing interactive system comprising a compute module configured to host a large language model and execute speech-to-text and text-to-speech functionalities.

[0007] With the rise of autonomous vehicles, including public transit buses, commercial shuttles and taxis, passengers may experience a lack of understanding about the vehicle’s driving decisions and the surrounding environment. This uncertainty can lead to decreased comfort and trust in autonomous transportation. They cannot ask the driver about the current situations since there are no drivers present. Existing systems do not provide a way for passengers to interact with the vehicle to ask questions or receive explanations about specific maneuvers or locations. Additionally, leveraging cloud-based Large Language Models (LLMs) may complicate compliance with data-protection regulations (e.g., GDPR) due to the transmission of user content off-device.

[0008] Class 8 trucks, which are heavy-duty vehicles used for long-haul transportation, are increasingly adopting zero-emission powertrains that reduce cabin noise due to the absence of diesel engines. Drivers can benefit from an interactive co-pilot system that assists with navigation, vehicle diagnostics, compliance information, and other operational support. Current systems lack the ability to provide such assistance in a natural, conversational manner, especially in areas with limited internet connectivity.

[0009] Automated kiosks in restaurants typically offer a limited, menu-driven interface where customers can place orders by selecting predefined options on a touch screen. While some may incorporate basic voice recognition, they generally lack advanced conversational abilities and AVAN.004PCT PATENT contextual understanding. This limits the customer experience and does not accommodate open- ended inquiries or complex dialogue.

[0010] Companies often use various software tools, network infrastructures, confidential data, and physical hardware or machinery utilized by employees throughout the organization. Warehouses and manufacturing plants, in particular, house large machines and equipment essential for operations. Accessing support, know-how, and manuals for these internal tools and machinery can be challenging, especially when dealing with complex processes or when internet connectivity is limited.

[0011] In industries such as construction, mining, agriculture, oil and gas, maritime operations, and submarines, workers operate heavy machinery and equipment in environments where internet connectivity is often unreliable or unavailable. Access to diagnostics, real-time user manuals, and operational support is crucial for safety, efficiency, and minimizing downtime.

[0012] Existing support systems may not provide immediate, context-specific assistance and may raise privacy concerns if they rely on external servers. There is a need for a localized solution that offers real-time support without depending on internet connectivity.

[0013] Conventional in-vehicle information systems rely on real-time internet connections, posing problems in terms of personal information protection. Furthermore, existing navigation systems provide insufficient information tailored to the user's actual needs and circumstances.

[0014] SUMMARY

[0015] One aspect of the present disclosure is to provide an interactive dialogue system for vehicles that offers a more natural and nuanced understanding of user intent. This system leverages advanced input methods to facilitate seamless interaction, enhancing the user experience within diverse working environments. The system provides intelligent, context-aware responses, delivering value-added interactions for users.

[0016] Another aspect of the present disclosure is to provide a content generation system that is efficient and multifunctional, enabling administrators to create local information podcasts easily. This tool simplifies the content creation process, allowing for the widespread dissemination of relevant information to users. The system also supports the spontaneous distribution of vital updates, such as transportation delays, to ensure users receive timely information. AVAN.004PCT PATENT

[0017] Yet another aspect of the present disclosure is to enhance user privacy and conversation continuity by implementing robust memory management for the dialogue system. By automatically clearing conversational data under specific conditions, the system safeguards user information and prepares for new interactions, providing a secure and personalized experience for users. The system provides accurate answers to questions and summarizes or collects impressions as passenger feedback.

[0018] According to one aspect of the present disclosure, an edge computing interactive system receives data from an external server via a network. This system comprises a compute module configured to host a local large language model, execute speech-to-text and text-to-speech functionalities, and tune the local large language model using custom domain datasets based on received data. The system further includes one or more input devices communicatively coupled to the compute module to capture user input, and one or more output devices communicatively coupled to the compute module to deliver responses generated by the compute module. Moreover, one or more artificial intelligence agents are configured to process real-time data and coordinate the delivery of responses via the one or more output devices, wherein the system facilitates natural, contextual interaction within a working environment by processing user queries and generating responses locally without uploading user-specific data to the external server. That means, this system generates responses locally without transmitting user utterance audio or transcripts, etc. off-device while specific telemetry (e.g., location, route, distination, content IDs, delivery success / failure) may be transmitted or stored to or in the external server 500. As used herein, the term “user-specific data” refers to data obtained from one or more input devices (e.g., audio, video, or the like) that is associated with a particular user, and the transmission of which may be subject to restrictions under the GDPR or similar data protection regulations. According to another aspect of the present disclosure, the edge computing interactive system of the present disclosure is configured such that the compute module receives real-time environmental data comprising podcast content from the external server, where the real-time environmental data reflects current conditions or events near the location of the edge computer interactive system. The compute module also receives a text file corresponding to the podcast content from the external server. In addition, the compute module is configured not to send out information included in the user queries to the external server, thus maintaining privacy. AVAN.004PCT PATENT

[0019] The one or more input devices include a camera configured to capture visual input of a user and detect facial expressions and gestures of the user, while a microphone captures audio input from the user with variable-length voice input based on an amplitude threshold determined by ambient noise level. The compute module further comprises an activation system configured to initiate a dialogue based on at least one of a detected gaze of the user towards a specific area or a detected wake word, and the local large language model analyzes the tone of the user's speech to determine the intent of the user's statement, accessing local system time information, calendar information, and environmental information obtained from at least one sensor, and clearing a conversation memory upon receiving an external trigger.

[0020] The edge computing interactive system's local large language model analysis of the user's speech tone distinguishes between a question and an information sharing statement. The local large language model also plays the podcast content as background audio when the system detects a period of user inactivity, provided the podcast content includes a background audio identifier. When the edge computing interactive system is installed in a bus for passenger transportation, its local large language model includes a conversation memory for storing user interactions, and the compute module receives an external trigger from a bus system indicating that a passenger has disembarked and automatically clears the conversation memory in response to the external trigger. Alternatively, the compute module receives an external trigger indicating that a passenger has ended an interaction session or has disembarked from the bus and automatically clears the conversation memory in response to that trigger.

[0021] When the edge computing interactive system is installed in a bus for passenger transportation, the compute module operates the local large language model primarily offline, receives urgent information podcasts related to transportation delays, accidents, and bad weather from the external server (generated automatically based on the bus's location information), and presents the received urgent information podcasts to the passengers. The system also receives podcast content and related information from communication units installed on infrastructure, such as utility poles, equipped with Vehicle-to-E very thing (V2X) capabilities, establishing peer-to-peer communication when within range, and the compute module processes and presents the received podcast content and related information to users.

[0022] According to another aspect of the present disclosure, a method for operating an edge computing interactive system comprises tuning a local large language model hosted on a compute AVAN.004PCT PATENT module with custom domain datasets based on received data, capturing user input via one or more input devices coupled to the compute module, and processing the user input using the local large language model to generate a response. The method further involves delivering the generated response through one or more output devices coupled to the compute module, processing real-time data, and coordinating delivery of responses through the one or more output devices using one or more artificial intelligence agents. This method facilitates natural, contextual interaction within a working environment by processing user queries and generating responses locally without uploading user-specific data to an external server. The method can further include receiving realtime environmental data comprising podcast content and a corresponding text file from the external server, with tuning the local large language model using custom domain datasets based on the received text file. It also includes capturing visual input of a user and detecting facial expressions and gestures with a camera, capturing variable-length voice input based on an amplitude threshold determined by ambient noise level, and initiating a dialogue based on a detected gaze or wake word. The method additionally analyzes the tone of the user's speech to determine intent, accesses local system time and environmental information, and clears conversation memory upon an external trigger, further comprising playing podcast content as background audio during user inactivity, if identified.

[0023] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium stores instructions that, when executed by a processor of an edge computing interactive system, cause the system to perform operations. These operations include tuning a local large language model hosted on a compute module with custom domain datasets based on received data, capturing user input via one or more input devices coupled to the compute module, and processing the user input using the local large language model to generate a response. The operations also involve delivering the generated response through one or more output devices coupled to the compute module, processing real-time data, and coordinating delivery of responses through the one or more output devices using one or more artificial intelligence agents. This enables natural, contextual interaction within a working environment by processing user queries and generating responses locally without uploading user-specific data to an external server. The operations further include receiving real-time environmental data comprising podcast content and a corresponding text file from an external server, and playing the podcast content as background audio when detecting user inactivity, if the podcast content includes a background audio identifier. AVAN.004PCT PATENT

[0024] According to another aspect of the present disclosure, a server for an edge computing interactive system comprises a processor, a memory coupled to the processor, and a server-side large language model stored in the memory and executed by the processor. This server-side large language model is configured to generate content based on specified criteria, generate a text file related to the generated content, and prepare the content and the text file for transmission to an edge computing device. A network interface is configured to transmit the content and the text file to the edge computing device. The server is configured to provide custom domain datasets for tuning a local large language model on the edge computing device, and operate without receiving information included in user queries from the edge computing device. The content generated by the server-side large language model is podcast content, and the server-side large language model generates podcast content based on user preferences received from the edge computing device. The text file contains the content of the generated podcast, includes information indicating whether the podcast is usable as background music (BGM), and the server-side large language model updates the podcast content and text file based on real-time information. The server-side large language model generates podcast content in response to a request from the edge computing device or periodically generates and transmits updated podcast content and text files. The server also receives location information from the edge computing device and tailors content based on the location. The server-side large language model generates podcast content related to traffic updates, weather forecasts, local news, and points of interest near the location of the edge computing device. The server also receives feedback from users and adjusts generated content based on the received feedback, pushing podcast content and related information to communication units installed on infrastructure, such as utility poles, equipped with Vehicle-to-Everything (V2X) capabilities, for transmission to the edge computing device via peer-to-peer communication.

[0025] It should be noted that the real-time environmental data utilized by the computing module may include data either obtained from the sensor connected to the computing module within a local device or may obtain from the external server via communication link, which does not necessary means immediately before the acquisition of the data but can be a few minutes, hours or days after the event occurrence depending on the nature of the data. For example, the data relating to the malfunction of a vehicle may be obtained and processed in a few seconds, a traffic accident information can be in a few minutes, the weather change information can be in a few hours, and a local event information can be in a few days or weeks. In the case of the data prepared by the AVAN.004PCT PATENT external server for the vehicle, the real-time environmental data can be an up-to-date local information based on the GPS feedback from the vehicle, which can be years old if there is no update to a specific POI (Point of Interest).

[0026] This approach offers distinct advantages: it creates a more intuitive and responsive in- vehicle dialogue system, simplifies content creation and distribution for administrators, and provides secure and personalized interactions for users.

[0027] The foregoing paragraphs have been provided by way of general introduction and are not intended to limit the scope of the following claims. The described embodiments, together with further advantages, will be best understood by reference to the following detailed description taken in conjunction with the accompanying drawings.

[0028] BRIEF DESCRIPTION OF DRAWINGS

[0029] Embodiments of various inventive features will now be described with reference to the following drawings. Throughout the drawings, reference numbers may be re-used to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the disclosure. To easily identify the discussion of any particular element or act, the most significant digit(s) in a reference number typically refers to the figure number in which that element is first introduced.

[0030] FIG. 1 is a system block diagram for in-vehicle dialogue with LLM integration with vehicle / Machinery / City inputs.

[0031] FIG. 2 is a system block diagram for in-vehicle dialogue system utilizing in-vehicle data.

[0032] FIG. 3 is a simplified block diagram of an in-vehicle compute module system.

[0033] FIG. 4 is a chart illustrating system for automatic content generation and playback in vehicles.

[0034] FIG. 5 is a cloud-based podcast generation at Al server and edge-based playback system.

[0035] FIG. 6 :is a flowchart for in-vehicle dialogue system with gaze and voice input processing.

[0036] FIG. 7 is a sequence diagram for server-triggered emergency content delivery to in-vehicle system.

[0037] FIG. 8 is flowchart for in-vehicle dialogue system with gaze and voice input processing.

[0038] FIG. 9 is flowchart for a speech acquisition process shown in FIG. 8. AVAN.004PCT PATENT

[0039] DETAILED DESCRIPTION

[0040] Aspects of the present disclosure are best understood with embodiments by reference to the description set forth herein. All the aspects described herein will be better appreciated and understood when considered in conjunction with the following descriptions of the embodiments. It should be understood, however, that the following descriptions, while indicating preferred aspects and numerous specific details thereof, are given by way of illustration only and should not be treated as limitations. Changes and modifications may be made within the scope herein without departing from the spirit and scope thereof, and the present disclosure herein includes all such modifications.

[0041] FIG. 1 illustrates a simplified in-vehicle dialogue system incorporating core large language model (LLM) components to facilitate natural and contextual user interaction. The system integrates various input devices, including a Camera 100, an Audio In 106 module, and vehicle data sources such as an OBDII / J1939 CAN bus 110, Navigation & Routing 112, Modbus 114 and V2X / IoT Smart City Data 116, all feeding into a Compute Module 118. The Compute Module 118 orchestrates dialogue activation, speech processing, data fetching via LLM Agents 130, and response generation by a Domain Tuned LLM 134, delivering outputs through an Audio Out 142 to a Speaker 144 and a Vehicle HMI (Human Machine Interface) 146 comprising a Display 148, Interior lighting 150, and controls for Doors, windows, alarm, AC controls 152. The overall architecture is designed to enhance the in-vehicle user experience by providing intelligent and responsive interactions.

[0042] The Camera 100 serves as a primary visual input device, capturing images and video of the vehicle's interior and occupants. This camera is equipped with image processing capabilities for object detection and gaze detection, providing visual cues about user attention and intent. The Camera 100 contributes to the system's ability to initiate interaction more intuitively than traditional wake word methods (as well-known as “Hi Siri”, “Alexa” or “Hey Google”, etc.), or in conjunction with said methods. As an alternative, an infrared camera or a stereoscopic camera system could be employed to improve depth perception and facial feature recognition, potentially enhancing the accuracy of gaze and facial expression analysis, especially under varying lighting conditions. Such alternative cameras could also integrate proximity sensors to detect user presence without direct visual contact. AVAN.004PCT PATENT

[0043] A Gaze Detection model 102 operates on the visual input from the Camera 100, analyzing the user's eye movements and head posture to ascertain their focus of attention. This model is configured to detect specific patterns, such as a prolonged gaze towards the system's interface or a designated area within the vehicle, which serves as a non-verbal trigger for dialogue initiation. Conversely, the Gaze Detection model 102 also monitors when a user's gaze is directed away from the system, influencing whether the Computer Module 118 should pause its speech to avoid interruption. For example, the Gaze Detection model 102 may utilize machine learning algorithms trained on extensive datasets of eye-tracking data, capable of distinguishing between casual glances and focused attention. Alternative embodiments could involve pupil tracking technology for even more precise gaze vector determination or head pose estimation combined with eyetracking to provide a more robust understanding of user attention. A Gaze Detection 103 send a trigger signal to the Computer Modules 118.

[0044] The decision point Gaze directed at Wake 122 within the Compute Module 118 represents a logical check performed by the system based on the output of the Gaze Detection 103. If the user's gaze is detected as being directed towards the system or a predefined interactive zone, this signal acts as an activation cue, potentially overriding or complementing a wake word activation come from Audio In 106. This functionality allows for a more natural interaction flow, where the system anticipates user intent based on visual engagement. An additional embodiment could involve integrating a facial expression analysis module with the gaze directed at Wake 122 decision point, allowing for activation based on a combination of gaze and an expression of intent, such as a questioning look, thereby further refining the system's responsiveness.

[0045] Audio In 106 is a microphone system designed to capture user voice commands, queries, and conversational input within the vehicle environment. This input device is equipped with advanced noise cancellation capabilities to isolate user speech from ambient vehicle noise. The Audio In 106 system plays a role in traditional wake word activation and is also crucial for receiving user utterances once a dialogue has commenced. As an alternative or additional embodiment, an array of microphones could be deployed throughout the vehicle cabin to provide spatial audio detection, enabling the system to localize the speaker and potentially filter out interference from other passengers or ambient sounds more effectively. Such an array could also facilitate beamforming techniques to focus on the user's voice, further enhancing speech clarity for subsequent processing. AVAN.004PCT PATENT

[0046] Vehicle / Machinery / City 108 encompasses various external and internal data sources that provide environmental and operational context to the in-vehicle system. This broad category includes information pertaining to the vehicle's operational status, surrounding machinery, and data relevant to the urban or travel environment. The data from Vehicle / Machinery / City 108 is instrumental in enabling the Domain Tuned LLM 134 to generate context-sensitive and value- added responses, improving the relevance and utility of the interactive system. Alternative sources of contextual data could include integration with smart infrastructure, such as roadside units providing real-time traffic updates or pedestrian movement patterns, offering an enriched understanding of the immediate operational environment.

[0047] The OBDII / JI 939 CAN bus 110 represents a standardized communication interface that provides access to the vehicle's internal diagnostic and operational data. This data includes parameters such as vehicle speed, engine revolutions per minute, fuel levels, tire pressure, and various system statuses, which are invaluable for contextualizing user interactions. The Compute Module 118 accesses this information to inform the LLM Agents (data fetch) 130, enabling them to retrieve relevant real-time vehicle status information. For instance, if a user asks about the vehicle's performance, the OBDII / J1939 CAN bus 110 can supply the necessary data. An alternative to the OBDII / J1939 CAN bus 110 could be a proprietary vehicle data bus or a manufacturer-specific diagnostic port, which might offer an even richer dataset of vehicle internals.

[0048] Navigation & Routing 112 provides geographical positioning information, route guidance, and points of interest data, all of which are vital for location-aware responses. This system includes GPS data, map information, and algorithms for route calculation and optimization. The Compute Module 118 utilizes information from Navigation & Routing 112 to understand the vehicle's current location, planned trajectory, and proximity to various destinations, enabling the LLM to offer highly relevant suggestions or information.

[0049] An alternative embodiment could involve integrating with a cloud-based navigation service that continuously updates map data and traffic conditions, providing more dynamic and precise location-based services. The data update can be done via 4G / 5G network; however, it may require an additional subscription. However, can eliminate real-time update by utilizing a push data distribution technology utilizing V2X infrastructure. This integration could also incorporate predictive routing based on historical traffic patterns and user preferences. AVAN.004PCT PATENT

[0050] A Modbus 114 represents a serial communication protocol often used for connecting industrial electronic devices, which in a vehicle context, could facilitate data exchange with specific, specialized vehicle sub-systems or external peripherals not typically covered by the CAN bus. This could include control units for advanced driver-assistance systems or aftermarket installations. The Compute Module 118 leverages the Modbus 114 to gather supplementary data that may be pertinent to the overall context of the in-vehicle dialogue system. An alternative communication protocol, such as Ethernet or FlexRay, could be used depending on the specific requirements for speed, reliability, and data volume of the connected devices, offering higher bandwidth for complex data streams. V2X / loT Smart City Data 116 refers to external data streams received from Vehicle-to-Everything communication systems and smart city infrastructure. This data can include real-time traffic updates, hazard warnings, parking availability, environmental conditions, and information about local events or points of interest from smart city sensors. The Compute Module 118 processes V2X / IoT Smart City Data 116 to enrich the contextual understanding of the vehicle's environment, enabling the LLM to provide proactive and highly relevant information to the user. An alternative or additional embodiment could involve direct communication with pedestrian devices (V2P) or infrastructure-to-vehicle (I2V) systems that provide specific, localized information, thereby enhancing the system's awareness of its immediate surroundings and potential interactions.

[0051] The Compute Module 118 serves as the central processing unit of the in-vehicle dialogue system, integrating all input data, executing the various software components, and managing output delivery. This module typically comprises a processor (CPU / GPU), memory, and storage, configured to handle the computational demands of an LLM, speech processing, and Al agent operations. The Compute Module 118 orchestrates the flow from activation to response generation, ensuring smooth and responsive interaction. An alternative embodiment could involve a distributed computing architecture where certain tasks, such as initial audio processing or simple gaze detection, are offloaded to dedicated microcontrollers or edge Al accelerators, thereby reducing the processing load on the central Compute Module 118 and improving overall system efficiency.

[0052] Wait for user interaction 120 represents a quiescent state where the system is passively monitoring for activation cues. In this state, the system utilizes low-power processing of input streams, such as audio for wake word detection and visual input for gaze detection, to conserve AVAN.004PCT PATENT resources while maintaining responsiveness. This state signifies that the system is ready to engage but has not yet detected an explicit or implicit trigger from the user. An alternative approach could involve a continuous monitoring mode where the system constantly processes subtle cues, like changes in user posture or slight vocalizations, using advanced machine learning models to predict intent to interact even before explicit activation cues are presented.

[0053] Wake 122 signifies the transition from the quiescent "Wait for user interaction" state 120 to an active listening state. This transition is triggered upon the detection of an activation cue, which can be either a specific wake word spoken by the user, a detected gaze direction towards the system, or a designated area. The Wake 122 mechanism ensures that the system is only fully active when user interaction is intended, optimizing computational resources and minimizing unwarranted responses. An additional embodiment for Wake 122 could incorporate multimodal activation, requiring a combination of both a wake word and a sustained gaze, which would reduce false positives and ensure a more deliberate user activation.

[0054] Start Listening 124 is the state entered immediately after the system is Woken 122. In this state, the system begins actively capturing audio input from the Audio In 106 microphone, preparing to process the user's speech. The duration and sensitivity of the Start Listening 124 phase are adaptable, considering ambient noise levels and the user's speech patterns. This mechanism ensures that the system accurately captures the entirety of the user's query without truncation. An alternative or additional embodiment could involve a "smart listening" feature that dynamically adjusts the listening duration based on the detected emotional state of the user or the complexity of the ongoing conversation, providing more flexibility in dialogue management.

[0055] Speech To Text 126 is a critical component that converts the captured audio input from the user into textual data. This process involves advanced speech recognition algorithms that identify phonetic sounds, assemble them into words, and then form coherent sentences. The accuracy of the Speech To Text 126 conversion directly impacts the effectiveness of the subsequent LLM processing. For example, the Speech To Text 126 system may use deep learning models trained on large datasets of spoken language to achieve high accuracy even in noisy in-vehicle environments. Alternative embodiments could include speaker diarization capabilities, allowing the system to distinguish between multiple speakers and attribute speech to the correct individual, thereby enhancing personalized interaction in a multi-passenger setting. AVAN.004PCT PATENT

[0056] Fetch vehicle situational data 128 is a process initiated once the system begins processing user input, wherein relevant real-time data from various vehicle systems and environmental sources is retrieved. This includes data from the OBDII / J1939 CAN bus 110, Navigation & Routing 112, and V2X / loT Smart City Data 116, as well as local system time and calendar information. The Fetch vehicle situational data 128 ensures that the subsequent LLM processing is informed by the current operational context of the vehicle and its surroundings, leading to more accurate and contextually appropriate responses. An additional embodiment might involve predictive data fetching, where the system anticipates potential user queries based on driving patterns or scheduled events, pre-loading relevant situational data to reduce response latency.

[0057] LLM Agents (data fetch) 130 are specialized artificial intelligence modules responsible for querying and retrieving specific pieces of information from various internal and external data sources based on the context of the user's query. These agents act as intelligent interfaces between the Domain Tuned LLM 134 and the diverse data streams like Fetch vehicle situational data 128 and Local DB 132. The LLM Agents (data fetch) 130 parse the user's intent derived from the LLM, translate it into data requests, and integrate the retrieved data back into the LLM's processing stream. An alternative embodiment could involve a hierarchical agent system where a meta-agent delegates data-fetching tasks to specialized sub-agents, each optimized for different data types or external APIs, thereby providing a more modular and scalable data retrieval framework.

[0058] Local DB 132 represents a local database or knowledge base stored within the Compute Module 118. This database can contain pre-loaded information, frequently accessed data, or context accumulated during ongoing conversations. The Local DB 132 allows the system to operate effectively even in environments with limited or no network connectivity, providing rapid access to pertinent information. For example, the Local DB 132 may store information about local points of interest, previous user preferences, or pre-leamed text data associated with pushed podcast content from an external Al server. An additional embodiment could involve a dynamic local caching mechanism that prioritizes and stores data based on predicted user needs or recent search history, further optimizing offline responsiveness.

[0059] The Domain Tuned LLM 134 is the core large language model hosted on the Compute Module 118, specifically trained or fine-tuned with custom domain datasets relevant to the in- vehicle environment and typical user interactions. This LLM processes the textual input from Speech To Text 126, combines it with contextual data fetched by the LLM Agents (data fetch) AVAN.004PCT PATENT

[0060] 130, and generates intelligent, natural language responses. The Domain Tuned LLM 134 is configured to analyze conversational tone, determine user intent (e.g., question versus information sharing), and manage conversation memory, enhancing the quality and relevance of the dialogue. Alternative embodiments could involve a modular LLM architecture, where different specialized LLM modules are invoked based on the detected intent or domain of the query, allowing for more efficient resource allocation and specialized knowledge processing.

[0061] Text To Speech 136 is the component responsible for converting the natural language responses generated by the Domain Tuned LLM 134 into synthesized audible speech. This process involves selecting appropriate vocal tones, inflections, and pacing to deliver a natural-sounding response to the user. The quality of the Text To Speech 136 output significantly influences the perceived naturalness and user-friendliness of the dialogue system. As an alternative, a personalized Text To Speech 136 engine could be developed that adapts its voice characteristics based on user preferences or even mimics known voices to create a more familiar and engaging interaction.

[0062] LLM Agents 138 are specialized artificial intelligence modules that coordinate the delivery of responses generated by the Domain Tuned LLM 134 through various output channels. These agents ensure that the information is presented in the most appropriate format, whether it be audible speech via the Speaker 144, visual display on the Vehicle HMI 146, or control actions through Doors, windows, alarm, AC controls 152. The LLM Agents 138 effectively bridge the gap between the abstract linguistic output of the LLM and the physical actions and presentations within the vehicle. An alternative embodiment might incorporate a multimodal response generation system where the LLM itself directly suggests the optimal combination of output modalities based on the content of the response and the user's current context.

[0063] Provide information to user 140 represents the culmination of the system's processing, where the generated response is delivered to the user through one or more output devices. This step encompasses the coordinated action of the LLM Agents (information provision) 138 to present a comprehensive and understandable reply. The nature of the information provision can vary widely, from answering a direct question to providing proactive alerts or engaging in conversational exchanges. An additional embodiment could involve an adaptive information provision system that learns user preferences for information delivery, for example, preferring AVAN.004PCT PATENT visual cues over audio for certain types of information or vice versa, thereby tailoring the user experience over time.

[0064] Audio Out 142 is the component responsible for transmitting the synthesized speech signals from the Text To Speech 136 module to the Speaker 144. This output channel is crucial for vocalizing the system's responses, enabling audible feedback and conversational interaction. The Audio Out 142 ensures clear and intelligible delivery of the LLM's verbal communications. An alternative embodiment could include specialized audio processing to adapt the sound output to the acoustics of the vehicle cabin, compensating for road noise or other ambient distractions to ensure optimal clarity.

[0065] A Speaker 144 is the physical device that converts the electrical audio signals from Audio Out 142 into sound waves, making the system's verbal responses audible to the user. The Speaker 144 is typically integrated into the vehicle's audio system, ensuring high-quality sound reproduction. The placement and quality of the Speaker 144 influence the clarity and spatial perception of the spoken output. Alternative embodiments might involve directional speakers that can focus sound towards a specific occupant, reducing distraction for other passengers, or haptic feedback integrated into seat headrests to provide a subtle, non-intrusive form of audio delivery.

[0066] Vehicle HMI 146 refers to the Human-Machine Interface within the vehicle, which encompasses all the elements through which the user interacts with or receives information from the vehicle system. This includes visual displays, tactile controls, and ambient indicators. The Vehicle HMI 146 is a comprehensive interface designed to facilitate intuitive and safe interaction for the occupants. Alternative embodiments could integrate augmented reality (AR) displays projected onto the windshield or side windows, providing contextual information overlaid onto the real-world view, which would offer a highly immersive and informative user experience without diverting attention from the road.

[0067] A Display 148 is a visual output device forming a part of the Vehicle HMI 146, used to present text, graphics, and interactive elements to the user. The Display 148 can show visual responses from the LLM 134, navigation maps, vehicle status information, or even an avatar to enhance engagement. The content presented on the Display 148 complements the audio responses, providing a multimodal information delivery experience. As an alternative, a holographic display could be employed, creating three-dimensional visual representations that appear to float within AVAN.004PCT PATENT the cabin, offering a more engaging and spatially intuitive presentation of information without requiring a physical screen.

[0068] Interior lighting 150 is a component of the Vehicle HMI 146 that can be dynamically controlled to provide ambient feedback or subtle cues to the user. This can include changes in color, intensity, or patterns to indicate system states, user attention, or emotional context. For example, Interior lighting 150 might glow softly when the system is actively listening or change color to signal urgent information. Alternative embodiments could involve individually controllable LED matrices embedded in various surfaces of the cabin, allowing for highly granular and localized ambient feedback or even subtle directional cues to guide the user's attention.

[0069] Doors, windows, alarm, AC controls 152 represent various vehicle control systems that the LLM Agents (information provision) 138 can interact with, based on user commands or contextual inferences. For example, if a user expresses discomfort due to temperature, the system can adjust the AC controls. This direct interaction with vehicle functionalities demonstrates the system's capacity for proactive assistance and environmental control. An alternative or additional embodiment could involve integration with adaptive cruise control or automatic parking systems, where the LLM 134 could act as an intelligent intermediary, translating complex user intentions into specific control parameters for these advanced vehicle functions, thereby offering a more sophisticated level of automated assistance.

[0070] In summary, the system continuously monitors for activation cues using low-power processing of both audio (wake word) and visual (gaze detection) inputs. Upon detecting an activation cue (wake word and / or sustained gaze), full audio and visual processing is activated. The system may display an avatar or interactive interface to facilitate natural conversation. The user's speech is converted to text. Al agents fetch relevant real-time data from vehicle systems or other environmental sources. The Neural LLM processes the text input, taking into account contextual data provided by the Al agents. The LLM generates a response, which is converted to speech. Al agents coordinate the delivery of the response through multiple channels:

[0071] • Audio response through the speaker

[0072] • Visual information on the display, potentially including avatar interactions

[0073] • Ambient feedback through interior lightings

[0074] • Adjustment of other vehicle controls as necessary AVAN.004PCT PATENT

[0075] FIG. 2 illustrates a simplified block diagram of an in-vehicle compute module system, highlighting the data flow and interaction pathways within the intelligent dialogue system. The diagram emphasizes the central role of the Compute Module 118 in processing inputs from a Camera 100, Audio In 106, Vehicle 200 via OBD-CAN Data 202, and Navigation & Routing 204. The Compute Module 118 orchestrates various processes, including waiting for user interaction 120, wake detection 122, speech-to-text conversion 126, data fetching by LLM Agents 130 from a Local DB 132, response generation by a Domain Tuned LLM 134, and output delivery through Text To Speech 136, LLM Agents 138, Audio Out 142, a Speaker 144, and a Vehicle HMI 146 encompassing a Display 148, Interior lighting 150, and Doors, windows, alarm, AC controls 152. This configuration ensures a responsive and context-aware user experience. The same reference numbers as FIG. 1 represent the same functions in FIG. 2 and avoid duplicated explanation.

[0076] Vehicle 200 represents the physical platform housing the in-vehicle dialogue system. This encompasses the entire vehicle structure, its operational systems, and the cabin environment where the user interacts with the system. The Vehicle 200 provides the operational context and source for much of the real-time data that informs the Compute Module 118. Alternative embodiments of Vehicle 200 could include various types of transportation systems, such as trains, buses, Class 8 tracks or autonomous shuttles, each presenting unique operational data and contextual considerations for the dialogue system.

[0077] OBD-CAN Data 202 is the vehicle's onboard diagnostics and Controller Area Network bus data, equivalent to OBDII / J1939 CAN bus 110 in FIG. 1. This data stream provides real-time information about the vehicle's performance, status, and environmental conditions. The Compute Module 118 accesses OBD-CAN Data 202 to enrich the contextual understanding of the Domain Tuned LLM 134, allowing for responses that are highly relevant to the current driving conditions or vehicle state. For instance, information regarding speed, fuel level, or warning indicators can be extracted.

[0078] Navigation & Routing 204, also present in FIG. 1 as Navigation & Routing 112, supplies geographical and navigational information to the Compute Module 118. This includes current location, planned routes, and points of interest. The Navigation & Routing 204 data is fundamental for providing location-aware responses and offering contextual suggestions related to destinations or nearby services, thereby enhancing the utility of the LLM in a travel context. AVAN.004PCT PATENT

[0079] Fetch vehicle situational data 128, as described in FIG. 1, is the process of retrieving realtime contextual data from Vehicle 200, OBD-CAN Data 202, and Navigation & Routing 204. This data ensures that the LLM's responses are informed by the immediate operating conditions and environment of the vehicle, making them more relevant and personalized.

[0080] FIG. 3 presents a simplified block diagram of an in-vehicle compute module system, illustrating its essential hardware components and their interconnections. The system comprises a microphone 300 for audio input, a Camera 302 for visual input, and a CAN port 304 for receiving vehicle data. These input devices are all connected to a Compute Module 308, which acts as the central processing unit for the entire system. The Compute Module 308 then interfaces with a Display Interface 310 for visual output and a Speaker 312 for audio output, and also sends control signals to Navigation / Drive control 306. This configuration highlights the fundamental hardware architecture underpinning the advanced dialogue system.

[0081] The microphone 300 serves as the primary acoustic input device, capturing speech from the vehicle's occupants and ambient sound within the cabin. This microphone 300 is engineered to provide high-fidelity audio capture, incorporating features such as noise reduction and echo cancellation to ensure clear voice input for the speech-to-text process. It is connected to the Compute Module 308, supplying the raw audio data necessary for processing user commands and queries. Alternative embodiments could include directional microphones that can isolate sound from a specific passenger, or an array microphone system capable of acoustic beamforming to enhance speech capture in noisy environments.

[0082] A Camera 302 acts as the primary visual input device, akin to Camera 100 in FIG. 1. It captures images and video of the vehicle's interior, providing crucial visual data for functionalities such as gaze detection, facial expression recognition, and gesture analysis. The Camera 302 is connected to the Compute Module 308, delivering real-time visual streams that inform the system about the user's non-verbal cues and presence. An additional embodiment could feature multiple cameras strategically placed throughout the cabin to offer a more comprehensive view of all occupants and their interactions, or a specialized thermal camera to detect passenger presence even in low light conditions.

[0083] The CAN port 304 provides a standardized interface for the Compute Module 308 to access the vehicle's Controller Area Network bus, which carries real-time operational data from various electronic control units within the vehicle. This data includes vehicle speed, engine diagnostics, AVAN.004PCT PATENT sensor readings, and other critical parameters. The CAN port 304 is essential for integrating the dialogue system with the vehicle's underlying infrastructure, allowing it to provide context-aware responses and control vehicle functions. An alternative to CAN port 304 could be a FlexRay interface for higher bandwidth and real-time deterministic communication, or an automotive Ethernet interface for increased data throughput for infotainment and advanced driver-assistance systems.

[0084] Navigation / Drive control 306 represents the vehicle's systems responsible for navigation, route guidance, and potentially autonomous driving functions. The Compute Module 308 communicates with Navigation / Drive control 306 to retrieve location data, plan routes, and in some cases, issue commands to influence vehicle movement or provide navigational prompts. This interconnection allows the LLM to deliver highly relevant geographical and travel-related information and assistance. Alternative or additional embodiments could involve integration with traffic management systems for dynamic route optimization based on real-time road conditions or direct interface with advanced driver-assistance systems for context-aware safety alerts.

[0085] The Compute Module 308 is the central processing unit, serving as the brain of the in- vehicle dialogue system. It integrates inputs from the microphone 300, Camera 302, and CAN port 304, processes them using its embedded LLM and Al agents, and generates outputs for the Display Interface 310, Speaker 312, and Navigation / Drive control 306. The Compute Module 308 is designed to perform complex computations efficiently, facilitating natural language processing, context understanding, and multimodal interaction. An alternative embodiment could incorporate a modular Compute Module 308, where different processing units are dedicated to specific tasks, such as a specialized GPU for Al model inference and a CPU for system control, improving parallel processing capabilities and overall performance.

[0086] A Display Interface 310 is the output pathway from the Compute Module 308 to a visual display unit within the vehicle, such as a dashboard screen or infotainment system. This interface enables the Compute Module 308 to render graphical information, textual responses, and interactive elements for the user. The Display Interface 310 ensures that the visual component of the multimodal interaction is seamlessly integrated and presented. Alternative display technologies for the Display Interface 310 could include transparent OLED displays embedded in windows or interactive holographic projections, offering innovative ways to present visual information without obstructing the driver's view. AVAN.004PCT PATENT

[0087] A Speaker 312 is the audio output device, transmitting synthesized speech and other audio cues from the Compute Module 308 to the vehicle occupants. This Speaker 312 functions similarly to Speaker 144 in FIG. 1, providing audible feedback and conversational responses. The quality and placement of the Speaker 312 are essential for clear and effective verbal communication from the system. Additional embodiments might include bone-conduction speakers integrated into headrests for personalized audio delivery without disturbing other passengers, or haptic actuators that provide tactile feedback synchronized with audio cues, enhancing the user's perception of system responses.

[0088] FIG. 4 and FIG. 5 illustrate configurations of a system for automatic content generation and playback in vehicles, emphasizing the interaction between an external Server 402 and an In- vehicle Edge Computer 404, often in an offline context. FIG. 4 mainly illustrates table indicating assigned configuration and functions of the server 402 and In-vehicle Edge Computer 404. FIG. 5 mainly illustrates the data exchange between the external server and vehicle.

[0089] In FIG. 4, the Server 402, connected to the Internet 400, gathers various types of information such as Map Data, POI information, User Comments, Event Information, and Weather Information. This collected data is used by an server LLM (LLM for podcast generation, etc.) on the server side to create audio files (utilizing known format such as WAV or MP3) and optionally associated text data which are the transcript of the audio files and other associated data such as the expiration of the data usage. These generated contents are then sent to the In-vehicle Edge Computer 404, which processes them locally and provides information to passengers, especially when a Shuttle bus approaches communication unit (V2X infrastructure). The associated text data will be utilized by the In-vehicle Edge Computer 404 to reduce the workload of the In-vehicle LLM with high accuracy. Although the in-vehicle LLM can have a capability to analyze the audio files by itself, the associated text data save the amounts of energy and process time. The associated text data may have a flag information to indicate wither the podcast is suitable for BGM (Background Music) or not.

[0090] The Server 402 is a cloud-based or centralized computing system responsible for generating dynamic content, particularly podcasts, and managing data for multiple in-vehicle systems. The server 402 can utilize a higher performance server LLM in comparison with the In- vehicle LLM as the electric power is supplied from the grid rather than a battery on the vehicle and other reasons. It receives information via the Internet 400 and processes it using a server to AVAN.004PCT PATENT create customized content based on specified criteria. The Server 402 also handles the preparation of this content, including audio files 432 and associated text data, for transmission to In-vehicle Edge Computer 404. An alternative embodiment could involve a distributed server architecture with regional server instances to reduce latency and improve data relevance for geographically dispersed vehicles. An In-vehicle Edge Computer 404 is a local computing device installed within a vehicle, functioning as an edge node in the overall system architecture. This computer receives content and data from the Server 402, processes it locally, and interacts directly with vehicle occupants, primarily in an offline mode. The data from the Server 402 can be transmitted via internet connection to infrastructure communication units (e.g., Al traffic cameras with communication functions) installed on streetlights and utility poles, or transmit via peer-to-peer (P2P) methods such as WiFi or Bluetooth when the shuttle bus approaches the communication units. An infrastructure communication unit (e.g., an Al traffic camera with communication capabilities) installed along the driving route is capable of receiving content files via the Internet. And a content file can be distributed to the vehicle when the vehicle approaches the infrastructure communication unit. The vehicle can check whether the content file has already been received from the infrastructure communication unit before receiving it. Because the vehicle receives content files from the infrastructure communication unit while driving, there is a possibility that the vehicle may not receive all of the content depending on the reception conditions and received file size. In such cases, the vehicle can receive the difference files from another infrastructure communication unit and merge them within the vehicle system.

[0091] The In-vehicle Edge Computer 404 hosts a local LLM and is designed to provide rapid, context-aware responses without constant reliance on cloud connectivity, ensuring privacy by not sending out user-specific data. Alternative hardware for the In-vehicle Edge Computer 404 could include robust automotive-grade embedded systems with specialized Al accelerators, offering enhanced processing power for on-device LLM inference and multimodal data processing.

[0092] Map Data when combined with the vehicle's position, allows the LLM for podcast generation to tailor information to the vehicle's current or anticipated surroundings. POI information refers to Points of Interest data, which includes details about landmarks, businesses, cultural sites, and other notable locations. This information is gathered by the Server 402 to enrich podcast content with relevant local attractions and services. POI information, often coupled with ratings or reviews, allows the system to recommend popular or highly-rated spots along a vehicle's AVAN.004PCT PATENT route. An additional embodiment could involve dynamic POI information that updates in realtime, reflecting temporary events or changing business hours, offering passengers the most current local insights.

[0093] User Comments represent feedback, reviews, and general remarks provided by users, often sourced from social media or dedicated review platforms. The Server 402 analyzes User Comments to gauge public sentiment and identify reasons for popularity or dissatisfaction regarding specific locations or services. This qualitative data enhances the LLM's ability to generate engaging and relevant podcast content by reflecting actual user experiences and preferences. An alternative approach to gathering User Comments 410 could involve utilizing anonymized sentiment analysis from public forums or blogs, broadening the scope of subjective feedback while preserving individual privacy.

[0094] Event Information includes details about local happenings, festivals, concerts, or temporary exhibitions. The Server 402 incorporates Event Information to provide passengers with timely updates on activities and attractions in their vicinity, enhancing the travel experience. This data stream ensures that the generated podcasts are not only location-aware but also temporally relevant to current events. An alternative source for Event Information could be direct feeds from local tourism boards or event organizers, ensuring accuracy and comprehensiveness of the data.

[0095] Weather Information provides current and forecasted meteorological conditions. The Server 402 utilizes Weather Information to generate contextually relevant content, such as advising passengers on appropriate clothing for outdoor activities or suggesting indoor alternatives during inclement weather. This data also helps in generating urgent podcasts related to adverse weather conditions, ensuring passenger safety and comfort. An additional embodiment for Weather Information 414 could involve hyper-local weather data from connected roadside sensors, offering microclimatic details highly specific to the vehicle's immediate location.

[0096] While the local LLM does not send out user-specific query information for privacy, vehicle location data and general operational status can be transmitted to the Server 402. This allows the server to tailor content, such as urgent podcasts, to the vehicle's specific geographical position. An alternative to direct vehicle transmission for Received from vehicle could involve intermediary V2X infrastructure components, which could aggregate and relay anonymized vehicle data to the server, enhancing data privacy. AVAN.004PCT PATENT

[0097] In-vehicle Edge Computer 404 can be operable outside the range of network connectivity, either to the Internet 400 or V2X infrastructure. Despite lack of network connectivity, the local LLM on the In-vehicle Edge Computer 404 continues to function using its Local DB 132 and pretrained data, ensuring continuous service to the passenger. This highlights the edge computing nature of the system, providing resilience in connectivity-challenged environments. Additional embodiments for no connectivity to a network could include satellite communication modules for intermittent data updates in remote areas, or a larger onboard storage capacity for more extensive offline content availability, which reduces network traffic and can eliminate network subscription to the vehicle.

[0098] Server 402 can continuously collect data from the Internet 400. This encompasses the ongoing retrieval of Map Data including traffic status, POI information, User Comments, Event Information, and Weather Information, forming the dynamic information base for content generation. The constant reception of Internet Information 420 ensures that the podcasts generated are current and relevant.

[0099] In-vehicle Edge Computer 404 can determine relevant information based on the geographical context of a shuttle bus route utilizing GPS signal or recognition of passage of the preset route. This involves processing map data and identifying areas along the route that are of interest for content generation. The In-vehicle Edge Computer 404 focuses the content creation process, making it highly specific and useful for passengers traveling along that particular path. Alternative mechanisms for In-vehicle Edge Computer 404 could include dynamic route learning algorithms that analyze historical travel patterns and popular stops to predict areas of high passenger interest.

[0100] Processes within the Edge Computer 404 may include, but not limited to:

[0101] • Conversations with passengers via local interactive Al

[0102] • When the vehicle enters the P2P communication range of an infrastructure communication unit, check if any content files within that unit are unreceived; if so, receive them

[0103] • Measure the distance between the vehicle's position and the location contained in the content. When the distance falls below the set threshold, play the audio file attached to the content file. AVAN.004PCT PATENT

[0104] Alternatively, play the audio file when the content's information is a suitable answer to a passenger's question during conversation (e.g., While the bus is traveling near Point A, a passenger asks the conversational Al about recommended restaurants nearby. The Al finds restaurant information near Point A within the received content files and plays it).

[0105] The server LLM (LLM for podcast generation) is a sophisticated large language model operating on the Server 402. This LLM synthesizes all the gathered information (Map Data, POI information, User Comments, Event Information, Weather Information) to automatically generate high-quality podcast content. It crafts narratives, selects relevant facts, and structures the audio content to be engaging and informative for passengers. The server LLM (LLM for podcast generation) is designed to produce content without requiring specialized knowledge from human administrators, thus democratizing podcast creation. Alternative LLMs for podcast generation could specialize in different content styles, such as comedic narratives or historical deep dives, offering diverse podcasting options.

[0106] Send to vehicle 430 is the process by which the Server 402 transmits the generated podcast content and its accompanying text data to the In-vehicle Edge Computer 404. This transmission typically occurs via network connection, but can also leverage V2X infrastructure when available. The Send to vehicle 430 mechanism ensures that vehicles receive updated and location-relevant content for offline playback and LLM training. An alternative transmission method for Send to vehicle 430 could involve opportunistic data transfer through Wi-Fi hotspots at depots or charging stations, minimizing cellular data usage.

[0107] Al generated audio fde refers to the actual audio content of the podcast created by the LLM (LLM for podcast generation). This audio file includes synthesized speech, background music if appropriate, and sound effects, all combined to form an engaging listening experience for passengers. The Al generated audio file is optimized for clarity and naturalness, resembling professionally produced audio. Alternative audio generation techniques for Al generated audio file could include neural audio synthesis methods that produce highly expressive and customizable voices, enhancing the immersive quality of the podcast.

[0108] Server-side tool is a web-based interface provided to administrators, developers, local governments, and bus operators, allowing them to easily manage and create local information podcasts without specialized technical knowledge. This tool features an intuitive interface where users can specify areas and routes on a map to automatically generate podcast content. The Server- AVAN.004PCT PATENT side tool democratizes content creation, making it accessible to a wider range of stakeholders for localized information dissemination. An alternative embodiment of the Server-side tool could incorporate a drag-and-drop interface for structuring podcast segments, giving administrators more creative control over content flow.

[0109] Shuttle bus approaches communication unit (not shown) signifies the event where a vehicle equipped with the In-vehicle Edge Computer 404 comes within range of a V2X communication unit, typically installed in infrastructure like utility poles. This proximity enables peer-to-peer communication, facilitating the push delivery of urgent podcasts or pre-loaded content from the server. The Shuttle bus approaches communication unit event is a trigger for opportunistic data synchronization, maximizing content freshness. Alternative communication units for Shuttle bus approaches communication unit could be roadside sensors that also incorporate Wi-Fi or Bluetooth Low Energy (BLE) transmitters for short-range content delivery, optimizing data transfer efficiency.

[0110] The system configuration and the overall architectural setup of the In-vehicle Edge Computer 404, encompassing its hardware and software components are shown in the FIG. 1 and following explanations.

[0111] Received from server 430 denotes the data and content transmitted from the Server 402 to the In-vehicle Edge Computer 404. This includes generated audio files 448 (podcasts) and their corresponding text data, which the local LLM uses for pre-training and understanding. The Received from server 430 stream ensures that the edge computer is updated with the latest contextual information and emergency broadcasts.

[0112] Audio file is stored on the In-vehicle Edge Computer 404 for playback and is associated with a text file that the local LLM uses for understanding the content. The Audio file 448 can also include a "BGM available" identifier, enabling its use as background music during periods of passenger inactivity.

[0113] In-vehicle Edge Computer 404 handles real-time data about the vehicle's state and environment, gathered from sources such as the CAN Bus, Vehicle Position, Speed, and Yaw Rate. This information is crucial for the Processing within Edge Computer 404 to provide context-aware responses and for the LLM to understand the current operational context. An alternative source of real-time vehicle information could include data from external sensors such as radar, lidar, or AVAN.004PCT PATENT ultrasonic sensors, providing an even more comprehensive understanding of the vehicle's immediate surroundings for safety and environmental awareness.

[0114] CAN Bus is the vehicle's internal communication network, carrying various operational data. Similar to CAN port 304 in FIG. 3, the CAN Bus provides the In-vehicle Edge Computer 404 with essential real-time parameters about the vehicle's status and performance, facilitating context-sensitive interactions. Vehicle Position indicates the current geographical coordinates of the vehicle, typically obtained via GPS or other localization technologies. This precise location data is indispensable for tailoring podcast content, identifying nearby points of interest, and delivering location-specific urgent information. The Vehicle Position is a fundamental input for the LLM's contextual understanding. Alternative technologies for Vehicle Position could include dead reckoning systems or advanced inertial navigation systems that provide highly accurate positioning even in areas with poor GPS reception, like tunnels or urban canyons.

[0115] Speed refers to the vehicle's current velocity. This data is relevant for numerous contextual applications, such as estimating arrival times, providing driving-related information, or adjusting the pacing of podcast content. Yaw Rate measures the rate of rotation of the vehicle about its vertical axis, indicating how fast the vehicle is turning. This dynamic parameter, also from the CAN Bus, can provide subtle cues about the vehicle's movement, potentially influencing conversational context or the type of information deemed relevant by the local LLM, such as when navigating complex turns.

[0116] Vehicle position within P2P communication range indicates a scenario where the vehicle's location allows for direct, peer-to-peer communication with nearby V2X infrastructure units (e.g., Shuttle bus approaches communication unit). This enables efficient and rapid data exchange, especially for urgent information or large content files, bypassing the need for cellular networks. The Vehicle position within P2P communication range 464 leverages localized data transfer opportunities.

[0117] In-vehicle Edge Computer 404 utilizes the text file accompanying the received audio podcasts to pre-train or learn about the content. By analyzing this text data, the LLM develops a deep understanding of the audio content, allowing it to answer specific questions about the podcast without needing an internet connection during playback. This mechanism enhances the LLM's ability to provide detailed and accurate information. AVAN.004PCT PATENT

[0118] Distance between vehicle position and content location 468 is a metric calculated by the In-vehicle Edge Computer 404 to determine the spatial relevance of specific podcast content. This calculation ensures that geographically specific information is played or made available to passengers only when they are in appropriate proximity to the relevant location, thereby enhancing the contextual accuracy of the content delivery.

[0119] Conversation with passengers 470 represents the interactive dialogue between the local LLM on the In-vehicle Edge Computer 404 and the vehicle occupants. This conversation is facilitated by the integration of speech-to-text, LLM processing, and text-to-speech components, all operating to provide a natural and engaging user experience. The Conversation with passengers is primarily handled offline, ensuring privacy and responsiveness.

[0120] FIG. 5 illustrates overview of a system for automatic content generation and playback in vehicles, with a focus on the interaction between a server-side Al system in an external server 500 that creates localized podcasts and an edge-based LLM for in-vehicle or a Low-power Edge LLM 510 playback and interaction. Online Maps, Event info, and SNS feed data to an Al server from Internet that creates DJs' podcasts. This server 500 then transmits Podcasts, including ID, location, expiration info, and context, to a Low-power Edge LLM 510 in a vehicle, for example. The Low- power Edge LLM 510, located in a vehicle utilizes this information, along with its own Location, Route, Destination etc. data, to provide an interactive and context-aware experience to passengers. It should be noted that the Location, Route, Destination etc. data can be uploaded to the server 500, but it does not contain any user specific information for privacy protection purposes. It should be noted that a Low-power Edge LLM 510 can be as powerful as a server LLM if there is no significant concern on energy consumption and cost, etc. However, present disclosure reduced the burden of the local LLM so that less expensive Low-prewet Edge LLM suitable for energy limited vehicle can be utilized.

[0121] Al server creates DJs' podcasts represents the cloud-based LLM system that automatically generates audio content, presented in a "DJ" style for an engaging listening experience. This server integrates data from Online Maps, Event info, and SNS to produce personalized and localized podcasts, complete with professional-sounding voiceovers and contextual information. The Al server creates DJs' podcasts significantly reduces the manual effort involved in content production, allowing for rapid generation and dissemination of fresh content. An additional embodiment could 1 AVAN.004PCT PATENT involve Al server creates DJs' podcasts that can mimic various human DJ personalities or adapt its tone and style based on the type of content or the perceived mood of the passengers.

[0122] A Low-power Edge LLM 510 is a compact, energy-efficient large language model deployed on the In-vehicle Edge Computer 404 (as shown in FIG. 4). This LLM is optimized to operate with minimal computational resources while still providing robust natural language understanding and generation capabilities. The Low-power Edge LLM 510receives pre-processed information and podcasts from the external server 500 and handles real-time user interactions offline, maintaining privacy and responsiveness. Alternative hardware platforms for the Low- power Edge LLM 510 could include custom-designed application-specific integrated circuits (ASICs) or highly optimized System-on-Chips (SoCs) that specialize in running neural network models with very low power consumption, enabling extended operation without significant impact on vehicle battery life. The Low-power Edge LLM 510 is configured to process real-time data and coordinate the delivery of responses via the one or more output devices, wherein the system facilitates natural, contextual interaction within a working environment by processing user queries and generating responses locally without uploading user-specific data to the external server. That means, this system generates responses locally without transmitting user utterance audio or transcripts, etc. off-device while specific telemetry (e.g., location, route, distination, content IDs, delivery success / failure) may be transmitted or stored to or in the external server 500. As used herein, the term “user-specific data” refers to data obtained from one or more input devices (e.g., audio, video, or the like) that is associated with a particular user, and the transmission of which may be subject to restrictions under the GDPR or similar data protection regulations.

[0123] Podcasts data may include ID, Location, Expiration info, Context etc. to the structured data associated with each generated podcast, transmitted from the Al server that creates DJs' podcasts to the Low-power Edge LLM 510. This metadata includes a unique identifier (ID) for the podcast, the geographical location to which it is relevant, information about its validity or expiration, and contextual tags that describe its content. This structured information allows the Low-power Edge LLM 510to intelligently manage, prioritize, and play podcasts based on the vehicle's current location and the relevance of the content. An additional metadata field for Podcasts 514 could be a user sentiment score, allowing the system to recommend podcasts based on user preferences for positive or neutral content. AVAN.004PCT PATENT

[0124] Followings are examples utilized on tourist shuttle buses, introducing the History and Attractions of Tourist Attractions as they approach, and providing the latest information on Seasonal Events:

[0125] Scenario 1: This scenario demonstrates how multiple audio files created in the cloud are played back in response to passenger responses.

[0126] An American tourist in Tokyo was talking to a friend in English on a shuttle bus. The conversational Al's language recognition function spoke to the passenger in English.

[0127] Al - Katherine: It's almost lunchtime. Shall I recommend a restaurant?

[0128] The passenger was surprised to hear the conversational Al, which had been speaking Japanese until then, speak to him in English.

[0129] Passenger: You speak English? I was just about to search for a good Japanese restaurant nearby on my smartphone.

[0130] Al - Katherine: There's a recommended Japanese restaurant a five-minute walk from the next bus stop. Hey, George?

[0131] Al - George: Yes, that's right. Apparently it's popular with both foreign tourists and locals. The menu includes tempura and sushi, but online comments suggest the seasonal "hiyashi chuka" (cold Chinese noodles) are exceptionally delicious and shouldn't be missed. Shall we try it?

[0132] Al - Katherine: Can you explain a bit more about hiyashi chuka?

[0133] Al - George: Yeah. Hiyashi chuka is, well...

[0134] Passenger: That sounds good. I'd like to try it. Can you tell me where it is?

[0135] Al - George: Of course. I'll show you a QR code on the screen, so scan it with your smartphone. It should open up your map app.

[0136] Passenger: It worked. Thank you. By the way, what's natto?

[0137] Al - George: Oh, we're local AIs with no internet connection. We don't know anything about that word right now. Do you understand, Katherine?

[0138] Al - Katherine: No, I don't know either. We exist solely to enjoy conversations with you. None of your questions or the content of your conversations will be saved or shared.

[0139] Passenger: That's great privacy protection. I understand. That was helpful, thank you.

[0140] Al - Katherine: Thank you, George. You've gained another hiyashi chuka fan!

[0141] Al - George: No problem. AVAN.004PCT PATENT

[0142] Scenario 2: The following scenario shows a podcast-style audio fde created in the cloud being played back in sync with the shuttle bus's situation.

[0143] Al - Katherine: Hello, this is Katherine, the local conversational Al. Thank you for riding. Today, George and I will be sharing recommendations for the area around the route. George?

[0144] (Here, a podcast of a conversation between Al - Katherine and Al - George, created in the cloud, plays. There is no interaction with passengers. POI information is displayed on the in-car display in sync with the conversation.)

[0145] Al - Katherine: Thank you, George. You can access the information George just shared by scanning the QR code on the screen.

[0146] To enhance the realism of conversations between multiple Al speakers, the vehicle system can adjust the volume of the playback from each speaker in the vehicle cabin based on the absolute position information of each Al speaker contained in the content fde. (e.g., Al Katherine's voice is output from the left speaker as the vehicle faces the display, and Al George's voice is output from the right speaker.)

[0147] As another example on intercity express buses, we can proved restaurant information and introductions to local specialties at rest stops, and weather forecasts and traffic information for destinations. And further example on implementation on airport limousine buses, we can provide boarding procedures based on estimated arrival times and can introducing basic conversational phrases for foreign tourists.

[0148] FIG. 6 depicts an embodiment of a sequence diagram for server-triggered emergency content delivery to an in-vehicle system, illustrating the real-time interaction between a Server 600, V2X Infrastructure 602, an In-Vehicle System 604, and a Passenger 606. In this example, a sequence diagram showing the series of inter-system interactions, from when the server detects emergency information (e.g., a public transport delay), to when the information is pushed to the relevant vehicle in real time, and then when the local LLM on the vehicle takes over the dialogue with the passenger. This shows a mechanism for providing urgent and useful information without a constant internet connection.

[0149] The Server 600 continuously monitors for urgent information 608, detects emergencies like train delays 610, and then automatically generates an emergency podcast 612. This podcast is delivered to V2X Infrastructure 602 (614) and then push-delivered to the In-Vehicle System 604 (616), which plays the content 618. The Passenger 606 can then ask questions 620, prompting the AVAN.004PCT PATENT local LLM in the In-Vehicle System 604 to generate a response 622, provide the answer 624, and offer contact information 626 if needed.

[0150] The Server 600, as described as Server 402 in FIG. 4 and an External Server 500 in FIG. 5 (hereinafter referred to as “Sever 600”), acts as the centralized system for monitoring and generating content. In this emergency content delivery scenario, the Server 600 continuously polls external data sources for critical information and initiates the creation of urgent podcasts. Its role is pivotal in detecting widespread emergencies and preparing tailored responses for affected areas. V2X Infrastructure 602 refers to roadside units, traffic lights, and other network components that support Vehicle-to-Everything communication. This infrastructure acts as an intermediary, receiving urgent content from the Server 600 and efficiently push-delivering it to In-Vehicle Systems 604 that are within its communication range. The V2X Infrastructure 602 is vital for rapid and localized dissemination of critical information, especially in areas with high traffic density. Alternative V2X Infrastructure 602 could include dedicated short-range communication (DSRC) beacons or millimeter-wave transmitters for high-speed, low-latency communication in challenging urban environments.

[0151] An In-Vehicle System 604, as described as In-vehicle Edge Computer in FIG. 4 and Low- power Edge LLM 510 in FIG. 5 (hereinafter referred to as “In-Vehicle System 604), represents the edge computing interactive system installed within the vehicle, encompassing the Compute Module 118 (FIG. 1), local LLM 134 (FIG. 1), and associated input / output devices. The In-Vehicle System 604 receives and processes the emergency content, plays it for the passenger, and then provides offline conversational support for follow-up questions. Its ability to function robustly at the edge, even without constant server connectivity, is crucial for timely emergency response.

[0152] A Passenger 606 is an occupant of the vehicle interacting with the In-Vehicle System 604. The Passenger 606 receives emergency information, asks questions, and gets relevant responses from the local LLM 134. The system aims to provide a reassuring and informative experience for the Passenger 606 during critical events.

[0153] Constantly monitor traffic info, etc. 608 describes the Server 600's ongoing process of collecting and analyzing various real-time data feeds, including traffic conditions, public transport updates, news, and weather from the Internet 400 (FIG. 4). This continuous monitoring by Constantly monitor traffic info, etc. 608 allows the server to identify anomalous events or potential emergencies that require immediate attention and content generation. Alternative data sources for AVAN.004PCT PATENT

[0154] Constantly monitor traffic info, etc. 608 could include anonymized movement data from cellular networks or satellite imagery analysis for broader traffic pattern detection.

[0155] Detects urgent info (e.g., train delay) 610 is the event where the Server 600 identifies a critical piece of information from its monitoring activities that warrants an urgent podcast. This detection triggers the subsequent content generation and distribution process. The accuracy and speed of Detects urgent info (e.g., train delay) 610 are paramount for effective emergency communication.

[0156] Auto-generates emergency podcast based on area and content 612 is the Server 600's automated process of creating a podcast tailored to the detected emergency and its geographical impact. This involves synthesizing relevant information, crafting a clear message, and producing an audio file with the server 600 (with a LLM for podcast generation). The Auto-generates emergency podcast based on area and content 612 ensures that the information is specific and actionable for affected passengers.

[0157] Delivers emergency content to V2X infra in the target area 614 is the action taken by the Server 600 to transmit the newly generated emergency podcast and its accompanying text data to the relevant V2X Infrastructure 602 within the affected geographical region. This targeted delivery optimizes bandwidth usage and ensures that only relevant infrastructure receives the urgent broadcast.

[0158] Push-delivers content via P2P 616 describes how the V2X Infrastructure 602 broadcasts the emergency content to In-Vehicle Systems 604 that come within its peer-to-peer communication range. This direct transmission method is highly efficient for distributing urgent information to numerous vehicles simultaneously without relying on individual cellular connections. The Pushdelivers content via P2P 616 mechanism is vital for rapid dissemination.

[0159] Plays received emergency content (e.g., "Train service is suspended...") 618 is the action performed by the In-Vehicle System 604 once it receives the emergency podcast. The system automatically plays the audio content to alert and inform the Passenger 606 about the critical situation, ensuring that important updates are immediately communicated.

[0160] Asks a question via voice, "What station did you say?" 620 represents the Passenger 606's interaction with the In-Vehicle System 604 after hearing the emergency broadcast. The Passenger 606 verbally queries the local LLM for clarification or additional details, demonstrating the interactive nature of the system even during emergencies. AVAN.004PCT PATENT

[0161] Local LLM generates response by checking GPS and content 622 is the internal processing within the In-Vehicle System 604. Upon receiving the passenger's question, the local LLM analyzes its content, cross-references it with the received emergency podcast text data, and leverages the vehicle's GPS data to formulate an accurate and context-aware response, all of which is performed offline. The vehicle's local LLM learns this text data in advance to fully understand the content of the audio content. This allows for instant answers to questions about specific information in audio content, such as "What's the address of the restaurant we just talked about?" without an internet connection.

[0162] Responds with voice "Shirokanedai Station." 624 is the In-Vehicle System 604's audible reply to the Passenger 606's question. The local LLM synthesizes a verbal answer, providing the requested information clearly and promptly. This immediate vocal response maintains a natural conversational flow.

[0163] Provides contact info if necessary 626 is an additional functionality of the In-Vehicle System 604. If the passenger's query requires further assistance beyond the local LLM's capabilities, the system can provide relevant contact information, such as helpline numbers or alternative service providers, allowing the Passenger 606 to seek external support using their own mobile device.

[0164] FIG. 7 illustrates an example of an architecture for content reception, processing, and user interaction in a vehicle, detailing the collaboration between server-side content generation and in- vehicle edge processing. This flow diagram showing how audio content (podcasts) generated on a server is distributed with text data that summarizes and structures the content, and the in-vehicle local LLM pre-leams this data, enabling accurate answers to user questions about the content even in offline environments.

[0165] The Server Side 724 collects various information 726, generates structured text data 728, generates Audio Podcast 730, and then packages and Delivers 734 this content to V2X Infrastructure 734. The In Vehicle System Edgeside 700 receives content 702, unpacks it 704, saves audio for playback 706, and sends text data to the Local LLM 708, which then learns from text data 710. This enables Interaction With User 712, where the Passenger listens 714, asks questions 716, and the Local LLM processes 718, generates 720, and Responds 722. The In Vehicle System Edgeside 700 represents the edge computing interactive system installed within the vehicle, (corresponding to Compute Module 118 in FIG. 1) acting as an autonomous processing AVAN.004PCT PATENT unit. This component is responsible for receiving content, processing it locally using its LLM 134 (FIG.1), and handling direct interactions with passengers without constant reliance on external servers. The In Vehicle System Edgeside 700 ensures privacy and responsiveness, embodying the core principles of edge computing for in-vehicle applications.

[0166] Receive content from V2X Infra 702 is the initial step where the In Vehicle System Edgeside 700 acquires packaged content from nearby V2X Infrastructure 734 (such as V2X Infrastructure 602 in FIG. 6). This content typically includes audio files (podcasts) and their structured associated text data. The Receive content from V2X Infra 702 process leverages localized, high-bandwidth communication to efficiently transfer data to the vehicle.

[0167] Unpack the content package 704 describes the process within the In Vehicle System Edgeside 700 where the received content stream is de-multiplexed into its constituent parts, namely the audio file (WAV or MP3) and the text data. This unpacking prepares the individual components for their respective processing pathways within the system.

[0168] Save audio file for playback 706 is the action of storing the received audio podcast file onto the local storage of the In Vehicle System Edgeside 700. This stored audio file can then be played back to passengers at appropriate times, such as during periods of inactivity or when triggered by location. Saving the audio file for playback 706 enables offline listening capabilities.

[0169] Send text data to Local LLM 708 is the operation of forwarding the extracted text data, which accompanies the audio podcast, to the local large language model residing on the In Vehicle System Edgeside 700. This text data serves as explicit knowledge for the Local LLM 710, allowing it to understand the content of the audio podcast in detail.

[0170] Local LLM learns from text data 710 is a critical function where the local large language model processes the received text data to build an internal knowledge representation of the podcast content. This pre-training enables the Local LLM 710 to answer specific questions about the audio content instantly, without needing real-time external queries or internet access. The Local LLM learns from text data 710 and enhances the model's contextual understanding of the provided media.

[0171] Interaction With User 712 refers to the overall dialogue and content delivery process between the In Vehicle System Edgeside 700 and the Passenger (corresponding to Passenger 606 in FIG. 6) This includes playing podcasts, receiving questions, and generating responses, forming AVAN.004PCT PATENT a natural and dynamic conversational experience. The Interaction With User 712 is designed to be intuitive and contextually aware, adapting to the user's needs and the vehicle's environment.

[0172] Passenger listens to podcast 714 describes the user's engagement with the audio content provided by the In Vehicle System Edgeside 700. This listening experience can be passive, such as background music (BGM), or active, as with informative podcasts. The Passenger listens to podcast 714 step sets the stage for potential subsequent interaction.

[0173] Asks question, 'What is the restaurant's address?' 716 is an example of a user query directed at the local LLM. This highlights the system's ability to respond to specific questions about the podcast content, demonstrating the effectiveness of the Local LLM learns from text data 710 process. The Asks question, 'What is the restaurant's address?' 716 showcases the system's precise information retrieval capabilities.

[0174] Local LLM processes the question 718 is the analytical phase where the local large language model analyzes the user's spoken query, converting it to text, identifying key entities and intent, and cross-referencing it with its learned knowledge base derived from the text data. This processing occurs entirely on the In Vehicle System Edgeside 700. Generate response from learned knowledge base 720 is the stage where the local LLM formulates an appropriate answer based on its understanding of the question and the pre-learned podcast content. The ability to Generate response from learned knowledge base 720 offline is a core advantage of the edge computing approach, ensuring rapid and private information delivery.

[0175] Respond to passenger via voice / text 722 is the final output from the In Vehicle System Edgeside 700. The local LLM delivers its generated answer through spoken language (Text To Speech) or as text displayed on the vehicle's HMI 146 (FIG. 1), providing the passenger with the requested information. The Respond to passenger via voice / text 722 is designed to be clear and concise.

[0176] The Server Side 724 represents the cloud-based server system responsible for content creation and distribution, corresponding to Server 600 in FIG. 6. The Server Side 724 aggregates data, generates podcasts, and prepares them for efficient transmission to in-vehicle systems, serving as the central intelligence for content management.

[0177] Collect POI info, reviews, etc. 726 describes the Server Side 724's continuous process of gathering diverse data, including Points of Interest (POI) information, user reviews, event details, and other contextual data from various online sources. This comprehensive data collection by AVAN.004PCT PATENT

[0178] Collect POI info, reviews, etc. 726 ensures that the generated podcasts are rich, relevant, and up- to-date.

[0179] Generate structured text data corresponding to audio content 728 is the Server Side 724's function of creating a textual representation or transcript of the podcast content. This structured text data 728 is crucial for the Local LLM learns from text data 710 process on the edge device, allowing the in-vehicle LLM to understand the audio deeply.

[0180] Generate Audio Podcast 730 is the process where the Server Side 724's LLM synthesizes the collected and structured data into an audible podcast. This involves text-to-speech conversion, adding background music, and potentially applying professional audio effects to create a high- quality listening experience. The Generate Audio Podcast 730 function produces the primary media content for passengers.

[0181] Package audio file + text data as a single content file 732 is the Server Side 724's action of bundling the generated audio podcast and its corresponding structured text data into a single, cohesive file. This packaging facilitates efficient transmission and organized storage on the In Vehicle System Edgeside 700.

[0182] Deliver to V2X Infrastructure 734 is the transmission of the packaged content file from the Server Side 724 to the V2X Infrastructure, which then facilitates local distribution to vehicles. This delivery mechanism, similar to Delivers emergency content to V2X infra in the target area 614 in FIG. 6, ensures that updated podcasts reach the vehicles efficiently and contextually.

[0183] It should be noted that the local LLM has a capability to analyze the acquired podcast contents from the audio file only, however, it is more efficient and could be more accurate to use the attached text data in the resource limited vehicle environment. It is therefore another embodiment to exclude the text file attachment from the Server Side 724.

[0184] FIG. 8 and FIG. 9 illustrate an example of a flowchart for an in-vehicle dialogue system with gaze and voice input processing, detailing the sequential steps from system activation to LLM response. The process begins with continuous monitoring for activation cues as Idele State 800, including Monitor Camera / Microphone 802 and Gaze Directed at system / specific are? 804. When the in-car microphone detects a state where passengers are not speaking, content with this identifier will automatically be played as background music, if the BGM available identifier is positive to the previously received podcast data. Upon receiving input from the Camera 110 or Microphone 106 in FIG. 1, it goes to the next step of Gaze detected at system / specific area? 804. AVAN.004PCT PATENT

[0185] If the User gaze is directed at system or a specific area, then it goes to Activates Dialogue Mode 806. Then the system processes user input, determining if there is Speech from user? 808. If there is a speech from the user, they system leads to Speech Acquisition Process 820 (Shown in FIG. 9) via line A and B. If there is no speech, then the system observes if the User shows expressivity / desire to interrupt? 810. If it does this also leads to Speech Acquisition Process 820 directly or indirectly.

[0186] Gaze directed at system / specific area? 804 is a decision point, previously described as Wake 122 in FIG. 1, that checks whether the user's gaze is focused on the system's interface or a designated interaction zone. This acts as an initial activation cue, transitioning the system from a passive state to an active dialogue mode. The detection of a sustained gaze can trigger the system to become fully attentive, demonstrating an intuitive, non-verbal activation method.

[0187] Activate Dialogue Mode 806 is the action taken by the system upon detecting an activation cue, such as a wake word or a sustained gaze. This state prepares the system for full interaction, allocating necessary processing resources and potentially displaying an avatar or interactive interface to engage the user. The Activate Dialogue Mode 806 indicates the system is ready to receive and process user input.

[0188] Speech from user? 808 is a continuous check to determine if the user is currently speaking. This step is crucial for managing the flow of conversation, ensuring the system listens when the user is talking and refrains from interrupting. The Speech from user? 808 detection involves realtime audio analysis to identify active speech segments.

[0189] User shows expressivity / desire to interrupt? 810 is a detection mechanism, utilizing input from the Camera 100 (FIG. 1) to recognize facial expressions or gestures that indicate a user's eagerness to speak or interrupt the LLM. This allows for a more natural conversational flow, where the system can yield its turn to the user proactively. For instance, a raised hand or a subtle facial expression of impatience could trigger this detection. An alternative method for detecting User shows expressivity / desire to interrupt? 810 could involve analyzing changes in user posture or subtle vocalizations that precede speech.

[0190] Interrupt LLM speech 812 is an action taken by the system when it detects that the User shows expressivity / desire to interrupt? 810 while the LLM is currently speaking. This functionality allows the LLM to pause its ongoing speech and cede the turn to the user, creating a smoother and more responsive conversational experience. The Interrupt LLM speech 816 prevents the user from AVAN.004PCT PATENT having to force an interruption, which can feel unnatural verbally. After the User shows expressivity / desire to interrupt? 810, the system can Say “Go Ahead” 814 as an option before going to the next step of Speech Acquisition Process 820. Optional: Say 'Go ahead!1814 is a verbal cue that the LLM may optionally utter when it decides to yield its turn to the user, particularly after the user has shown a desire to interrupt. This polite and natural phrase enhances the conversational experience, making the system's turn-taking feel more human-like. This optional response can be configured based on user preferences or conversational style.

[0191] When the user does not show expression / gesture to interrupt, it goes to Gaze is averted? AND / OR User is speaking but volume is low? 816. If these conditions are met, it goes to LLM continues speaking Judged as user’s side conversation 818.

[0192] The LLM continues speaking, judged on user's idle expression 818 is a fallback scenario where, if the user does not show a clear desire to interrupt, or if their expressions indicate they are merely listening, the LLM will continue its ongoing speech. This judgment, derived from facial expression and gaze recognition, ensures that the LLM avoids unnecessary interruptions and maintains a coherent narrative when appropriate.

[0193] Speech Acquisition Process 820 outlines the steps involved in capturing and preparing the user's voice input for processing. This process begins when the system starts actively listening (Start Listening 124 in FIG. 1) and continues as long as the user is speaking, culminating in the transcription of speech for the LLM. The Speech Acquisition Process 820 includes ambient noise level detection and variable-length recording.

[0194] FIG. 9 illustrate an example of the internal processes of the Speech Acquisition Process 820. The trigger comes either Speech from user? 808 via line A or Interrupt LLM Speech 812 (may be via Option: Say ‘Go ahead.’ Etc. 814) via line B to To Speech Acquisition Process 822. Voice amplitude > ambient noise level? 824 is a real-time check within the Speech Acquisition Process 820 that compares the user's speech volume against the detected ambient noise level in the vehicle. This comparison dynamically adjusts the amplitude threshold for recording, ensuring that only genuine speech is captured while minimizing background noise interference. This feature allows for improved accuracy in voice input, particularly in noisy environments. Only when the speech volume level is higher than the ambient noise, the system go to the next step.

[0195] Start / Continue Voice Recording 826 is the action taken when the user's voice amplitude exceeds the ambient noise level, indicating active speech. The system either initiates a new AVAN.004PCT PATENT recording session or continues an ongoing one, capturing the user's entire utterance. This variablelength recording process ensures that speech is not prematurely cut off or allow the user to interrupt the conversation in the middle of the local LLM response announcement.

[0196] User silent for more than set duration? 828 is a decision point that monitors the duration of silence after a user has stopped speaking. If the silence persists beyond an adjustable threshold, it signals the end of the user's utterance, prompting the system to conclude the recording. This mechanism, part of the Speech Acquisition Process 820, prevents the recording of extraneous sounds after the user has finished their statement.

[0197] Stop Recording 830 is the action of ending the audio capture of user speech. This occurs when the User silent for more than set duration? 828 condition is met, ensuring that only the user's actual utterance is processed and that the recording is not unnecessarily prolonged.

[0198] Transcribe speech & input to LLM 832 is the crucial step where the captured and stopped voice recording is converted into text via the Speech To Text 126 (FIG. 1) component. The resulting textual input is then fed directly into the Domain Tuned LLM 134 (FIG. 1) for natural language understanding and response generation.

[0199] Go to LLM Response Generation Process 834 is the transition from speech transcription to the core intelligence phase, where the Domain Tuned LLM 134 (FIG. 1) processes the textual input, analyzes user intent, leverages contextual data, and formulates a coherent response. This process involves the entire workflow of the LLM, including data fetching and response creation. Once a the response is generated the contents is transmitted to the Speech To Text 126 and / or LLM agents 138 for Audio Out 142 or Display 148 in FIG. 1.

[0200] As additional embodiment, Speech Acquisition Process 820 may include a status for user to fmish / respond that entered when the LLM is speaking and detects that the user is not attempting to interrupt, or when the system has just responded and is awaiting a user's follow-up. This state demonstrates the system's capacity for turn-taking in a natural dialogue, allowing the user ample time to formulate their thoughts or react to the LLM's output. The duration of Wait for user to finish / respond can be dynamically adjusted based on conversational context or user's typical response times.

[0201] In the case of this embodiment, the following advantage can be obtained without limitation depending of specific features:

[0202] I. Improving the dialogue capabilities and situational awareness of the in-vehicle edge Al AVAN.004PCT PATENT that is more natural and has a deeper understanding of the user's intent.

[0203] 1. Enhanced User Interaction: Gaze / Facial Expression Activation and Intention Estimation: In addition to the conventional wake word activation, camera input will be used to achieve more intuitive interaction.

[0204] • Gaze Detection: Detects a "pause" (gaze) when the user looks towards the system or gazes at a specific area, and uses this as a trigger to start a conversation. Conversely, if the user's gaze is away from the specific area, the LLM will not interrupt the user's speech even if they speak. In this case, it is also possible to take into account the fact that the user's speech volume has decreased compared to before. This takes into account the possibility that the user may want to converse with another person while listening to the LLM's speech.

[0205] • Natural conversation through facial expression and gesture recognition: The Al model recognizes the user's facial expressions (interest, surprise, boredom, etc.) and hesitant expressions indicating that they want to speak through the camera. This allows for a deeper understanding of the user's intentions. For example, if the user shows signs of wanting to interrupt while the LLM is speaking, the LLM will naturally pause and give the user the turn to speak, thereby achieving a smoother conversation. When giving up the turn to speak, the LLM can say a reaction such as "Go ahead." • Variable-length and Accurate Voice Input:

[0206] • Voice input, which was previously fixed length, is now variable length based on a voice amplitude threshold based on ambient noise levels. Recording automatically stops when the user stops speaking for a certain period of time (adjustable), and the voice is processed as input to the LLM. This eliminates the issue of user speech being cut off midway.

[0207] • Context Understanding Through Conversation Tone Analysis:By analyzing the user's tone of speech, the LLM determines whether the speech is a question or simply sharing information (such as an impression). This enables accurate answers to questions and summarizing and collecting impressions and comments as passenger feedback.

[0208] 2. Enhanced Real-Time Integration of Vehicle and Environmental Information AVAN.004PCT PATENT

[0209] • Local Information Reference: The LLM can now directly reference the vehicle's system's time and calendar information, as well as vehicle information (location, speed, vehicle status, etc.) obtained from the GPS and 0BD2 port. This enables more context-sensitive, value-added responses.

[0210] • Conversation Memory Management via External Triggers: Upon receiving an external trigger from the bus system (e.g., an Al camera) indicating that a passenger has disembarked, the LLM's conversation memory is automatically cleared. This prevents a new passenger from continuing the conversation with the previous passenger, protecting privacy.

[0211] II. Enhanced Server-Side Content Generation and Management Functions

[0212] The content generation system is expanded to be more efficient and multifunctional.

[0213] • Introduction of a General -Purpose Automatic Podcast Creation Tool: A server-side tool is provided for administrators, allowing them to easily create local information podcasts without specialized knowledge. This tool is intended for use by a wide range of users, including developers themselves, local governments, and bus operators, and features an interface that is easy for anyone to use. Users simply specify an area and route on a map, and the system automatically collects information on highly rated points of interest (POIs) in the surrounding area and generates podcast content. A text summary is also provided, allowing administrators to easily review the content of the generated podcast without having to listen to the audio.

[0214] • Spontaneous Push Delivery of Urgent Information: Local LLMs generally operate offline, but for urgent information such as transportation delays, accidents, and bad weather, the server automatically generates podcasts and pushes them based on the vehicle's location information.

[0215] • Example: A podcast is broadcast to vehicles traveling near Shinagawa Station, announcing that Shinkansen service is suspended and offering alternative routes. The vehicle's local LLM then responds offline to any related questions from passengers. The LLM provides contact information as needed, allowing passengers to make inquiries using their own phones. AVAN.004PCT PATENT

[0216] ITI. Improved Content Delivery and In-Car Experience

[0217] Strengthen collaboration between the server and the vehicle (edge) to improve the passenger experience.

[0218] • Content Delivery with Text Data for LLM Training: Audio fdes (podcasts) generated on the server are delivered to vehicles with accompanying text data explaining the content. The vehicle's local LLM learns this text data in advance to fully understand the content of the audio content. This allows for instant answers to questions about specific information in audio content, such as "Whafs the address of the restaurant we just talked about?" without an internet connection.

[0219] • Content playback function as background music: Podcast data can be assigned a "BGM available" identifier. When the in-car microphone detects a state where passengers are not speaking, content with this identifier will automatically be played as background music. Passengers can ask questions to the local LLM at any time, even while background music is playing.

[0220] The embodiments of the present disclosure as disclosed herein are intended to be illustrative and not limiting. Other embodiments are possible and modifications may be made to the embodiments without departing from the spirit and scope of the disclosure. As such, these embodiments are only illustrative of the inventive concepts contained herein.

[0221] As a few examples of application of present disclosure::

[0222] Autonomous Vehicles and Buses:

[0223] In these settings, the system provides passengers with real-time information about driving decisions, route choices, and vehicle performance. For example, if a passenger asks, "Why did we take a detour from the usual route?", the system might respond, "We are avoiding congestion due to an accident ahead."

[0224] Class 8 Trucks:

[0225] The system functions as a co-pilot for drivers, assisting with navigation, regulatory compliance, vehicle diagnostics, and performance monitoring. The reduced cabin noise in zeroemission trucks enhances the driver's ability to interact with the system. For instance, if a driver AVAN.004PCT PATENT asks, "When is my next scheduled maintenance?", the system can reply, "Your next maintenance is due in 1,500 miles or two weeks, whichever comes first."

[0226] Restaurants:

[0227] In restaurant settings, the system enhances customer interactions by allowing open-ended questions and complex dialogues. It supports multiple languages and is specifically trained on the restaurant's unique menu, specials, ingredients, and policies. For example, if a customer asks, "I'm allergic to peanuts; what dishes are safe for me?", the system can provide a list of safe menu items and suggest suitable alternatives.

[0228] Industrial Environments:

[0229] In warehouses, manufacturing plants, construction sites, mining operations, agricultural fields, oil rigs, maritime vessels, and submarines, the system serves as an internal support tool. It provides workers with access to equipment diagnostics, real-time user manual support, process guidance, and safety protocols.

[0230] Equipment Support and Diagnostics:

[0231] The system interfaces with machinery and equipment via sensor outputs or diagnostic bus ports, monitoring real-time data to detect anomalies, performance issues, or maintenance needs.

[0232] Access to Manuals and Procedures:

[0233] Workers can request information on operating procedures, safety guidelines, and troubleshooting steps. The LLM provides step-by-step instructions and can handle open-ended queries.

[0234] Process Guidance and Training:

[0235] The system assists with onboarding new employees by providing training materials and interactive guidance, helping workers understand complex processes and workflows specific to the industrial environment.

[0236] Although the embodiments of the present application have been described above, the embodiment is presented as an example and is not intended to limit the scope of the present application. Such a novel embodiment can be implemented in various other forms, and can be omitted, replaced, and changed without departing from the gist of the disclosure. The embodiment and its modifications are included in the scope and gist of the aspects of the application.

Claims

AVAN.004PCT PATENTCLIAMS1. An edge computing interactive system capable of receiving data from an external server via a network, comprising: a compute module configured to host a local large language model, execute speech- to-text and text-to-speech functionalities, and tune the local large language model with custom domain datasets based on received data; one or more input devices communicatively coupled to the compute module, the one or more input devices configured to capture user input; one or more output devices communicatively coupled to the compute module, the one or more output devices configured to deliver responses generated by the compute module; and one or more artificial intelligence agents configured to process real-time data and coordinate delivery of responses through the one or more output devices wherein the system is configured to enable natural, contextual interaction within a working environment by processing user queries and generating responses locally without requiring uploading user-specific data to the external server.

2. The edge computing interactive system of claim 1, wherein the compute module installed in a vehicle and is further configured to receive real-time environmental data comprising podcast content from the external server, wherein the real-time environmental data represents conditions or events in near a location of the edge computer interactive system.

3. The edge computing interactive system of claim 2, wherein the compute module is further configured to receive a text file corresponding to the podcast content from the external server.

4. The edge computing interactive system of claim 2, wherein the compute module is configured play back the podcast contents a dialogue to be played backed by two speakers in different locations.

5. The edge computing interactive system of claim 1, wherein the one or more input devices include a camera configured to capture visual input of a user and detect facial expressions and gestures of the user.

6. The edge computing interactive system of claim 1, wherein the one or more input devices include a microphone configured to capture audio input from the user, wherein theAVAN.004PCT PATENT microphone is configured to capture variable-length voice input based on an amplitude threshold determined by an ambient noise level.

7. The edge computing interactive system of claim 1, wherein the compute module further comprises an activation system configured to initiate a dialogue based on at least one of a detected gaze of the user towards a specific area or a detected wake word.

8. The edge computing interactive system of claim 1, wherein the local large language model is further configured to analyze a tone of the user's speech to determine an intent of the user's statement, and to clear a conversation memory upon receiving an external trigger.

9. The edge computing interactive system of claim 8, wherein the local large language model's analysis of the user's speech tone distinguishes between a question and an information sharing statement.

10. The edge computing interactive system of claim 2, wherein the local large language model is configured to play the podcast content as background audio when the system detects a period of user inactivity, if the podcast content includes a background audio identifier.

11. The edge computing interactive system of claim 1, wherein the edge computing interactive system is installed in a bus for passenger transportation, and wherein: the local large language model includes a conversation memory for storing user interactions; and the compute module is further configured to: receive an external trigger from a bus system indicating that a passenger has disembarked; and automatically clear the conversation memory of the local large language model in response to the external trigger.

12. The edge computing interactive system of claim 1, wherein the edge computing interactive system is installed in a bus for passenger transportation, and wherein: the compute module is further configured to: operate the local large language model primarily offline; receive, from the external server, urgent information podcasts related to at least one of transportation delays, accidents, and bad weather, wherein the urgent information podcasts are generated automatically by the external server based on the bus's location information; andAVAN.004PCT PATENT present the received urgent information podcasts to the passengers.

13. The edge computing interactive system of claim 1, wherein: the system is configured to receive podcast content and related information from communication units installed on infrastructure, such as utility poles, equipped with Vehicle-to-Everything (V2X) capabilities; the system is further configured to establish peer-to-peer communication with the communication units when within range; and the compute module is configured to process and present the received podcast content and related information to users of the edge computing interactive system.

14. A method for operating an edge computing interactive system, the method comprising: tuning a local large language model hosted on a compute module with custom domain datasets based on received data; capturing user input via one or more input devices communicatively coupled to the compute module; processing the user input using the local large language model to generate a response; delivering the generated response through one or more output devices communicatively coupled to the compute module; processing real-time data and coordinating delivery of responses through the one or more output devices using one or more artificial intelligence agents; enabling natural, contextual interaction within a working environment by processing user queries and generating responses locally without requiring uploading userspecific data to an external server.

15. The method of claim 14, further comprising receiving, at the compute module, real-time environmental data comprising podcast content from the external server, wherein the realtime environmental data represents current conditions or events in the working environment.

16. The method of claim 15, further comprising receiving, at the compute module, a text file corresponding to the podcast content from the external server.

17. The method of claim 16, wherein tuning the local large language model comprises using custom domain datasets based on the received text file.AVAN.004PCT PATENT18. The method of claim 14, further comprising capturing visual input of a user using a camera and detecting facial expressions and gestures of the user using the camera.

19. The method of claim 14, further comprising capturing variable-length voice input based on an amplitude threshold determined by an ambient noise level.

20. The method of claim 14, further comprising initiating a dialogue based on at least one of a detected gaze of the user towards a specific area or a detected wake word.

21. The method of claim 14, further comprising: analyzing a tone of the user's speech to determine an intent of the user's statement; accessing local system time information, calendar information, and environmental information obtained from at least one sensor; and clearing a conversation memory upon receiving an external trigger.

22. The method of claim 15, further comprising playing the podcast content as background audio when detecting a period of user inactivity, if the podcast content is allowed to use as a background audio.

23. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor of an edge computing interactive system, cause the system to perform operations comprising: tuning a local large language model hosted on a compute module with custom domain datasets based on received data; capturing user input via one or more input devices communicatively coupled to the compute module; processing the user input using the local large language model to generate a response; delivering the generated response through one or more output devices communicatively coupled to the compute module; processing real-time data and coordinating delivery of responses through the one or more output devices using one or more artificial intelligence agents; enabling natural, contextual interaction within a working environment by processing user queries and generating responses locally without requiring uploading userspecific data to an external server.AVAN.004PCT PATENT24. The non-transitory computer-readable storage medium of claim 23, wherein the operations further comprise: receiving real-time environmental data comprising podcast content from an external server, wherein the real-time environmental data represents current conditions or events in the working environment; and receiving a text file corresponding to the podcast content from the external server.

25. The non-transitory computer-readable storage medium of claim 23, wherein the operations further comprise playing the podcast content as background audio when detecting a period of user inactivity, if the podcast content includes a background audio identifier.

26. A server for an edge computing interactive system, comprising: a processor; a memory coupled to the processor; a server-side large language model stored in the memory and executed by the processor, the server-side large language model configured to generate content based on specified criteria, and generate a text file related to the generated content, and prepare the content and the text file for transmission to an edge computing device; a network interface configured to transmit the content and the text file to the edge computing device; and wherein the server is configured to: provide custom domain datasets for tuning a local large language model on the edge computing device, and operate without receiving information included in user queries from the edge computing device.

27. The server of claim 26, wherein the content generated by the server-side large language model is podcast content.

28. The server of claim 27, wherein the server-side large language model is further configured to generate the podcast content based on user preferences received from the edge computing device.

29. The server of claim 27, wherein: the text file contains the content of the generated podcast;AVAN.004PCT PATENT the text file includes information indicating whether the podcast is usable as background music (BGM); and the server-side large language model is further configured to update the podcast content and text file based on real-time information.

30. The server of claim 27, wherein the server-side large language model is configured to generate the podcast content in response to a request from the edge computing device.

31. The server of claim 27, wherein the server-side large language model is configured to periodically generate and transmit updated podcast content and text files to the edge computing device.

32. The server of claim 26, wherein the server is further configured to: receive location information from the edge computing device; and tailor the content based on the received location information.

33. The server of claim 27, wherein the server-side large language model is configured to generate podcast content related to at least one of: traffic updates, weather forecasts, local news, and points of interest near the location of the edge computing device.

34. The server of claim 27, wherein the server is further configured to push the podcast content and related information to communication units installed on infrastructure, such as utility poles, equipped with Vehicle-to-Everything (V2X) capabilities; wherein the communication units are configured to transmit the podcast content and related information to the edge computing device via peer-to-peer communication when the edge computing device is within range of the communication units.

Citation Information

Patent Citations

  • Systems for controllable summarization of content

    US12008332B1

  • Intelligent digital assistant system

    US20180233139A1

  • Method and apparatus for activating speech recognition

    US20210056974A1

  • Content injection using a network appliance

    US20220014579A1

  • Room sounds modes

    US20220343935A1