An adaptive audio-visual assistant system

WO2026190525A1PCT designated stage Publication Date: 2026-09-17DEWAN MOHAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/062713
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-13
Filing Date
2025-12-11
Publication Date
2026-09-17

Smart Images

  • Figure IB2025062713_17092026_PF_FP_ABST
    Figure IB2025062713_17092026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an adaptive audio-visual assistant system (100) configured to provide real-time, context-aware, multilingual, and user-adaptive guidance on a portable electronic device (142) The system (100) comprises a voice processing module (102) configured to process speech input, and a context detection unit (104) configured to fuse sensor data from a touch sensor (126), an inertial measurement unit (IMU) (128), system APIs (130), and a front-facing camera (150) to generate a structured digital context packet. A guidance generation unit (106) processes the packet to generate a digital guidance instruction set, rendered synchronously by an audio-visual coordination unit (108) on a display (136) and a speaker (138). A speech-style adaptation unit (110) and a multilingual translation module (116) generate personalized multilingual output. An offline operation unit (112), a dynamic control unit (114), an on-demand activation controller (122), and a hardware accelerator (124) enable continuous, low-latency, and resource-efficient operation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AN ADAPTIVE AUDIO-VISUAL ASSISTANT SYSTEM

[0002] FIELD

[0003] The present disclosure generally relates to the field of artificial intelligence. More particularly, the present disclosure relates to an adaptive audio-visual assistant system.

[0004] DEFINITION

[0005] As used in the present disclosure, the following terms are generally intended to have the meaning as set forth below, except to the extent that the context in which they are used indicates otherwise.

[0006] Angle of Arrival (AoA): The term “Angle of Arrival (AoA) ” refers to the estimated direction from which a speech signal reaches the microphone array, used to focus beamforming on the user's voice.

[0007] Augmented Reality (AR) Techniques: The term “Augmented Reality (AR) Techniques” refers to the methods that overlay visual elements like highlights, animations, or cues onto a device interface or environment to enhance user guidance.

[0008] Beamforming: The term “beamforming” refers to a signal -processing technique that combines inputs from multiple microphones to enhance sound arriving from a desired direction while suppressing noise from other directions.

[0009] Context-Aware Assistant: The term “Context-Aware Assistant” refers to an assistant capable of recognizing the user’s current activity or device state and providing relevant, realtime assistance based on that context.

[0010] Fundamental Frequency (F0): The term “Fundamental Frequency (F0)” refers to the lowest frequency component of a speech waveform, perceived as the pitch of the user's voice.

[0011] Hidden Markov Model (HMM): The term “Hidden Markov Model (HMM)” refers to a probabilistic model used to represent sequences of observable signals generated by hidden states, commonly used for speech-to-text conversion in automatic speech recognition.Inertial Measurement Unit (IMU): The term “Inertial Measurement Unit (IMU)” refers to a motion-sensing component that measures acceleration, rotation, and device orientation to detect gestures and physical movements.

[0012] Matrix-Multiplication Cores: The term “matrix-multiplication cores” refers to specialized hardware units optimized for performing large-scale matrix operations used in neural network processing.

[0013] Multilingual: The term “Multilingual” refers to the ability of a system to understand, process, and respond in multiple languages or dialects, with automatic detection and translation capability.

[0014] Named Entity Recognition (NER): The term “Named Entity Recognition (NER) ” refers to a natural language processing method that identifies and classifies key information in text, such as commands, objects, or user intents.

[0015] Neural Inference: The term “neural inference ” refers to the process of executing a trained neural network model on new input data to generate outputs such as predictions, classifications, or guidance instructions.

[0016] Offline Operation: The term “Offline Operation” refers to the capability of the system to function fully without internet connectivity by using locally stored data and resources within the device.

[0017] Phoneme Duration: The term “phoneme duration” refers to the temporal length of individual speech sounds, which affects the rhythm and pacing of synthesized speech.

[0018] Prosody Parameters: The term “prosody parameters” refers to acoustic features of speech, such as pitch, rhythm, stress, and intonation, used to shape natural, expressive spoken output.

[0019] Speech-Style Adaptation: The term “Speech-Style Adaptation” refers to the ability of the system to adjust its speech tone, phrasing, and rhythm to match the user’s natural communication style for a personalized interaction.

[0020] Statistical Machine Translation (SMT) Model: The term “Statistical Machine Translation (SMT) Model ” refers to a language translation method that predicts the most likely translated text using statistical patterns learned from bilingual datasets.The above definitions are in addition to those expressed in the art.

[0021] BACKGROUND

[0022] The background information herein below relates to the present disclosure but is not necessarily prior art.

[0023] The rapid increase of electronic devices, such as mobile phones, laptops, personal computers, and other portable or interactive systems, has transformed how users interact with technology in daily life. These devices integrate a wide range of complex features, such as email management, file organisation, camera operation, and shortcut navigation, that often require users to learn intricate interfaces and settings to achieve desired outcomes.

[0024] However, existing assistance systems that aim to simplify such interactions exhibit several limitations. Many of these systems are restricted to a single language or a limited set of languages, creating accessibility barriers for users with diverse linguistic preferences. Conventional assistants also rely heavily on continuous internet connectivity, making them unreliable or entirely non-functional in offline environments such as remote areas or during network disruptions, where users may still require guidance for basic operations like capturing photos or locating downloaded files.

[0025] Furthermore, conventional systems lack robust context-awareness and audiovisual integration. They typically deliver generic or pre-scripted responses that do not adapt to the specific feature or function being used at a given moment, such as adjusting a camera angle, navigating file directories, or executing application shortcuts. This results in user frustration, extended learning curves, and reduced task efficiency, particularly for individuals who are new to a device or who benefit from visual reinforcement alongside verbal guidance.

[0026] In addition, many known assistants operate in a resource-intensive manner, consuming significant processing power and battery life due to the absence of adaptive control or timebound activation mechanisms. Their limited configurability further prevents users from selectively enabling or disabling assistance based on proficiency, task complexity, or preference.

[0027] Therefore, there is felt a need for an adaptive audio-visual assistant system to alleviate the above-mentioned drawbacks.OBJECTS

[0028] Some of the objects of the present disclosure, which at least one embodiment herein satisfies, are as follows:

[0029] An object of the present disclosure is to provide an adaptive audio-visual assistant system. Another object of the present disclosure is to provide a system that delivers synchronized audio and visual guidance outputs to the user.

[0030] Still another object of the present disclosure is to provide a system that offers real-time, context-aware assistance through synchronized audio and visual guidance to enhance user interaction with electronic devices.

[0031] Yet another object of the present disclosure is to provide a system that operates seamlessly in offline environments without dependence on network connectivity.

[0032] Still another object of the present disclosure is to provide a system that adapts its speech tone, phrasing, and delivery style based on individual user speech patterns and linguistic characteristics.

[0033] Yet another object of the present disclosure is to provide a system that supports a plurality of languages and dialects to ensure accessibility across diverse user groups.

[0034] Still another object of the present disclosure is to provide a system that incorporates dynamic control functionality enabling selective activation, suspension, or reactivation of the assistant based on user proficiency or task completion.

[0035] Yet another object of the present disclosure is to provide a system that integrates audiovisual guidance with feedback and learning capabilities for a personalized user experience.

[0036] Still another object of the present disclosure is to provide a system that optimizes power and processing resources through adaptive operation and time-limited activation.

[0037] Yet another object of the present disclosure is to provide a system that supports gesture-based and voice-based triggering mechanisms for intuitive on-demand activation.

[0038] Still another object of the present disclosure is to provide a system that is compatible across multiple platforms.Other objects and advantages of the disclosure will be more apparent from the following description, which is not intended to limit the scope of the disclosure.

[0039] SUMMARY

[0040] The present disclosure envisages an adaptive audio-visual assistant system for a portable electronic device. The system comprises a context detection unit configured to generate a structured digital context packet by fusing real-time data streams and active application state data from system -level application programming interfaces (APIs). A guidance generation unit communicatively coupled with the context detection unit and configured to process the structured digital context packet through a pre-trained neural network model to generate a digital guidance instruction set. An audio-visual coordination unit communicatively coupled with the guidance generation unit, and configured to command a concurrent rendering of a visual output on a touch-sensitive display and a speech output on a speaker of the portable electronic device, based on the digital guidance instruction set. A speech-style adaptation unit communicatively coupled with the guidance generation unit and the audio-visual coordination unit, and configured to process the textual command string and generate a speech output to match a user-specific profile stored in a user profile database. An offline operation unit stored in a non-volatile storage medium of the portable electronic device and operatively coupled to the guidance generation unit and the speech-style adaptation unit, the offline operationconfigured to enable full system operation in an offline mode. A dynamic control unit communicatively coupled to the context detection unit and the guidance generation unit, the dynamic control unit is configured to optimize power consumption by implementing a state machine.

[0041] In an embodiment, the system further comprises a visual assistance module operatively coupled to the guidance generation unit and the audio-visual coordination unit, the visual assistance module configured to generate a visual overlay instruction set by mapping the target UI element identifier to absolute screen coordinates and rendering a semi-transparent graphical highlight using an augmented reality (AR) technique that overlays a specific user interface element on the display.

[0042] In an embodiment, the system further comprises a hardware accelerator, comprising a neural processing unit (NPU) with dedicated matrix multiplication cores, operatively coupled to the system via a high-speed bus, wherein the hardware accelerator is configured to:• execute inference computations for the pre-trained neural network model in the guidance generation unit; and

[0043] • perform real-time digital signal processing for the prosody model in the speech-style adaptation unit.

[0044] In an embodiment, the context detection unit comprises:

[0045] • a feature detection engine configured to parse a system event log and a frame buffer to identify the active application and a specific feature in use; and • a gesture recognition extension, implementing a convolutional neural network (CNN), configured to process image frames from a front-facing camera to interpret dynamic hand gestures for system control.

[0046] In an embodiment, the real-time data streams comprise touch events detected by a touch sensor and motion vectors captured by an inertial measurement unit (IMU).

[0047] In an embodiment, the guidance generation unit is communicatively coupled with the context detection unit via an inter-process communication (IPC) bus.

[0048] In an embodiment, the digital guidance instruction set generated by the guidance generation unit comprises a textual command string and a target user interface (UI) element identifier. In an embodiment, the speech-style adaptation unit is configured to apply a prosody model that dynamically adjusts fundamental frequency (FO) contours, phoneme duration, and spectral features of the speech output.

[0049] In an embodiment, the offline operation unit further maintains a local encrypted database comprising multilingual acoustic models, a library of visual asset templates, and parameters for the pre-trained neural network model and performs bidirectional synchronization with a cloud server upon detection of a secure network connection.

[0050] In an embodiment, the state machine of the dynamic control unit selectively disables the guidance generation unit for specific device functions upon determining that a user proficiency score exceeds a predefined threshold, the user proficiency score being derived from a success rate of historical user interactions.In an embodiment, the system further comprises a voice processing module operatively coupled to a multi-microphone array of the portable electronic device and the context detection unit, the voice processing module configured to:

[0051] • perform beamforming and noise suppression on raw audio input signals; • convert the enhanced audio signals into text through a hidden Markov model (HMM)-based automatic speech recognition (ASR) engine; and • perform named entity recognition (NER) on the text to extract user intent and update the structured digital context packet.

[0052] In an embodiment, the system further comprises a multilingual translation module communicatively coupled between the guidance generation unit and the speech-style adaptation unit, the multilingual translation module configured to:

[0053] • detect a user’s preferred language by analyzing a system locale setting and a language n-gram distribution from historical user input;

[0054] • translate the textual command string from the digital guidance instruction set into the preferred language through a statistical machine translation (SMT) model; and

[0055] • provide the translated textual command string to the speech-style adaptation unit.

[0056] In an embodiment, the system further comprises a user feedback and learning module communicatively coupled to the dynamic control unit, the user feedback and learning module configured to:

[0057] • record timestamped user interaction events, including task completion times and error rates, in a structured log file;

[0058] • calculate the user proficiency score using a moving average filter applied to the success rate over a sliding window of the most recent N interactions; and • transmit the updated user proficiency score to the dynamic control unit to adjust its activation policy.In an embodiment, the dynamic control unit further comprises a time-limited assistance controller configured to:

[0059] • initiate a countdown timer upon activation of the guidance generation unit;

[0060] and

[0061] • trigger a low-power interrupt to a central processing unit (CPU) of the portable electronic device to disable the guidance generation unit when the countdown timer reaches zero without a subsequent user interaction event from the context detection unit.

[0062] In an embodiment, the system further comprises an on-demand activation controller operatively coupled to the gesture recognition extension and the dynamic control unit, the on-demand activation controller configured to:

[0063] • monitor for a predefined reactivation trigger signature from a voice keyword or a specific hand gesture; and

[0064] • upon verification of the trigger signature, send a system call to the dynamic control unit to re-enable the guidance generation unit.

[0065] In an embodiment, the predefined reactivation trigger signature is derived from a reference data set stored in the non-volatile storage medium, the reference data set comprising:

[0066] • a specific audio waveform pattern for the voice keyword, obtained during an initial user enrollment phase; and

[0067] • a specific sequence of spatial coordinates representing the hand gesture, captured and stored as a template by the gesture recognition extension.

[0068] In an embodiment, the predefined threshold for the user proficiency score is a configurable parameter stored in the user profile database, settable by at least one of: a system administrator via a configuration file, or the user through an accessibility settings menu of the portable electronic device.

[0069] In an embodiment, the pre-trained neural network model of the guidance generation unit is a supervised learning model, initially trained on a curated dataset of annotated user interactions prior to its deployment on the portable electronic device, wherein the curated datasetcomprises a plurality of data pairs, each pair consisting of an input context vector derived from historical sensor and application state data, and a corresponding labeled output vector representing a validated guidance instruction.

[0070] In an embodiment, the portable electronic device comprises:

[0071] • a system-on-chip (SoC) comprising a central processing unit (CPU) and a graphics processing unit (GPU);

[0072] • a non-volatile storage medium;

[0073] • a touch-sensitive display;

[0074] • a multi-microphone array, an audio codec, and a speaker;

[0075] • a sensor hub comprising a touch sensor, an inertial measurement unit (IMU), and the front-facing camera; and

[0076] • the adaptive audio-visual assistant system.

[0077] The present disclosure further envisages a method for providing adaptive audio-visual assistance on a portable electronic device, the method comprises the steps of:

[0078] • receiving, at a context detection unit, real-time data streams and active application state data from system-level application programming interfaces (APIs) and generating a structured digital context packet based on the data;

[0079] • processing, at a guidance generation unit, the structured digital context packet through a pre-trained neural network model to generate a digital guidance instruction set;

[0080] • commanding, via an audio-visual coordination unit, a concurrent rendering of a visual output on a touch-sensitive display and a speech output on a speaker of the portable electronic device, based on the digital guidance instruction set;

[0081] • processing, at a speech-style adaptation unit, a textual command string of the digital guidance instruction set to generate a speech output matching a userspecific profile stored in a user profile database;• enabling, via an offline operation unit, full system operation in an offline mode using data stored in a non-volatile storage medium; and • optimizing, via a dynamic control unit, power consumption by implementing a state machine configured to manage activation of the guidance generation unit. In an embodiment, the method further comprises generating the structured digital context packet comprises sampling raw sensor signals from a touch sensor, an inertial measurement unit (IMU), and a front-facing camera, and pre-processing the sampled signals by applying digital filtering, feature extraction, and data fusion.

[0082] In an embodiment, the digital guidance instruction set comprises a textual command and a target user interface (UI) element identifier.

[0083] In an embodiment, the method further comprises generating, via a visual assistance module, a visual overlay instruction set by mapping the target UI element identifier to screen coordinates.

[0084] In an embodiment, the method further comprises processing the textual command through the speech-style adaptation unit comprises applying a prosody model to synthesize a speech waveform with user-adapted acoustic characteristics.

[0085] BRIEF DESCRIPTION OF THE ACCOMPANYING DRAWING

[0086] An adaptive audio-visual assistant system, of the present disclosure will now be described with the help of the accompanying drawing, in which:

[0087] Figure 1 illustrates a block diagram of the adaptive audio-visual assistant system, in accordance with an embodiment of the present disclosure; and

[0088] Figures 2A-2B illustrate a flowchart of a method for providing adaptive audio-visual assistance on a portable electronic device, in accordance with an embodiment of the present disclosure.

[0089] LIST OF REFERENCE NUMERALS

[0090] 100 System

[0091] 102 Voice Processing Module104 Context Detection Unit

[0092] 104a Feature Detection Engine

[0093] 104b Gesture Recognition Extension

[0094] 106 Guidance Generation Unit

[0095] 108 Audio-visual Coordination Unit 108a Visual Assistance Module

[0096] 110 Speech-style Adaptation Unit

[0097] 112 Offline Operation Unit

[0098] 114 Dynamic Control Unit

[0099] 116 Multilingual Translation Module

[0100] 118 User Feedback and Learning Module 120 Time-limited Assistance Controller 122 On-Demand Activation Controller 124 Hardware Accelerator

[0101] 124a Neural Processing Unit (NPU)

[0102] 126 Touch Sensor

[0103] 128 Inertial measurement unit (IMU)

[0104] 130 System-level APIs

[0105] 132 Inter-process communication (IPC) bus 134 Graphics processing unit (GPU)

[0106] 136 Touch-sensitive display138 Speaker

[0107] 140 User profile database

[0108] 142 Portable electronic device

[0109] 144 Central processing unit (CPU)

[0110] 146 Non-volatile storage medium

[0111] 148 Multi-microphone array

[0112] 150 Front-facing camera

[0113] 152 ASR engine (HMM -based)

[0114] 154 High-speed bus

[0115] DETAILED DESCRIPTION

[0116] Embodiments, of the present disclosure, will now be described with reference to the accompanying drawing.

[0117] Embodiments are provided so as to thoroughly and fully convey the scope of the present disclosure to the person skilled in the art. Numerous details, are set forth, relating to specific components, and methods, to provide a complete understanding of embodiments of the present disclosure. It will be apparent to the person skilled in the art that the details provided in the embodiments should not be construed to limit the scope of the present disclosure. In some embodiments, well-known processes, well-known apparatus structures, and well-known techniques are not described in detail.

[0118] The terminology used, in the present disclosure is only for the purpose of explaining a particular embodiment and such terminology shall not be considered to limit the scope of the present disclosure. As used in the present disclosure, the forms "a,” "an," and "the" may be intended to include the plural forms as well, unless the context clearly suggests otherwise. The terms “including,” and “having,” are open-ended transitional phrases and therefore specify the presence of stated features, elements and / or components, but do not forbid the presence or addition of one or more other features, elements, components, and / or groups thereof.The terms first, second, third, etc., should not be construed to limit the scope of the present disclosure, as the aforementioned terms may be only used to distinguish one element, component, region, layer or section from another component, region, layer or section. Terms such as first, second, third, etc., when used herein, do not imply a specific sequence or order unless clearly suggested by the present disclosure.

[0119] The widespread use of portable electronic devices such as smartphones, laptops, and tablets has introduced complex features that often challenge users due to intricate interfaces and limited in-built assistance. Existing systems typically support only a few languages, depend on continuous internet connectivity, and lack real-time context-awareness or audiovisual integration, resulting in generic guidance, poor adaptability, and user frustration. Moreover, conventional assistants operate with high resource consumption and offer limited flexibility in activation or personalization, reducing efficiency and accessibility.

[0120] To address the issues of the existing systems and methods, the present disclosure envisages a multilingual adaptive audio-visual assistant system (hereinafter referred to as “system (100)”). The system (100) will now be described with reference to Figure 1.

[0121] The present disclosure relates to an adaptive audio-visual assistant system (100) configured to provide real-time, context-aware, multilingual, and user-adaptive assistance on a portable electronic device (142). The system (100) is configured to deliver synchronized audio and visual outputs, to analyse multimodal sensor data, to generate personalised speech output, and to operate seamlessly in offline environments. The system (100) further incorporates adaptive learning, gesture-based and voice-based triggering, and resource-optimised activation to enhance user interaction efficiency and accessibility.

[0122] In an embodiment, the system (100) is implemented as software, firmware, hardware circuitry, or any combination thereof, integrated within a portable electronic device (142) such as a smartphone, tablet, smartwatch, or computing device.

[0123] The system (100) comprises a plurality of operatively connected modules including a voice processing module (102), a context detection unit (104), a guidance generation unit (106), an audio-visual coordination unit (108), a speech-style adaptation unit (110), an offline operation unit (112), a dynamic control unit (114), a multilingual translation module (116), a user feedback and learning module (118), a time-limited assistance controller (120), an on-demand activation controller (122), and a hardware accelerator (124) comprising a neuralprocessing unit (124a). The components communicate via an inter-process communication (IPC) bus (132), a high-speed bus (154), or internal processor buses, as applicable.

[0124] The portable electronic device (142) housing the adaptive audio-visual assistant system (100) comprises a system-on-chip (SoC) incorporating a central processing unit (CPU) (144) and a graphics processing unit (GPU) (134), configured to execute application logic, graphical rendering, neural inference tasks, and audio-visual output coordination. The portable electronic device (142) further comprises a non-volatile storage medium (146) configured to store operating system files, application data, encrypted model parameters, multilingual acoustic libraries, visual asset templates, and user profile information required for operation of the system (100). The device (142) additionally includes a touch-sensitive display (136) configured to receive user touch inputs and to render visual overlays, highlights, and augmented reality (AR) visual assets generated by the system (100). A multi-microphone array (148), an audio codec, and a speaker (138) are provided and are configured to capture high-quality speech input and to reproduce synthesized speech output generated by the system (100). The portable electronic device (142) further comprises a sensor hub including a touch sensor (126), an inertial measurement unit (IMU) (128), and a front-facing camera (150), each configured to supply real-time sensor data pertaining to user interaction, motion, and gestures. These hardware components collectively support the operation of the adaptive audio-visual assistant system (100), enabling context detection, guidance generation, audiovisual synchronisation, and real-time user assistance across various operational environments. The voice processing module (102) is operatively coupled to the multi-microphone array (148) of the portable electronic device (142) and is configured to receive raw audio signals generated in response to user speech. The voice processing module (102) is configured to perform beamforming to spatially filter the received audio signals and enhance the dominant speech source, and is further configured to apply noise suppression and echo cancellation to remove ambient noise, reverberation, and device -generated audio artefacts.

[0125] In an embodiment, the voice processing module (102) is configured to convert the enhanced audio signals into a text output using a hidden Markov model (HMM)-based automatic speech recognition (ASR) engine (152).

[0126] The voice processing module (102) is additionally configured to process the recognized text using a named entity recognition (NER) engine to extract semantic features, user intent,command keywords, and contextual linguistic parameters. The processed text and extracted intent data generated by the voice processing module (102) are configured to be transmitted to the context detection unit (104) for integration into the structured digital context packet. In an embodiment, the voice processing module (102) is configured to utilize adaptive beamforming techniques that dynamically adjust microphone directionality based on the estimated angle of arrival (Ao A) of the user’s speech, thereby isolating speech from competing background sources.

[0127] In another embodiment, the voice processing module (102) is configured to select languagespecific acoustic models from the offline operation unit (112) based on a language preference determined by the multilingual translation module (116).

[0128] In yet another embodiment, the voice processing module (102) is configured to generate confidence scores for recognized words and to transmit these confidence values to the guidance generation unit (106), thereby enabling context-aware fallback strategies such as requesting clarification or simplifying subsequent guidance instructions.

[0129] The context detection unit (104) is operatively coupled to the voice processing module (102), the touch sensor (126), the inertial measurement unit (IMU) (128), the system-level application programming interfaces (APIs) (130), and the front-facing camera (150). The context detection unit (104) is configured to continuously acquire real-time interaction data comprising touch events, motion vectors, application state information, and gesture imagery. The context detection unit (104) is further configured to fuse the acquired multimodal data into a structured digital context packet through digital filtering, time-alignment, sensor fusion, and contextual annotation. The structured digital context packet is transmitted to the guidance generation unit (106) via an inter-process communication (IPC) bus (132).

[0130] In an embodiment, the structured digital context packet generated by the context detection unit (104) is configured to represent the user’s current operational context, including the active application, the specific user interface element in focus, the detected physical gesture, and the semantic intent contributed by the voice processing module (102).

[0131] In an embodiment, the context detection unit (104) comprises a feature detection engine (104a) configured to parse system event logs, frame buffer metadata, and UI hierarchy data obtained from the system APIs (130) to identify the active application and a specific featureor function currently in use. The feature detection engine (104a) is configured to extract interface -level attributes such as button identifiers, menu selections, camera mode states, or file navigation paths and to incorporate these attributes into the structured digital context packet.

[0132] In an embodiment, the context detection unit (104) further comprises a gesture recognition extension (104b) configured to interpret dynamic gestures by processing sequential image frames acquired from the front-facing camera (150). The gesture recognition extension (104b) implements a convolutional neural network (CNN) model configured to detect motion trajectories, hand shapes, and positional variations associated with predefined gesture patterns. The gesture recognition extension (104b) is further configured to generate a gesture signature that is transmitted to the on-demand activation controller (122) for triggering or reactivating the system (100) based on user-performed gestures.

[0133] In an embodiment, the context detection unit (104) is further configured to apply temporal smoothing filters and a confidence-weighting mechanism that prioritises stable and high-certainty signals when generating the structured digital context packet.

[0134] In another embodiment, the context detection unit (104) is configured to track user interaction sequences across time, enabling recognition of compound actions such as multi-step navigation, long -press gestures, or rotational device movements. In yet another embodiment, the context detection unit (104) is configured to identify error-prone user behaviours, such as repeated incorrect taps or unstable device motion, and to annotate the structured digital context packet with error flags to facilitate corrective guidance by the guidance generation unit (106).

[0135] In an embodiment, the context detection unit (104) is further configured to obtain, via the system-level APIs (130), application metadata including feature descriptions, permission states, operational capabilities, and user-accessible functions associated with applications installed on the portable electronic device (142). The context detection unit (104) is configured to analyse the metadata and generate an application-capability map describing the functional depth of each installed application. The application-capability map is configured to be incorporated into the structured digital context packet and supplied to the guidance generation unit (106) for generating application-specific or feature-level assistance prompts.The guidance generation unit (106) is operatively coupled to the context detection unit (104) via the inter-process communication (IPC) bus (132) and is configured to receive the structured digital context packet. The guidance generation unit (106) is configured to process the packet using a pre-trained neural network model stored on the portable electronic device (142) to generate a digital guidance instruction set comprising a textual command string and a target user interface (UI) element identifier. The guidance instruction set is configured to be transmitted to the audio-visual coordination unit (108), the multilingual translation module (116), the speech-style adaptation unit (110), and the visual assistance module (108a).

[0136] In an embodiment, the guidance generation unit (106) comprises an Al-based guidance module (106a) configured to execute neural inference using the hardware accelerator (124) to reduce latency.

[0137] In another embodiment, the guidance generation unit (106) is configured to adjust the level of detail in its instructions based on the user's proficiency score supplied by the dynamic control unit (114).

[0138] In yet another embodiment, the guidance generation unit (106) is configured to incorporate error indicators from the structured digital context packet to generate corrective or stabilising guidance prompts.

[0139] In a further embodiment, the guidance generation unit (106) is configured to analyse the application-capability map transmitted by the context detection unit (104) and to generate feature-specific guidance for an active application. The guidance generation unit (106) is configured to determine whether the user is attempting to access a feature that is advanced, nested within multiple interface layers, or unfamiliar, and to generate corresponding instructions that assist the user in locating, activating, or operating such features.

[0140] In another embodiment, the guidance generation unit (106) is configured to compare functional capabilities of an installed application with those of a newly requested or newly downloaded application by evaluating feature descriptors, operational metadata, and usage data obtained via the system-level APIs (130). The guidance generation unit (106) is further configured to generate a digital guidance instruction set indicating whether the newly considered application provides additional functionality, lacks requisite features, or duplicates capabilities of the installed application, thereby facilitating informed installation or deletion decisions by the user.In yet another embodiment, the guidance generation unit (106) is configured to generate recommendations relating to application management by analysing storage availability, application redundancy, and usage patterns derived from the user feedback and learning module (118). The guidance generation unit (106) is configured to generate a digital guidance instruction suggesting deletion of underutilised or redundant applications or warning the user when installation of a new application may exceed available storage or replicate existing functionality.

[0141] In still another embodiment, the guidance generation unit (106) is configured to detect usage of an artificial intelligence (Al) tool through application state data and system API interactions, and to prompt the user to specify an intended purpose of using the Al tool. Based on the received purpose, the guidance generation unit (106) is configured to identify one or more Al tools installed on the portable electronic device (142), or available through online services, that are best suited for the intended purpose, and to generate a digital guidance instruction recommending the preferred Al tool or workflow.

[0142] The audio-visual coordination unit (108) is communicatively coupled to the guidance generation unit (106) and is configured to receive the digital guidance instruction set comprising the textual command string and the target UI element identifier. The audio-visual coordination unit (108) is configured to perform concurrent rendering of:

[0143] i. visual output on the touch-sensitive display (136) via the graphics processing unit (GPU) (134), and

[0144] ii. speech output through an audio codec to the speaker (138) of the portable electronic device (142).

[0145] The audio-visual coordination unit (108) is configured to synchronise the timing of the visual and auditory outputs to ensure coherent multimodal guidance delivery.

[0146] In an embodiment, the audio-visual coordination unit (108) comprises a visual assistance module (108a) configured to map the target UI element identifier to absolute display coordinates and to generate a visual overlay instruction set. The visual assistance module (108a) is further configured to render semi-transparent graphical highlights or augmented reality (AR) techniques onto the touch-sensitive display (136) to visually reinforce the guidance instructions.In another embodiment, the audio-visual coordination unit (108) is configured to adjust output intensity, animation style, or overlay size based on user settings or accessibility preferences.

[0147] The speech-style adaptation unit (110) is communicatively coupled to the guidance generation unit (106), the multilingual translation module (116), and the audio-visual coordination unit (108). The speech-style adaptation unit (110) is configured to receive the textual command string and to process the text using a prosody model configured to dynamically adjust fundamental frequency (F0), phoneme duration, pacing, and spectral characteristics of synthesized speech. The speech-style adaptation unit (110) is further configured to generate a speech signal tailored to the linguistic and acoustic preferences of the user, as defined in a user-specific linguistic profile stored in the user profile database (140). The generated speech signal is configured to be transmitted to the audio-visual coordination unit (108) for synchronized playback with visual overlays.

[0148] In an embodiment, the speech-style adaptation unit (110) is configured to utilize prosody parameters that adapt over time based on user interaction patterns received from the user feedback and learning module (118).

[0149] In another embodiment, the speech-style adaptation unit (110) is configured to employ hardware acceleration through the neural processing unit (124a) to perform real-time digital signal processing for prosody transformation, thereby reducing computational load on the central processing unit (CPU) (144).

[0150] In yet another embodiment, the speech-style adaptation unit (110) is configured to switch between formal, casual, concise, or expressive speaking styles based on predefined user settings or contextual cues detected in the structured digital context packet.

[0151] The multilingual translation module (116) is operatively connected to the guidance generation unit (106) and the speech-style adaptation unit (110) and is configured to translate textual guidance instructions into a user-preferred language prior to speech synthesis. The multilingual translation module (116) is configured to receive the textual command string generated by the guidance generation unit (106), the textual command representing the system-generated instruction intended for audio-visual delivery. Upon receiving the textual command string, the multilingual translation module (116) is configured to analyse thesystem locale settings of the portable electronic device (142) as well as historical user linguistic patterns obtained from interaction logs to determine the user’s preferred language. The multilingual translation module (116) is further configured to translate the textual command string into the identified language using a statistical machine translation (SMT) model stored within the non-volatile storage medium (146). The multilingual translation module (116) is configured to generate a translated textual output that preserves semantic accuracy and contextual relevance with respect to the structured digital context packet originally processed by the guidance generation unit (106). The translated textual command string is thereafter transmitted to the speech-style adaptation unit (110), which applies prosody and user-specific linguistic features for generating the final speech signal.

[0152] In an embodiment, the multilingual translation module (116) is configured to support a plurality of languages and dialects, including regionally adapted vocabulary sets, enabling real-time switching between languages without disrupting the continuity of interaction.

[0153] In another embodiment, the multilingual translation module (116) is configured to store language models, translation rules, and bilingual dictionaries within a local encrypted database maintained by the offline operation unit (112), thereby enabling full translation functionality during offline operation.

[0154] In yet another embodiment, the multilingual translation module (116) is configured to update its translation models through periodic cloud synchronization facilitated by the offline operation unit (112) whenever a secure network connection is available.

[0155] The offline operation unit (112) is stored within the non-volatile storage medium (146) of the portable electronic device (142) and is configured to ensure uninterrupted operation of the adaptive audio-visual assistant system (100) when network connectivity is unavailable. The offline operation unit (112) is configured to provide local access to multilingual acoustic models, visual asset templates, translation resources, prosody parameters, and neural network inference parameters required by the guidance generation unit (106) and the speech-style adaptation unit (110).

[0156] In operation, the offline operation unit (112) is configured to supply the guidance generation unit (106) with pre-stored neural network parameters, language models, and visual instruction templates necessary for generating the digital guidance instruction set during offlineconditions. The offline operation unit (112) further provides the speech-style adaptation unit (110) with locally stored acoustic models and prosodic configuration datasets that enable the unit to synthesize natural-sounding speech without external server support. The offline operation unit (112) thereby acts as a local repository that maintains consistency of system behaviour across both online and offline states.

[0157] In an embodiment, the offline operation unit (112) is configured to maintain all stored models and templates in an encrypted database to ensure data security and integrity.

[0158] In another embodiment, the offline operation unit (112) is configured to perform bidirectional synchronization with a cloud server upon detecting a secure network connection, wherein updated translation models, speech datasets, UI templates, or neural parameters are downloaded and existing user-specific learning data are uploaded for long-term personalization.

[0159] In yet another embodiment, the offline operation unit (112) is configured to maintain version control for stored assets, ensuring rollback capability and stable system performance following updates.

[0160] The dynamic control unit (114) is communicatively coupled to the context detection unit (104), the guidance generation unit (106), the user feedback and learning module (118), the time-limited assistance controller (120), and the on-demand activation controller (122). The dynamic control unit (114) is configured to regulate, optimise, and manage the operational state of the guidance generation unit (106) based on real-time interaction data, user proficiency metrics, and system resource conditions.

[0161] During operation, the dynamic control unit (114) is configured to receive the updated user proficiency score from the user feedback and learning module (118). The dynamic control unit (114) is further configured to compare the proficiency score with a predefined threshold stored in the user profile database (140). When the proficiency score exceeds the threshold, the dynamic control unit (114) is configured to selectively disable the guidance generation unit (106) for the corresponding device functions, thereby conserving processing resources and preventing unnecessary guidance prompts. The dynamic control unit (114) also receives inactivity notifications or countdown expirations from the time-limited assistance controller (120) and is configured to transition the guidance generation unit (106) into a suspended or low-power state accordingly.The dynamic control unit (114) is further configured to process reactivation signals received from the on-demand activation controller (122). Upon receiving a valid trigger signature derived from a recognised gesture or voice keyword, the dynamic control unit (114) is configured to re-enable the guidance generation unit (106) and restore full system operation. In an embodiment, the dynamic control unit (114) implements a state machine configured to shift between active, semi-active, low-power, and suspended states based on contextual data, resource availability, and user behaviour profiles.

[0162] In another embodiment, the dynamic control unit (114) is configured to provide manual override options wherein a user may explicitly enable or disable guidance for selected applications through device accessibility settings.

[0163] The user feedback and learning module (118) is communicatively coupled to the dynamic control unit (114) and is configured to record timestamped user interaction data, including task completion times, error occurrences, and dismissed suggestions. The user feedback and learning module (118) is further configured to compute a user proficiency score using a moving average of successful interactions and to transmit the updated score to the dynamic control unit (114) for adjusting activation behaviour of the guidance generation unit (106). The time-limited assistance controller (120) is operatively coupled to the dynamic control unit (114) and the CPU (144). The time-limited assistance controller (120) is configured to initiate a countdown timer each time the guidance generation unit (106) becomes active and to trigger a low-power interrupt to disable the guidance generation unit (106) when no user interaction is detected before the timer expires. The time-limited assistance controller (120) is configured to conserve processing resources and prevent prolonged unnecessary system activity.

[0164] The on-demand activation controller (122) is operatively coupled to the gesture recognition extension (104b) and the dynamic control unit (114). The on-demand activation controller (122) is configured to monitor for predefined reactivation triggers such as a specific voice keyword or a designated hand gesture. The on-demand activation controller (122) is further configured to verify the detected trigger against reference templates stored in the non-volatile storage medium (146) and, upon successful verification, send a system call to the dynamic control unit (114) to re-enable the guidance generation unit (106).The hardware accelerator (124) is operatively coupled to the adaptive audio-visual assistant system (100) through a high-speed bus (154) configured to support rapid transfer of model parameters, feature vectors, intermediate tensors, and audio-visual processing data. The hardware accelerator (124) comprises a neural processing unit (NPU) (124a) having dedicated matrix-multiplication cores, vector processing units, and parallel compute engines optimized for executing deep learning workloads with high throughput and low power consumption.

[0165] The hardware accelerator (124) is configured to receive from the guidance generation unit (106) an input corresponding to the structured digital context packet and to perform neural inference using the pre-trained artificial neural network model. The NPU (124a) is configured to process multi-dimensional matrices associated with convolution operations, attention mechanisms, or recurrent computations and to generate an output tensor representing the digital guidance instruction set. This enables the guidance generation unit (106) to operate with significantly reduced computational load on the central processing unit (CPU) (144), thereby improving overall system efficiency.

[0166] The hardware accelerator (124) is further configured to support the speech-style adaptation unit (110) by executing real-time digital signal processing (DSP) operations required for prosody modelling. The hardware accelerator (124) offloads these audio signal manipulations from the CPU (144), resulting in faster synthesis and lower latency during spoken output generation.

[0167] In an embodiment, the digital signal processing (DSP) operations include applying transformations to adjust fundamental frequency (F0) contours, phoneme duration, spectral shaping coefficients, and amplitude envelopes associated with the synthesized speech signal. In an embodiment, the hardware accelerator (124) is further configured to operate in both online and offline conditions, enabling local execution of all inference and prosody computations without requiring cloud resources. By reducing execution time, energy consumption, and model latency, the hardware accelerator (124) enhances the real-time performance of the assistant system (100) and ensures smooth operation even on resource-constrained portable electronic devices (142).

[0168] Figures 2A-2B illustrate a flowchart of a method for providing adaptive audio-visual assistance on a portable electronic device, in accordance with an embodiment of the presentdisclosure. The order in which method (200) is described is not intended to be construed as a limitation, and any number of the described method steps may be combined in any order to implement method (200), or an alternative method. Furthermore, method (200) may be implemented by processing resource or computing device(s) through any suitable hardware, non-transitory machine-readable medium / instructions, or a combination thereof. The method (200) comprises the following steps:

[0169] At step 202, the method (200) includes receiving, at a context detection unit (104), realtime data streams and active application state data from system-level application programming interfaces (APIs) (130) and generating a structured digital context packet based on the data.

[0170] At step 204, the method (200) includes processing, at a guidance generation unit (106), the structured digital context packet through a pre-trained neural network model to generate a digital guidance instruction set.

[0171] At step 206, the method (200) includes commanding, via an audio-visual coordination unit (108), a concurrent rendering of a visual output on a touch-sensitive display (136) and a speech output on a speaker (138) of the portable electronic device (142), based on the digital guidance instruction set.

[0172] At step 208, the method (200) includes processing, at a speech-style adaptation unit (110), a textual command string of the digital guidance instruction set to generate a speech output matching a user-specific profile stored in a user profile database (140). At step 210, the method (200) includes enabling, via an offline operation unit (112), full system operation in an offline mode using data stored in a non-volatile storage medium (146).

[0173] At step 212, the method (200) includes optimizing, via a dynamic control unit (114), power consumption by implementing a state machine configured to manage activation of the guidance generation unit (106).

[0174] In an embodiment, the step of generating the structured digital context packet comprises sampling raw sensor signals from a touch sensor, an inertial measurement unit (IMU), and afront-facing camera, and pre-processing the sampled signals by applying digital filtering, feature extraction, and data fusion.

[0175] In an embodiment, the digital guidance instruction set comprises a textual command and a target user interface (UI) element identifier.

[0176] In an embodiment, the method (200) further comprises the step of generating, via a visual assistance module, a visual overlay instruction set by mapping the target UI element identifier to screen coordinates.

[0177] In an embodiment, the method further comprises the step of processing the textual command through the speech-style adaptation unit (110), which comprises applying a prosody model to synthesize a speech waveform with user-adapted acoustic characteristics.

[0178] In an operative configuration, the adaptive audio-visual assistant system (100) receives multimodal user inputs through a multi-microphone array (148), a touch sensor (126), an inertial measurement unit (IMU) (128), system APIs (130), and a front-facing camera (150). The voice processing module (102) is configured to transform raw speech into intent-rich text, while the context detection unit (104) is configured to fuse sensor data into a structured digital context packet representing the user’s current interaction state. The guidance generation unit (106) processes the packet using a pre-trained neural network to generate a digital guidance instruction set comprising a textual command string and a target UI element identifier. The multilingual translation module (116) and the speech-style adaptation unit (110) are configured to translate and synthesize the command into user-specific multilingual speech. Concurrently, the audio-visual coordination unit (108), assisted by the visual assistance module (108a), is configured to render synchronized visual overlays on the display (136) and speech output via the speaker (138). The offline operation unit (112) supplies locally stored models to ensure full functionality without network connectivity, while the dynamic control unit (114), supported by the user feedback and learning module (118), the time-limited assistance controller (120), and the on-demand activation controller (122), is configured to manage activation, suspension, and reactivation of guidance based on user proficiency, inactivity, or predefined triggers. The hardware accelerator (124) is configured to execute neural inference and prosody processing to achieve low-latency, resource-efficient, real-time operation across varying usage conditions.Advantageously, the adaptive audio-visual assistant system (100) enables real-time, multimodal guidance by combining sensor fusion, neural inference, and synchronized audiovisual rendering to enhance user interaction on a portable electronic device (142). The system (100) operates reliably in both online and offline environments through locally stored models, while dynamically adapting its speech style, language, and instructional detail to individual user behaviour. Further, the system (100) optimizes power and compute resources through proficiency-based activation control, time-limited assistance, and on-demand reactivation, thereby ensuring efficient and personalised assistance with minimal system overhead.

[0179] The foregoing description of the embodiments has been provided for purposes of illustration and is not intended to limit the scope of the present disclosure. Individual components of a particular embodiment are generally not limited to that particular embodiment but are interchangeable. Such variations are not to be regarded as a departure from the present disclosure, and all such modifications are considered to be within the scope of the present disclosure.

[0180] The disclosed system (100) will now be explained with the help of hypothetical non-limiting anecdotal examples as stated herein below.

[0181] In an exemplary scenario, an elderly user operating a portable electronic device (142) opens a mobile banking application and struggles to locate the “fund transfer” option. The context detection unit (104) detects repeated incorrect taps, while the guidance generation unit (106) receives a structured digital context packet indicating confusion in navigating the interface. The audio-visual coordination unit (108) displays a highlighted overlay on the correct menu item and the speech-style adaptation unit (110) generates slow, clearly articulated speech guidance suited to the user’s profile. The user successfully completes the task with real-time, synchronized audio-visual assistance.

[0182] In another exemplary scenario, a teenage user launches a newly installed mathematics learning application. The context detection unit (104) identifies that the user is opening the app for the first time, and the guidance generation unit (106) produces beginner-level instructions explaining how to navigate chapters, quizzes, and progress reports. When the user attempts a gesture-controlled feature, the gesture recognition extension (104b) interprets the hand movement and triggers on-screen visual aids. The system (100) adapts overtime asthe user’s proficiency increases, reducing the level of guidance as the teenager becomes familiar with the features.

[0183] In yet another exemplary scenario, a working professional intends to download a second PDF reader because the existing one appears to lack a document-signing feature. The system -level APIs (130) reveal that the installed app already includes the required feature, and the guidance generation unit (106) compares feature sets of both applications using metadata and app-capability data. The system (100) generates an audio-visual explanation that the existing PDF reader supports document signing and displays a highlighted path showing where the feature is located. The user avoids installing a redundant app, demonstrating the system’s intelligent app-comparison and smart installation assistance capability.

[0184] In a further exemplary scenario, a young child opens an Al-powered drawing application and attempts to use it to “make a school project chart.” The context detection unit (104) captures this intent through voice input, and the guidance generation unit (106) queries application metadata to identify whether the opened tool is suitable for academic-style layouts. Based on this purpose, the system recommends another child-friendly Al tool better suited for creating labelled diagrams and charts. The system provides a simple voice prompt and a highlighted screen path to switch applications, demonstrating the assistant’s purpose-aware Al tool recommendation capability.

[0185] In still another exemplary scenario, a senior corporate employee uses the device (142) on an airplane without network connectivity to prepare for an upcoming board presentation. While reviewing offline documents such as quarterly financial reports and a draft slide deck, the offline operation unit (112) provides local access to language models, visual assets, and neural network parameters, enabling the guidance generation unit (106) to deliver accurate task guidance such as, helping the user locate specific charts within a PDF or reorganize slides within the presentation application. The speech-style adaptation unit (110) generates concise, professional speech output aligned with the user’s preferences, and the dynamic control unit (114) conserves power by disabling guidance for familiar operations like scrolling or zooming. Even in a fully offline environment, the system (100) continues to operate with full functionality, enabling uninterrupted productivity during travel.

[0186] TECHNICAL ADVANCEMENTSThe present disclosure described hereinabove has several technical advantages including, but not limited to, the realization of an adaptive audio-visual assistant system that:

[0187] • fuses multimodal sensor data to generate a precise real-time context;

[0188] • produces on-device neural guidance instructions with low latency;

[0189] • delivers synchronized audio and AR-based visual overlays;

[0190] • adapts speech style using user-specific prosody profiles;

[0191] • performs full multilingual operation in offline mode;

[0192] • optimizes system resources through proficiency-based and time-limited control;

[0193] • enables intuitive gesture- or voice-based activation;

[0194] • accelerates inference using a dedicated hardware NPU; and

[0195] • continuously learns from user interactions for personalized assistance.

[0196] The aspect herein and the various features and advantageous details thereof are explained with reference to the non-limiting embodiments in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein may be practiced and to further enable those of skill in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.

[0197] The foregoing description of the specific embodiments so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and / or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognizethat the embodiments herein can be practiced with modification within the spirit and scope of the embodiments as described herein.

[0198] The use of the expression “at least” or “at least one” suggests the use of one or more elements or ingredients or quantities, as the use may be in the embodiment of the disclosure to achieve one or more of the desired objects or results.

[0199] Any discussion of devices, articles or the like that has been included in this specification is solely for the purpose of providing a context for the disclosure. It is not to be taken as an admission that any or all of these matters form a part of the prior art base or were common general knowledge in the field relevant to the disclosure as it existed anywhere before the priority date of this application.

[0200] While considerable emphasis has been placed herein on the components and component parts of the preferred embodiments, it will be appreciated that many embodiments can be made and that many changes can be made in the preferred embodiments without departing from the principles of the disclosure. These and other changes in the preferred embodiment as well as other embodiments of the disclosure will be apparent to those skilled in the art from the disclosure herein, whereby it is to be distinctly understood that the foregoing descriptive matter is to be interpreted merely as illustrative of the disclosure and not as a limitation.

Claims

CLAIMS:

1. An adaptive audio-visual assistant system (100) for a portable electronic device (142), said system (100) comprising:• a context detection unit (104) configured to generate a structured digital context packet by fusing real-time data streams and active application state data from system-level application programming interfaces (APIs) (130); • a guidance generation unit (106) communicatively coupled with said context detection unit (104) and configured to process the structured digital context packet through a pre-trained neural network model to generate a digital guidance instruction set;• an audio-visual coordination unit (108) communicatively coupled with said guidance generation unit (106), and configured to command a concurrent rendering of a visual output on a touch-sensitive display (136) and a speech output on a speaker (138) of the portable electronic device (142), based on the digital guidance instruction set;• a speech-style adaptation unit (110) communicatively coupled with said guidance generation unit (106) and said audio-visual coordination unit (108), and configured to process the textual command string and generate a speech output to match a user-specific profile stored in a user profile database (140);• an offline operation unit (112) stored in a non-volatile storage medium (146) of the portable electronic device (142) and operatively coupled to said guidance generation unit (106) and said speech-style adaptation unit (110), said offline operation unit (112) configured to enable full system operation in an offline mode; and• a dynamic control unit (114) communicatively coupled to said context detection unit (104) and said guidance generation unit (106), said dynamic control unit (114) configured to optimize power consumption by implementing a state machine.. The system (100) as claimed in claim 1, further comprising a visual assistance module (108a) operatively coupled to said guidance generation unit (106) and said audiovisual coordination unit (108), said visual assistance module (108a) configured to generate a visual overlay instruction set by mapping the target UI element identifier to absolute screen coordinates and rendering a semi-transparent graphical highlight using an augmented reality (AR) technique that overlays a specific user interface element on the display (136).

3. The system (100) as claimed in claim 1, further comprising a hardware accelerator (124), comprising a neural processing unit (NPU) (124a) with dedicated matrix multiplication cores, operatively coupled to the system (100) via a high-speed bus (154), wherein said hardware accelerator (124) is configured to:• execute inference computations for the pre-trained neural network model in said guidance generation unit (106); and• perform real-time digital signal processing for the prosody model in said speech-style adaptation unit (110).

4. The system (100) as claimed in claim 1, wherein said context detection unit (104) comprises:• a feature detection engine (104a) configured to parse a system event log and a frame buffer to identify the active application and a specific feature in use; and • a gesture recognition extension (104b), implementing a convolutional neural network (CNN), configured to process image frames from a front-facing camera (150) to interpret dynamic hand gestures for system control.

5. The system (100) as claimed in claim 1, wherein the real-time data streams comprise touch events detected by a touch sensor (126) and motion vectors captured by an inertial measurement unit (IMU) (128).

6. The system (100) as claimed in claim 1, wherein said guidance generation unit (106) is communicatively coupled with said context detection unit via an inter-process communication (IPC) bus (132).

7. The system (100) as claimed in claim 1, wherein the digital guidance instruction set generated by said guidance generation unit comprising a textual command string and a target user interface (UI) element identifier.

8. The system (100) as claimed in claim 1, wherein said speech-style adaptation unit (110) is configured to apply a prosody model that dynamically adjusts fundamental frequency (F0) contours, phoneme duration, and spectral features of the speech output.

9. The system (100) as claimed in claim 1, wherein said offline operation unit (112) further maintains a local encrypted database comprising multilingual acoustic models, a library of visual asset templates, and parameters for the pre-trained neural network model and performs bidirectional synchronization with a cloud server upon detection of a secure network connection.

10. The system (100) as claimed in claim 1, wherein the state machine of said dynamic control unit selectively disables said guidance generation unit (106) for specific device functions upon determining that a user proficiency score exceeds a predefined threshold, said user proficiency score being derived from a success rate of historical user interactions.

11. The system (100) as claimed in claim 1, further comprising a voice processing module (102) operatively coupled to a multi-microphone array (148) of the portable electronic device (142) and said context detection unit (104), said voice processing module (102) configured to:• perform beamforming and noise suppression on raw audio input signals;• convert the enhanced audio signals into text through a hidden Markov model (HMM)-based automatic speech recognition (ASR) engine (152); and • perform named entity recognition (NER) on the text to extract user intent and update the structured digital context packet.

12. The system (100) as claimed in claim 1, further comprising a multilingual translation module (116) communicatively coupled between said guidance generation unit (106)and said speech-style adaptation unit (110), said multilingual translation module (116) configured to:• detect a user’s preferred language by analyzing a system locale setting and a language n-gram distribution from historical user input;• translate the textual command string from the digital guidance instruction set into the preferred language through a statistical machine translation (SMT) model; and• provide the translated textual command string to said speech-style adaptation unit (110).

13. The system (100) as claimed in claim 1, further comprising a user feedback and learning module (118) communicatively coupled to said dynamic control unit (114), said user feedback and learning module (118) configured to:• record timestamped user interaction events, including task completion times and error rates, in a structured log file;• calculate the user proficiency score using a moving average filter applied to the success rate over a sliding window of the most recent N interactions; and • transmit the updated user proficiency score to said dynamic control unit (114) to adjust its activation policy.

14. The system (100) as claimed in claim 1, wherein said dynamic control unit (114) further comprises a time-limited assistance controller (120) configured to:• initiate a countdown timer upon activation of said guidance generation unit (106); and• trigger a low-power interrupt to a central processing unit (CPU) (144) of the portable electronic device (142) to disable said guidance generation unit (106) when the countdown timer reaches zero without a subsequent user interaction event from said context detection unit (104).

15. The system (100) as claimed in claim 1, further comprising an on-demand activation controller (122) operatively coupled to said gesture recognition extension (104b) and said dynamic control unit (114), said on-demand activation controller (122) configured to:• monitor for a predefined reactivation trigger signature from a voice keyword or a specific hand gesture; and• upon verification of the trigger signature, send a system call to said dynamic control unit (114) to re-enable said guidance generation unit (106).

16. The system (100) as claimed in claim 15, wherein the predefined reactivation trigger signature is derived from a reference data set stored in the non-volatile storage medium (146), the reference data set comprising:• a specific audio waveform pattern for the voice keyword, obtained during an initial user enrollment phase; and• a specific sequence of spatial coordinates representing the hand gesture, captured and stored as a template by the gesture recognition extension (104b).

17. The system (100) as claimed in claim 1, wherein the predefined threshold for the user proficiency score is a configurable parameter stored in the user profile database (140), settable by at least one of: a system administrator via a configuration file, or the user through an accessibility settings menu of the portable electronic device.

18. The system (100) as claimed in claim 1, wherein the pre-trained neural network model of said guidance generation unit (106) is a supervised learning model, initially trained on a curated dataset of annotated user interactions prior to its deployment on the portable electronic device (142), wherein the curated dataset comprises a plurality of data pairs, each pair consisting of an input context vector derived from historical sensor and application state data, and a corresponding labeled output vector representing a validated guidance instruction.

19. A portable electronic device (142) comprising:• a system-on-chip (SoC) comprising a central processing unit (CPU) (144) and a graphics processing unit (GPU) (134);• a non-volatile storage medium (146);• a touch-sensitive display (136);• a multi-microphone array (148), an audio codec, and a speaker (138);• a sensor hub comprising a touch sensor (126), an inertial measurement unit (IMU) (128), and the front-facing camera (150); and• the adaptive audio-visual assistant system (100) as claimed in any one of claims 1 to 12.

20. A method (200) for providing adaptive audio-visual assistance on a portable electronic device (142), said method (200) comprising:• receiving, at a context detection unit (104), real-time data streams and active application state data from system-level application programming interfaces (APIs) (130) and generating a structured digital context packet based on said data;• processing, at a guidance generation unit (106), the structured digital context packet through a pre-trained neural network model to generate a digital guidance instruction set;• commanding, via an audio-visual coordination unit (108), a concurrent rendering of a visual output on a touch-sensitive display (136) and a speech output on a speaker (138) of the portable electronic device (142), based on the digital guidance instruction set;• processing, at a speech-style adaptation unit (110), a textual command string of the digital guidance instruction set to generate a speech output matching a user-specific profile stored in a user profile database (140);• enabling, via an offline operation unit (112), full system operation in an offline mode using data stored in a non-volatile storage medium (146); andoptimizing, via a dynamic control unit (114), power consumption by implementing a state machine configured to manage activation of the guidance generation unit (106).

21. The method (200) as claimed in claim 20, wherein generating the structured digital context packet comprises sampling raw sensor signals from a touch sensor (126), an inertial measurement unit (IMU) (128), and a front-facing camera (150), and pre-processing the sampled signals by applying digital filtering, feature extraction, and data fusion.

22. The method (200) as claimed in claim 20, wherein the digital guidance instruction set comprises a textual command and a target user interface (UI) element identifier.

23. The method (200) as claimed in claim 22, further comprising generating, via a visual assistance module (108a), a visual overlay instruction set by mapping the target UI element identifier to screen coordinates.

24. The method (200) as claimed in claim 20, wherein processing the textual command through the speech-style adaptation unit (110) comprises applying a prosody model to synthesize a speech waveform with user-adapted acoustic characteristics.