Assistive system for task guidance using subvocalized commands, visual analysis, and biosensor data
The assistive system integrates visual analysis and biosensor data with subvocalized commands to enhance task guidance for visually impaired individuals, improving recognition accuracy and user interaction, enabling daily tasks and sports participation.
Patent Information
- Application Number
- PCT/IB2025/056642
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-04
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-08
AI Technical Summary
Existing assistive technologies for visually impaired individuals face challenges in providing comprehensive and context-aware assistance, struggling with real-time processing of complex environments, accurate object identification, and intuitive user interfaces that do not rely heavily on visual cues, and lack integration of multiple input modalities and context-appropriate feedback.
An assistive system that integrates visual analysis, biosensor data, and subvocalized commands using a neural language model to provide comprehensive and intuitive task guidance, enhancing object and individual recognition, and enabling personalized feedback through a combination of image-capture devices, biosensors, and subvocalized commands.
Improves accuracy in object and individual recognition, offers more natural and intuitive user interaction, and provides context-aware assistance, enabling visually impaired individuals to perform daily tasks and engage in sports activities with enhanced safety and independence.
Smart Images

Figure IB2025056642_08012026_PF_FP_ABST
Abstract
Description
ASSISTIVE SYSTEM FOR TASK GUIDANCE USING SUBVOCALIZED COMMANDS, VISUAL ANALYSIS, AND BIOSENSOR DATACROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to Indian Provisional Application No. IN202411051251 , which was filed on July 4, 2024. The above stated Patent Application(s) are hereby incorporated herein by reference in their entirety.FIELD OF INVENTION
[0002] The present disclosure relates to assistive technologies for visually impaired individuals, and more particularly to an electronic device that provides task guidance using subvocalized commands, visual analysis, and biosensor data.BACKGROUND
[0003] Assistive technologies for visually impaired or disabled individuals have evolved significantly in recent years, leveraging advancements in computer vision, natural language processing, and wearable sensors. The assistive technologies aim to enhance the independence and quality of life for users by providing information about their surroundings and assisting with daily tasks.
[0004] Typical solutions in this domain include screen readers, text-to-speech systems, and object recognition applications. Some approaches utilize cameras to capture visual information, while others employ various sensors to detect obstacles or provide navigation cues. However, existing systems often face challenges in provision of comprehensive and context-aware assistance. Many solutions struggle with real-time processing of complex environments, accurate object identification in diverse settings, or intuitive user interfaces that do not rely heavily on visual cues. Additionally, the integration of multiple input modalities and the generation of natural, context-appropriate feedbackremain areas for improvement in assistive technologies for the visually impaired or disabled individuals.
[0005] Further limitations and disadvantages of conventional and traditional approaches will become apparent to one of skill in the art, through comparison of described systems with some aspects of the present disclosure, as set forth in the remainder of the present application and with reference to the drawings.SUMMARY
[0006] An electronic device and method to provide task guidance using subvocalized commands, visual analysis, and biosensor data is provided substantially as shown in, and / or described in connection with, at least one of the figures, as set forth more completely in the claims.
[0007] These and other features and advantages of the present disclosure may be appreciated from a review of the following detailed description of the present disclosure, along with the accompanying figures in which like reference numerals refer to like parts throughout.BRIEF DESCRIPTION OF FIGURES
[0008] FIG. 1 is a diagram illustrates a network environment for task guidance using subvocalized commands, visual analysis, and biosensor data, in accordance with at least one embodiment of the disclosure.
[0009] FIG. 2 is a block diagram of an electronic device for provision of task guidance using subvocalized commands, visual analysis, and biosensor data, in accordance with at least one embodiment of the disclosure.
[0010] FIG. 3 is a diagram that illustrates an execution pipeline for provision of task guidance using subvocalized commands, visual analysis, and biosensor data, in accordance with at least one embodiment of the disclosure.
[0011] FIG. 4A and FIG. 4B are diagrams that collectively illustrate an exemplary execution pipeline for detection of objects using visual noise cancellation technique, in accordance with at least one embodiment of the disclosure.
[0012] FIG. 5 is a diagram that illustrates an exemplary execution pipeline for visual text / sign recognition, in accordance with at least one embodiment of the disclosure.
[0013] FIG. 6A and FIG. 6B are diagrams that collectively illustrate an exemplary execution pipeline for facial recognition, in accordance with at least one embodiment of the disclosure.
[0014] FIG. 7A and FIG. 7B are diagrams that collectively illustrate an exemplary execution pipeline for movie narration assistance, in accordance with at least one embodiment of the disclosure.
[0015] FIG. 8A and FIG. 8B are diagrams that collectively illustrate an exemplary execution pipeline for room assistance, in accordance with at least one embodiment of the disclosure.
[0016] FIG. 9A and FIG. 9B are diagrams that collectively illustrate an exemplary execution pipeline for detection of infrared rays, in accordance with at least one embodiment of the disclosure.
[0017] FIG. 10 is a block diagram of an exemplary audio processing system, in accordance with at least one embodiment of the disclosure.
[0018] FIG. 11 is an exemplary diagram of a convolution block of the audio processing system of FIG. 10, in accordance with at least one embodiment of the disclosure.
[0019] FIG. 12 is a diagram that illustrates a waveform graph between audio and EMG signal patterns, in accordance with at least one embodiment of the disclosure.
[0020] FIG. 13 is a diagram that illustrates an exemplary implementation of the assistive device in sports assistance, in accordance with at least one embodiment of the disclosure.
[0021] FIG. 14 is a diagram that illustrates an exemplary flowchart of a method of providing task guidance using subvocalized commands, visual analysis, and biosensor data, in accordance with at least one embodiment of the disclosure.DETAILED DESCRIPTION
[0022] The present disclosure relates to an assistive system for task guidance using subvocalized commands, visual analysis, and biosensor data. The following described implementation may be found in an electronic device and method that may be configured to provide task guidance using subvocalized commands, visual analysis, and biosensor data. The electronic device may include circuitry. The circuitry may be configured to capture, via a set of image-capture devices, visual information of an environment associated with a user. The circuitry may be further configured to determine first information indicative of at least one of an object or an individual in the environment, based on the captured visual information. The circuitry may be further configured to receive sensor data from a set of biosensors associated with the user. The circuitry may be further configured to determine a subvocalized command of the user, based on the received sensor data and apply a neural language model to the determined first information and the determined subvocalized command. The circuitry may be further configured to determine, based on the application of the neural language model, second information corresponding to at least one of an audio feedback or a haptic feedback for the user. The circuitry may be further configured to render the determined second information to the user. The second information may include instructions to guide the user to perform a task associated with the determined subvocalized command.
[0023] Existing assistive technologies, especially for visually impaired individuals, may often face challenges in provision of comprehensive and context-aware assistance. Many current solutions struggle with real-time processing of complex environments, accurate object identification in diverse settings, and intuitive user interfaces that do not rely heavilyon visual cues. Additionally, the integration of multiple input modalities and the generation of natural, context-appropriate feedback remain areas for improvement. These limitations can result in suboptimal user experiences and reduced independence for visually impaired individuals, which highlights the need for a more advanced and integrated assistive system.
[0024] The disclosed assistive system combines visual analysis, biosensor data, and subvocalized commands to provide a more comprehensive and intuitive task guidance solution for visually impaired users. Unlike traditional systems that may rely on a single input modality, the disclosed assistive system may integrate multiple data sources to create a more accurate understanding of the user's environment and intentions. The assistive system utilizes advanced neural language models to process the combined input and generate context-appropriate feedback. Thus, the assistive system may improve accuracy in object and individual recognition, more natural and intuitive user interaction through subvocalized commands, and the ability to provide personalized task guidance based on the user's specific needs and environment.
[0025] The assistive system may assist users in performing tasks corresponding to day-to-day activities such as moving from one room to another, walking on road, picking required things, and the like. The tasks may also correspond to sports activity. The assistive system may play a crucial role in enabling both normal and disabled users to participate in sports activities by enhancing their physical capabilities and ensuring safety. For the disabled users, the assistive system aid the users to engage in activities that would otherwise be challenging or impossible. For example, an eye-assistive device may aid a visually impaired user in sports such as baseball, cricket, tennis, and the like. The assistive system not only facilitate physical activity but also promote inclusivity, enabling disabled athletes to participate alongside their able-bodied peers. For normal users, the assistive device may enhance performance and prevent injuries. For instance, the assistive devicemay include a voice-assistance that may guide the users hence drawing attention of the user to important information, thereby enhancing performance users. By leveraging the assistive device, both normal and disabled users may achieve full potential in sports, fostering a more inclusive and competitive environment.
[0026] FIG. 1 is a diagram illustrates a network environment for task guidance using subvocalized commands, visual analysis, and biosensor data, in accordance with at least one embodiment of the disclosure. With reference to FIG. 1 , there is shown an exemplary network environment 100. The network environment 100 includes an electronic device 102, a set of image-capture devices 104, a set of biosensors 106, a neural language model 108, a facial recognition model 110, a motion tracking model 112, a server 114, a database 116, a communication network 118, and a user device 120. The electronic device 102 may be associated with the set of image-capture devices 104, the set of biosensors 106, the neural language model 108, the facial recognition model 110, and the motion tracking model 112. The electronic device 102 may also be connected to the server 114, a device hosting the database 116, or the user device 120, via the communication network 118. The set of image-capture devices 104 may be associated with an environment surrounding a user and may capture visual information 104A of the environment. In an embodiment, the set of image-capture devices 104 and the set of biosensors 106 may be associated with the user device 120.
[0027] The electronic device 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to capture visual information from the set of image-capture devices 104. The electronic device 102 may further be configured to determine first information indicative of at least one of an object or an individual in the environment, based on the captured visual information. The electronic device 102 may further be configured to receive sensor data from the set of biosensors 106 associated with a user. The electronic device 102 may further be configured to determine a subvocalized command of the user,based on the received sensor data. The electronic device 102 may further be configured to apply the neural language model 108 to the determined first information and the determined subvocalized command. The electronic device 102 may further be configured to determine, based on the application of the neural language model 108, second information corresponding to at least one of an audio feedback or a haptic feedback for the user. The electronic device 102 may further be configured to render the determined second information to the user. The second information may include instructions to guide the user to perform a task associated with the determined subvocalized command. As used herein, the term “subvocalized command” refers to a silent, internal command issued by the user that does not include audible speech. Instead, the subvocalized commands may be detected through the physiological signals generated when the user engages in subvocalization, that may be the act of internally articulating words or commands without producing sound.
[0028] Examples of the electronic device 102 may include, but are not limited to, a digital media player (DMP), a micro-console, a TV tuner, a digital media streamer, a media extender / regulator, a smart TV, a gaming console, a digital media hub, a computer workstation, a mainframe computer, a handheld computer, a smart appliance, a plug-in device, a mobile phone, a smart phone, a tablet computer, a personal computer, a wearable device, an augmented reality / virtual reality (AR / VR) device, and / or any other computing device, such as, a consumer electronic device.
[0029] The electronic device 102 may store the neural language model 108, the facial recognition model 110, and the motion tracking model 112 or may be remotely connected to another system (such as the server 114) that hosts the neural language model 108, the facial recognition model 110, and the motion tracking model 112. When hosted on another system, the electronic device 102 may send instructions to control training or inference ofthe neural language model 108, the facial recognition model 110, and the motion tracking model 112 via remote calls (e.g., application programming interface (API) calls).
[0030] The set of image-capture devices 104 may include suitable logic, circuitry, interfaces, and / or code that may be configured to capture visual information of an environment surrounding the user. In an instance, the set of image-capture devices 104 may include, but not limited to, cameras, infrared sensors, or depth sensors that capture the visual information 104A of the environment surrounding the user. The visual information 104A may be data or details captured from the physical space I environment around a user. For example, the visual information 104A may be captured using a visual sensor device such as cameras, sensors, or the human eye. The visual information 104A may include elements like objects, people, spatial arrangements, colors, lighting, movement, and any visual cues present in the environment. In an embodiment, some image-capture devices 104 may be integrated in the electronic device 102 and / or the server 114. In one embodiment, some image-capture devices of the set of image-capture devices 104 may be installed on a wearable device adapted to be worn by the user. The wearable device may include headgear, helmet, glasses, clothing, accessories, smart watch, VR headsets, and the like. In another embodiment, some image-capture devices of the set of image-capture devices 104 may be installed in nearby surroundings of the environment. For instance, a camera may be mounted on a wall of a room associated with the user, such that the camera may capture images of objects and individuals in the user's vicinity. In another instance, an infrared sensor may be integrated on an entrance door to detect heat signature of the user. In another instance, a depth sensor may provide spatial information about the environment.
[0031] The set of biosensors 106 may include suitable logic, circuitry, interfaces, and / or code that may be configured to monitor the sensor data associated with the user. The set of biosensors 106 may include a plurality of sensors including, but not limited to,a surface electromyography (sEMG) sensor, a brain-computer interface (BCI) sensor, an electroencephalography (EEG) sensor, an electrocardiography (ECG) sensor, an eye tracker, a voice recognition sensor, a galvanic skin response (GSR) sensor, an electrodermal activity (EDA) sensor, a face detector, a motion sensor, and the like. For example, the sensor data may include data related to, but not limited to, muscle movements, brain, neurons, voluntary / involuntary response, heart, eyes, voice, skin response, dermal response, facial features, motion associated with the body of the user, and the like. At least one sensor of the set of biosensors 106 may be placed around the user, such that the sensor may generate the sensor data associated with the user, and further monitor the user attributes such as the muscle movements, or the subvocalized command, brain, neurons, voluntary / involuntary response, heart, eyes, voice, skin response, dermal response, or facial features, based on the sensor data. In an instance, sEMG sensor may detect muscle movements associated with subvocalization of the user. The sEMG sensor may be placed at or around a throat or face of the user to monitor subtle muscle movements that occur when the user thinks about speaking without actually producing audible sounds. The captured subtle muscle movements may allow the electronic device 102 to interpret the user's intended commands without a requirement of vocal output. In one embodiment, the set of biosensors 106 may transmit the sensor data directly to the electronic device 102. In another embodiment, based on a location of the user to be remote, the set of biosensors 106 may transmit the sensor data to the server 114. In such a case, the electronic device 102 may receive the sensor data from the server 114.
[0032] The neural language model 108 may include suitable logic, circuitry, interfaces, and / or code that may be configured to process the visual information and biosensor data to generate appropriate responses or instructions for the user. In an instance, the neural language model 108 may process the visual information and the subvocalized commandto determine second information corresponding to at least one of an audio feedback or a haptic feedback for the user. The second information may include instructions to guide the user to perform a task associated with the determined subvocalized command. For example, the neural language model 108 may be a Large Language Model (LLM) that may understand, generate, and process human language such as text content. The neural language model 108 may be trained on datasets corresponding to visual data and sensor data. The neural language model 104B may perform tasks like interpretation of text, generation of coherent content, translation of languages, answering questions, and engaging in conversations using a deep learning technique. The deep learning technique may include a neural networks and attention mechanisms to determine the context information and predict text sequences associated with the text content.
[0033] The neural language model 108 may be a neural network that may be a computational network or a system of artificial neurons, arranged in a plurality of layers, as nodes. The plurality of layers of the neural network may include an input layer, one or more hidden layers, and an output layer. Each layer of the plurality of layers may include one or more nodes (or artificial neurons). Outputs of all nodes in the input layer may be coupled to at least one node of hidden layer(s). Similarly, inputs of each hidden layer may be coupled to outputs of at least one node in other layers of the neural network. Outputs of each hidden layer may be coupled to inputs of at least one node in other layers of the neural network. Node(s) in the final layer may receive inputs from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined from hyper-parameters of the neural network. Such hyper-parameters may be set before or after training the neural network on the training dataset.
[0034] The neural network may include electronic data, which may be implemented as, for example, a software component of an application executable on the electronic device 102. The neural network may rely on libraries, external scripts, or other logic / instructionsfor execution by a processing device, such as the electronic device 102. The neural network may rely on code and routines to enable a computing device, such as the electronic device 102 to perform one or more operations, such as render the second information to the user. The second information may include instructions to guide the user to perform the task associated with the determined subvocalized command. In some embodiments, the neural network may be implemented using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the neural network may be implemented using a combination of hardware and software.
[0035] Examples of the neural language model 108 may include, but are not limited to, a Bidirectional Encoder Representations from Transformers (BERT) model, a Generative Pre-trained Transformer (GPT) model and its variants, Text-to-Text Transfer Transformer (T5) models, Long Short-Term Memory (LSTM) networks, attention-based models, multimodal transformer models, fine-tuned large language models, encoder-decoder architectures, hierarchical neural language models, and reinforcement learning-based language models.
[0036] The facial recognition model 110 may include suitable logic, circuitry, interfaces, and / or code that may be configured to analyze images captured by the set of imagecapture devices 104 to identify individuals in the user's environment. In an embodiment, the facial recognition model 110 may identify faces of individuals based on an analysis and comparison of patterns using facial features of the individuals. The facial recognition model 110 may detect a face within an image or video frame. Once the face is detected, the facial recognition model 110 may extract key features such as distance between eyes, shape of cheek bones, contour of lips, depth of eye sockets, and the like. The key features may then be converted into a mathematical representation known as a faceprint. The facialrecognition model 110 may leverage deep learning techniques to compare the faceprint against a database of known faces to find a match or confirm an identity of corresponding individual.
[0037] For example, the facial recognition model 110 may be a neural network that may be a computational network or a system of artificial neurons, arranged in a plurality of layers, as nodes. The plurality of layers of the neural network may include an input layer, one or more hidden layers, and an output layer. Each layer of the plurality of layers may include one or more nodes (or artificial neurons). Outputs of all nodes in the input layer may be coupled to at least one node of hidden layer(s). Similarly, inputs of each hidden layer may be coupled to outputs of at least one node in other layers of the neural network. Outputs of each hidden layer may be coupled to inputs of at least one node in other layers of the neural network. Node(s) in the final layer may receive inputs from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined from hyper-parameters of the neural network. Such hyper-parameters may be set before or after training the neural network on the training dataset.
[0038] The neural network may include electronic data, which may be implemented as, for example, a software component of an application executable on the electronic device 102. The neural network may rely on libraries, external scripts, or other logic / instructions for execution by a processing device, such as the electronic device 102. The neural network may rely on code and routines to enable a computing device, such as the electronic device 102 to perform one or more operations, such as detection of face of individuals. In some embodiments, The neural network may be implemented using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an applicationspecific integrated circuit (ASIC). Alternatively, in some embodiments, the neural network may be implemented using a combination of hardware and software.
[0039] Examples of the facial recognition model 110 may include, but are not limited to, Convolutional Neural Networks (CNNs), VGGFace2 networks, Deep Face Recognition models, FaceNet, DeepFace, ArcFace, Siamese networks, 3D face recognition models, attention-based facial recognition models, multi-modal face recognition systems, finetuned large vision models for face recognition, and deep metric learning approaches for facial recognition. The facial recognition model 110 may also incorporate feature extraction techniques, landmark detection techniques, and face verification methods to enhance recognition accuracy and robustness across various environmental conditions and user demographics.
[0040] The motion tracking model 112 may monitor movements of objects or individuals in the captured visual information. The motion tracking model 112 may detect an object or an individual, followed by continuous tracking of position of the detected object / individual over time. The process of tracking may involve identification of key points or features on the individual, such as joints in case of human motion tracking, and calculation of movement trajectories of the key points or features. The identified key points or features may be processed to understand the dynamics of the motion of the individual.
[0041] Examples of the motion tracking model 112 may include, but are not limited to, optical flow techniques, Kalman filters, particle filters, Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN) based tracking models. The motion tracking model 112 may also incorporate Long Short-Term Memory (LSTM) networks for temporal motion analysis, Multiple Object Tracking (MOT) techniques, feature-based tracking methods, and template matching techniques. Deep learning-based object tracking models, Simultaneous Localization and Mapping (SLAM) for motion and environment tracking, pose estimation networks, and 3D skeleton tracking models may be also utilized. The motion tracking model 112 may include gesture recognition techniques, motion vector analysis techniques, Time-of-Flight (ToF) based tracking systems, and InertialMeasurement Unit (IMU) based motion tracking. Additionally, fusion techniques that combine data from multiple sensors for robust tracking and attention-based tracking models for focusing on relevant motion patterns may be employed in the motion tracking model 112.
[0042] The server 114 may include suitable logic, circuitry, interfaces, and / or code that may be configured to enable the electronic device 102 to provide task guidance using subvocalized commands, visual analysis, and biosensor data. For example, the server 114 may analyze volume of sensor data, such as, visual information from the set of imagecapture devices 104 and biosensor data from the set of biosensors 106. The server 114 may employ advanced computer vision techniques, such as convolutional neural networks to process and analyze the visual information in real-time or near real-time. Similarly, the server 114 may apply various machine learning models on the biosensor data to determine insights from the biosensor data. The server 114 may be configured to store and manage user profiles and preferences. The server 114 may maintain and / or host a database (such as the database 116) of individual user settings, such as preferred feedback modalities or personalized object recognition priorities. The preferred feedback modalities may refer to the types or forms of feedback that a user prefers to receive from the electronic device 102 or the user device 120. The feedback modalities may include visual cues (e.g., notifications or highlights), auditory signals (e.g., spoken information, messages, or sounds), haptic feedback (e.g., vibrations), or other forms of sensory communication. The personalized object recognition priorities may refer to the customization of how the electronic device 102 may prioritizes identification or interaction with specific objects in the user's environment based on the preferences of the user. For example, a user may configure the electronic device 102 to recognize certain objects, like personal belongings, workplace tools, or home appliances, over less relevant items. For example, the server 114 may store user's frequently visited locations and customize navigation assistance onthe database 116 or on a memory associated with the server 114, based on such historical data. In some cases, the server 114 may host more computationally intensive machine learning models that may not be feasible to run on the electronic device 102 itself. This may include large-scale natural language processing models to interpret complex subvocalized commands or advanced facial recognition models capable of identification of individuals in challenging lighting conditions. The server 114 may provide cloud-based storage for historical data and environmental maps. This functionality may help to build and maintain detailed three-dimensional (3D) maps of frequently visited areas, which may be used to enhance navigation assistance and object localization. For instance, the server 114 may store and update a map of a user's home, including the locations of furniture and potential obstacles.
[0043] The server 114 may send software updates and model refinements to the electronic device 102. For example, the server 114 may push regular updates to improve the performance of the neural language model 108, the facial recognition model 110, and the motion tracking model 112 based on aggregated user data and new research findings. The server 114 may enable multi-user data aggregation to improve overall system performance. The data aggregation may allow the electronic device 102 to learn from the experiences of multiple users, that potentially may improve object recognition accuracy in diverse environments or enhancing the interpretation of subvocalized commands across different accents and languages.
[0044] The server 114 may be implemented as a cloud server and may execute operations through web applications, cloud applications, HTTP requests, repository operations, file transfer, and the like. Other example implementations of the server 114 may include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, a machine learning server (enabled withor hosting, for example, a computing resource, a memory resource, and a networking resource), or a cloud computing server.
[0045] In at least one embodiment, the server 114 may be implemented as a plurality of distributed cloud-based resources by use of several technologies that are well known to those ordinarily skilled in the art. A person with ordinary skill in the art will understand that the scope of the disclosure may not be limited to the implementation of the server 114 and the electronic device 102, as two separate entities. In certain embodiments, the functionalities of the server 114 may be incorporated in its entirety or at least partially in the electronic device 102, without a departure from the scope of the disclosure. In certain embodiments, the server 114 may host the database 116. Alternatively, the server 114 may be separate from the database 116 and may be communicatively coupled to the database 116.
[0046] The database 116 may include suitable logic, interfaces, and / or code that may be configured to store information such as user profiles, historical data, or reference images for object and facial recognition. The database 116 may be derived from data off a relational or non-relational database, or a set of comma-separated values (csv) files in conventional or big-data storage. The database 116 may be stored or cached on a device, such as a server (e.g., the server 114) or the electronic device 102. The device storing the database 116 may be configured to receive commands or instructions from the electronic device 102 or the server 114. In response, the device associated with the database 116 may be configured to retrieve and provide required information from the stored information.
[0047] In some embodiments, the database 116 may be hosted on a plurality of servers stored at the same or different locations. The operations of the database 116 may be executed using hardware including, but not limited to, a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some other instances,the database 116 may be implemented using software.
[0048] The communication network 118 may include a communication medium through which the electronic device 102, the server 114, the database 116, and the user device 120 may communicate with one another. The communication network 118 may be one of a wired connection or a wireless connection. Examples of the communication network 118 may include, but are not limited to, the Internet, a cloud network, Cellular or Wireless Mobile Network (such as Long-Term Evolution and 5thGeneration (5G) New Radio (NR)), satellite communication system (using, for example, low earth orbit satellites), a Wireless Fidelity (Wi-Fi) network, a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). Various devices in the network environment 100 may be configured to connect to the communication network 118 in accordance with various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of a Transmission Control Protocol and Internet Protocol (TIP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Zig Bee, EDGE, IEEE 802.11 , light fidelity (Li-Fi), 802.16, IEEE 802.11 s, IEEE 802.11 g, multi-hop communication, wireless access point (AP), device to device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.
[0049] The user device 120 may include a user-interface through which the user may interact with the electronic device 102, send images / videos or audio and feed commands and instructions to the electronic device 102. The user may be a person who may be authorized to use or operate the electronic device 102 or the server 114. The user device 120 may be fixed at a place or may be portable. Examples of the user device 120 may include, but not limited to, a smartphone, a wearable device, a personal computer, an admin terminal of a server, or a display device.
[0050] In operation, the electronic device 102 may be configured to capture, via the setof image-capture devices 104, visual information of an environment associated with a user. In an instance, the visual information may include objects and individuals surrounding the user. The set of image-capture devices 104 may include, but not limited to, cameras, infrared sensors, or depth sensors that capture visual information of the environment surrounding the user. Details related to capture of visual information are further provided, for example, in FIG. 3 (at 302).
[0051] The electronic device 102 may be configured to determine first information indicative of at least one of an object or an individual in the environment, based on the captured visual information. The electronic device 102 may analyze the captured visual information and may detect an object or an individual relatable with the user. The electronic device 102 may determine the first information based on the detection. Details related to determination of first information are further provided, for example, in FIG. 3 (at 304).
[0052] The electronic device 102 may be configured to receive sensor data from the set of biosensors 106 associated with the user. The set of biosensors 106 may include a plurality of sensors including, but not limited to, a sEMG sensor, a BCI sensor, an EEG sensor, an ECG sensor, an eye tracker, a voice recognition sensor, a GSR sensor, an EDA sensor, a face detector, a motion sensor, and the like. The first sensor data may include data related to, but not limited to, muscle movements, brain, neurons, voluntary / involuntary response, heart, eyes, voice, skin response, dermal response, facial features, motion associated with the body of the user, and the like. Details related to reception of sensor data are further provided, for example, in FIG. 3 (at 306).
[0053] The electronic device 102 may be configured to determine a subvocalized command of the user, based on the received sensor data from the set of biosensors 106 associated with the user. The set of biosensors 106 may include the sEMG sensors that may be configured to detect muscle movements associated with the subvocalized command of the user. The electronic device 102 may be configured to detect surfaceelectromyography (sEMG) signals based on the received sensor data. The detected sEMG signals correspond to the subvocalized command of the user. Further, the electronic device 102 may convert the detected sEMG signals into textual content and synthesize an audible speech from the converted textual content. For an instance, sEMG sensor may detect muscle movements associated with subvocalization of the user. The sEMG sensor may be placed throat or face of the user to capture subtle muscle movements that occur when the user thinks about speaking without actually producing audible sounds. The captured subtle muscle movements may allow the electronic device 102 to interpret the user's intended commands without requiring vocal output. The sEMG signals, received from the sEMG sensor, may correspond to the subvocalized command of the user. The electronic device 102 may analyze the detected sEMG signals to determine an intended emotional tone associated with subvocalized command of the user. The electronic device 102 may further adjust characteristics of the synthesized audible speech based on the determined intended emotional tone. Details related to determination of subvocalized command are further provided, for example, in FIG. 3 (at 308).
[0054] The electronic device 102 may be configured to apply the neural language model 108 on the determined first information and the determined subvocalized command. The electronic device 102 may process the determined first information and subvocalized command through the neural language model 108 using natural language processing techniques. Details related to application of the neural language model are further provided, for example, in FIG. 3 (at 310).
[0055] The electronic device 102 may be configured to determine, based on the application of the neural language model 108, second information corresponding to at least one of an audio feedback or a haptic feedback for the user. The neural language model 108 may process the visual information and subvocalized command to determine the second information. The second information may include instructions to guide the user toperform a task associated with the determined subvocalized command. In an embodiment, the task may correspond to day-to-day activities of the user. Alternatively, the task may correspond to a sports activity such as, baseball, cricket, badminton, and the like, in which the user participates. Details related to determination of second information are further provided, for example, in FIG. 3 (at 312).
[0056] The electronic device 102 may be configured to render the determined second information to the user. The electronic device 102 may control the user device 120 to render the second information on the user device 120. Details related to rendering of second information are further provided, for example, in FIG. 3 (at 314).
[0057] Unlike prior solutions that may rely solely on voice commands or simple gesture recognition, the disclosed technique integrates multiple input modalities including visual information, biosensor data, and subvocalized commands. This approach offers advantages such as improved accuracy in interpreting user intentions, enhanced privacy for the user as commands may not be audible to others, and the ability to provide more context-aware and personalized assistance. The combination of visual analysis and biosensor data allows the system to understand both the user's environment and their intended actions, enabling more precise and relevant guidance.
[0058] FIG. 2 a block diagram of an electronic device for provision of task guidance using subvocalized commands, visual analysis, and biosensor data, in accordance with at least one embodiment of the disclosure. FIG. 2 is described in conjunction with elements from FIG. 1. With reference to FIG. 2, there is shown an exemplary block diagram 200 of the electronic device 102. The electronic device 102 may include circuitry 202, a memory 204, a network interface 206, and an input / output (I / O) device 208. The I / O device 208 may include a display device 208A. The network interface 206 may connect the electronic device 102 with the server 114, the database 116, and the user device 120, via the communication network 118.
[0059] The circuitry 202 may include suitable logic, circuitry, and / or interfaces that may be configured to execute program instructions associated with different operations to be executed by the electronic device 102. The operations may include, for instance, capturing visual information of environment associated with a user, first information determination, sensor data reception, subvocalized command determination, neural language model application, second information determination, second information rendering, and the like. The circuitry 202 may include one or more processing units, which may be implemented as a separate processor. In an embodiment, the one or more processing units may be implemented as an integrated processor or a cluster of processors that perform the functions of the one or more specialized processing units, collectively. The circuitry 202 may be implemented based on a number of processor technologies known in the art. Examples of implementations of the circuitry 202 may be an X86-based processor, a Graphics Processing Unit (GPU), a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or a combination thereof.
[0060] The memory 204 may include suitable logic, circuitry, interfaces, and / or code that may be configured to store one or more instructions to be executed by the circuitry 202. The one or more instructions stored in the memory 204 may be configured to execute the different operations of the circuitry 202 (and / or the electronic device 102). The memory 204 may be further configured to store the neural language model 108, the facial recognition model 110, and the motion tracking model 112. Examples of implementation of the memory 204 may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Hard Disk Drive (HDD), a Solid-State Drive (SSD), a CPU cache, and / or a Secure Digital (SD) card.
[0061] The network interface 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communication between the electronic device 102 and various devices of the network environment (such as, the server 114, the database 116, and the user device 120), via the communication network 118. The network interface 206 may be implemented by use of various known technologies to support wired or wireless communication of the electronic device 102 with the communication network 118. The network interface 206 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuitry.
[0062] The network interface 206 may be configured to communicate via wireless communication with networks, such as the Internet, an Intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a plurality of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), 5thGeneration (5G) New Radio (NR), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11 a, IEEE 802.11 b, IEEE 802.11 g or IEEE 802.11 n), voice over Internet Protocol (VoIP), light fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), a protocol for email, instant messaging, and a Short Message Service (SMS).
[0063] The I / O device 208 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the captured visual information, receive the sensor data, and render the second information. For example, the I / O device 208 may receive images or videos and the sensor data from the user device 120. The I / O device 208 may be furtherconfigured to render the second information on the user interface of a user device, such as the user device 120. Examples of the I / O device 208 may include, but are not limited to, a display (e.g., a touch screen), a keyboard, a mouse, a joystick, a microphone, or a speaker. Examples of the I / O device 208 may further include braille I / O devices, such as, braille keyboards and braille readers.
[0064] The display device 208A may include suitable logic, circuitry, and interfaces that may be configured to display or render the second information. In some embodiments, the display device 208A may be a touch screen which may enable a user to provide a userinput via the display device 208A. The display device 208A may be realized through several known technologies such as, but not limited to, at least one of a Liquid Crystal Display (LCD) display, a Light Emitting Diode (LED) display, a plasma display, or an Organic LED (OLED) display technology, or other display devices. In accordance with an embodiment, the display device 208A may refer to a display screen of a head mounted device (HMD), a smart-glass device, a see-through display, a projection-based display, an electro-chromic display, or a transparent display. Various operations of the circuitry 202 are described further, for example, in FIG. 3.
[0065] FIG. 3 is a diagram that illustrates an execution pipeline for provision of task guidance using subvocalized commands, visual analysis, and biosensor data, in accordance with at least one embodiment of the disclosure. FIG. 3 is described in conjunction with elements from FIG. 1 and FIG. 2. With reference to FIG. 3, there is shown an exemplary execution pipeline 300 that may include operations for task guidance using subvocalized commands, visual analysis, and biosensor data. In some cases, the execution pipeline 300 may be implemented by the circuitry 202 of the electronic device 102. The execution pipeline 300 may include operations from 302 to 314. The operations 302 to 314 may be performed by any suitable system, apparatus, or device, such as the example electronic device 102 of FIG. 1 or the circuitry 202 of FIG. 2. Although illustratedwith discrete blocks, the steps and operations associated with one or more of the blocks of the execution pipeline 300 may be divided into additional blocks, combined into fewer blocks, or eliminated, depending on the particular implementation.
[0066] At 302, visual information may be captured. The circuitry 202 may be configured to acquire the captured visual information including high-resolution imagery and depth data from the environment surrounding the user. In an embodiment, the circuitry 202 may acquire the captured visual information based on application of the multi-modal sensing approach on the set of image-capture devices 104. The multi-modal sensing approach may involve utilizing a combination of RGB cameras, infrared sensors, and depth cameras from the set of image-capture devices 104, each of which may contribute unique data to create a comprehensive environmental model. For instance, RGB cameras may capture color and texture information, infrared sensors may detect heat signatures for improved object detection in low-light conditions, and depth cameras may provide precise spatial information for accurate distance estimation and object localization. The circuitry 202 may be configured to synchronize multiple camera feeds with sub-millisecond precision, which may ensure temporal alignment of data from different sensors.
[0067] To address varying environmental conditions, the circuitry 202 may implement adaptive exposure control and high dynamic range (HDR) imaging techniques, to automatic adjust for scenarios that range from dimly lit indoor spaces to bright outdoor environments. Real-time image stabilization may be achieved through a combination of hardware-based optical image stabilization (OIS) and software-based electronic image stabilization (EIS), using gyroscopic sensors and predictive techniques to compensate for user movement and vibrations. Additionally, the circuitry 202 may employ wide-angle lenses with fields of view exceeding degrees to capture a broader perspective, which may enhance situational awareness for the user. In scenarios where the user navigates through a crowded urban environment, the expanded field of view may allow for simultaneousdetection of approaching or nearby pedestrians, nearby vehicles, and relevant signage. The circuitry 202 may also incorporate advanced noise reduction techniques such as temporal averaging and spatial filtering to improve image quality in challenging conditions, such as low-light or high-motion environments. The circuitry 202 may utilize on-device artificial intelligence (Al) processors to perform real-time scene segmentation, which may allow rapid identification and prioritization of relevant environmental features to optimize subsequent operations.
[0068] At 304, first information may be determined. The circuitry 202 may be configured determine the first information indicative of at least one of objects or an individual in the environment, based on the captured visual information. For example, the first information may include, but not limited to an object Identification for the object recognition or detection, individual user recognition, or spatial context determination. The circuitry 202 may be configured to process the captured visual information using advanced computer vision techniques and machine learning models. The processing may include implementation of a cascade of convolutional neural networks for object detection, segmentation, and classification. The circuitry 202 may utilize transfer learning techniques to adapt pre-trained models for specific environmental contexts to improving accuracy and reduce computational overhead. For instance, in a crowded urban environment, the circuitry 202 may employ a specialized pedestrian detection model fine-tuned on city street scenes to accurately identify and track individuals in complex and dynamic settings. The facial recognition model 110 may employ a combination of landmark detection, texture analysis, and deep metric learning to identify individuals with high precision, even under difficult conditions such as partial occlusion or varying poses.
[0069] In an example, the circuitry 202 may incorporate multi-modal fusion techniques, based on a combination of RGB image data with depth information from infrared sensors to enhance object detection and scene understanding in low-light conditions. Alternativeapproaches may include the use of attention mechanisms in neural networks to focus on salient features, or the implementation of graph neural networks to model spatial relationships between objects in the scene. The circuitry 202 may also employ real-time instance segmentation techniques, such as Mask R-CNN, to not only detect objects but also precisely delineate their boundaries, which may be crucial for tasks like obstacle avoidance or fine-grained object manipulation guidance. The circuitry 202 may also utilize temporal information through recurrent neural networks or 3D convolutions to track objects and predict their trajectories, to enable more robust and anticipatory assistance for the user.
[0070] At 306, sensor data may be received. The circuitry 202 may be configured to receive the sensor data from set of biosensors 106 associated with the user. For example, the circuitry 202 may acquire and process high-fidelity electromyographic signals from the set of biosensors 106. The processing may involve implementation of adaptive noise cancellation techniques to isolate subtle muscle activations associated with subvocalization from background physiological signals. The circuitry 202 may utilize advanced signal processing techniques, such as, wavelet decomposition or empirical mode decomposition, to extract relevant features from the raw electromyographic signals. Additionally, the circuitry 202 may be configured to handle multiple sensor channels simultaneously, which may allow redundancy and improved signal quality through sensor fusion techniques. For example, in a noisy environment like a busy street, the circuitry 202 may employ a combination of surface EMG sensors placed on the user's throat and jaw, along with bone conduction sensors on the temples, to capture subvocalized commands more accurately. In an instance, the circuitry 202 may implement real-time adaptive filtering using a least mean squares (LMS) algorithm to continuously adjust to changing noise conditions.
[0071] The circuitry 202 may use various machine learning techniques, such as support vector machines (SVM) or deep neural networks, to classify and interpret the processed EMG signals, to enable more robust command recognition even in difficult conditions. In scenarios where traditional EMG sensors may be impractical, such as during physical activities, the circuitry 202 may utilize alternative biosensing modalities like capacitive sensing through smart textiles integrated into the user's clothing. The circuitry 202 may also be capable of dynamic adjustment of its sampling rate and resolution based on the detected signal quality and user activity level, which may optimize power consumption while a high signal fidelity may be maintained.
[0072] At 308, a subvocalized command may be determined. The circuitry 202 may be configured to determine a subvocalized command of the user, based on the received sensor data. The circuitry 202 may be configured to interpret the processed EMG signals using a combination of pattern recognition techniques and deep learning models. This may involve implementing advanced recurrent neural networks (RNNs) or long short-term memory (LSTM) networks to capture complex temporal dependencies in the muscle activation patterns associated with subvocalization. The circuitry 202 may employ transfer learning techniques to adapt pre-trained models to individual users, to account for variations in muscle physiology, subvocalization styles, and even potential speech impediments. For instance, the circuitry 202 may use a model pre-trained on a large dataset of subvocalized commands and fine-tune the model on a specific user's data to improve accuracy and response time.
[0073] In an example, the circuitry 202 may implement context-aware command interpretation based on environmental cues, user history, and current task context to disambiguate similar subvocalized commands. The interpretation may involve integration of data from other sensors, such as accelerometers or gyroscopes, to understand the user's current activity and refine command interpretation. The circuitry 202 may alsoemploy ensemble methods, based on a combination of outputs from multiple models (e.g., convolutional neural networks for spatial features and LSTMs for temporal features) to improve robustness and accuracy in diverse environments. In a specific scenario, if a user subvocalizes "navigate" in a crowded street, the circuitry 202 may interpret the subvocalization command as a request for navigation assistance, based on the user's location, previous routes, and current obstacles, objects, or individuals detected by other sensors. Alternatively, the same subvocalized command in a quiet room may be interpreted as a request to browse through a digital interface. The circuitry 202 may also incorporate adaptive noise cancellation techniques to filter out EMG artifacts caused by non-speech related muscle movements, ensuring clean signal processing even in dynamic environments or during physical activities.
[0074] In some embodiment, the circuitry 202 may be configured to detect the sEMG signals based on the received sensor data. The detected sEMG signals may correspond to the subvocalized command of the user. Further, the circuitry 202 may be configured to convert the detected sEMG signals into textual content and synthesize the audible speech from the converted textual content.
[0075] In some embodiments, the circuitry 202 may be configured to analyze the detected sEMG signals to determine an intended emotional tone of the user. The analysis may include processing the amplitude, frequency, and patterns of muscle activations associated with different emotional states. The circuitry 202 may utilize machine learning techniques trained on datasets of sEMG signals correlated with various emotions to accurately classify the user's intended emotional tone. Based on the determinate intended emotional tone, the circuitry 202 may adjust characteristics of the synthesized audible speech, such as pitch, speed, volume, and intonation, to reflect the user's emotional state. The adjustments may enable the synthesized speech to convey not only the literal contentof the user's subvocalized commands but also the emotional context, potentially enhancing the naturalness and expressiveness of the communication.
[0076] At 310, a neural language model may be applied to the first information and the subvocalized command. The circuitry 202 may be configured to apply the neural language model 108 to the first information and the subvocalized command. The circuitry 202 may be configured to process the determined first information and subvocalized command through the neural language model 108 using natural language processing techniques. The application of the neural language model 108 may involve implementation of transformer-based architectures, such as BERT or GPT, to generate contextually appropriate responses. The circuitry 202 may utilize multi-modal fusion techniques to integrate visual information with linguistic inputs, which may enable more comprehensive scene understanding and response generation. For example, when a user subvocalizes "identify" while looking at a crowded street scene, the circuitry 202 may combine object detection results with the command to provide a detailed description of the people and vehicles present in the street scene.
[0077] The neural language model 108 may employ few-shot learning or meta-learning approaches to quickly adapt to new scenarios or user-specific language patterns without extensive retraining. Additionally, the circuitry 202 may incorporate attention mechanisms to focus on relevant parts of the street scene, which may improve the accuracy and relevance of generated responses. In scenarios where real-time processing is crucial, such as in case of navigation at a busy intersection, the circuitry 202 may implement model compression techniques like knowledge distillation or quantization to reduce latency and maintain high accuracy. The neural language model 108 may also utilize transfer learning from large pre-trained models, fine-tuned on domain-specific datasets to enhance performance in specialized contexts, such as medical terminology for hospital navigation or technical jargon for workplace assistance. The circuitry 202 may implement continuallearning techniques to incrementally update the neural language model 108 based on user interactions, to gradually improve personalization over time.
[0078] At 312, second information may be determined. The circuitry 202 may be configured to determine, based on the application of the neural language model 108, the second information corresponding to at least one of an audio feedback or a haptic feedback for the user. The circuitry 202 may be configured to generate multi-modal feedback such as the audio feedback or the haptic feedback tailored to the user's needs and environmental context. The generation of the multi-modal feedback may involve implementation of a decision-making model based on multiple factors, such as user preferences, environmental hazards, task complexity, user's current emotional state, and real-time sensor data to determine the optimal feedback modality. For example, the decision-making model may be an approach implemented by the electronic device 102 to evaluate multiple factors and determine the most suitable feedback modality for the user. The decision-making model integrates diverse inputs, such as user preferences (individual choices or desired feedback forms), environmental hazards (safety risks in the surroundings), task complexity (the difficulty or intricacy of the user's activity), user's emotional state (current mood or cognitive condition), and real-time sensor data (dynamic information captured from the user's biosensors). For audio feedback, the circuitry 202 may utilize advanced text-to-speech synthesis with dynamic prosody modulation to indicate urgency, emphasis, or emotional nuances. The text-to-speech synthesis may include pitch, speed, and volume adjustment based on an importance of spoken information or stress level of the user. For example, when a user is guided through a crowded street, the electronic device 102 may use a calm, steady voice for general directions and may switch to a more urgent tone to warn about an approaching vehicle.
[0079] For haptic feedback, the circuitry 202 may implement vibrotactile patterns that encode spatial information or complex instructions through intricate variations in intensity,frequency, duration, and location of vibrations. The implementation may include creation of a "haptic language" where different vibration patterns represent specific objects or actions. For instance, the haptic feedback may include a series of short, intense vibrations moving from left to right across the user's back may indicate the direction and speed of an approaching object. The circuitry 202 may also incorporate adaptive feedback mechanisms that may learn from user responses over time, to adjust the intensity and frequency of haptic feedback based on the user's sensitivity and preferences.
[0080] In some implementations, the circuitry 202 may utilize a hybrid approach, based on a combination of audio and haptic feedback synergistically. For example, for a user to locate a specific object, the circuitry 202 may generate audio descriptions of the object's appearance while using haptic feedback to guide the user's hand towards the object's location. Here, the audio descriptions may be generated based on the captured visual information. Additionally, the circuitry 202 may incorporate the situational awareness to adjust the haptic feedback based on the user's environment. For example, the situational awareness may be automatically switching to haptic-only mode in quiet settings like libraries or using bone conduction audio in noisy environments to ensure clarity of instructions.
[0081] The circuitry 202 may be configured to detect frequently-visited places based on the captured visual information and further determine first navigation assistance to the user for the detected frequently visited places. The second information may include the determined first navigation assistance information for the detected frequently-visited places. For example, the circuitry 202 may be configured to identify and store locations that the user frequently visits based on the visual information captured over time. This may allow the circuitry 202 to build a personalized map of important places for the user, such as their home, workplace, or favorite stores. The circuitry 202 may use accumulated knowledge of the personalized map to provide customized first navigation assistance,offering familiar landmarks and routes tailored to the user's regular habits. When navigating to or near these frequently-visited places, the circuitry 202 may include specific, contextual information (i.e., first navigation information) as part of the second information provided to the user, which may enhance their spatial awareness and confidence in familiar environments.
[0082] The circuitry 202 may be configured to detect and track objects in a virtual environment associated with the user, and further determine second navigation assistance information for the virtual environment, based on the detected and tracked objects. The second information may include the determined second navigation assistance information for the virtual environment. For example, the circuitry 202 may be configured to detect and track objects within virtual environments associated with the user, such as in virtual reality or augmented reality settings. The detection capability may allow the circuitry 202 to provide real-time assistance and guidance in digital spaces, similar to how the circuitry 202 functions in physical environments. Based on the detected and tracked virtual objects, the circuitry 202 may determine the second navigation assistance information specifically tailored for the virtual environment. Such a virtual navigation assistance may include directions, object descriptions, and spatial relationships within the digital space, which may enhance the user's ability to interact with and navigate through virtual worlds. The second information rendered to the user may incorporate this virtual environment-specific guidance, enabling a seamless assistive experience across both physical and digital realms.
[0083] The circuitry 202 may be configured to learn user movement patterns based on the captured visual information and the received sensor data, and further determine third navigation assistance information for the environment, based on the learned movement patterns. The second information may include the determined third navigation assistance information for the environment. For example, the circuitry 202 may be configured toanalyze and learn the user's movement patterns by processing the captured visual information and the received sensor data over time. This learning process may involve identifying recurring routes, preferred paths, and common behaviors in various environments. Based on these learned movement patterns, the circuitry 202 may determine third navigation assistance information tailored to the user's habitual movements and preferences. This personalized navigation assistance may include anticipatory guidance, such as proactively alerting the user to upcoming familiar turns or potential obstacles on frequently traversed routes. The second information provided to the user may incorporate this learned, pattern-based navigation assistance, which may offer a more intuitive and personalized guidance experience that aligns with the user's established routines and preferences.
[0084] At 314, second information may be rendered. The circuitry 202 may be configured to render the second information through various modalities tailored to the user's needs and environmental context. The second information may include instructions to guide the user to perform a task associated with the determined subvocalized command. For example, the circuitry 202 may implement a multi-modal feedback system that combines audio, haptic, and visual content. For audio feedback, the circuitry 202 may utilize advanced text-to-speech synthesis with dynamic prosody modulation to convey urgency, emphasis, or emotional nuances. For example, when a user is being guided through a crowded street, the circuitry 202 may use a calm, steady voice for general directions but switch to a more urgent tone to warn about an approaching vehicle. Haptic feedback may be delivered through a network of vibration motors strategically placed on the user's body, to create complex spatial patterns for indication of direction, distance, or object type. In a scenario where the user is being navigated in a supermarket, the circuitry 202 may use different vibration patterns to distinguish between product categories, with different intensities indicative of a proximity of the user to a product.
[0085] The circuitry 202 may also implement adaptive feedback mechanisms that learn from user responses over time, to adjust the intensity and frequency of cues based on the user's sensitivity and preferences. In some implementations, the circuitry 202 may utilize bone conduction audio technology to provide clear audio feedback in noisy environments without impeding the user's ability to hear ambient sounds. For users with varying degrees of visual impairment, the circuitry 202 may incorporate high-contrast visual displays or augmented reality overlays to enhance residual vision capabilities. The rendering process may also take into account the user's cognitive load, which may dynamically adjust the complexity and frequency of feedback to prevent information overload in difficult environments.
[0086] In some cases, the circuitry 202 may be further configured to detect and classify objects in the environment using computer vision techniques and provide audio feedback to the user by describing the detected and classified objects. The process of detection and classification may involve using machine learning models trained on large datasets of objects to recognize and categorize items in the user's environment. For example, the circuitry 202 may detect a chair and provide audio feedback such as "There is a wooden chair about two meters in front of you."
[0087] The disclosed techniques utilize a combination of RGB cameras, infrared sensors, and depth cameras to capture high-resolution imagery and depth data from the user's environment. The multi-modal approach may provide a more comprehensive understanding of the surroundings, enabling the circuitry 202 to offer more accurate and context-aware assistance to visually impaired users. By incorporating subvocalized command detection through sEMG sensors, the disclosed electronic device 102 may allow for discreet and hands-free user input. The electronic device 102 may enhance the user experience by enabling more natural and intuitive interactions with the assistive system, particularly in public or social settings where vocal commands may be less desirable.
[0088] For example, the circuitry 202 may employ natural language model or a multimodal fusion to generate contextual information associated with the audio / haptic feedback. The generation of contextual information may result in more personalized and relevant guidance for the user, taking into account factors such as user preferences, environmental hazards, and task complexity to determine the optimal feedback modality. The execution pipeline 300 may be configured to operate continuously, processing visual and sensor information in real-time. This may allow the electronic device 102 to adapt quickly to changes in the environment or new user commands, which may potentially enhance the user's ability to navigate and interact with their surroundings more effectively and safely. The disclosed electronic device 102 may implement sophisticated vibration patterns to convey rich environmental and navigational information. By using a library of standardized haptic primitives that may be combined and modulated, the circuitry 202 may provide more nuanced and informative tactile feedback to users, potentially improving their spatial awareness and navigation abilities.
[0089] FIG. 4A and FIG. 4B are diagrams that illustrate an exemplary execution pipeline for detection of objects using visual noise cancellation technique, in accordance with at least one embodiment of the disclosure. FIG. 4A and FIG. 4B are described in conjunction with elements from FIG. 1 , FIG. 2, and FIG. 3. With reference to FIG. 4A and FIG. 4B, there is shown an exemplary execution pipeline 400. The execution pipeline 400 is executed by the circuitry 202 of the electronic device 102, or the user device 120. In an exemplary embodiment, the execution pipeline 400 includes a user input 402, a decision node associated with ad hoc request (in 404 of Fig. 4), a microphone 406, a camera I infrared (IR) sensor 408, an ultrasonic sensor 410, a data interpretation system 412, a preprocessing system 414, a convolutional neural network 416, a detection layer system 418, a non-max suppression system 420, a post processing system 422, an object detection system 424, a filtering system 426, a visualization system 428, an integrationsystem 430, and haptic devices 432. The execution pipeline 400 illustrates the process of detection of objects using visual noise cancellation technique in the network environment 100. The various operations of the execution pipeline 400 may be performed by the circuitry 202 of the electronic device 102.
[0090] In some cases, the user input 402 may initiate data acquisition process through interaction with the electronic device 102. The decision node associated with ad hoc request (in 404 of Fig. 4) may determine the type of data to be collected (whether ad hoc or not) based on the user input 402 or environmental conditions. For example, if the user input 402 may be the subvocalized command related to object identification (i.e., in case of an ad hoc request), then the electronic device 102 may activate the camera (or IR) sensor 408 and ultrasonic sensor 410 for visual and spatial data collection. Otherwise, the electronic device 102 may activate the microphone 406 to capture an audio command .
[0091] The microphone 406 may be a part of the set of biosensors 106 and may capture audio data, including ambient sounds and the user's voice commands. The audio data may be crucial to understand the user's environment and intentions. For instance, the microphone 406 may detect traffic sounds, indicative of a busy street environment, which may prompt the circuitry 202 to prioritize obstacle detection and navigation assistance.
[0092] The camera / IR sensor 408 may be a part of the set of image-capture devices 104, that captures high-resolution visual information of the environment associated with the user. In some cases, the camera / IR sensor 408 may include multiple IR sensors disposed at different locations and multiple cameras with different focal lengths or fields of view to provide comprehensive visual coverage. The camera (or IR) sensor 408 may work in conjunction with the ultrasonic sensor 410 to provide depth information and enhance spatial awareness.
[0093] The ultrasonic sensor 410 may emit high-frequency sound waves and measure the time taken for the waves to bounce back, to provide accurate distance measurementsof nearby objects. The ultrasonic sensor 410 may be particularly useful for detection of changes in elevation or terrain in the environment. For example, the ultrasonic sensor 410 may detect a curb or a step, allowing the circuitry 202 to provide audio or haptic feedback user input 402 that is descriptive of these changes in elevation to the user input 402.
[0094] The data interpretation system 412 may receive inputs that include the captured audio data from the microphone 406, and the visual and spatial data collection from the camera (or IR) sensor 408 and the ultrasonic sensor 410. The circuitry 202 may process the multi-modal data to create a comprehensive understanding of the user's environment. In some cases, the data interpretation system 412 may utilize the neural language model 108, the facial recognition model 110, and the motion tracking model 112 stored in the memory 204 of the electronic device 102 to analyze the data.
[0095] The preprocessing system 414 may preprocess the captured visual information to prepare the raw image data for analysis by performance of tasks such as noise reduction, color correction, and image resizing on the captured visual information. The preprocessing system 414 may enhance the quality of the received inputs, improving the accuracy of subsequent processing stages.
[0096] The convolutional neural network (CNN) 416 may be a deep learning model configured to extract relevant features from the preprocessed virtual information. In some cases, the convolutional neural network 416 may be trained on large datasets of objects, text, and faces to enable accurate recognition and classification. The CNN 416 may be particularly useful to detect and recognize text or signs in the environment, which may allow the electronic device 102 to provide audio feedback to the user input 402 descriptive of the detected text or signs.
[0097] The detection layer system 418 may analyze the features extracted by the CNN 416 to identify specific objects, individuals, or environmental elements. The detection layer system 418 may work in conjunction with the facial recognition model 110 to perform facialrecognition on individuals in the environment. The circuitry 202 may be configured to provide audio feedback to the user input 402 identifying the recognized individuals, enhancing social interactions for visually impaired users.
[0098] The non-max suppression system 420 may refine detection results (identified objects and recognized faces) by elimination of redundant or overlapped detections, which may ensure that each object or feature is identified only once. The non-max suppression system 420 may improve the accuracy and efficiency of object detection process.
[0099] The post processing system 422 may further refine the detection results, based on application of additional filters or adjustments to improve the accuracy of object identification and localization. The post processing system 422 may detect objects in motion within the environment. Further, the post processing system 422 may allow the circuitry 202 to provide real-time audio or haptic feedback to the user input 402. The haptic feedback may be description of the motion of detected objects.
[0100] The object detection system 424 may consolidate the processed information to create a comprehensive map of the user's environment, including the location, type, and characteristics of the user's environment. Further, the comprehensive map of the user environment may include the location, type, and characteristics of the detected objects. In some embodiments, the object detection system 424 may utilize detected objects. In other embodiments, the object detection system 424 may utilize data from the ultrasonic sensor 410 to enhance depth perception and create to enhance depth perception and create accurate 3D maps of the environment.
[0101] The filtering system 426 may prioritize and categorize the detected objects based on their relevance to the user's current task or navigation needs. For example, the filtering system 426 may highlight potential obstacles or points of interest that are most relevant to the user's current location and direction of movement.
[0102] The visualization system 428 may create a digital representation of the environment may be used for internal processing “or”, in some cases, , for generating augmented reality (AR) overlays for the user with partial vision.
[0103] The integration system 430 may combine the processed visual information from other sensors, such as the microphone 406 and the ultrasonic sensor 410, to create a multi-modal representation of the environment. This integrated data may be used to generate more accurate and context-aware feedback for the user input 402.
[0104] The haptic devices 432 may receive signals from the integration system 430 and provide audio feedback or tactile feedback, respectively to the user input 402. In some embodiments, the haptic devices 432 may use advanced vibration patterns to convey spatial information or alert the user to changes in elevation or terrain.
[0105] The process of detection of objects using visual noise cancellation technique may enable the electronic device 102 to detect frequently visited places based on the captured visual information. The circuitry 202 may be configured to provide navigation assistance to the user input 402 for these detected frequently visited places, enhancing the user's ability to navigate familiar environments independently.
[0106] In some embodiments, the set of image-capture devices 104 may include LiDAR sensors for more accurate depth perception and 3D mapping. The LiDAR sensors may emit laser pulses and measure the time taken for the pulses to return, creating highly detailed 3D maps of the environment. The additional depth information may enhance the system's ability to detect changes in elevation or terrain and provide more accurate navigation assistance.
[0107] The combination of multiple sensors and sophisticated processing systems may allow for comprehensive environmental analysis and interpretation. The multi-modal approach may enable the electronic device 102 to provide a detailed, context-awareassistance to visually impaired users, which may enhance the ability of the users to navigate and interact with their surroundings safely and independently.
[0108] FIG. 5 is a diagram that illustrates an exemplary execution pipeline for visual text / sign recognition, in accordance with at least one embodiment of the disclosure. FIG. 5 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, and FIG. 4B. With reference to FIG. 5, there is shown an exemplary execution pipeline 500. The execution pipeline 500 is executed by the circuitry 202 of the electronic device 102, or the user device 120. In an exemplary emebodiment. The execution pipeline 500 includes a user input 502, an image acquisition system 504, images / videos 506, a text / sign board 508, a mask extraction module 522, a speech conversion system 544, and an output device 546. The various operations of the execution pipeline 500 may be performed by the circuitry 202 of the electronic device 102.
[0109] As shown in the execution pipeline 500, the user input 502 may be interaction of the user with the image acquisition system 504. The image acquisition system 504 may be part of the set of image-capture devices 104 associated with the electronic device 102. In some embodiments, the image acquisition system 504 may capture images / videos 506 (and, for example, the text / sign board 508 information) from the environment surrounding the user input 502. The captured images / videos 506 and the text / sign board 508 may correspond to the captured visual information.
[0110] The captured visual information may undergo frame annotation 510, where key features and objects within each frame of the images / videos 506 may be labeled. The circuitry 202 may perform the frame annotation 510. The process may enhance the accuracy of subsequent object detection and classification. Based on the annotation, the circuitry 202 may perform frame preprocessing 512 on each frame, which may involve noise reduction, color correction, and contrast enhancement to optimize the image quality for analysis.
[0111] The preprocessed frames may undergo augmentation 514, where various transformations may be applied to increase the diversity of the training data and improve the robustness of object detection models. The circuitry 202 may perform the augmentation 514. The process of the augmentation 514 may include techniques such as random rotations, flips, or changes in brightness and contrast. The augmented frames may be subject to resizing 516 to ensure consistent input dimensions for the neural networks, followed by splitting 518 to divide the images into smaller segments for more efficient processing. The circuitry 202 may perform the resizing 516 and the splitting 518.
[0112] The circuitry 202 may perform normalization and anchoring 520 on the split frames, where pixel values of the frames may be standardized and reference points may be established for object localization. The circuitry 202 may further process the normalized frames using the mask extraction module 522 (which may correspond to a mask R-CNN Region Proposal Network (RPN) and feature extraction and detection module), which may utilize advanced architectures such as EfficientDet or DETR (DEtection TRansformer) for object detection and segmentation. The advanced architectures may offer improved accuracy and efficiency compared to traditional YOLO or Mask R-CNN approaches.
[0113] The circuitry 202 may perform image subtraction 524 to isolate objects of interest from the background. The image subtraction 524 may then be followed by the generation of an RPN proposal 526 by the circuitry 202, where potential regions containing objects may be identified. The circuitry 202 may perform classification of region proposals to specific classes 528 to assign categories to the detected objects.
[0114] Based on the classification, the circuitry 202 may perform Region-of-lnterest (Rol) pooling and box (bounding box) generation 530, where features from the Rol may be extracted and resized to a fixed dimension. The circuitry 202 may perform refinement and classification 532 post the Rol pooling and box generation 530, where initial detections may be further refined to improve accuracy. The circuitry 202 may further extract classscores from Rol (as in 534 of the fig. 5), based on assignment of confidence values to each detected object and its classification.
[0115] The circuitry 202 may further refine the class scores with threshold values ( as in 536 of fig. 5). The refined class score may result in output objects with annotations and positions (as in 538 of fig. 5). At the (text detection) decision box 540, the circuitry 202 may determine whether any text or signs are present in the processed images. If any text or signs are detected to be present, the circuitry 202 may send detected text as parameters 542 for further processing. The circuitry 202 may further pass the detected text parameters through a speech conversion system 544, that may generate audio descriptions of the detected text. In some embodiments, the circuitry 202 may be configured to generate audio descriptions of visual content in movies or television programs based on the captured visual information. The audio descriptions may be provided to the user input 502 through the output device 546, which may be part of the I / O device 208 of the electronic device 102. The output device 546 may include speaker, headphone, recorder, amplifier, and the like. If no text is detected at the decision box 540, the circuitry 202 may proceed directly to TTS (Text-to-Speech) output 548, where other relevant information about the detected objects and environment may be converted to audio feedback for the user input 502.
[0116] FIG. 6A and FIG. 6B are diagrams that collectively illustrate an exemplary execution pipeline for facial recognition, in accordance with at least one embodiment of the disclosure. FIG. 6A and FIG. 6B are described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, FIG. 4B, and FIG. 5. With reference to FIG. 6A and FIG. 6B, there is shown an exemplary execution pipeline 600. The execution pipeline 600 may include blocks 602 to 650. The various operations of the execution pipeline 600 may be performed by any suitable system, apparatus, or device, such as the example electronic device 102 of FIG. 1 , or the circuitry 202 of FIG. 2.
[0117] At 602, a camera, which may be part of the set of image-capture devices 104, may capture visual information of the environment associated with the user. At 604, the circuitry 202 may preprocess the captured information. The image preprocessing may involve techniques such as noise reduction, contrast enhancement, and color correction. The image preprocessing may be crucial for improvement of the quality of input data for subsequent analysis, for instance in challenging light conditions.
[0118] At 606, a training dataset may be generated or accessed. The dataset may be continuously updated based on the user's experiences and the audio feedback or the haptic feedback from the user. The continuous update may allow the circuitry 202 to adapt to individual needs and preferences over time. The circuitry 202 may apply the trained / pre- trained model to the preprocessed image to further process the image. The processing may involve using the training dataset to train machine learning models to recognize various objects, faces, and text in different environments.
[0119] At 608, frames associated with the image may be split. The circuitry 202 may split the frames associated with the image. The frames may be split based on sampling of the frames. The process of splitting may enable parallel processing of different parts of the image, which may potentially reduce overall computation time. For example, on analysis of a complex scene like multiple people who stand at a busy intersection, different regions of the image may be processed simultaneously to quickly identify features of faces of the people, obstacles, traffic, and the like.
[0120] At 610, blobs may be generated. The circuitry 202 may generate the blobs. The generated blobs, or Binary Large Objects, may represent areas of interest within the frames associated with the image. The process of blob generation may involve identification of contiguous regions of pixels that may represent distinct objects or features in the image. In a scenario where the user explores a museum, the blob generation mayhelp to isolate individual exhibits or artwork from the background, which may facilitate more focused analysis and description.
[0121] At 612, the generated blobs may be fed to a deep learning model for facial recognition tasks, for instance, a Visual Geometry Group (VGG) Face2 Network. The circuitry 202 may feed the blobs to the deep learning model. In case of social interactions, the VGG Face2 Network may allow the circuitry 202 to identify familiar faces or detect the presence of people in the user's vicinity. For instance, in a workplace setting, the circuitry 202 may help the user to recognize colleagues or identify when someone may approach the user.
[0122] At 614, output from the deep learning model may undergo face normalization and bounding box construction. The circuitry 202 may perform the face normalization and the bounding box construction. The process may standardize the detected faces and create defined areas around the faces for further analysis.
[0123] At 616, weak predictions may be filtered from the normalization and bounding box construction. The weak prediction may be a low-confidence detection made by the electronic device 102 during the analysis or decision-making process. The weak predictions may be considered unreliable because the predictions may not meet the confidence thresholds or fall below the standards required for accurate results. The weak predictions may arise due to ambiguous input data, noise, or limitations in the detection techniques. The filtering of the weak predictions may eliminate low-confidence detections to improve overall accuracy. The filtering of weak predictions may involve application of confidence thresholds or ensemble methods to retain only the most reliable detections. In a practical application, such as navigation through a crowded public space, the filtering may help to provide the user with information about clearly identified obstacles or individuals while false alarms may be minimized.
[0124] At 618, bounding boxes may be updated and confidence scores for the detected faces may be generated. The circuitry 202 may update the bounding boxes and generate the confidence scores. Information related to updated bounding boxes and confidence scores may be crucial for providing accurate spatial information to the user. Bounding boxes are rectangular regions that encapsulate detected faces within the image, which may provide precise spatial localization. As the circuitry 202 processes subsequent frames or refines its initial detections, these bounding boxes may be adjusted to fit the facial contours or account more accurately for movement. Confidence scores associated with each detected face indicate a level of certainty in its identification. The generation of the confidence scores may be based on various factors such as facial feature clarity, pose, lighting conditions, and match quality against known facial templates.
[0125] At 620, relative position and confidence of recognition may be found, which may provide spatial information about the detected faces within the environment. The circuitry 202 may determine the relative position and confidence of recognition. The finding of relative position and confidence of recognition may involve mapping the 2D image detections to 3D space, potentially using additional depth information from sensors like the ultrasonic sensor 410. In a scenario where the user may be reaching for an object, the 3D mapping may allow the circuitry 202 to provide precise guidance on the object's location and the best approach to grasp the object.
[0126] At 622, face verification and identification may be performed. The circuitry 202 may perform the face verification and identification, using the facial recognition model 110 stored in the memory 204 of the electronic device 102. For instance, when exploring a new environment, this feature may help to identify familiar landmarks or confirm the presence of specific objects the user may be searching for.
[0127] At 624, class identification and text description generation may be performed based on the application of the facial recognition model 110. The circuitry 202 may performclass identification and text description generation. The class identification and text description generation may involve creating natural language descriptions of the detected objects, their attributes, and spatial relationships. In a practical application, such as description of a painting in an art gallery, this feature may generate detailed descriptions of the artwork's composition, colors, and style, enhancing the user's appreciation and understanding. Based on the identified class and generated text description, information about the corresponding recognized individuals may be prepared for output to the user input 502.
[0128] At 626, it may be determined whether a face has been detected. The circuitry 202 may use the facial recognition model 110 to determine whether a face has been detected. If a face of an individual is detected, then text examination and scanning may be performed at 628. The text examination and scanning may involve analysis of any text associated with the detected face, such as name tags or nearby signs. Further, when the face of the individual may not be detected, then the circuitry 202 may proceed to 636 of the FIG. 6.
[0129] At 630, letter to sound conversion may be performed. The circuitry 202 may perform letter to sound conversion based on the detected text, preparing the letter for an audio output. This process may involve translation of the detected text into phonetic representations and applying text-to-speech techniques to generate natural-sounding pronunciations.
[0130] At 632, a speech generation module may convert the processed text and face recognition results into natural language descriptions. The circuitry 202 may perform may convert the processed text and face recognition results into natural language descriptions using the speech generation module. The speech generation module may employ text-to- speech techniques to produce clear and expressive audio feedback. For instance, whendescribing a complex scene, the speech generation module may vary tone and pacing to emphasize important elements or convey spatial relationships more effectively.
[0131] At 634, the generated audio may be transmitted as output through headphones, which may be part of the I / O device 208 of the electronic device 102. The circuitry 202 may render the generated audio through the speakers or headphones of the electronic device 102.
[0132] At 636, in case a face is not detected in the image, the circuitry 202 may activate / initiate an embedded device associated with the electronic device 102 to trigger ultrasonic sensor at 638. In an instance, the embedded device may include Raspberry Pi®. The initiation of the embedded device may involve initializing various hardware components and loading necessary software modules for ultrasonic sensing and data processing. In some cases, the initialization process may include self-diagnostic routines to ensure all sensors associated with the electronic device 102 are functioning correctly before beginning active sensing.
[0133] At 638, ultrasonic sensors may be triggered and ultrasonic waves may be transmitted from the ultrasonic sensors. The circuitry 202 may trigger the ultrasonic sensors. The ultrasonic sensors may be part of the set of biosensors 106 in the electronic device 102. Transmission parameters, such as frequency and pulse duration, of the ultrasonic waves may be dynamically adjusted based on the current environment. For example, for an open outdoor space, the lower frequency waves may be used for longer- range detection, while in a cluttered indoor environment, higher frequency waves may be used for more precise short-range sensing.
[0134] At 640, it may be determined whether the reflected echo waves are received. The circuitry 202 may determine whether the reflected echo waves are received. This determination process may include signal processing techniques to detect and isolate valid echo signals from background noise. For complex environments with multiple reflectivesurfaces like crowded urban streets, the circuitry 202 may use advanced techniques to separate overlapping echo signals and identify relevant obstacles or objects.
[0135] In case the reflected echo waves are received, data related to the reflected echo waves may be sent to the embedded device for analysis at 642. The circuitry 202 may send the data related to the reflected echo waves. The data transmission may occur in real-time, that may allow immediate processing and feedback generation. In some cases, the circuitry 202 may employ edge computing techniques to perform initial data analysis directly on sensors, which may reduce latency and bandwidth requirements for data transmission to the main processor. Further, in case the reflected echo waves are not received, the circuitry 202 may loop back to 638 for the ultrasonic sensing and the ultrasonic waves transmission.
[0136] At 644, a threshold value may be set, based on the data received by the embedded device or main processing unit. The circuitry 202 may set the threshold value. The threshold value may be dynamically adjusted based on the user's current activity, speed of movement, or environmental conditions. For instance, when the user is moving quickly, such as during a jog, the threshold value may be increased to provide earlier warnings about potential obstacles.
[0137] At 646, a distance (D) may be calculated based on the time taken for the ultrasonic waves to return. The circuitry 202 may calculate the distance. The calculation may incorporate factors such as air temperature and humidity, which may affect the speed of sound, to improve accuracy. In complex environments, the circuitry 202 may use multiple ultrasonic sensors and triangulation techniques to provide more precise 3D spatial information.
[0138] At 648, it may be determined whether the calculated distance is less than the threshold. The circuitry 202 may determine whether the calculated distance is less than the threshold. The comparison may not only consider absolute distance but also the rateof change in the distance, which may allow the circuitry 202 to predict potential collisions or rapidly approaching objects. For example, in a dynamic environment like a busy sidewalk, the circuitry 202 may prioritize warnings about fast-moving objects that are on a collision course with the user, even if the user are currently further away than stationary obstacles.
[0139] If the distance may be less than the threshold, then the distance may be indicative of the presence of a nearby object or obstacle. Further, the ultrasonic sensors may be triggered again, and ultrasonic waves may be transmitted for continuous monitoring at 650. Further, when the distance may be determined to be greater than the threshold, then the circuitry 202 may loop back to the 646 (for other epoch or iteration of the data).
[0140] In some embodiments, the circuitry 202 may be configured to detect and suppress ambient noise in the environment and enhance audio signals relevant to the user's task or navigation. The process may involve using adaptive noise cancellation techniques in conjunction with the ultrasonic sensing data to provide clearer audio feedback to the user input 502, especially in noisy environments.
[0141] The combination of visual processing, facial recognition, and ultrasonic sensing allows the electronic device 102 to provide comprehensive environmental awareness to the user input 502. Based on integration of the different sensing modalities, the electronic device 102 may offer more accurate and context-aware assistance, enhancing the user's ability to navigate and interact with their surroundings safely and independently.
[0142] FIG. 7A and FIG. 7B are diagrams that collectively illustrate an exemplary execution pipeline for movie narration assistance, in accordance with at least one embodiment of the disclosure. FIG. 7A and FIG. 7B are described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, FIG. 4B, FIG. 5, FIG. 6A, and FIG. 6B. With reference to FIG. 7A and FIG. 7B, there is shown an exemplary execution pipeline 700.The execution pipeline 700 may include operations that may be performed by any suitable system, apparatus, or device, such as the example electronic device 102 of FIG. 1 , or the circuitry 202 of FIG. 2.
[0143] With reference to FIG. 7A, at 702, a user may interact with the electronic device 102. The interaction may involve determination of subvocalized commands of the user based on sensor data captured by the set of biosensors 106. For example, the set of biosensors 106, such as sEMG sensors may be used to detect subtle muscle movements associated with subvocalization. For example, in a noisy environment such as a busy street, the user may subvocalize a command to identify nearby objects without having to speak aloud and thus maintain privacy and reduce interference from ambient noise.
[0144] At 704, the set of image-capture devices 104 (such as a camera I IR sensor) may capture visual information from the environment surrounding the user. The set of image-capture devices 104 may include a combination of RGB cameras, infrared sensors, and depth sensors to provide a comprehensive view of the environment. For instance, in a low-light situation such as a dimly lit movie hall, the infrared sensors may capture heat signatures of objects and people, while the depth sensors may provide accurate spatial information, enabling the electronic device 102 to create a detailed 3D map of the surroundings. In an instance, the circuitry 202 may check video frames associated with a movie (referred to as movie frames, herein) for scene change checkpoint.
[0145] At 706, it may be determined whether the entire scene is captured at the scene change checkpoint. The circuitry 202 may determine whether the entire scene is captured at the scene change checkpoint. In case the entire scene is determined to be captured, the control may move to 708.
[0146] At 708, a feature extraction module may process the captured visual information associated with the scene to identify key element such as features and anchors associated with the scene. The circuitry 202 may process the captured visual information using thefeature extraction module. The feature extraction module may employ computer vision techniques, such as convolutional neural networks (CNNs), to detect and classify objects, recognize text, and identify facial features.
[0147] At 710, a region proposal network (RPN) may be used to generate potential regions of interest proposals based on the extracted features and anchors. The circuitry 202 may generate potential Rol proposals using the RPN. For instance, in a crowded urban scene, the RPN may prioritize regions including the anchors, to ensure that the most critical information is processed first.
[0148] At 712, a region of interest (Rol) pooling may be performed based on the generate potential regions of interest proposals. The circuitry 202 may perform the Rol pooling. In an instance, features from the proposed regions of interest may be extracted and standardized for Rol pooling. The Rol pooling may enable more efficient processing of varied-sized objects and scenes.
[0149] At 714, a classification and regression module may analyze the pooled Rol to determine the labels and annotations associated with features and anchors of the pooled Rol. The circuitry 202 may analyze the pooled Rol. The classification and regression module may utilize machine learning models trained on diverse datasets to accurately categorize environmental elements. For example, in a workplace setting scene, these models may distinguish between different types of office equipment, identify office personnel, and recognize important documents or signage.
[0150] At 716, a non-max suppression module may refine and consolidate the labels and annotations based on elimination of redundant or overlapping information. The circuitry 202 may refine and consolidate the labels and annotations. This process aims to simplify the output by retaining only the most relevant and distinct detections, which is crucial for provision of clear and concise feedback to the user. The consolidation becomes particularly important in complex environmental scenes where multiple objects may bepresent. The consolidation may help to prevent information overload and enhance the user's ability to interpret the surroundings effectively.
[0151] At 718, a scene processing module may integrate the consolidated labels and annotations to create a comprehensive understanding of the environment of the movie scene and correspondingly generate a processed scene information. The circuitry 202 may integrate the consolidated labels and annotations. The scene processing module may consider spatial relationships between objects, temporal information, and contextual cues to generate a coherent representation of the scene. For instance, in a home environment scene, the scene processing module may not only identify individual objects but also understand their functional relationships, such as recognition of a dining area with a table, chairs, and nearby kitchen appliances.
[0152] At 720, the processed scene information may be converted into appropriate feedback for the user through headphones or haptic devices. The circuitry 202 may convert processed scene information into appropriate feedback. In some cases, the circuitry 202 may be configured to incorporate bone conduction technology for audio feedback. This may allow the user to receive audio information without blockage of their ears and may maintain their ability to hear environmental sounds directly.
[0153] In case the entire scene is not captured at the scene change checkpoint (at 706), control may move to 722. With reference to FIG. 7B, at 722, the preprocessing module may prepare preprocessed data for analysis based on execution of tasks such as noise reduction, image enhancement, data normalization, resizing, and formatting of the movie frame. The circuitry 202 may perform the preprocessing tasks of the preprocessing module. In some embodiments, the preprocessing module may employ adaptive techniques to adjust preprocessing parameters based on environmental conditions in the movie frame. For example, in a movie frame scenario, where the user moves from a brightly lit outdoor area to a dimly lit indoor space, the preprocessing module maydynamically adjust contrast and brightness levels to maintain optimal image quality for subsequent processing stages.
[0154] At 724, convolutional neural network (CNN) may extract relevant features from the preprocessed data. The circuitry 202 may extract relevant features from the preprocessed data using the CNN. The CNN may be configured to recognize a wide range of objects, textures, and patterns in the movie frame, and correspondingly generate high- level spatial information.
[0155] At 726, a detection layer module may analyze the high-level spatial information to generate detection results based on detection of specific objects, individuals, or environmental elements. The circuitry 202 may analyze the high-level spatial information. The detection layer module may employ advanced techniques such as anchor-free detection or feature pyramid networks to improve accuracy across different scales and object types. The detection layer module may generate bounding boxes around detected objects, individuals, or environmental elements. The detection layer module may also generate confidence scores for the detected objects, individuals, or environmental elements.
[0156] At 728, a non-max suppression module may refine the detection results based on elimination of redundant or low-confidence detections. The circuitry 202 may refine the detection results using the non-max suppression module. The non-max suppression module may use techniques such as soft-NMS (Non-Maximum Suppression) or adaptive thresholding to improve the accuracy of object localization. In a crowded desk scene, the non-max suppression module may help to provide clear information about distinct objects without confusion from overlapped detections of items on a crowded desk.
[0157] At 730, a post processing module may further refine the detection results, by applying contextual reasoning and temporal consistency checks. The circuitry 202 may further refine the detection results using the post processing module. The post processingmodule may refine bounding boxes and corresponding coordinates based on the application of contextual reasoning and temporal consistency checks. For example, in a library scene, the post processing module may prioritize detections of bookshelves and book spines and temporarily downplay other environmental details.
[0158] At 732, an object detection module may detect objects in the scene based on analysis of the refined bounding boxes and corresponding coordinates. The circuitry 202 may use the object detection module to detect objects in the scene using, for example, one or more deep learning models.
[0159] At 734, a filtering module may annotate images associated with the scene based on prioritization and categorization of the detected objects based on their relevance. The circuitry 202 may annotate images using the filtering module. The filtering module may employ user preference learning techniques to adapt its filtering criteria over time based on the user's behavior and feedback.
[0160] At 736, a visualization module may create a digital representation of the annotated images. In some cases, the circuitry 202 may be configured to detect and track objects in a virtual environment and provide audio or haptic feedback to the user for navigation within the virtual environment. This capability may be particularly useful for training scenarios or for previewing unfamiliar physical spaces in a safe, virtual context.
[0161] At 738, an integration module may combine objects with annotations in the annotated images with path planning and guidance data (received from different sources such as the server 114, user device 120, and the like) to create a multi-modal representation of the environment of the scene. The integration module may employ sensor fusion techniques and probabilistic models to resolve any conflicts or inconsistencies in the data from different sources. The circuitry 202 may combine objects with annotations using the integration module.
[0162] The headphones / haptic devices 720 may deliver the multi-modal representation of the environment of the scene to the user through audio or tactile feedback. In some embodiments, the haptic feedback mechanism may include a glovelike device with multiple vibration points for more nuanced directional and spatial information. The advanced haptic system may enable the communication of complex spatial relationships and object characteristics through touch.
[0163] In some embodiments, the electronic device 102 may also detect a plurality of audio streams in the environment. The electronic device 102 may further receive a user input from the user device 120, where the user input is indicative of a selection of an audio stream of interest from the plurality audio streams detected in the environment. The electronic device 102 may further amplify the selected audio stream of interest and output the amplified audio stream to the user. The electronic device 102 may further receive a user command. Based on the user command, the electronic device 102 may perform at least one of play-back the amplified audio stream, store the amplified audio stream, or generate a summary of the amplified audio stream.
[0164] For example, the circuitry 202 may detect multiple audio streams present in the user's environment. This feature may allow the circuitry 202 to identify and differentiate between various sound sources, such as conversations, background noise, or specific audio signals. The circuitry 202 may receive input from the user device 120, which may indicate the user's selection of a particular audio stream of interest from among the detected streams. Upon receipt of this selection, the circuitry 202 may amplify the chosen audio stream to enhance its clarity and prominence for the user. The circuitry 202 may then output the amplified audio stream to the user, potentially through headphones or other audio output devices. Additionally, the circuitry 202 may accept user commands for further processing of the amplified audio stream. Based on these commands, the circuitry 202 may offer functionalities such as playback of the amplified audio, storage for futurereference, or generation of a content summary, which may provide the user with flexible options to manage and interact with the selected audio information.
[0165] FIG. 8A and FIG. 8B are diagrams that collectively illustrate an exemplary execution pipeline for room assistance, in accordance with at least one embodiment of the disclosure. FIG. 8A and FIG. 8B are described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, FIG. 4B, FIG. 5, FIG. 6A, FIG. 7A, and FIG. 7B. With reference to FIG. 8A and FIG. 8B, there is shown an exemplary execution pipeline 800. The execution pipeline 800 may include operations that may be performed by any suitable system, apparatus, or device, such as the example electronic device 102 of FIG. 1 , or the circuitry 202 of FIG. 2.
[0166] At 802, visual information may be captured. A camera, which may be part of the set of image-capture devices 104, may capture visual information associated with a room in which the user is. In an instance, the camera may capture multiple frames per second to enable real-time analysis of dynamic environment of the room. The circuitry 202 may control the capture of the visual information.
[0167] At 804, the captured images undergo image preprocessing (in a similar manner as elaborated at 604 of FIG. 6A). The circuitry 202 may perform the image preprocessing. For instance, when the room is dimly lit, the image preprocessing techniques may be applied to extract maximum detail from low-light images of the room.
[0168] At 806, a training dataset may be generated or accessed (in a similar manner as elaborated at 606 of FIG. 6A). The dataset may be continuously updated based on the user's experiences and feedback, allowing the circuitry 202 to adapt to individual needs and preferences over time. In a practical application, the training dataset may be enriched with examples of common objects and their locations in the room.
[0169] At 808, the preprocessed frames may be split into smaller segments or regions for more efficient processing (in a similar manner as elaborated at 608 of FIG. 6A). The circuitry 202 may split the preprocessed frames into smaller segments or regions.
[0170] At 810, blob generation may be performed (in a similar manner as elaborated at 610 of FIG. 6A). The circuitry 202 may perform the blob generation.
[0171] At 812, the generated blobs may be fed into a VGG Face2 Network or a similar advanced facial recognition model (in a similar manner as elaborated at 612 of FIG. 6A). The circuitry 202 may feed the generated blobs into a facial recognition model.
[0172] At 814, weak predictions may be filtered out (in a similar manner as elaborated at 616 of FIG. 6A). The circuitry 202 may filter out weak predictions.
[0173] At 816, bounding boxes may be updated, and confidence scores may be generated for the detected objects or faces (in a similar manner as elaborated at 618 of FIG. 6A). The circuitry 202 may perform the bounding box update and the confidence score generation. For example, when a layout of a room may be described, the circuitry 202 may use the bounding box information to convey the size and position of furniture items relative to the user's location.
[0174] At 818, relative position and confidence of recognition for detected objects or faces may be determined (in a similar manner as elaborated at 620 of FIG. 6A). The circuitry 202 may determine relative position and confidence of recognition for detected objects or faces.
[0175] At 820, image verification and identification may be performed. The circuitry 202 may perform the image verification and identification. The process of image verification and identification may include cross-referencing of detected objects or faces with a database of known entities, potentially stored in the database 116 connected to the server 114.
[0176] At 822, class identification and text description may be generated (in a similar manner as elaborated at 624 of FIG. 6A). The circuitry 202 may perform the class identification and text description.
[0177] At 824, it may be determined whether an object has been detected. The circuitry 202 may determine whether an object has been detected. In case an object is detected in the room, control may proceed to 826 for text examination and scanning. The text examination and scanning may be particularly useful for reading text on signs, labels, or documents in the room. Further, in case, the object may not be detected in the room, then the process may be directed to 834 for initiation of the embedded device.
[0178] At 828, letter-to-sound conversion may be performed. The circuitry 202 may perform the letter-to-sound conversion. The letter-to-sound conversion may involve translation of the detected text into phonetic representations for audio output. In some cases, the process may consider the user's preferred language or dialect for more naturalsounding pronunciation.
[0179] At 830, a speech generation module may convert the processed text into natural language audio output (in a similar manner as elaborated at 632 of FIG. 6A). The circuitry 202 may convert the processed text into the natural language audio output.
[0180] At 832, the generated audio may be output through headphones, which may be part of the I / O device 208 of the electronic device 102. The circuitry 202 may control output of the generated audio. In some cases, the circuitry 202 may be configured to learn user movement patterns in the room based on the captured visual information and sensor data and provide personalized navigation assistance based on the learned movement patterns. For example, if the circuitry 202 determines that the user frequently pauses at certain locations in the room of their home or office, it may proactively provide more detailed information about such locations or offer reminders about tasks associated with the locations.
[0181] At 834, the embedded device or similar embedded system may be initiated (in a similar manner as elaborated at 636 of FIG. 6B). The circuitry 202 may initiate the embedded device.
[0182] At 836, ultrasonic sensors may be triggered to transmit ultrasonic waves (in a similar manner as elaborated at 638 of FIG. 6B). The circuitry 202 may trigger the ultrasonic sensors.
[0183] At 838, it may be determined whether reflected echo waves are received. The circuitry 202 may determine whether reflected echo waves are received (in a similar manner as elaborated at 640 of FIG. 6B).
[0184] In case the reflected echo waves are received, at 840, the circuitry 202 may send data related to the reflected echo waves to the embedded device or main processing unit for analysis (in a similar manner as elaborated at 642 of FIG. 6B). Further, in case the reflected echo waves are not received, at 840, the circuitry 202 may loop back to the 836 for that triggers the ultrasonic sensors to generate and transmit the ultrasonic wave.
[0185] At 842, a threshold value may be set for distance calculations (in a similar manner as elaborated at 644 of FIG. 6B). The circuitry 202 may set the threshold value.
[0186] At 844, distance (D) may be calculated based on the time taken for the ultrasonic waves to return (in a similar manner as elaborated at 646 of FIG. 6B). The circuitry 202 may calculate the distance.
[0187] At 846, it may be determined whether the calculated distance is less than the threshold. The circuitry 202 may determine whether the calculated distance is less than the threshold (in a similar manner as elaborated at 648 of FIG. 6B).
[0188] If D is less than the threshold, indicating the presence of a nearby object or obstacle, at 848, the circuitry 202 may send a signal to vibration motor sensors or haptic feedback devices. The haptic feedback may be configured to convey not just the presence of an obstacle, but also information about its size, direction, and rate of approach. Forinstance, a large, stationary obstacle directly ahead may be indicated by a strong, continuous vibration, while a smaller moving object, such as ball, approaching from the side may be signaled by a series of quick pulses increasing in intensity.
[0189] In some embodiments, the haptic feedback may be integrated with audio feedback to provide a multi-modal alert system. For example, while haptic feedback system provides immediate spatial awareness, audio feedback system may offer more detailed descriptions of the detected obstacles or environment. The combination of tactile and auditory feedback may enable user to build a more comprehensive mental map of the surroundings, enhancing ability of the user to navigate complex environments independently.
[0190] The integration of visual processing, ultrasonic sensing, and multi-modal feedback may provide a comprehensive system for environmental awareness and navigation assistance. Based on combination of different sensing and feedback modalities, the electronic device 102 may offer more accurate, context-aware, and personalized assistance, significantly enhancing the user's ability to interact with the surroundings safely and independently.
[0191] FIG. 9A and FIG. 9B are diagrams that collectively illustrate an exemplary execution pipeline for detection of infrared rays, in accordance with at least one embodiment of the disclosure. FIG. 9A and FIG. 9B are described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, FIG. 4B, FIG. 5, FIG. 6A, FIG. 7A, FIG. 7B, FIG. 8A, and FIG. 8B. With reference to FIG. 9A and FIG. 9B, there is shown an exemplary execution pipeline 900. The execution pipeline 900 may include operations that may be performed by any suitable system, apparatus, or device, such as the example electronic device 102 of FIG. 1 , or the circuitry 202 of FIG. 2.
[0192] At 902, an IR camera, which may be part of the set of image-capture devices 104, may capture visual information from the environment. In some cases, the IR cameramay be used in conjunction with other sensors to provide a comprehensive view of the surroundings. For example, in low-light conditions such as nighttime navigation, the IR camera may detect heat signatures of objects and individuals, complementing the data from traditional RGB cameras. The circuitry 202 may control the IR camera to capture the visual information.
[0193] At 904, the captured visual information may undergo image acquisition. The circuitry 202 may acquire image associated with the captured visual information. The image acquisition may involve techniques such as noise reduction, contrast enhancement, and image stabilization to improve the quality of the visual information. In some embodiments, the image acquisition process may employ adaptive techniques that adjust parameters based on current lighting conditions and movement speed of the electronic device 102.
[0194] At 906, based on the image acquisition, captured images and videos may be stored for further analysis. The circuitry 202 may store the captured images and videos in the memory 204 or the database 116. In some cases, the storage may involve implementation of efficient data compression techniques to optimize storage usage while image quality suitable for subsequent processing stages may also be maintained.
[0195] At 908, frame annotation may be performed on the stored images / videos. The circuitry 202 may perform the frame annotation. In the process of frame annotation, key features, objects, and regions of interest may be labelled within each frame. In some cases, the frame annotation may utilize pre-trained machine learning models to automatically identify and tag common elements in the environment, such as doors, stairs, or pedestrian crossings.
[0196] At 910, frame preprocessing may be performed. The circuitry 202 may perform the frame preprocessing. The frame preprocessing may involve operations such as resizing, color space conversion, and normalization to prepare the annotated frames forinput into more advanced analysis modules. In some cases, the preprocessing operation may also include data augmentation techniques to artificially increase the diversity of the input data, to improve the robustness of subsequent object detection and classification processes.
[0197] At 912, a Gaussian filter may be applied to the preprocessed frames. The circuitry 202 may apply the Gaussian filter. This filtering operation may help reduce noise and smooth the image, to improve the performance of edge detection and feature extraction in later stages. In some cases, the parameters of the Gaussian filter may be dynamically adjusted based on the current environmental conditions and the specific requirements of subsequent processing modules.
[0198] At 914, 916, and 918, augmentation, resizing, and splitting of the filtered frames may respectively be performed. The circuitry 202 may perform the augmentation, resizing, and splitting of the filtered frames. These operations may prepare the visual data for efficient processing by machine learning models used in later stages. For example, augmentation techniques such as random rotations or brightness adjustments may help improve the system's ability to recognize objects under varying conditions, while resizing and splitting the frames into smaller segments may allow for parallel processing and improved computational efficiency.
[0199] At 920, normalization and anchoring may be performed on the processed frame segments. The circuitry 202 may perform the normalization and anchoring. The normalization and anchoring may standardize the frame segments and establish reference points for object localization, potentially improving the accuracy and consistency of subsequent detection and classification processes.
[0200] At 922, a Fast R-CNN RPN, feature extraction, and detection system may be deployed. The circuitry 202 may control the Fast R-CNN RPN, feature extraction, anddetection system to perform multiple operations for visual analysis of the images / videos, which may include operations 924 to 936.
[0201] At 924, image subtraction may be performed to isolate moving objects or changes in the scene by comparing consecutive frames. The circuitry 202 may perform the image subtraction. In some cases, the image subtraction may be particularly useful in detection of dynamic elements in the environment, such as approaching vehicles or pedestrians.
[0202] At 926 and 928, a classifier module and a proposal module may work in tandem to identify regions of interest and recommend potential object locations within the frame. The circuitry 202 may control operations of the classifier module and the proposal module.
[0203] At 930, a CNN module may extract high-level features from the proposed regions, to enable more accurate object classification and recognition. The circuitry 202 may control operations of the CNN module. In some cases, the CNN module may employ transfer learning techniques, utilizing pre-trained models fine-tuned on domain-specific datasets to enhance performance in particular environments, such as indoor spaces or urban streets.
[0204] At 932 and 934, bounding box regression and Rol pooling may be performed to refine the object detections, adjust bounding box coordinates, and aggregate features to improve localization accuracy. The circuitry 202 may perform the bounding box regression and the Rol pooling.
[0205] At 936, anchoring images may establish reference points for object detection across different scales and aspect ratios, to improving accuracy of identification of objects of varying sizes and shapes. The circuitry 202 may anchor images.
[0206] The Fast R-CNN RPN, feature extraction, and detection system 922 may be feed an output associated with the visual analysis of the images / videos into a non-max suppression model 938 for post-processing. The circuitry 202 may control operations ofthe non-max suppression model 938. The non-max suppression model 938 may apply techniques such as non-maximum suppression to eliminate redundant or low-confidence detections, which may result in a cleaner and more accurate set of object identifications. In some cases, the non-max suppression model 938 may employ adaptive thresholding techniques that adjust based on the complexity of the current scene and the density of detected objects.
[0207] At 940, object detection may be performed. The circuitry 202 may perform object detection, which consolidates the processed information to create a comprehensive representation of the detected objects in the environment. The process of object detection may include operations such as, identification 942 to determine object categories, confidence scoring 944 to assess reliability of each detection, localization 946 for precise determination of object positions and tracking 948 to monitor object movements over time. The circuitry 202 may control the various operations associated with the object detection (i.e., 942 to 948).
[0208] At 950, objects with annotations and positions may be output. The circuitry 202 may generate output objects with annotations and positions as a structured representation.
[0209] At 952, it may be determined whether objects have been successfully detected. The circuitry 202 may determine whether objects have been successfully detected. If objects are detected, at 954, the circuitry 202 may send the parameters associated with the objects to a speech generation model (as described at 956). Examples of the parameters may include object type, position, elevation, shape, and movement.
[0210] At 956, the speech generation model may convert the extracted parameters into natural language descriptions, using context-aware language generation techniques to provide more relevant and easily understandable information to the user. The circuitry 202 may control the speech generation model to determine the natural language descriptions.
[0211] At 958, a speaker I headphones may output the natural language descriptions to the user. The circuitry 202 may control the I / O device 208 (which may include the speaker / headphones) to output an audio including the natural language descriptions to the user.
[0212] If the object is not detected at 952, control may pass to 960. With reference to FIG. 9B, at 960, the embedded device with distance sensor, which may serve as the central processing unit for the distance sensing system, may coordinate the operations of various sensors and processing components, enabling real-time distance measurements and obstacle detection. The circuitry 202 may control the embedded device.
[0213] At 962, the embedded device with distance sensor may initialized. Upon initialization, the embedded device module may perform self-diagnostics and calibration routines to ensure accurate sensor operation. In some cases, the initialization may involve adjusting sensor parameters based on current environmental conditions, such as temperature or humidity, which may affect ultrasonic wave propagation.
[0214] At 964, the set of biosensors 106 may emit high-frequency sound waves into the environment. In some cases, the electronic device 102 may employ multiple ultrasonic sensors arranged in an array configuration to provide wider coverage and enable more precise localization of detected objects.
[0215] At 966, the electronic device 102 may analyze the reflected ultrasonic waves to determine the presence and distance of objects in the environment. The process may involve sophisticated signal processing techniques to filter out noise and identify valid echo signals, particularly in complex environments with multiple reflective surfaces.
[0216] At 968, the electronic device 102 may establish distance thresholds for object detection and obstacle avoidance. In some cases, the thresholds may be dynamically adjusted based on the user's movement speed, the complexity of the current environment, or user-defined preferences for personal space.
[0217] At 970, the electronic device 102 may use the time-of-flight principle to compute precise distances to detected objects. In some cases, the distance calculation may incorporate additional data from other sensors, such as the set of image-capture devices 104, to refine distance estimates and provide more accurate spatial information.
[0218] At 972, the calculated distance information may be formatted and prepared for integration with data from other sensing modalities, such as the visual analysis system described in FIG. 9A. This multi-modal approach may enable more comprehensive and reliable environmental mapping and obstacle detection.
[0219] At 974, the electronic device 102 may compare the calculated distances to the thresholds, to determine whether detected objects pose potential obstacles or hazards to the user. If an object is detected within the threshold distance, control may proceed to haptic trigger at 976.
[0220] At 976, the electronic device 102 may activate haptic feedback devices to alert the user to the presence of nearby objects or obstacles. In some cases, the haptic feedback may be configured to convey not only the presence of an obstacle but also information about its distance, direction, and potential motion. For example, the intensity or frequency of vibrations may increase as the user approaches an obstacle, providing intuitive spatial awareness without relying on auditory or visual cues.
[0221] The integration of visual analysis and distance sensing methods in the electronic device 102 may provide a comprehensive system for environmental perception and navigation assistance. By combining data from multiple sensing modalities, including IR cameras, RGB cameras, and ultrasonic sensors, the electronic device 102 may offer more accurate and reliable object detection, distance measurement, and obstacle avoidance capabilities. This multi-modal approach may enable the electronic device 102 to adapt to a wide range of environmental conditions and user scenarios, potentially enhancing the independence and safety of visually impaired users in various settings.
[0222] FIG. 10 is a block diagram of an exemplary audio processing system, in accordance with at least one embodiment of the disclosure. FIG. 10 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, FIG. 4B, FIG. 5, FIG. 6A, FIG. 7A, FIG. 7B, FIG. 8A, FIG. 8B, FIG. 9A, and FIG. 9B. With reference to FIG. 10, there is shown a diagram 1000 representing exemplary audio processing system. The diagram 1000 includes an EMG signal input 1002, a transduction model 1004, a convolution block 1006, a transformer layer 1008, a WaveNet model 1010, and an audio output 1012. The EMG signal input 1002 may be transmitted to the transduction model 1004, which may contain the convolution block 1006 and the transformer layer 1008. The output of the transduction model 1004 may be connected to the WaveNet model 1010, which may produce the audio output 1012. In an embodiment, the audio processing system of diagram 1000 may be implemented by the electronic device 102.
[0223] The audio processing system may be configured to convert subvocalized speech, detected through EMG signals, into audible speech output. The audio processing system may enhance communication capabilities for individuals with speech impairments or in situations where silent communication is preferred. In some cases, the audio processing system may be integrated with the electronic device 102, utilizing the set of biosensors 106 for EMG signal acquisition.
[0224] The EMG signal input 1002 may receive electrical signals from surface electrodes placed on the user's throat or facial muscles. The signals may correspond to the subtle muscle movements associated with subvocalization. In some cases, the EMG signal input 1002 may employ advanced noise reduction techniques, such as adaptive filtering, to isolate the relevant muscle activity signals from background physiological noise. For example, EMG signal input 1002 may be at a frequency of 800 Hz.
[0225] The transduction model 1004 may process the raw EMG signals to extract meaningful features and patterns associated with speech. The convolution block 1006within the transduction model 1004 may apply a series of convolutional filters to the input signal, potentially identifying local patterns and temporal dependencies in the EMG data. In some cases, the convolution block 1006 may utilize dilated convolutions to capture long- range dependencies without significantly increasing computational complexity.
[0226] The transformer layer 1008 may further process the convoluted features, employing self-attention mechanisms to capture complex relationships between different parts of the input sequence. The layer may be particularly effective in modeling the contextual dependencies present in speech-related muscle activations. In some cases, the transformer layer 1008 may incorporate positional encodings to maintain information about the temporal order of the input signals.
[0227] The WaveNet model 1010 may generate the final audio output 1012 based on the processed EMG signals. This model may employ a series of dilated causal convolutions to generate high-quality audio waveforms. In some cases, the WaveNet model 1010 may be conditioned on the output of the transduction model 1004, which may allow the WaveNet model 1010 to produce speech sounds that closely match the intended subvocalized speech.
[0228] The audio output 1012 may represent the synthesized speech signal, which may be played through speakers or headphones connected to the electronic device 102. In some cases, the audio output 1012 may undergo additional post-processing operations, such as spectral shaping or formant adjustment, to enhance naturalness and intelligibility. For example, the audio output 1012 may be at a frequency of 16 kHz.
[0229] FIG. 11 is an exemplary diagram of a convolution block of the audio processing system of FIG. 10, in accordance with at least one embodiment of the disclosure. FIG. 11 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, FIG. 4B, FIG. 5, FIG. 6A, FIG. 7A, FIG. 7B, FIG. 8A, FIG. 8B, FIG. 9A, FIG. 9B, and FIG. 10. With reference to FIG. 11 , there is shown a diagram representing the convolution block 1006.In an instance, the convolution block 1006 may include a convolution module 1102, a convolution module 1104, and a convolution module 1106. The convolution module 1102 and the convolution module 1104 may be connected in parallel, with their outputs combined and fed into the convolution module 1106.
[0230] The convolution module 1102 may apply convolutions with a stride of 2, effectively downsampling the input during extraction of features. The convolution module 1102 may help to reduce the computational complexity of subsequent layers while important spatial information may be preserved. In some cases, the convolution module 1102 may employ different kernel sizes to capture multi-scale features from the input EMG signals. In an example, the convolution module 1102 may be a convolution neural network model with a width 3 and a stride 2.
[0231] The convolution module 1104 may use convolutions with a larger kernel width to capture broader contextual information from the input signals. The convolution module 1104 may be particularly useful for identifying longer-range patterns in the EMG data that correspond to specific phonemes or speech sounds. In some cases, the convolution module 1104 may incorporate residual connections to facilitate the flow of information across multiple layers. In an example, the convolution module 1104 may be a convolution neural network model with a width 3.
[0232] The convolution module 1106 may further process the combined output of the previous two modules 1102 and 1104 and apply additional feature extraction or dimensionality reduction operations. In some cases, the convolution module 1106 may adapt its parameters based on the characteristics of the input signal, to allow more flexible and robust feature extraction across different users or speaking styles. In an example, the convolution module 1102 may be a convolution neural network model with a width 1 and a stride 2.
[0233] FIG. 12 is a diagram that illustrates a waveform graph between audio and EMG signal patterns, in accordance with at least one embodiment of the disclosure. FIG. 12 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, FIG. 4B, FIG. 5, FIG. 6A, FIG. 7A, FIG. 7B, FIG. 8A, FIG. 8B, FIG. 9A, FIG. 9B, FIG. 10, and FIG. 11. With reference to FIG. 12, there is shown a diagram of a waveform graph 1200 that may include representations of audio waveforms from vocalized speech 1202A, EMG signals from silent speech 1202B, and processed output signals. The visualization of the waveform graph 1200 may help to understand a transformation of the subvocalized speech signals into the audible speech output. The waveform graph 1200 may represent an alignment between the vocalized speech and the EMG signals.
[0234] The audio processing system (implemented, for example, by the electronic device 102) of the disclosure may enable various applications beyond basic speech synthesis. For example, in a medical context, the audio processing system may assist in speech rehabilitation for individuals recovering from stroke or other conditions that affect speech production. By providing real-time auditory feedback based on subvocalized attempts at speech, the audio processing system may help patients relearn proper muscle coordination for speech production.
[0235] In professional settings, the audio processing system may facilitate silent communication in noise-sensitive environments. For instance, factory workers in loud industrial settings may use the system to communicate clearly without shouting or removing hearing protection. Similarly, military personnel may employ the system for covert communication during stealth operations.
[0236] The integration of the audio processing system with other components of the electronic device 102, such as the neural language model 108, may enable more sophisticated applications. For example, the audio processing system may learn to adapt its output based on the user's context, potentially adjusting vocabulary or speech style tomatch different social or professional situations. In some cases, the audio processing system may also incorporate real-time translation capabilities, converting subvocalized speech in one language to audible output in another.
[0237] The audio processing system may also benefit from continuous learning and adaptation. Based on analysis of the relationship between the user's subvocalized EMG patterns and their intended speech output over time, the audio processing system may improve its accuracy and naturalness. In some cases, this adaptation may involve fine- tuning the parameters of the transduction model 1004 or the WaveNet model 1010 based on user feedback or detected errors in the output speech.
[0238] To enhance privacy and security, the audio processing system may incorporate advanced encryption techniques for both the EMG signal processing and the generated audio output. This may prevent unauthorized access to the user's subvocalized speech content, which could contain sensitive information. In some cases, the audio processing system may also include user authentication features, such as, using analysis unique characteristics of the user's EMG patterns to ensure that only authorized individuals may use the device.
[0239] The development of the audio processing system may involve extensive training on diverse datasets, including EMG recordings from individuals with different accents, speaking styles, and physiological characteristics. The comprehensive training approach may help ensure that the system performs robustly across a wide range of users and usage scenarios. In some cases, the system may also incorporate transfer learning techniques, allowing it to quickly adapt to new users with minimal additional training data.
[0240] FIG. 13 is a diagram that illustrates an exemplary implementation of the assistive device in sports assistance, in accordance with at least one embodiment of the disclosure. FIG. 13 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, FIG. 4B, FIG. 5, FIG. 6A, FIG. 6B, FIG. 7A, FIG. 7B, FIG. 8A, FIG. 8B, FIG. 9A,FIG. 9B, FIG. 10, FIG. 11 , and FIG. 12. With reference to FIG. 13, there is shown an exemplary diagram 1300. The diagram 1300 includes a first player (batter) 1302 holding a bat 1304 and a second player (bailer) 1306 holding a ball 1308 in a baseball field 1310. In an instance, the batter 1302 may be visually impaired and the electronic device 102 may assist the batter 1302 in sports like baseball. The bailer 1302 throws the ball 1308 towards the batter 1302. The electronic device 102 may detect the incoming ball 1308, recognize trajectory of the ball 1308, and may provide immediate haptic feedback and / or audio feedback to the batter 1302. In similar manner, the electronic device 102 may recognize trajectory of the ball 1308 once it is hit by the batter 1302, and may provide haptic feedback and / or audio feedback to the bailer 1306. By providing the haptic feedback and / or audio feedback, the electronic device 102 may enable the batter 1302 and the bailer 1306 to react promptly. By bridging sensory gap, the electronic device 102 ensures equal participation and enjoyment in baseball for players with visual impairments, fostering inclusivity and accessibility in baseball and other sports.
[0241] FIG. 14 is a diagram that illustrates an exemplary flowchart of a method of providing task guidance using subvocalized commands, visual analysis, and biosensor data, in accordance with at least one embodiment of the disclosure. FIG. 14 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4A, FIG. 4B, FIG. 5, FIG. 6A, FIG. 7A, FIG. 7B, FIG. 8A, FIG. 8B, FIG. 9A, FIG. 9B, FIG. 10, FIG. 11 , FIG. 12, and FIG. 13. With reference to FIG. 14, there is shown an exemplary flowchart 1400. The flowchart 1400 operations from 1402 to 1416 and may be implemented by the electronic device 102 of FIG. 1 or the circuitry 202 of FIG. 2. The flowchart 1400 may start at 1402 and proceed to 1404.
[0242] At 1404, visual information of an environment associated with a user may be captured, via a set of image-capture devices. The circuitry 202 may capture the visual information of the environment associated with the user via the set of image-capturedevices 104. The circuitry 202 may be configured to acquire high-resolution imagery and depth data from the environment surrounding the user through a sophisticated multi-modal sensing approach. Details related to capturing of visual information are further provided, for example, in FIG. 3 (at 302).
[0243] At 1406, first information indicative of at least one of an object or an individual in the environment may be determined. The circuitry 202 may be configured to determine the first information indicative of at least one of an object or an individual in the environment, based on the captured visual information. The circuitry 202 may be configured to process the captured visual data using advanced computer vision techniques and machine learning models. Details related to determination of first information are further provided, for example, in FIG. 3 (at 304).
[0244] At 1408, sensor data may be received from a set of biosensors associated with the user. The circuitry 202 may be configured to receive sensor data from the set of biosensors 106 associated with the user. The set of biosensors 106 may include a plurality of sensors including, but not limited to, a sEMG sensor, a BCI sensor, an EEG sensor, an ECG sensor, an eye tracker, a voice recognition sensor, a GSR sensor, an EDA sensor, a face detector, a motion sensor, and the like. The first sensor data may include data related to, but not limited to, muscle movements, brain, neurons, voluntary / involuntary response, heart, eyes, voice, skin response, dermal response, facial features, motion associated with the body of the user, and the like. Details related to reception of sensor data are further provided, for example, in FIG. 3 (at 306).
[0245] At 1410, a subvocalized command of the user may be determined. The circuitry 202 may be configured to determine a subvocalized command of the user, based on the received sensor data. In an instance, sEMG sensor may detect muscle movements of the user, associated with subvocalization. The sEMG sensor may be placed throat or face of the user to capture subtle muscle movements that occur when the user thinks aboutspeaking without actually producing audible sounds. The captured subtle muscle movements may allow the electronic device 102 to interpret the user's intended commands without requiring vocal output. Details related to determination of subvocalized command are further provided, for example, in FIG. 3 (at 308).
[0246] At 1412, a neural language model may be applied on the determined first information and the determined subvocalized command. The circuitry 202 may be configured to apply the neural language model 108 on the determined first information and the determined subvocalized command. The electronic device 102 may process the determined first information and subvocalized command through the neural language model 108 using advanced natural language processing techniques. Details related to application of the neural language model are further provided, for example, in FIG. 3 (at 310).
[0247] At 1414, second information corresponding to at least one of an audio feedback or a haptic feedback for the user may be determined, based on the application of the neural language model. The circuitry 202 may be configured to determine, based on the application of the neural language model 108, second information corresponding to at least one of an audio feedback or a haptic feedback for the user. The neural language model 108 may process the visual information and subvocalized command to determine the second information. The second information may include instructions to guide the user to perform a task associated with the determined subvocalized command. Details related to determination of second information are further provided, for example, in FIG. 3 (at 312).
[0248] At 1416, the determined second information may be rendered to the user, wherein the second information may include instructions to guide the user to perform a task associated with the determined subvocalized command. The circuitry 202 may be configured to render the determined second information to the user. The second information may include instructions to guide the user to perform the task associated withthe determined subvocalized command. The electronic device 102 may control the user device 130, such as smart phone, television, LCD display, laptop, and the like, to render the second information on the user device 130. Details related to rendering of second information are further provided, for example, in FIG. 3 (at 314). Control may pass to end.
[0249] Although the flowchart 1400 is illustrated as discrete operations, such as 1402, 1404, 1406, 1408, 1410, 1412, 1414, and 1416, the disclosure is not so limited. Accordingly, in certain embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation without detracting from the essence of the disclosed embodiments.
[0250] Various embodiments of the disclosure may provide a non-transitory computer- readable medium and / or storage medium having stored thereon, computer-executable instructions by a machine and / or a computer to operate an electronic device (for example, the electronic device 102 of FIG. 1 ). Such instructions may cause the electronic device 102 to perform operations that may include capture of visual information of an environment associated with a user via a set of image-capture devices (for example, the set of imagecapture devices 104). The operations may further include determination of first information indicative of at least one of an object or an individual in the environment, based on the captured visual information. The operations may further include receipt of sensor data from a set of biosensors (for example, the set of biosensors 106) associated with the user and determination of a subvocalized command of the user based on the received sensor data. The operations may further include application of a neural language model (for example, the neural language model 108) on the determined first information and the determined subvocalized command. The operations may further include determination of second information indicative of at least one of an audio feedback or a haptic feedback for the user, based on the application of the neural language model 108. The operations may further include rendering the determined second information to the user, where the secondinformation may include instructions to guide the user to perform a task associated with the determined subvocalized command.
[0251] The present disclosure provides an electronic device (for example, the electronic device 102 of FIG. 1 ) that includes circuitry (for example, the circuitry 202 of FIG. 2) configured to capture visual information of an environment associated with a user via a set of image-capture devices (for example, the set of image-capture devices 104). The circuitry 202 may further be configured to determine first information indicative of at least one of an object or an individual in the environment, based on the captured visual information. The electronic device 102 may further receive sensor data from a set of biosensors (for example, the set of biosensors 106) associated with the user and determines a subvocalized command of the user based on the received sensor data. The circuitry 202 may further apply a neural language model (for example, the neural language model 108) on the determined first information and the determined subvocalized command. The circuitry 202 may determine second information indicative of at least one of an audio feedback or a haptic feedback for the user, based on the application of the neural language model 108. The circuitry 202 may render the determined second information to the user, where the second information may include instructions to guide the user to perform a task associated with the determined subvocalized command. In an embodiment, the task corresponds to at least one of a day-to-day activity or a sports activity in which the user participates.
[0252] In some embodiments, the set of image-capture devices 104 may include at least one of a camera, an infrared sensor, or a depth sensor. In some embodiments, the set of biosensors 106 may include surface electromyography (sEMG) sensors configured to detect muscle movements associated with subvocalization command of the user.
[0253] The electronic device 102 may be further configured to detect frequently-visited places based on the captured visual information and further determine first navigationassistance to the user for the detected frequently visited places. The second information may include the determined first navigation assistance information for the detected frequently-visited places.
[0254] The electronic device 102 may also be configured to detect and recognize text or signs in the environment based on the captured visual information. The electronic device 102 may further determine a description of the detected text or signs. The second information may include the determined description of the detected text or signs.
[0255] The electronic device 102 may also be configured to apply a facial recognition model on the captured visual information and determine an identification of the individual in the environment based on the application of the facial recognition model. The second information may include the identification of the individual.
[0256] The electronic device 102 may also be configured to apply a motion tracking model on at least one of the object or the individual in the environment, and determine a description of a motion of at least one of the object or the individual, based on the motion tracking model. The second information may include the determined description of the motion of at least one of the object or the individual.
[0257] The electronic device 102 may be configured to detect changes in elevation or terrain in the environment based on the captured visual information. The second information may include information related to the detected changes in elevation or terrain.
[0258] The electronic device 102 may be configured to detect and suppress an ambient noise in the environment, and further enhance audio signals relevant to the task associated with the determined subvocalized command, based on the detected and suppressed ambient noise.
[0259] The electronic device 102 may be configured to generate audio descriptions of visual content being played-back on a media rendering device in the environment. Theaudio descriptions may be generated based on the captured visual information. The second information may include the generated audio descriptions.
[0260] The electronic device 102 may be configured to detect and track objects in a virtual environment associated with the user, and further determine second navigation assistance information for the virtual environment, based on the detected and tracked objects. The second information may include the determined second navigation assistance information for the virtual environment.
[0261] The electronic device 102 may be configured to learn user movement patterns based on the captured visual information and the received sensor data, and further determine third navigation assistance information for the environment, based on the learned movement patterns. The second information may include the determined third navigation assistance information for the environment.
[0262] The electronic device 102 may be configured to detect and classify the object in the environment using a computer vision model. The second information may include information related to the detection and the classification of the object in the environment.
[0263] In some embodiments, the haptic feedback may include vibration patterns corresponding to different types of environmental information or user instructions.
[0264] In some embodiments, the electronic device 102 may be configured to detect a plurality of audio streams in the environment, receive a user input indicative of a selection of an audio stream of interest from the plurality audio streams detected in the environment, amplify the selected audio stream of interest, and output the amplified audio stream to the user. Based on a user command, the electronic device 102 may perform at least one of: play-back the amplified audio stream, store the amplified audio stream, or generate a summary of the amplified audio stream.
[0265] In some embodiments, the electronic device 102 may be configured to detect surface electromyography (sEMG) signals based on the received sensor data. Thedetected sEMG signals correspond to the subvocalized command of the user. The electronic device 102 may further convert the detected sEMG signals into textual content and synthesize an audible speech from the converted textual content.
[0266] In some embodiments, the electronic device 102 may be configured to apply the neural language model 108 on the converted textual content based on contextual information to refine the converted textual content. The audible speech may be synthesized using the refined textual content.
[0267] In some embodiments, the electronic device 102 may be configured to analyze the detected sEMG signals to determine an intended emotional tone of the user, and may further adjust characteristics of the synthesized audible speech based on the determined intended emotional tone.
[0268] The present disclosure may also be positioned in a computer program product, which comprises all the features that enable the implementation of the methods described herein, and which when loaded in a computer system is able to carry out these methods. Computer program, in the present context, means any expression, in any language, code or notation, of a set of instructions intended to cause a system with information processing capability to perform a particular function either directly, or after either or both of the following: a) conversion to another language, code or notation; b) reproduction in a different material form.
[0269] While the present disclosure is described with reference to certain embodiments, it will be understood by those skilled in the art that various changes may be made, and equivalents may be substituted without departure from the scope of the present disclosure. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the present disclosure without departure from its scope. Therefore, it is intended that the present disclosure is not limited to the embodimentdisclosed, but that the present disclosure will include all embodiments that fall within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:
1. An electronic device, comprising: circuitry configured to: capture, via a set of image-capture devices, visual information of an environment associated with a user; determine first information indicative of at least one of an object or an individual in the environment, based on the captured visual information; receive sensor data from a set of biosensors associated with the user; determine a subvocalized command of the user, based on the received sensor data; apply a neural language model on the determined first information and the determined subvocalized command; determine, based on the application of the neural language model, second information corresponding to at least one of an audio feedback or a haptic feedback for the user; and render the determined second information to the user, wherein the second information includes instructions to guide the user to perform a task associated with the determined subvocalized command.
2. The electronic device of claim 1 , wherein the set of image-capture devices comprises at least one of: a camera, an infrared sensor, or a depth sensor, and the set of biosensors comprises surface electromyography (sEMG) sensors configured to detect muscle movements associated with the subvocalized command of the user.
3. The electronic device of claim 1 , wherein the task corresponds to at least one of a day-to-day activity or a sports activity in which the user participates.
4. The electronic device of claim 1 , wherein the circuitry is further configured to: detect frequently-visited places in the environment based on the captured visual information; and determine first navigation assistance information for the detected frequently- visited places, wherein the second information includes the determined first navigation assistance information for the detected frequently-visited places.
5. The electronic device of claim 1 , wherein the circuitry is further configured to: detect and recognize text or signs in the environment based on the captured visual information; and determine a description of the detected text or signs, wherein the second information includes the determined description of the detected text or signs.
6. The electronic device of claim 1 , wherein the circuitry is further configured to: apply a facial recognition model on the captured visual information; and determine an identification of the individual in the environment based on the application of the facial recognition model, wherein the second information includes the identification of the individual.
7. The electronic device of claim 1 , wherein the circuitry is further configured to:apply a motion tracking model on at least one of the object or the individual in the environment; and determine a description of a motion of at least one of the object or the individual, based on the motion tracking model, wherein the second information includes the determined description of the motion of at least one of the object or the individual.
8. The electronic device of claim 1 , wherein the circuitry is further configured to: detect changes in elevation or terrain in the environment based on the captured visual information, wherein the second information includes information related to the detected changes in elevation or terrain.
9. The electronic device of claim 1 , wherein the circuitry is further configured to: detect and suppress an ambient noise in the environment; and enhance audio signals relevant to the task associated with the determined subvocalized command, based on the detected and suppressed ambient noise.
10. The electronic device of claim 1 , wherein the circuitry is further configured to: generate audio descriptions of visual content being played-back on a media rendering device in the environment, wherein the audio descriptions are generated based on the captured visual information, and the second information includes the generated audio descriptions.11 . The electronic device of claim 1 , wherein the circuitry is further configured to:detect and track objects in a virtual environment associated with the user; and determine second navigation assistance information for the virtual environment, based on the detected and tracked objects, wherein the second information includes the determined second navigation assistance information for the virtual environment.
12. The electronic device of claim 1 , wherein the circuitry is further configured to: learn user movement patterns based on the captured visual information and the received sensor data; and determine third navigation assistance information for the environment, based on the learned movement patterns, wherein the second information includes the determined third navigation assistance information for the environment.
13. The electronic device of claim 1 , wherein the circuitry is further configured to: detect and classify the object in the environment using a computer vision model, wherein the second information includes information related to the detection and the classification of the object in the environment.
14. The electronic device of claim 1 , wherein the haptic feedback comprises vibration patterns corresponding to different types of environmental information or user instructions.
15. The electronic device of claim 1 , wherein the circuitry is further configured to:detect a plurality of audio streams in the environment; receive a user input indicative of a selection of an audio stream of interest from the plurality audio streams detected in the environment; amplify the selected audio stream of interest; output the amplified audio stream to the user; and based on a user command, perform at least one of: play-back the amplified audio stream, store the amplified audio stream, or generate of a summary of the amplified audio stream.
16. The electronic device of claim 1 , wherein the circuitry is further configured to: detect surface electromyography (sEMG) signals based on the received sensor data, wherein the detected sEMG signals correspond to the subvocalized command of the user; convert the detected sEMG signals into textual content; and synthesize an audible speech from the converted textual content.
17. The electronic device of claim 16, wherein the circuitry is further configured to: apply the neural language model on the converted textual content based on contextual information to refine the converted textual content, wherein the audible speech is synthesized using the refined textual content.
18. The electronic device of claim 16, wherein the circuitry is further configured to: analyze the detected sEMG signals to determine an intended emotional tone of the user; andadjust characteristics of the synthesized audible speech based on the determined intended emotional tone.
19. A method, comprising: in an electronic device: capturing, via a set of image-capture devices, visual information of an environment associated with a user; determining first information indicative of at least one of an object or an individual in the environment, based on the captured visual information; receiving sensor data from a set of biosensors associated with the user; determining a subvocalized command of the user, based on the received sensor data; applying a neural language model on the determined first information and the determined subvocalized command; determining, based on the application of the neural language model, second information corresponding to at least one of an audio feedback or a haptic feedback for the user; and rendering the determined second information to the user, wherein the second information includes instructions to guide the user to perform a task associated with the determined subvocalized command.
20. A non-transitory computer-readable medium having stored thereon, computerexecutable instructions that when executed by an electronic device, causes the electronic device to execute operations, the operations comprising: capturing, via a set of image-capture devices, visual information of an environment associated with a user;determining first information indicative of at least one of an object or an individual in the environment, based on the captured visual information; receiving sensor data from a set of biosensors associated with the user; determining a subvocalized command of the user, based on the received sensor data; applying a neural language model on the determined first information and the determined subvocalized command; determining, based on the application of the neural language model, second information corresponding to at least one of an audio feedback or a haptic feedback for the user; and rendering the determined second information to the user, wherein the second information includes instructions to guide the user to perform a task associated with the determined subvocalized command.
Citation Information
Patent Citations
Audio navigation assistance
US20150211858A1
Intelligent trip prediction in autonomous vehicles
US20190186939A1
Conversational artificial intelligence system in a virtual reality space
US20230108256A1
Artificial intelligence assisted wearable
US20230267299A1
Using pattern analysis to provide continuous authentication
US20240073219A1