Low power always online listening artificial intelligence (AI) system
The dual-layer voice monitoring system, consisting of a low-power always-on listening module and a high-power response module, solves the energy consumption and context awareness issues of voice-activated systems during continuous monitoring, achieving more efficient battery use and more accurate user responses.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-04-07
AI Technical Summary
Existing voice activation systems consume a lot of energy when continuously monitoring user voice, which can easily deplete device battery resources. They also lack context awareness, leading to a decline in user experience.
A low-power always-on listening module (LPALM) is used to continuously monitor speech. Combined with a wake-up module and a high-power response module (HPRM), the high-power mode is activated only when context-awareness is required by analyzing verbal and non-verbal cues, thus realizing a two-layer speech monitoring system that balances performance and power consumption.
It effectively extends device battery life, improves context awareness, provides more accurate and immediate responses, reduces unnecessary high-power activation, and enhances the user experience.
Smart Images

Figure CN121816611A_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims priority to U.S. Nonprovisional Application No. 18 / 468,964, filed September 18, 2023, the entire contents of which are incorporated herein by reference. Background Technology
[0003] Typically, voice-activated computing systems interact with users through voice commands or queries. These systems rely on a combination of microphone hardware and natural language processing software to capture and interpret human speech. Upon receiving specific voice prompts or keywords, a voice-activated system can transition from a passive listening state to an active state, in which it can perform tasks, provide information, or perform various functions. Typically, such systems facilitate user interaction with devices or software applications by replacing or supplementing manual input methods such as typing or clicking. The capabilities of voice-activated computing systems range from simple task execution to more complex conversational interactions.
[0004] Conventional voice activation systems can initiate interaction with the user based on predefined voice prompts or keywords. Unlike continuous voice monitoring systems, these systems remain in a passive listening state until activated by specific vocal cues (such as "Hey Siri" or "Ok Google"). In response to the detected cue, the system can transition from its low-power, idle state to a more engaged mode, in which it can receive and process additional commands or queries. Typically, such systems are not context-aware and focus solely on performing specific tasks or answering questions based on immediate instructions or direct queries. Summary of the Invention
[0005] Various aspects include methods and processing systems for continuously or repeatedly monitoring speech to obtain preemptive or context-aware responses to user queries. In some aspects, the method may include: collecting ambient audio data in a low-power always-on listening mode operating on a processing system in a computing device, and storing the collected ambient audio data and an activation timestamp in an audio buffer as buffered audio data; determining a confidence score based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues; and activating a high-power listening mode in response to determining that the confidence score exceeds a threshold, and providing the final portion of the audio buffer to the high-power mode for immediate context-aware assistance.
[0006] In some aspects, determining a confidence score based on the results of analyzing buffered audio data for linguistic or non-linguistic cues can include determining a confidence score based on the results of analyzing buffered audio data for linguistic cues, including grammatical cues, semantic cues, sub-word cues, contextual cues, and co-occurrence cues. In other aspects, determining a confidence score based on the results of analyzing buffered audio data for linguistic or non-linguistic cues can include determining a confidence score based on the results of analyzing buffered audio data for non-linguistic cues, including prosodic cues, pitch cues, speech rate cues, volume cues, temporal pattern cues, and acoustic feature cues.
[0007] Some aspects may also include using a trained language model to identify linguistic or non-linguistic cues. Some aspects may also include switching to a periodic training mode in response to determining that the computing device can be connected to a stable power source. Some aspects may also include: retrieving buffered audio data and activation timestamps from an audio buffer; using the results of applying the retrieved audio data to a Large Language Model (LLM) to identify instances where activation of a high-power listening mode should have occurred but did not; labeling the identified instances and their corresponding timestamps; generating an updated machine learning model based on the labels for a low-power always-on listening mode; and replacing the current machine learning model for the low-power listening mode with the generated updated machine learning model.
[0008] Another aspect may include a computing device having a processing system configured with processor-executable instructions for performing various operations corresponding to the methods described above. Another aspect may include a non-transitory processor-readable storage medium having processor-executable instructions stored thereon, these instructions being configured to cause the processing system to perform various operations corresponding to the operations of the methods described above. Another aspect may include a computing device having various components for performing functions corresponding to the operations of the methods described above. Attached Figure Description
[0009] The accompanying drawings, which are incorporated herein and form part of this specification, illustrate exemplary embodiments of the claims and, together with the general description and detailed description given, serve to explain the features of this document.
[0010] Figure 1 This is a component block diagram illustrating example components that can be included in a computing device and configured to implement some implementation scheme in a system-in-package (SIP).
[0011] Figure 2 It is a component block diagram illustrating example components and operations in a system configured to implement some implementation schemes.
[0012] Figures 3 to 6This is a flowchart illustrating a method for implementing or operating a continuous voice monitoring artificial intelligence (AI) system according to some implementation schemes, which continuously listens to the user and analyzes their voice to obtain contextual cues and / or proactively initiates actions or generates responses.
[0013] Figure 7 This is a component block diagram illustrating an example computing device in the form of a laptop computer suitable for implementing some implementation schemes.
[0014] Figure 8 This is a component block diagram illustrating an example wireless communication device suitable for use with various implementation schemes. Detailed Implementation
[0015] Various embodiments will be described in detail with reference to the accompanying drawings. Where possible, the same reference numerals will be used throughout the drawings to refer to the same or similar parts. References to specific examples and embodiments are for illustrative purposes and are not intended to limit the scope of the claims.
[0016] Various implementations include methods and devices for providing a low-power always-on listening mode for AI-type assistant functions, such as smartphones, tablets, and personal computers, which identify when it is appropriate to activate the AI assistant without requiring the user to utter a specific trigger word or phrase. These implementations enable the use of machine learning to train the low-power always-on listening mode functionality over time to recognize when a user is speaking or in a manner that suggests the user can benefit from the AI assistant. Upon recognizing such a situation, the low-power always-on listening mode functionality can activate a high-power AI assistant mode, which then listens for and responds to the user. In some implementations, the low-power always-on listening mode uses machine learning to learn to activate the high-power AI assistant mode based on the context of the spoken words, plus the tone, rhythm, and other features of the user's speech.
[0017] In some implementations, the always-on listening mode functionality may include a memory buffer that continuously records sound (at least in response to the user's voice). This buffer may be accessible to a high-power AI assistant, and / or a portion of the words the user has already spoken (or a buffer pointer) stored in the buffer may be provided as part of activating the high-power AI assistant mode. In this case, the AI assistant receives data about the words and context that led to the activation, enabling the AI assistant to immediately provide assistance to the user without requiring the user to repeat the activation question or statement.
[0018] In some implementations, the low-power always-on listening mode functionality can perform periodic training, such as when a computing device is connected to power. During this training, the high-power AI assistant can process audio data in the always-on listening buffer to identify when the AI assistant should be activated (i.e., the user may have already used the assistance and must speak a keyword to trigger activation) but is not activated, and to identify when the AI assistant is unnecessarily activated (i.e., the user does not need or want the assistance). Since the AI assistant can be a trained Large Language Model (LLM) AI system, these instances can be identified based on the user's voice recorded in the buffer and the dialogue between the AI assistant and the user after activation. For each instance of missed or unnecessary activation, the AI assistant can provide a timestamped, determined conclusion to the corresponding sound in the buffer in a manner that enables machine learning of the context through the low-power always-on listening functionality.
[0019] The term "computing device" in this document may refer to any or all of the following: personal computer, laptop computer, tablet computer, user equipment (UE), smartphone, personal or mobile multimedia player, personal data assistant (PDA), handheld computer, wireless email receiver, cellular phone with multimedia internet support, gaming system (e.g., PlayStation). ™ Xbox ™ Nintendo Switch ™ Wearable devices (e.g., headphones, smartwatches, head-mounted displays, fitness trackers, etc.), media players (e.g., DVD players, ROKU...), etc. ™ Apple TV ™ Digital video recorders (DVRs), automotive displays, portable projectors, 3D holographic displays, and other similar devices including displays and programmable processing systems that can be configured to provide functionality for various implementation schemes.
[0020] The term "processing system" is used herein to refer to one or more processors, including multi-core processors, that are organized and configured to perform various computational functions. Various implementation methods can be implemented in one or more of the multiple processors within a processing system as described herein.
[0021] The term "System-on-a-Chip" (SoC) is used herein to refer to a single integrated circuit (IC) chip containing multiple resources or independent processors integrated on a single substrate. A single SoC may contain circuitry for digital, analog, mixed-signal, and radio frequency functions. A single SoC may include a processing system comprising any number of general-purpose or specialized processors (e.g., network processors, digital signal processors, modem processors, video processors, etc.), blocks of memory (e.g., ROM, RAM, flash memory, etc.), and resources (e.g., timers, voltage regulators, oscillators, etc.). For example, an SoC may include an application processor operating as the SoC's main processor, central processing unit (CPU), microprocessor unit (MPU), arithmetic logic unit (ALU), etc. An SoC may also include software for controlling the integrated resources and processors, as well as software for controlling peripheral devices.
[0022] The term "System-in-Package" (SIP) may be used herein to refer to a single module or package that contains multiple resources, computing units, cores, or processors on two or more IC chips, substrates, or SoCs. For example, a SIP may comprise a single substrate on which multiple IC chips or semiconductor dies are stacked in a vertical configuration. Similarly, a SIP may comprise one or more multi-chip modules (MCMs) on which multiple ICs or semiconductor dies are packaged into a single substrate. A SIP may also comprise multiple independent SoCs coupled together and packaged adjacently (e.g., on a single motherboard, in a single UE, or in a single CPU device) via high-speed communication circuitry. The proximity of SoCs facilitates high-speed communication and the sharing of memory and resources.
[0023] Voice-activated AI systems (such as Siri, Google Assistant, etc.) convert spoken language into text through a process called speech recognition. After interpreting the text, these systems match queries against a set of predefined algorithms or processes to provide an appropriate response or perform a specified task. These systems can also use machine learning models to interpret natural language and provide more accurate and / or more relevant answers to user queries. Many of these systems also include the ability to interact with other software and hardware components to allow users to control various aspects of their digital environment via verbal commands (e.g., turning on lights, televisions, etc.). While these systems have become increasingly adept at understanding and responding to a wide range of queries, they face numerous technical challenges and limitations (e.g., interpreting context, understanding accents, dialects, or complex requests).
[0024] Some implementations may include computing devices equipped with a continuous speech monitoring AI system that continuously or permanently listens to and analyzes human speech and / or other sensor data to preemptively identify user queries. Unlike conventional voice-activated systems that wait for specific prompts to initiate interaction, continuous speech monitoring AI systems can maintain continuous auditory observation of the user (e.g., they are constantly listening and analyzing) to better understand the user's context, emotional tone, and immediate needs. Continuous speech monitoring AI systems can also implement and / or use advanced natural language processing (NLP) algorithms to interpret spoken words, phrases, or sentences to obtain more accurate and / or context-aware responses, assistance, actions, or control.
[0025] Continuous voice monitoring AI systems face various technical challenges due to the high energy consumption associated with continuous monitoring using advanced LLM AI capabilities. For example, LLM AI systems may require significant processing and / or energy resources to continuously listen to users and analyze their speech for contextual cues and / or proactively initiate actions or generate responses. This can rapidly deplete the typically limited battery resources of user devices, rendering them unresponsive and / or otherwise degrading the user experience. Continuous voice monitoring AI systems can be configured to overcome these and other technical challenges.
[0026] In various implementations, a continuous voice monitoring AI system may include a low-power always-on listening module (LPALM) configured to maintain continuous auditory awareness or alertness without consuming excessive processing, memory, or battery resources of the user's computing system or device. In this way, the LPALM can operate on the computing device for extended periods without depleting the device's battery resources, causing the user device to become unresponsive, or otherwise negatively or perceptibly affecting the performance, functionality, or power consumption characteristics of the user device.
[0027] In some implementations, the computing system can be configured to implement a two-layer continuous speech monitoring system that provides context-sensitive, always-on listening AI capabilities while balancing various trade-offs between performance, functionality, and power consumption. The two-layer continuous speech monitoring system can be configured to operate primarily in a low-power mode and switch to a higher-power mode that provides more robust functionality only when contextual cues indicate that the user can benefit from immediate or more robust AI assistance.
[0028] In some implementations, a two-layer continuous speech monitoring system may include an LPALM, a wake-up module, a high-power response module (HPRM), and / or a periodic training module (PTM). The LPALM can operate in a low-power state or mode (e.g., a low-power always-on listening mode) to perform basic listening operations such as speaker identification, logging detected speech or speech features in a log, and / or scanning to ensure activation of more comprehensive sound or environmental triggers for listening and processing. For example, in some implementations, the LPALM may be configured to capture ambient sound, analyze the captured sound to obtain potential speech or environmental triggers, generate metadata based on the analysis results, store the metadata in a log, and / or transmit the metadata to the wake-up module.
[0029] The wake-up module can be configured to monitor information collected by LPALM triggers and activate or invoke HPRM for deeper interaction and analysis in response to the detection of a sound or environmental trigger that ensures more comprehensive listening and processing. For example, the wake-up module can be configured to receive metadata from LPALM, evaluate the metadata and compare it to a set of predefined triggers, use data collected from additional modules (e.g., visual, motion, etc.) and / or multiple sources (e.g., a combination of user feedback data and Large Language Model (LLM) feedback data) to determine or calculate a confidence or probability value for an immediate action or response, and compare the confidence or probability value to a threshold. The wake-up module can perform a "quick action" in response to determining that the confidence or probability value exceeds the threshold. Alternatively, the wake-up module can activate HPRM in response to determining that the confidence or probability value does not exceed the threshold.
[0030] HPRM can be a more computationally intensive module, using more power / energy but offering better performance or more robust functionality (e.g., full speech recognition, complex natural language processing, etc.). HPRM can be configured to operate in standby mode until it receives an activation signal (e.g., a specific message, flag, etc.) from a wake-up module. In response to receiving the activation signal, HPRM can transition from standby to active mode, allocate or prioritize computational and memory resources for its operation, record and analyze deep speech data, perform advanced analysis using natural language processing (NLP) algorithms, linguistic cues, and / or non-linguistic cues, generate LLM input queries (e.g., single input values or strings for LLM components) using the results of the advanced analysis, and pass the LLM input queries to the LLM components. In some implementations, the LPALM, wake-up module, HPRM, and / or other components in the system can receive and use output from the LLM components to generate fine-grained and comprehensive output for the user and / or as feedback for refining or fine-tuning the operation of components in the system. In some implementations, the computing system can update supervisory data and / or adaptively tune the wake-up module and / or HPRM based on output, results, and user feedback (e.g., for future interactions). While the HPRM is operating in an active state, the computing system can also monitor battery and resource usage. The computing system can re-enter standby mode in response to determining that available resources are below a threshold level.
[0031] In some implementations, the wake-up module may include, generate, or use multi-source data technologies and / or supervised data including user feedback data and LLM feedback data. Each type of data or feedback data can provide unique, specific, exact, or different insights. The wake-up module can merge these datasets into a centralized supervised dataset and use this centralized supervised dataset to dynamically and intelligently determine whether to activate HPRM.
[0032] In some implementations, the wake-up module may include dynamic decision-making capabilities. For example, the wake-up module may use supervised data to calculate a confidence or probability value that represents the likelihood that an expected immediate action, a suggested response, is correct, and / or that the action or response will be acceptable to the user. In some implementations, the wake-up module may be configured to perform a “quick” action or reaction in response to determining that the confidence or probability value exceeds a threshold, which avoids activation of high-power modules. By avoiding activation of high-power modules, the wake-up module can improve the performance and power consumption characteristics of the computing system.
[0033] In some implementations, the continuous speech monitoring system may include or use speech recognition to add a layer of specificity. For example, the system may be configured such that LPALM is activated only for recognized speech, which can enhance system security and / or context relevance by eliminating unwanted activations.
[0034] In some implementations, the waker module can be configured to determine whether to activate the HPRM based on linguistic cues and / or non-linguistic cues. Examples of linguistic cues that can be generated or determined by the system include grammatical cues (e.g., word and phrase arrangement, etc.), semantic cues (e.g., the meaning of individual words, phrases, or sentences, etc.), sub-word cues (e.g., morphemes and other units smaller than words, etc.), contextual cues (e.g., surrounding words, etc.), and co-occurrence cues (e.g., words that frequently appear together in a given context, such as "bread and butter," etc.). Examples of non-linguistic cues that can be generated or used by the system include prosodic cues (e.g., rhythm, stress patterns, and intonation of speech, etc.), pitch cues (e.g., variations in speech frequency, etc.), speech rate cues, volume cues (e.g., loudness or softness of speech, etc.), temporal pattern cues (e.g., pauses or hesitations in speech, etc.), and acoustic feature cues (characteristics used as identifiers of different speakers, emotional states, etc.).
[0035] In some implementations, the wake-up module can be configured to implement a multi-layered and / or multi-modal solution that activates the LPALM and / or HPRM based on any combination of information collected on the device, such as audio, visual, motion, etc. By collecting and using information from multiple modules, the wake-up module can better identify, analyze, or interpret the often complex layers of human interaction and environmental context. This, in turn, allows the wake-up module to make better or smarter decisions about whether or when to activate the LPALM and / or HPRM or perform a quick action.
[0036] In some implementations, HPRM can be configured to use linguistic and non-linguistic cues to generate a single input sequence for LLM to obtain a more refined and / or comprehensive output that better matches the user's request.
[0037] In some implementations, the computing system can be configured to initialize or load the LPALM, HPRM, audio buffer, external memory storage, wake-up criteria (e.g., phrase, tone), and privacy settings (e.g., complete speech log and features). The LPALM can continuously listen to audio, record detected audio in the audio buffer, analyze incoming audio to identify the registered speaker, examine or analyze context indicators of the audio (e.g., tone, rhythm), determine if the context indicators match the wake-up criteria, transmit a wake-up signal to the HPRM in response to determining that the indicators match the wake-up criteria (e.g., via a wake-up module, etc.), transfer the audio buffer to external memory when capacity is approaching, apply privacy settings (e.g., feature extraction, etc.), and / or invoke the PTM in response to determining that the system is connected to a power source. The HPRM can be configured to activate upon receiving a wake-up signal from the LPALM, access the audio buffer to retrieve the stored audio that caused the wake-up signal, perform complete speech recognition on the retrieved audio to establish a context, process the context and incoming audio through an NLP engine, and generate appropriate responses or actions. HPRM can return to standby or low-power state and signal LPALM to resume its operation. PTM can be configured to review logs stored in external memory, identify points where LPALM missed or erroneously activated, update the wake-up criteria based on this analysis, fine-tune the algorithm or model used by LPALM using the updated wake-up criteria and full speech recognition, and store the updated algorithms and criteria for future use.
[0038] In some implementations, the computing system can be configured to include a low-power always-on listening mode (corresponding to LPALM), a high-power listening mode (corresponding to HPRM), and machine learning models for both the low-power and high-power modes.
[0039] In some implementations, the low-power always-on listening mode can be intelligent because it wakes up a higher-power AI system (e.g., HPRM) based on the context of the spoken words (plus tone, rhythm, etc.) rather than just key phrases or words. When operating in low-power always-on listening mode, the computing device can continuously capture audio data and store it in an audio buffer, which can be a scrolling buffer of a predefined size. Similarly, while in low-power always-on listening mode, the computing device can analyze the buffered audio data to obtain the spoken words, tone, rhythm, etc., use a trained language model to recognize the context, tone, rhythm, etc., determine or calculate a confidence score / value based on the model's output, and compare the confidence score / value to a predefined threshold. In response to determining that the confidence score / value exceeds the predefined threshold, the computing system can activate a high-power listening mode and provide the last portion of the audio buffer to the high-power mode for immediate context-aware assistance.
[0040] In some implementations, the low-power always-on listening mode may include a buffer (i.e., an always-on recording buffer, etc.) and, when the high-power listening mode or the higher-power AI system is activated, provides a portion of the spoken words stored in the buffer, so that the higher-power mode / system is aware of the words and context that led to the activation and can therefore be activated immediately (instead of asking "How can I help?" etc.).
[0041] When operating in high-power listening mode, the computing device can receive buffered audio data, perform advanced natural language processing on the received buffer and any additional spoken words, and generate responses or perform actions based on the processed audio data.
[0042] In some implementations, the computing system may be configured to include a periodic training mode (corresponding to PTM). In some implementations, the computing system may be configured to enter the periodic training mode in response to determining that the device is connected to a stable power source (e.g., not operating from battery, etc.). When operating in the periodic training mode, the computing device may retrieve stored audio data and activation timestamps from an audio buffer, apply the collected audio data to a high-power mode LLM, identify instances where high-power listening mode activation should have occurred but did not (and / or vice versa, or when it was incorrectly activated, etc.), tag the identified instances and their corresponding timestamps, update the machine learning model used for low-power listening functionality based on the tags, and replace the current low-power listening model with the newly trained model to improve context awareness.
[0043] The implementation scheme can provide technical solutions to overcome various technical challenges faced by existing and conventional AI systems. For example, the implementation scheme can provide context-sensitive, always-on listening AI capabilities while balancing trade-offs between the performance and power consumption characteristics of the user's computing device. The implementation scheme can maintain continuous auditory perception or alertness without consuming excessive processing, memory, or battery resources of the computing system. Furthermore, unlike conventional solutions that rely on specific "wake words" to activate the AI system, the implementation scheme can activate a robust AI system based on a fine set of criteria, such as the context of the spoken words, the tone and / or rhythm of certain phrases or intonations indicating that the user is seeking interaction with the AI system. Additionally, some implementation schemes may include options for storing the complete voice log or only extracting specific features based on the user's data privacy settings, and the stored data can be transferred to external storage for future use or analysis when the internal buffer reaches its capacity. When connected to a power source, the HPRM can perform a review of the stored voice logs and adjust its activation criteria based on that data (which allows the system to learn and improve rapidly over time). Furthermore, the system can perform supervised training, and feedback from HPRM (which can perform full speech recognition and semantic analysis) can be used to fine-tune LPALM's AI model and / or decision-making algorithm.
[0044] Various implementation schemes can be implemented on multiple single-processor and multi-processor computer systems, including system-on-a-chip (SOC) or system-in-package (SIP) systems. Figure 1 Example computing systems or SIP 100 architectures that can be used in mobile computing devices implementing continuous voice monitoring AI systems according to various implementation schemes are illustrated.
[0045] refer to Figure 1 The illustrated example SIP 100 includes two SOCs 102 and 104, a clock 106, a voltage regulator 108, and a wireless transceiver 166. The first SOC 102 and the second SOC 104 can communicate via an interconnect bus 150. Various processors 110, 112, 114, 116, 118, 121, and 122 can be interconnected with each other and to one or more memory elements 120, system components and resources 124, and a thermal management unit 132 via an interconnect bus 126, which may include advanced interconnects such as high-performance network-on-chip (NOC). Similarly, processor 152 can be interconnected to a power management unit 154, a millimeter-wave transceiver 156, memory 158, and various additional processors 160 via an interconnect bus 164. These interconnect buses 126, 150, and 164 may include arrays of reconfigurable logic gates and / or implement bus architectures (e.g., CoreConnect, AMBA, etc.). Communication can be provided by advanced interconnect components such as NOC.
[0046] In various implementation schemes, any or all of the processors 110, 112, 114, 116, 121, and 122 in the system can operate as the main processor, central processing unit (CPU), microprocessor unit (MPU), arithmetic logic unit (ALU), etc. of the SoC. One or more of the coprocessors 118 can operate as the CPU.
[0047] In some implementations, the first SOC 102 may operate as a central processing unit (CPU) of a mobile computing device, which executes instructions by performing arithmetic, logic, control, and input / output (I / O) operations specified by instructions from software applications. In some implementations, the second SOC 104 may operate as a dedicated processing unit. For example, the second SOC 104 may operate as a dedicated 5G processing unit responsible for managing high-capacity, high-speed (e.g., 5Gbps) and / or ultra-high frequency short-wavelength (e.g., 28GHz millimeter-wave spectrum) communications.
[0048] The first SOC 102 may include a digital signal processor (DSP) 110, a modem processor 112, a graphics processor 114, an application processor 116, one or more coprocessors 118 (e.g., vector coprocessors, CPUCP, etc.) connected to one or more of these processors, memory 120, a deep processing unit (DPU) 121, an artificial intelligence processor 122, system components and resources 124, an interconnect bus 126, one or more temperature sensors 130, a thermal management unit 132, and a thermal power envelope (TPE) component 134. The second SOC 104 may include a 5G modem processor 152, a power management unit 154, an interconnect bus 164, multiple millimeter-wave transceivers 156, memory 158, and various additional processors 160, such as application processors, packet processors, etc.
[0049] Each processor 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, 160 may include one or more cores, and each processor / core may perform operations independently of other processors / cores. For example, the first SOC 102 may include a processor running a first type of operating system (e.g., FreeBSD, LINUX, OS X, etc.) and a processor running a second type of operating system (e.g., MICROSOFT WINDOWS 11). Additionally, any or all of processors 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, 160 may be included as part of a processor cluster architecture (e.g., synchronous processor cluster architecture, asynchronous or heterogeneous processor cluster architecture, etc.).
[0050] Processors 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, and 160, or any or all of them, can operate as the CPU of a mobile computing device. Additionally, processors 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, and 160, or any or all of them, can be included as one or more nodes in one or more CPU clusters. A CPU cluster can be a group of interconnected nodes (e.g., processing cores, processors, SOCs, SIPs, computing devices, etc.) configured to work in a coordinated manner to perform computational tasks. Each node can run its own operating system and contains its own CPU, memory, and storage devices. Tasks assigned to the CPU cluster can be divided into smaller tasks, which are distributed across the nodes for processing. Nodes can work together to complete a task, with each node handling a portion of the computation. The results of the computations from each node can be combined to produce a final result. CPU clusters are particularly useful for tasks that can be parallelized and executed concurrently. This allows CPU clusters to complete tasks much faster than a single high-performance computer. Furthermore, because CPU clusters consist of multiple nodes, they are generally more reliable and less prone to failure than a single high-performance component.
[0051] The first SOC 102 and the second SOC 104 may include various system components, resources, and custom circuitry for managing sensor data, analog-to-digital conversion, wireless data transmission, and performing other specialized operations such as decoding data packets and processing encoded audio and video signals for presentation in a web browser. For example, the system components and resources 124 of the first SOC 102 may include power amplifiers, voltage regulators, oscillators, phase-locked loops, peripheral bridges, data controllers, memory controllers, system controllers, access ports, timers, and other similar components for supporting processors and software clients running on computing devices. The system components and resources 124 may also include circuitry for interfacing with peripheral devices such as cameras, electronic displays, wireless communication devices, external memory chips, etc.
[0052] The first SOC 102 and / or the second SOC 104 may also include input / output modules (not illustrated) for communicating with external resources such as clock 106, voltage regulator 108, and wireless transceiver 166 (e.g., cellular wireless transceiver, Bluetooth transceiver, etc.). External resources (e.g., clock 106, voltage regulator 108, wireless transceiver 166) may be shared by two or more internal SOC processors / cores.
[0053] In addition to the example SIP 100 discussed above, various implementations can be implemented in a wide variety of computing systems, which may include a single processor, multiple processors, multi-core processors, or any combination thereof.
[0054] Figure 2 Example components that may be included in a system configured to implement a continuous voice monitoring AI system, according to the implementation scheme, are illustrated. References Figure 1 and Figure 2 System 200 (e.g., SIP 100, SOC 102, 104, etc.) may include a multi-layer continuous speech monitoring subsystem 201, which includes a periodic training module (PTM) 202, a low-power always-on listening module (LPALM) 204, a wake-up module 206, a high-power response module (HPRM) 208, an audio buffer 210, a wake-up criteria (e.g., phrase, intonation) component 212, and a privacy settings (e.g., complete speech log and features) component 214. In some embodiments, system 200 may also include an external memory storage device 220 and / or a large language model (LLM) component 240. In some embodiments, the audio buffer may include low-power (LP) speech logs and high-power (HP) speech logs.
[0055] System 200 may activate PTM 202 and / or enter a periodic training mode in response to determining that the device is connected to a stable power source (e.g., not operating from battery, etc.). In some implementations, LPALM 204 may activate or invoke PTM 202 in response to determining that System 200 is connected to a power source. PTM 202 may be configured to receive or retrieve user feedback and LLM feedback, retrieve stored audio data and activation timestamps from audio buffer 210, apply the collected audio data to HPRM 208, identify instances where HPRM activation should have occurred but did not (and / or vice versa, or when it was incorrectly activated, etc.), tag the identified instances and their corresponding timestamps, update the machine learning model of LPALM 204 based on the tags, and replace the current low-power listening machine learning model with the newly trained model to improve context awareness.
[0056] LPALM 204 can be configured to maintain a continuous state of auditory perception or alertness without placing an excessive burden on system resources, such as processing, memory, or battery resources. LPALM 204 can achieve this by operating primarily in a low-power state for basic sensory operations, such as hearing. These operations may include capturing ambient sounds, analyzing the captured ambient sounds to obtain potential sound or environmental triggers, and generating metadata based on the analysis results. LPALM 204 can store the metadata locally or transmit the metadata to a wake-up module 206, which can use the metadata in conjunction with other contextual information, such as visual and motion data, to intelligently determine whether to activate the HPRM.
[0057] LPALM 204 can continuously capture audio and store it in an audio buffer, which can be a predefined-sized scrolling buffer or a sliding window. LPALM 204 can use a machine learning model to understand the context, pitch, and rhythm of spoken words. LPALM 204 can determine or calculate a confidence score based on the model's output. LPALM 204 or waker module 206 can activate HPRM 208 in response to determining that the confidence score exceeds a predefined threshold. Activating HPRM 208 can cause system 201 to transition to a higher-power operating state equipped to provide immediate context-aware assistance. Furthermore, to further improve performance and resource utilization, LPALM 204 can move its audio buffer 210 to external memory storage device 220 when approaching capacity, and / or invoke PTM 202 when a stable power source is detected.
[0058] In some implementations, LPALM 204 can be configured to store data or activate HPRM by identifying registered speakers and based on whether the current speaker is an identified registered speaker, thereby providing additional security and context relevance and / or reducing unnecessary system activations. In some implementations, LPALM 204 can be integrated with PTM 202 to improve its machine learning model. PTM 202 can be triggered when the device is connected to a stable power source, thereby ensuring that the machine learning model of LPALM 204 is updated periodically without negatively impacting the device's performance or power consumption characteristics.
[0059] The waker module 206 can be configured to act as an intermediary agent in communication with LPALM 204 and HPRM 208. The waker module 206 can evaluate the data and metadata collected by LPALM 204 to determine whether to invoke the more resource-intensive HPRM 208 for AI or complex auditory processing tasks. In response to receiving metadata from LPALM 204, the waker module 206 can perform an evaluation or comparison of the metadata with a set of predefined sound or environmental triggers. As part of these operations, the waker module 206 can also collect, use, or merge additional data streams, such as visual or motion information. This multi-layered approach enhances the waker module 206's ability to evaluate the complexity of human communication, human interaction, and environmental context.
[0060] In some implementations, the wake-up module 206 can be configured to determine its action using confidence or probability values calculated based on various metrics. These values can be compared to thresholds to determine whether immediate activation of HPRM 208 or some other action is warranted. The wake-up module 206 can perform a shortcut action in response to determining that the confidence or probability value exceeds a predetermined threshold, effectively circumventing the need to activate HPRM 208.
[0061] In some implementations, the waker module 206 can be configured to use a centralized supervised dataset to merge user feedback and LLM feedback, and to refine its decision-making process over time. This can also allow for more granular activation criteria, including linguistic cues (e.g., grammatical, semantic, contextual, etc.) and non-linguistic cues (e.g., pitch, prosody, acoustic features, etc.). The waker module 206 can be configured to intelligently determine whether to activate LPALM 204 or HPRM 208 (or perform a shortcut action) based on multimodal input and complex contextual considerations, balancing a trade-off between accuracy and effectiveness.
[0062] The HPRM 208 can be configured to perform more complex processing operations, such as NLP and full speech recognition. The HPRM 208 can operate in standby mode and await an activation signal from the wake-up module 206. Upon receiving the activation signal, the HPRM 208 transitions from standby to active mode and allocates computational and memory resources for its tasks. The HPRM 208 can record and analyze deep speech data, employing a complex mixture of linguistic and non-linguistic cues. The HPRM 208 can use advanced NLP algorithms to generate single input values or strings for the LLM 240. The HPRM 208 can further refine the output from the LLM 240 to provide a more detailed and user-centric response to the user. This multi-level approach helps ensure that the output presented to the user is consistent with the user's needs and the broader context of interaction.
[0063] In some implementations, the HPRM 208 can be configured to monitor its resource usage, determine whether the current resource usage or availability in the computing device has fallen below a specified threshold, and re-enter standby mode to conserve power in response to determining that resource usage or availability has fallen below the threshold. In some implementations, the HPRM 208 can be configured to review logs and refine its activation criteria when connected to a power source, which can facilitate rapid machine learning and improvement over time.
[0064] In some implementations, HPRM 208 can be configured to fine-tune the performance of LPALM 24 using stored algorithms and standards. For example, HPRM 108 can review logs to identify instances where LPALM 204 or waker module 206 may have missed activating or incorrectly activated HPRM 108, and update wake-up standard 212 accordingly.
[0065] Figures 3 to 6 These are flowcharts illustrating methods 300, 400, 500, and 600 for implementing or operating a continuous voice monitoring AI system according to some implementation schemes. This continuous voice monitoring AI system continuously listens to the user and analyzes their speech to obtain contextual cues and / or proactively initiates actions or generates responses. (Reference) Figures 1 to 6Methods 300, 400, 500, and / or 600 may be executed in a computing device by a processing system including one or more processors (e.g., 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, 160, etc.), components, or subsystems discussed herein. Components for performing the operations in methods 300, 400, 500, and / or 600 may include the processing system, which includes one or more of processors 110, 112, 114, 116, 118, 121, 122, 121, 122, 152, 160, and other components described herein. Furthermore, one or more processors of the processing system may be configured with software or firmware to perform some or all of the operations in methods 300, 400, 500, and / or 600. To cover alternative configurations implemented in the various implementation schemes, the hardware implementing any or all of methods 300, 400, 500 and / or 600 is referred to herein as the “processing system”.
[0066] refer to Figures 1 to 6 In box 302, the processing system can load machine learning models into the LPALM and HPRM to facilitate continuous speech analysis and response actions. Loading machine learning models into the LPALM allows it to perform basic speech recognition tasks with reduced or minimal energy consumption. Loading machine learning models into the HPRM allows it to perform more complex, resource-intensive actions when activated. Examples of machine learning models that can be loaded into the LPALM are decision trees or lightweight neural networks designed to efficiently recognize words or phrases. The HPRM can be configured with more complex models, such as Long Short-Term Memory (LSTM) networks or Transformer models, which are better suited for understanding context, detecting emotions in user speech, etc. These models contribute to the implementation or operation of continuous speech monitoring systems by providing the necessary capabilities for real-time analysis of audio input. The LPALM can use its machine learning models to sift through audio with reduced or minimal energy consumption, passing relevant data to the HPRM, which can then use its more complex models to make better or more refined interpretations or decisions. In some implementations, this hierarchical processing arrangement can allow the system to make context-aware decisions and perform tasks more effectively, such as identifying user needs, suggesting actions, or initiating workflows based on the analyzed speech.
[0067] In box 304, the processing system can set initial privacy settings and wake-up criteria. The processing system can use the initial privacy settings to determine how to manage the captured audio data and / or allow the multi-layer continuous speech monitoring system to balance user data privacy with the use of personal or private data to perform its functions. For example, the system can default to storing only specific extracted features instead of the complete speech log, thereby reducing the amount of personal data stored and processed. These settings can be particularly beneficial in systems where privacy is a strong concern. Setting wake-up criteria (e.g., specific phrases, tone quality, rhythm, etc.) can allow the system to better guide LPALM, reduce false alarms, and better identify relevant information that triggers HPRM for further analysis and action. Wake-up criteria can be granular and include contextual cues that provide improvements over traditional systems that typically rely on specific “wake words.” For example, enhanced wake-up criteria can allow the system to maintain context awareness while using less computational and energy resources.
[0068] In some implementations, the PTM and / or HPRM can fine-tune privacy settings and wake-up criteria over time. For example, the HPRM can analyze logs to identify situations where the LPALM may have failed to activate or erroneously activated, and modify the wake-up criteria accordingly to enhance system performance. These adjustments can allow for more dynamic and context-aware systems and improve upon some limitations of conventional solutions and systems, such as rigidity of activation cues or high computational loads that may have negative or user-perceptible impacts on device performance or user privacy.
[0069] In determination box 306, the processing system determines whether a stable power source is available. The processing system can use any of a variety of known techniques to determine the availability of a stable power source. Hardware-based sensors and software algorithms typically work together to detect the current power state of a device. For example, voltage levels and current can be monitored to assess whether the device is connected to a power outlet, rather than relying on battery power. The system can also use application programming interface (API) calls to the operating system to query the power state, which can return information about whether the device is operating at AC power and, if so, information about the stability of the connection.
[0070] In response to determining the presence of a stable power source (i.e., determining box 306 = "Yes"), in box 308, the processing system can activate the Periodic Training Module (PTM) and / or begin operating in Periodic Training Mode. As described above, the PTM can be configured to perform various resource-intensive tasks, such as processing stored audio data and activation timestamps, updating machine learning models, receiving or retrieving feedback, applying collected audio data to the HPRM, etc. These operations may require higher levels of power, and performing such tasks while operating solely on battery power can quickly deplete the device's battery life. Therefore, the system can be configured to trigger the PTM only when the device is connected to a stable power source. This allows the system to perform processor-intensive or power-intensive operations without excessively consuming the device's resources.
[0071] In box 310, the processing system can retrieve audio data and activation timestamps from an audio buffer. The audio buffer can be a temporary storage device where incoming audio data and its corresponding activation timestamps, important for both real-time and post-event analysis, are stored for a specified time period or until processed. Furthermore, the processing system can participate in periodic training activities. During these training cycles, the processing system can refine the machine learning model using the audio data and activation timestamps from the audio buffer. The timestamps can serve as a temporal guide to align audio data with specific events, thus providing a contextual basis for analyzing the model's performance. Additionally, by examining the retrieved audio data and timestamps, the processing system can better determine whether the LPALM or HPRM failed to activate when it should have and / or was activated when it shouldn't have. The processing system can use such observations during training for a relabeling process to further enhance the model's performance in context recognition and responsiveness. Moreover, in real-time operation, audio data and activation timestamps may be required for on-the-fly processing. For example, upon receiving a wake-up signal, the HPRM can access the audio buffer to obtain the audio that led to the activation event. This data may be important for performing deep speech recognition tasks and generating appropriate responses or actions.
[0072] In box 312, the processing system can generate updated machine learning models based on stored audio data. For example, the processing system can retrieve and evaluate audio data stored in an audio buffer and associated activation timestamps to identify specific instances where the system failed to activate when it should have or was activated when it should not have, label the identified instances and their corresponding timestamps, generate a training dataset, and use the labeled dataset to adjust or train an existing machine learning model. In some implementations, training operations may include minimizing a loss function using gradient descent, backpropagation, or other optimization techniques, quantifying the difference between the model's predictions and actual results, guiding the model toward improved performance, and generating relevant features that contribute to the learning process using feature extraction methods, which may include applying Fourier transforms or Mel-frequency cepstral coefficients (MFCCs) to the audio data. Identifying MFCCs is a feature extraction method used for speech and audio analysis, particularly for AI systems. MFCCs are a compact representation of the audio spectrum.
[0073] In box 314, the processing system can replace the existing machine learning model in LPALM with an updated machine learning model. That is, after training is complete, the processing system can replace the existing machine learning model with a newly trained and / or updated machine learning model. These new machine learning models can provide better context awareness and responsiveness based on recent data. By continuously updating its machine learning model, the system can better adapt to changing conditions and nuances in user interactions.
[0074] After replacing the existing machine learning model with the updated machine learning model in box 314, or in response to determining that no stable power source is available (i.e., determining box 306 = "No"), the processing system can begin to execute the operation of method 400.
[0075] In box 402, the processing system can repeatedly or continuously capture sensor data (e.g., ambient audio data, etc.). This continuous data capture can allow the system to maintain an up-to-date context of its surrounding environment and / or can facilitate a rapid response to relevant stimuli.
[0076] In box 404, the processing system can store the captured data in an audio buffer, which in some implementations can be a rolling buffer or a buffer memory with a predefined sliding window. A rolling buffer (or circular buffer) is a data structure used to store information such that when the buffer reaches its capacity limit, the latest data overwrites the oldest data. An audio buffer can operate on a "first-in, first-out" (FIFO) principle, where when the buffer is full, the oldest data is removed to make room for new data. This is typically used when receiving a continuous stream of data but only being interested in the most recent data. Similarly, a predefined sliding window can be a fixed-size segment of data that moves across a large dataset in a predefined manner. Unlike a rolling buffer, data in a sliding window is typically not overwritten. Instead, a "window" through which the dataset is viewed slides along the dataset. This is typically used in analyses that require considering relationships between consecutive or nearly consecutive data points over a specific time period. Once the window reaches its capacity, it slides forward to analyze the next segment, which may have some overlap with the previous segment. When the captured data is stored in an audio buffer (e.g., a scrolling buffer or one with a predefined sliding window), the system can add a time dimension, such as timestamps, to the data. The system can analyze not only individual data points but also patterns or sequences that may unfold over a period of time.
[0077] In box 406, the processing system can analyze the captured data (e.g., ambient sound, etc.) for triggering. For example, the processing system can analyze data stored in an audio buffer for spoken words, pitch, rhythm, etc., using a trained language model to identify context, pitch, rhythm, etc., to identify a specific trigger, which could be a specific sound, sound sequence, or other related sensory pattern. The system can use such analysis to determine whether and how to perform further actions, such as activating or waking up a high-power module.
[0078] In box 408, the processing system can generate metadata based on the results of the analysis. The generated metadata can summarize the key characteristics of the captured data. In some implementations, the metadata may include tags, scores, or other descriptors to simplify the content and make it easier to process quickly.
[0079] In box 410, the processing system can transmit metadata to the waker module. See below for reference. Figure 5 In further detail, the wake-up module can use metadata to determine whether the conditions for waking up the higher-power processing module have been met.
[0080] In determination box 412, the processing system can determine whether the audio buffer (e.g., a scrolling buffer or a predefined sliding window buffer) is close to its storage capacity. That is, the processing system can check to determine whether the audio buffer data storage device is close to its capacity limit.
[0081] In response to determining that the audio buffer is nearing capacity (i.e., determining box 412 = "Yes"), in box 414, the processing system can transfer the data in the audio buffer to an external memory storage device. This helps ensure that real-time data capture can continue without interruption, while the stored data remains available for future analysis or training.
[0082] In box 502, the processing system can perform monitoring to receive metadata from the LPALM. The metadata may include analysis of sensor information (such as ambient audio information) that has previously been captured and evaluated by the LPALM.
[0083] In box 504, the processing system can monitor to receive additional contextual data collected from additional modules (e.g., vision, motion, etc.) and / or multiple sources (e.g., a combination of user feedback data and LLM feedback data, etc.).
[0084] In box 506, the processing system can identify and evaluate triggers based on received metadata and received additional contextual data. A trigger can be a word, phrase, or other indicator that prompts the system to take a specific action.
[0085] In box 508, the processing system can determine or calculate a confidence or probability score for the identified trigger and / or immediate action or response. The score can be numerically represented to characterize or indicate the likelihood that the identified trigger is a reasonable action prompt.
[0086] In determination box 510, the processing system can determine whether the determined score exceeds a first threshold. In response to determining that the determined score does not exceed the first threshold (i.e., determination box 510 = "No"), in box 502, the processing system can continue monitoring to receive metadata from LPALM.
[0087] In response to determining that the determined score exceeds a first threshold (i.e., determination box 510 = "Yes"), in determination box 512, the processing system may determine whether the determined score exceeds a second threshold.
[0088] In response to determining that the determined score exceeds the second threshold (i.e., determining box 512 = "Yes"), in box 514, the processing system may perform a quick action and / or generate a quick response.
[0089] In response to determining that the determined score does not exceed the second threshold (i.e., determination box 512 = "No"), the processing system can activate HPRM in box 516.
[0090] In box 602, the processing system can begin operating in standby mode. Standby mode can be an intermediate state that keeps the system alert to specific triggers without fully engaging its resource-intensive modules. In this mode, the system can limit its activity to basic tasks, such as monitoring activation signals or specific conditions that require it to transition to a more active state. Operating in standby mode can be particularly beneficial for mobile devices or other battery-powered operating systems where battery life is critical. Furthermore, operating in standby mode reduces the risk of unnecessary activation and data processing. This can be particularly relevant for systems designed to balance robust performance with the constraints of limited computing resources.
[0091] In box 604, the processing system can monitor for activation signals from the wake-up module. As part of its function, the processing system can monitor activation signals from the wake-up module to facilitate transitions from low-resource-consumption states to more active and resource-intensive states. The wake-up module can perform an initial evaluation of sensor data (such as audio or visual cues) to determine if conditions have been met to ensure attention from computationally more expensive modules in the system. By focusing on detecting specific triggers or conditions, the wake-up module can act as an initial filter to reduce the computational burden on the entire system. Therefore, by monitoring activation signals in box 604, the system can maintain responsiveness to user input or environmental changes without consuming excessive computational resources.
[0092] In determination box 606, the processing system can determine whether an activation signal has been received. In response to determining that no activation signal has been received (i.e., determination box 606 = "No"), in boxes 602 and 604, the processing system can continue operating in standby mode and monitor for activation signals from the wake-up module. In response to determining that an activation signal has been received (i.e., determination box 606 = "Yes"), in box 608, the processing system can switch to operating in active mode (or a higher power state, etc.). That is, when an activation signal is detected, the processing system can transition from its standby or low-power state to a more active state, thereby allocating more resources to complex tasks such as natural language processing or full speech recognition. This allows the system to provide more advanced functions only when they are likely needed, thus improving system performance, functionality, and efficiency.
[0093] In box 610, the processing system can allocate or reserve resources for complex processing tasks. By dedicating resources only when they might be used for complex tasks, the system can save energy and extend the device's battery life. This is particularly relevant for applications where the system spends significant amounts of time waiting to be activated in a low-power or standby state. These operations also help ensure that the system has sufficient computing power and memory to successfully execute complex and / or resource-intensive tasks, such as those involved in natural language processing or high-resolution image recognition.
[0094] In box 612, the processing system can retrieve, record, generate, and / or analyze deep speech data. Deep speech data can include robust contextual information that the processing system can use to understand a user's intent or emotions. For example, subtle differences in pitch, speed, or even background noise can provide valuable insights into the context in which the user's command was given, which in turn can facilitate the generation of more accurate and / or context-aware responses. Furthermore, detailed analysis of the speech data can help develop more sophisticated NLP algorithms. Over time, the system can refine its understanding of spoken language, idiomatic phrases, regional accents, or even individual user speech patterns, thereby improving its ability to interact in a more natural and intuitive way.
[0095] Furthermore, recording and storing speech data can be beneficial for retrospective analysis, especially for training. For example, if a system misinterprets a command or fails to act, the processing system can analyze the stored data to understand what went wrong and how to refine the algorithm. Similarly, deep speech data can be used to train machine learning models to adapt to new speech patterns or forms, allowing the system to maintain a degree of flexibility. Deep speech data can also provide auxiliary benefits such as supporting multimodal interaction (e.g., combining speech with visual cues) or enabling more advanced features, such as speech biometrics, to enhance security.
[0096] In box 614, the processing system can perform advanced analysis using NLP techniques, linguistic cues, nonverbal cues, and more. By using advanced NLP techniques and various cues, the processing system can better understand and respond to user commands in a granular and context-rich manner, enhancing its operational efficiency and the quality of user interaction. As part of the operations in box 614, the processing system can analyze linguistic cues such as grammar, semantics, sub-word elements, context, and co-occurrence. The processing system can also analyze nonverbal cues such as prosody, pitch, volume, speech rate, temporal patterns, and acoustic features. In some implementations, the processing system can also analyze the output from the LLM generated by other components, along with the results and user feedback, for adaptive tuning or refinement of future interactions.
[0097] In box 616, the processing system can generate LLM input queries based on the results of advanced analytics. For example, the processing system can use the results of advanced analytics to generate a single input value or string that serves as a concise yet comprehensive representation of the user's intent, context, and emotional state, which is then input into the LLM component.
[0098] In box 618, the processing system can pass the generated LLM input query to the LLM component. The generated LLM input query can be a concise representation passed to the LLM component. Providing the LLM with well-crafted input queries allows the LLM component to generate more accurate and context-aware responses.
[0099] Various implementation plans (including but not limited to the above references) Figures 1 to 6 The described implementation scheme can be used in a wide variety of wireless devices and computing systems (including laptop computers 700, examples of which are shown in...). Figure 7 Implemented in the example (see below). Reference Figures 1 to 7 The laptop computer 700 may include a processor 702 coupled to volatile memory 704 and a disk drive 706 containing mass non-volatile memory (such as flash memory). The laptop computer 700 may include a touchpad touch surface 708 serving as a pointing device for the computer and thus capable of receiving drag, scroll, and tap gestures. Additionally, the laptop computer 700 may have one or more antennas 710 for transmitting and receiving electromagnetic radiation, connectable to a wireless data link, and / or a cellular transceiver 712 coupled to the processor 702. The computer 700 may also include a BT transceiver 714 coupled to the processor 702, a compact disc (CD) drive 716, a keyboard 718, and a display 720. Other configurations of the computing device may include a computer mouse or trackball, as is well known, coupled to the processor (e.g., via a Universal Serial Bus (USB) input), which may also be used in various implementations.
[0100] Figure 8 This is a component block diagram of a computing device 800 suitable for use with various implementation schemes. (Reference) Figures 1 to 8 Various implementation schemes can be implemented on a variety of computing devices (examples of which are available in 800). Figure 8Implemented on a smartphone (example provided). The computing device 800 may include a first SoC 102 coupled to the second SoC 104. The first SoC 102 and the second SoC 104 may be coupled to internal memory 816, a display 812, and a speaker 814. The first SoC 102 and the second SoC 104 may also be coupled to at least one subscriber identity module (SIM) 840 and / or a SIM interface, which may store information supporting a first 5G NR subscription and a second 5G NR subscription, which support services on a 5G non-standalone (NSA) network.
[0101] The computing device 800 may include an antenna 804 for transmitting and receiving electromagnetic radiation, which may be connected to a wireless transceiver 166 coupled to one or more processors in a first SOC 102 and / or a second SOC 104. The computing device 800 may also include a menu selection button or rocker switch 820 for receiving user input.
[0102] The computing device 800 also includes a sound codec (CODEC) circuit 810 that digitizes sound received from a microphone into data packets suitable for wireless transmission and decodes the received sound data packets to generate an analog signal for use with a speaker to produce sound. Furthermore, one or more of the processor in the first circuit 102 and the second circuit 104, the wireless transceiver 166, and the CODEC 810 may include a digital signal processor (DSP) circuit (not shown separately).
[0103] The processor or processing unit discussed in this application can be any programmable microprocessor, microcomputer, or one or more multiprocessor chips that can be configured via software instructions (applications) to perform a variety of functions, including those described in the various embodiments. In some computing devices, multiple processors may be provided, such as one processor within a first circuit dedicated to wireless communication functions and another processor within a second circuit dedicated to running other applications. Software applications may be stored in memory and then accessed and loaded into the processor. The processor may include internal memory sufficient to store application software instructions.
[0104] Specific implementation embodiments are described in the following paragraphs. While some of the specific implementation embodiments described below are in the form of exemplary methods, further exemplary implementations may include: example methods discussed in the following paragraphs implemented by a computing device including a processor configured (e.g., using processor-executable instructions) to perform the methods of the following specific embodiments; example methods discussed in the following paragraphs implemented by a computing device including components for performing the methods of the following specific embodiments; and example methods discussed in the following paragraphs may be implemented as non-transitory processor-readable storage media storing processor-executable instructions configured to cause the processor of the computing device to perform the operations of the methods of the following specific embodiments.
[0105] Example 1: A method for continuously monitoring speech to obtain a preemptive or context-aware response to a user query, the method comprising: collecting ambient audio data in a low-power always-on listening mode operating on a processing system in a computing device, and storing the collected ambient audio data and an activation timestamp in an audio buffer as buffered audio data; determining a confidence score based on the result of analyzing the buffered audio data for linguistic or non-linguistic cues; and activating a high-power listening mode in response to determining that the confidence score exceeds a threshold and providing the last portion of the audio buffer to the high-power mode for immediate context-aware assistance.
[0106] Example 2: According to the method in Example 1, the confidence score is determined based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues, including analyzing the buffered audio data for linguistic cues such as grammatical cues, semantic cues, sub-word cues, contextual cues, and co-occurrence cues.
[0107] Example 3: The method according to any one of Examples 1 or 2, wherein determining the confidence score based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues includes determining the confidence score based on the results of analyzing the buffered audio data for non-linguistic cues including prosodic cues, pitch cues, speech rate cues, volume cues, time pattern cues, and acoustic feature cues.
[0108] Example 4: The method according to any one of Examples 1 to 3, the method further includes using a trained language model to identify linguistic cues or non-linguistic cues.
[0109] Example 5: The method according to any one of Examples 1 to 4, the method further includes switching to a periodic training mode in response to determining that the computing device is connected to a stable power source.
[0110] Example 6: The method according to any one of Examples 1 to 5, the method further includes: retrieving the buffered audio data and activation timestamp from the audio buffer; using the result of applying the retrieved audio data to a Large Language Model (LLM) to identify instances in which activation of the high-power listening mode should have occurred but did not; marking the identified instances and their corresponding timestamps; generating an updated machine learning model for the low-power always-on listening mode based on the markings; and replacing the current machine learning model for the low-power listening mode with the generated updated machine learning model.
[0111] As used in this application, the terms "component," "module," "system," etc., are intended to include computer-related entities such as, but not limited to, hardware, firmware, combinations of hardware and software, software, or software being executed, configured to perform specific operations or functions. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a computing device and the computing device itself can be referred to as a component. One or more components may reside within a process and / or a thread of execution, and components may reside on a processor or core and / or be distributed across two or more processors or cores. Furthermore, these components may execute on various non-transitory computer-readable media on which various instructions and / or data structures are stored. Components may communicate via local and / or remote processes, function or procedure calls, electronic signals, data packets, memory read / write, and other known network, computer, processor, and / or process-related communication methods.
[0112] A variety of different memory types and memory technologies are available or conceivable in the future, and any or all of these different memory types and memory technologies can be included and used in systems and computing devices implementing various implementation schemes. Such memory technologies / types may include non-volatile random access memory (NVRAM), such as magnetoresistive RAM (M-RAM), resistive random access memory (ReRAM or RRAM), phase-change random access memory (PC-RAM, PRAM, or PCM), ferroelectric RAM (F-RAM), spin-transfer torque magnetoresistive random access memory (STT-MRAM), and 3D-XPOINT memory. Such memory technologies / types may also include non-volatile or read-only memory (ROM) technologies, such as programmable read-only memory (PROM), field-programmable read-only memory (FPROM), and one-time programmable non-volatile memory (OTP NVM). Such memory technologies / types may also include volatile random access memory (RAM) technologies, such as dynamic random access memory (DRAM), double data rate (DDR) synchronous dynamic random access memory (DDR SDRAM), static random access memory (SRAM), and pseudo static random access memory (PSRAM). Systems and computing devices implementing various embodiments may also include or use electronic (solid-state) non-volatile computer storage media, such as flash memory. Each of the memory technologies mentioned above includes, for example, elements suitable for storing instructions, programs, control signals, and / or data for use in or for use in: advanced driver assistance systems (ADAS) of vehicles, system-on-a-chip (SOC), or other electronic components. Any references to terms and / or technical details relating to individual memory types, interfaces, standards, or memory technologies are for illustrative purposes only and are not intended to limit the scope of the claims to a particular memory system or technology, unless expressly stated in the language of the claims.
[0113] The various embodiments illustrated and described are provided merely as examples illustrating the various features of the claims. However, the features shown and described with respect to any given embodiment are not necessarily limited to the associated embodiment and may be used or combined with other embodiments shown and described. Furthermore, the claims are not intended to be limited to any one of the exemplary embodiments. For example, one or more operations in the operation of the method may substitute for or combine with one or more operations of the method.
[0114] The foregoing method descriptions and process flowcharts are provided as illustrative examples only and are not intended to require or imply that the operations of the various embodiments must be performed in the given order. As those skilled in the art will appreciate, the operations in the foregoing embodiments can be performed in any order. Words such as “afterward,” “then,” “next,” etc., are not intended to restrict the order of operations; these words are only used to guide the reader through the description of the method. Furthermore, any reference to singular claim elements (e.g., references using the articles “a,” “an,” or “described”) should not be construed as limiting that element to the singular.
[0115] The various exemplary logic blocks, modules, circuits, and algorithmic operations described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and operations have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. While those skilled in the art may implement the described functionality in different ways for each specific application, such implementation decisions should not be construed as departing from the scope of the claims.
[0116] Hardware for implementing the various exemplary logic units, logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or executed using a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (TCUASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic unit, discrete hardware component, or any combination thereof designed to perform the functions described herein. While the general-purpose processor may be a microprocessor, in alternative embodiments, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Alternatively, some operations or methods may be performed by circuitry specific to a given function.
[0117] In one or more embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, such functionality may be stored as one or more instructions or code on a non-transitory computer-readable medium or a non-transitory processor-readable medium. The operation of the methods or algorithms disclosed herein may be embodied in a processor-executable software module that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium may be any storage medium accessible by a computer or processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to store object program code in the form of instructions or data structures and is accessible by a computer. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, the operation of a method or algorithm may reside as a single line of code and / or instruction, or any combination or set of code and / or instructions, on a non-transitory processor-readable medium and / or computer-readable medium that may be incorporated into a computer program product.
[0118] The above description of the disclosed embodiments is provided to enable any person skilled in the art to implement or use the claims. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without departing from the scope of the claims. Therefore, this disclosure is not intended to be limited to the embodiments shown herein, but should be granted the broadest scope consistent with the following claims and the principles and novel features disclosed herein.
Claims
1. A computing device, the computing device comprising: Processing system, the processing system being configured to: Ambient audio data is collected in low-power always-on listening mode, and the collected ambient audio data and activation timestamp are stored in an audio buffer as buffered audio data. The confidence score is determined based on the results of analyzing the buffered audio data in response to linguistic or non-linguistic cues. as well as In response to determining that the confidence score exceeds a threshold, a high-power listening mode is activated and the last portion of the audio buffer is provided to the high-power mode for immediate context-aware assistance.
2. The computing device of claim 1, wherein the processing system is configured to determine the confidence score based on the results of analyzing the buffered audio data for linguistic cues including grammatical cues, semantic cues, sub-word cues, contextual cues, and co-occurrence cues, thereby determining the confidence score based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues.
3. The computing device of claim 1, wherein the processing system is configured to determine the confidence score based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues, including determining the confidence score based on the results of analyzing the buffered audio data for non-linguistic cues including prosodic cues, pitch cues, speech rate cues, volume cues, temporal pattern cues, and acoustic feature cues.
4. The computing device of claim 1, wherein the processing system is further configured to use a trained language model to identify linguistic cues or non-linguistic cues.
5. The computing device of claim 1, wherein the processing system is further configured to switch to a periodic training mode in response to determining that the computing device is connected to a stable power source.
6. The computing device of claim 1, wherein the processing system is further configured to: Retrieve the buffered audio data and activation timestamp from the audio buffer; The results of applying the retrieved audio data to a large language model (LLM) are used to identify instances where high-power listening patterns should have occurred but did not. The instance is marked and its corresponding timestamp is displayed. Based on the tags, an updated machine learning model is generated for the low-power always-on listening mode; and Replace the current machine learning model in the low-power listening mode with the generated updated machine learning model.
7. A method for continuously monitoring speech to obtain a preemptive or context-aware response to a user query, the method comprising: Ambient audio data is collected in a low-power always-on listening mode operating on the processing system in the computing device, and the collected ambient audio data and activation timestamps are stored in an audio buffer as buffered audio data. The confidence score is determined based on the results of analyzing the buffered audio data in response to linguistic or non-linguistic cues. as well as In response to determining that the confidence score exceeds a threshold, a high-power listening mode is activated and the last portion of the audio buffer is provided to the high-power mode for immediate context-aware assistance.
8. The method of claim 7, wherein determining the confidence score based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues includes determining the confidence score based on the results of analyzing the buffered audio data for linguistic cues including grammatical cues, semantic cues, sub-word cues, contextual cues, and co-occurrence cues.
9. The method of claim 7, wherein determining the confidence score based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues includes determining the confidence score based on the results of analyzing the buffered audio data for non-linguistic cues including prosodic cues, pitch cues, speech rate cues, volume cues, temporal pattern cues, and acoustic feature cues.
10. The method of claim 7, further comprising using a trained language model to identify linguistic or non-linguistic cues.
11. The method of claim 7, further comprising switching to a periodic training mode in response to determining that the computing device is connected to a stable power source.
12. The method according to claim 7, further comprising: Retrieve the buffered audio data and activation timestamp from the audio buffer; The results of applying the retrieved audio data to a large language model (LLM) are used to identify instances where high-power listening patterns should have occurred but did not. The instance is marked and its corresponding timestamp is displayed. A machine learning model for updating the low-power always-on listening mode is generated based on the tags; as well as Replace the current machine learning model in the low-power listening mode with the generated updated machine learning model.
13. A non-transitory computer-readable storage medium having processor-executable software instructions stored thereon, the processor-executable software instructions being configured to cause a processing system in a computing device to perform operations for continuously monitoring speech to obtain a preemptive or context-aware response to a user query, the operations comprising: In low-power always-on listening mode operation, ambient audio data is collected, and the collected ambient audio data and activation timestamp are stored in an audio buffer as buffered audio data. The confidence score is determined based on the results of analyzing the buffered audio data in response to linguistic or non-linguistic cues. as well as In response to determining that the confidence score exceeds a threshold, a high-power listening mode is activated and the last portion of the audio buffer is provided to the high-power mode for immediate context-aware assistance.
14. The non-transitory computer-readable storage medium of claim 13, wherein the stored processor-executable software instructions are configured to cause the processor to perform operations such that determining the confidence score based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues includes determining the confidence score based on the results of analyzing the buffered audio data for linguistic cues including grammatical cues, semantic cues, sub-word cues, contextual cues, and co-occurrence cues.
15. The non-transitory computer-readable storage medium of claim 13, wherein the stored processor-executable software instructions are configured to cause the processor to perform operations such that determining the confidence score based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues includes determining the confidence score based on the results of analyzing the buffered audio data for non-linguistic cues including prosodic cues, pitch cues, speech rate cues, volume cues, time pattern cues, and acoustic feature cues.
16. The non-transitory computer-readable storage medium of claim 13, wherein the stored processor-executable software instructions are configured to cause the processor to perform operations, said operations further comprising using a trained language model to identify linguistic or non-linguistic cues.
17. The non-transitory computer-readable storage medium of claim 13, wherein the stored processor-executable software instructions are configured to cause the processor to perform operations, the operations further comprising switching to a periodic training mode in response to determining that the computing device is connected to a stable power source.
18. The non-transitory computer-readable storage medium of claim 13, wherein the stored processor-executable software instructions are configured to cause the processor to perform an operation, said operation further comprising: Retrieve the buffered audio data and activation timestamp from the audio buffer; The results of applying the retrieved audio data to a large language model (LLM) are used to identify instances where high-power listening patterns should have occurred but did not. The instance is marked and its corresponding timestamp is displayed. A machine learning model for updating the low-power always-on listening mode is generated based on the tags; as well as Replace the current machine learning model in the low-power listening mode with the generated updated machine learning model.
19. A computing device, the computing device comprising: This component is used to collect ambient audio data in low-power always-on listening mode and store the collected ambient audio data and activation timestamp in an audio buffer as buffered audio data. A component for determining a confidence score based on the results of analyzing the buffered audio data for linguistic or non-linguistic cues; as well as A component for activating a high-power listening mode in response to determining that the confidence score exceeds a threshold and providing the last portion of the audio buffer to the high-power mode for immediate context-aware assistance.
20. The computing device of claim 19, wherein the component for determining the confidence score based on the result of analyzing the buffered audio data for linguistic or non-linguistic cues includes a component for determining the confidence score based on the result of analyzing the buffered audio data for linguistic cues including grammatical cues, semantic cues, sub-word cues, contextual cues, and co-occurrence cues.
21. The computing device of claim 19, wherein the component for determining the confidence score based on the result of analyzing the buffered audio data for linguistic or non-linguistic cues includes a component for determining the confidence score based on the result of analyzing the buffered audio data for non-linguistic cues including prosodic cues, pitch cues, speech rate cues, volume cues, temporal pattern cues, and acoustic feature cues.
22. The computing device of claim 19, further comprising a component for identifying linguistic or non-linguistic cues using a trained language model.
23. The computing device of claim 19, further comprising components for switching to a periodic training mode in response to determining that the computing device is connected to a stable power source.
24. The computing device of claim 19, further comprising: Components for retrieving the buffered audio data and activating the timestamp from the audio buffer; A component for identifying instances in which high-power listening patterns should have occurred but did not, using the results of applying the retrieved audio data to a large language model (LLM); A component used to mark the identified instance and its corresponding timestamp; Components for generating an updated machine learning model based on the tags for the low-power always-on listening mode; as well as A component used to replace the current machine learning model in low-power listening mode with the generated updated machine learning model.