Voice-based user authentication
Through the two-stage user authentication process, keyword detection and similarity comparison are used to solve the problem of increased user authentication delays based on voice, and a more efficient and accurate authentication process is achieved.
Patent Information
- Application Number
- CN202280100851.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-05-16
AI Technical Summary
During voice-based user authentication, processing of long-lasting speech may lead to an increase in delay, affecting the accuracy and efficiency of authentication.
A two-stage user authentication process is adopted, first determining the status of the user equipment by detecting keywords, then determining the user identity based on similarity comparison, and using the first and second thresholds to optimize the authentication process.
Reduces the latency of user authentication and improves the accuracy and efficiency of authentication, especially when processing speech for a longer duration.
Smart Images

Figure CN120019356A_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to processing audio data.For example, aspects of the present disclosure relate to systems and techniques for providing improvements (eg, latency reduction) for speech-based (eg, text-independent) user authentication (also referred to as user verification). Background Art
[0002] Electronic devices (e.g., mobile devices and other electronic devices) can communicate audio (e.g., speech or voice) and data packets via a wireless network. Such devices may also provide additional functionality via one or more applications, such as capturing images using a digital still camera, capturing videos using a digital video camera, recording data (e.g., audio, image data, video, etc.) using a digital recorder, outputting audio using an audio player (e.g., streaming music or music files, book content, etc.), and / or other functionality. Some electronic devices may be configured to process speech or voice input for various purposes. For example, a speech recognition application such as a virtual digital assistant of an electronic device may translate a spoken speech command into a function or action to be performed by one or more other applications (e.g., an audio file player, etc.) of the device. In some cases, an electronic device may perform user authentication or verification to authenticate / verify and identify a user based on voice or speech characteristics to determine whether the user is an authorized user of the device.
[0003] In some cases, there may be latency issues when performing user authentication / verification based on voice or speech. For example, when processing speech with a longer duration, the user authentication or verification application may provide more accurate user authentication / verification results. However, when processing speech with a longer duration, the user authentication or verification application may experience more latency. Summary of the invention
[0004] In some examples, systems and techniques for authenticating a user of an electronic device using voice input (e.g., using text-independent speech analysis) are described. These systems and techniques can reduce the latency associated with voice-based user authentication.
[0005] According to at least one example, a method for processing audio is provided. The method includes: obtaining first audio information from a user using an audio sensor of a user device; determining whether the first audio information includes audio corresponding to a detected keyword, the detected keyword configuring the user device to receive or process one or more commands from the user; based on the first audio information including the audio corresponding to the detected keyword, determining the similarity between the first audio information corresponding to the detected keyword and a model of an authenticated user; and determining whether to authenticate the user as the authenticated user based on a first comparison of the similarity between the first audio information and the model of the authenticated user and a first threshold.
[0006] In another example, a device for processing audio is provided, the device comprising at least one memory and at least one processor coupled to the at least one memory. The at least one processor is configured to: obtain first audio information from a user using an audio sensor of a user device; determine whether the first audio information includes audio corresponding to a detected keyword, the detected keyword configuring the user device to receive or process one or more commands from the user; based on the first audio information including the audio corresponding to the detected keyword, determine the similarity between the first audio information corresponding to the detected keyword and a model of an authenticated user; and determine whether to authenticate the user as the authenticated user based on a first comparison of the similarity between the first audio information and the model of the authenticated user and a first threshold.
[0007] In another example, a non-transitory computer-readable medium having instructions stored thereon is provided, which, when executed by one or more processors, causes the one or more processors to: obtain first audio information from a user using an audio sensor of a user device; determine whether the first audio information includes audio corresponding to a detected keyword, the detected keyword configuring the user device to receive or process one or more commands from the user; determine a similarity between the first audio information corresponding to the detected keyword and a model of an authenticated user based on the first audio information including the audio corresponding to the detected keyword; and determine whether to authenticate the user as the authenticated user based on a first comparison of the similarity between the first audio information and the model of the authenticated user with a first threshold.
[0008] In another example, a device for processing audio is provided. The device includes: a component for obtaining first audio information from a user using an audio sensor of a user device; a component for determining whether the first audio information includes audio corresponding to a detected keyword, the detected keyword configuring the user device to receive or process one or more commands from the user; a component for determining a similarity between the first audio information corresponding to the detected keyword and a model of an authenticated user based on the first audio information including the audio corresponding to the detected keyword; and a component for determining whether to authenticate the user as the authenticated user based on a first comparison of the similarity between the first audio information and the model of the authenticated user with a first threshold.
[0009] In some aspects, the apparatus is a mobile device, is part of a mobile device and / or includes a mobile device (e.g., a mobile phone and / or a mobile handset and / or a so-called "smart phone" or other mobile device), an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device or a mixed reality (MR) device, a head mounted device (HMD) device, a vehicle or a computing system, device or component of a vehicle, a wearable device (e.g., a network-connected watch or other wearable device), a wireless communication device, a camera, a personal computer, a laptop computer, a server computer, another device, or a combination thereof. In some aspects, the apparatus includes one or more cameras for capturing one or more images. In some aspects, the apparatus also includes a display for displaying one or more images, notifications and / or other displayable data. In some aspects, the apparatus may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof and / or other sensors).
[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.
[0011] The foregoing and other features and aspects will become more apparent upon reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Illustrative aspects of the present application are described in detail below with reference to the following drawings:
[0013] Figure 1 A conceptual diagram of a voice input device incurring a delay due to delayed authentication of the voice input.
[0014] Figure 2 is a block diagram of a voice input device that reduces latency by improving authentication of voice input according to some aspects of the present disclosure.
[0015] Figure 3 is a flow chart of a process for reducing latency performed by a voice input device according to some aspects of the present disclosure.
[0016] Figure 4A and Figure 4B is a timing diagram illustrating different voice authentication scenarios according to aspects of the present disclosure.
[0017] Figure 5 is an illustration of a speaker device including voice input functionality for authenticating a user according to some aspects of the present disclosure.
[0018] Figure 6 is an illustration of a mobile device including voice input functionality for authenticating a user according to some aspects of the present disclosure.
[0019] Figure 7 is an illustration of an automated cleaning device 700 including voice input functionality for authenticating a speaker in accordance with some aspects of the present disclosure.
[0020] Figure 8 is a flow chart illustrating an example of a method 800 for processing audio data in accordance with certain aspects of the present disclosure.
[0021] Fig. 9 is a diagram illustrating an example of a system for implementing certain aspects described herein. DETAILED DESCRIPTION
[0022] Certain aspects of the present disclosure are provided below. Some of these aspects may be applied independently, and some of them may be applied in combination, which is obvious to those skilled in the art. In the following description, specific details are set forth for explanation purposes to provide a thorough understanding of various aspects of the application. However, it will be apparent that various aspects may be implemented without these specific details. Each of the drawings and descriptions is not intended to be restrictive.
[0023] The following description provides only example aspects and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of the example aspects will provide a description that can be used to implement the example aspects to those skilled in the art. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the essence and scope of the present application as set forth in the appended claims.
[0024] The following description provides only exemplary aspects and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of exemplary aspects will provide an enabling description for implementing aspects of the present disclosure to those skilled in the art. It should be understood that various changes may be made to the function and arrangement of elements without departing from the essence and scope of the present application as set forth in the appended claims.
[0025] The terms "exemplary" and / or "example" are used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" and / or "example" is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term "aspects of the disclosure" does not require that all aspects of the disclosure include the discussed feature, advantage, or mode of operation.
[0026] As previously described, an electronic device may be configured to receive audio input (e.g., speech or voice input) and intelligently process the audio input to perform one or more functions, such as controlling the device, causing the device to output audio content (e.g., music content, book content, etc.), controlling an auxiliary device or system (such as a lighting system) connected (e.g., wirelessly connected) to the voice input device, and the like. Such electronic devices may be referred to as voice input devices. In some aspects, applications utilized by the voice input device to process audio input may be referred to as voice assistants. Examples of voice input devices include mobile phones, XR devices (e.g., VR devices, AR devices, and / or MR devices), vehicles or systems, devices, or components of vehicles, tablet computers, televisions (TVs), external TV input devices (e.g., Roku, etc.), and the like. TM etc.), smart speakers, laptops, desktop computers, or any other suitable electronic device.
[0027] In some cases, the voice input device can be placed in a low power state. When in a low power state, the voice input device can monitor the environment to find speech related to keywords. For example, the voice input device can be configured to wake up and transition to a higher power state after detecting keywords in the speech input. In some cases, the voice input device can provide feedback (e.g., audio feedback, visual feedback such as using a display, one or more lights or other visual feedback, etc.) to indicate that the voice assistant of the voice input device is active. For example, after the user says the keyword configured to activate the voice assistant of the voice input device, the voice input device can light up a lamp integral with the device and / or provide an audio output to indicate that the voice input device is waiting for additional input (e.g., referred to as a command) from the user. If a command (e.g., voice or speech command) is received from the user, the voice assistant can activate one or more applications (e.g., music applications, applications for controlling auxiliary devices or systems, etc.). In some cases, the initial speech input can include keywords and commands, in which case, if the keywords are identified, the voice input device (e.g., voice assistant) can process the keywords and then (e.g., after entering a higher power state) process the commands.
[0028] The voice input device may also be configured to verify or authenticate the user from which the voice or speech input is received, which may be referred to as user verification or user authentication. User verification or authentication is the process of verifying that a user corresponds to a registered identity (e.g., a user profile) of a voice input device and / or voice assistant. Once verified or authenticated, the voice input device and / or voice assistant may enable the user to engage in activities authorized by the user, such as accessing one or more applications (e.g., a music application, an application for controlling an auxiliary device or system, etc.).
[0029] Authenticating a user using voice or speech input that is independent of one or more predefined keywords is a complex process and may require a significant amount of processing power. In addition, authenticating a user based on voice or speech input that includes a detected keyword and a subsequent voice or speech command may result in significant latency. For example, when processing speech with a longer duration (e.g., a keyword and subsequent command), a user authentication or verification application can provide more accurate user authentication / verification results. However, processing such longer duration voice or speech input may cause the user authentication or verification application to experience more latency. As described in reference Figure 1As further described, the entire verification / authentication process using a voice input device may take a significant amount of time (e.g., 500 milliseconds (ms), 1 second, etc.), resulting in an application (e.g., a music application, an application for controlling an auxiliary device or system, etc.) not receiving the user verification / authentication result and the corresponding command to be processed until an even longer period of time (e.g., 2 seconds, 3 seconds, 4 seconds, etc.). Such delays may result in noticeable and inconvenient delays for the user.
[0030] In some aspects, systems, devices, processes (also referred to as methods), and computer-readable media (collectively referred to herein as "systems and techniques") are described for improving speech-based user authentication or verification (e.g., text-independent user authentication and verification). The systems and techniques may perform a two-stage user authentication or verification process, which in some cases may be a text-independent user authentication or verification process. For example, the systems and techniques may determine whether the obtained first audio information includes audio corresponding to a detected keyword (e.g., a keyword previously detected as a valid keyword), the detected keyword configuring the user device to receive or process one or more commands from the user. In one aspect, the first audio information may correspond to a detected keyword associated with the user device (e.g., a text-independent keyword created by a user of the user device), and based on the first audio information including audio corresponding to the detected keyword, the systems and techniques may determine the similarity between the first audio information corresponding to the keyword and a model of an authenticated user. In one illustrative aspect, the model is trained by an authenticated user.
[0031] After determining the similarity between the first audio information corresponding to the detected keyword and the model of the authenticated user, the systems and techniques may determine whether to authenticate the user as the authenticated user based on a first comparison of the similarity between the first audio information and the model of the authenticated user with a first threshold value. For example, if the similarity is greater than the first threshold value, the systems and techniques may authenticate the user as the authenticated user at this time without inputting a command or query to the systems and techniques.
[0032] In another exemplary aspect, when the similarity between the first audio information and the model of the authenticated user is less than a first threshold, the system and technology may obtain a second audio information that follows the first audio information and includes a command or query. The system and technology may determine the similarity between the second audio information (and in some cases, the combination of the first audio information and the second audio information) and the model of the authenticated user. The system and technology determine whether to authenticate the user as an authenticated user based on a second comparison of the similarity between the second audio information (and in some cases, the combination of the first audio information and the second audio information) and the model of the authenticated user with a second threshold different from the first threshold. In some cases, the similarity may be based on the portion of the second audio information with the longest duration (e.g., based on a timer). For example, the system and technology may use a portion of a command or query, such as two seconds of a command or query, and authenticate the user as an authenticated user while continuing to enter the command or query. In some aspects, the system and technology reduce the latency of user authentication (e.g., based on authenticating the user based only on the detected keywords included in the first audio information) and may provide visual feedback in some cases to improve voice input capabilities.
[0033] In one illustrative example, a keyword detection engine of the system may receive an audio sample (e.g., pulse code modulation (PCM) data from one or more microphones) as input and may determine whether a target keyword is included in the audio input. In some cases, a trained neural network keyword detection model may be used to determine whether the audio data includes a keyword. The audio sample is determined to include a keyword, and the audio sample may be stored in a detected keyword buffer. In addition, if a keyword is detected, the system may use the audio sample in the detected keyword buffer to start the first stage of a two-stage text-independent user verification process. In the first stage, the system may compare features extracted from the detected keyword audio sample with a registered / registered user model to determine whether the keyword is uttered by a target / registered user (which may be referred to as an authorized user). For example, if the similarity between the keyword audio sample and the model (e.g., a user voice confidence score) is higher than the above-mentioned first threshold, the system may determine with high confidence that the user is an authorized user using only the keyword audio sample. In such a case, the system may stop the verification / authentication process and may begin transmitting subsequent information (e.g., commands) to an upper layer such as a client application. Using only keyword audio samples can greatly reduce the user verification / authentication process, thereby reducing the end-to-end latency of voice activation, because the system does not need to wait for the audio command to make a decision on the keyword and the authorized user. However, if the user voice confidence score is not high enough (e.g., the similarity is less than the first threshold), the system may not have enough confidence to confirm whether it is the authorized user who is speaking. The system can then proceed to obtain subsequent command speech audio samples (when available) and perform the second stage of the two-stage text-independent user verification process.
[0034] In some cases, rather than utilizing text-independent user verification or authentication, a voice activation system may use keyword-dependent user verification or authentication. For example, initial voice activation may use the same keyword-only audio sample to concurrently perform keyword detection and keyword-dependent user verification. Such keyword-dependent voice activation systems do not use command audio samples (e.g., audio samples including commands that appear after an audio sample including a keyword). Such voice activation systems may require that, during the enrollment phase, user enrollments must be the same keyword from the same user repeated a specific number of times (e.g., five times) to create a user voice model. During the detection phase, the keyword buffer audio data may be processed for keyword detection and also for user verification.
[0035] In some cases, the user verification system can be extended to support text-independent user verification or authentication. In such cases, during the registration phase, the user registration can use random speech samples (excluding keywords) from the same target user (e.g., five commands / sentences read by the target user) to create a user voice model. During the detection phase, keywords and commands can be used for user verification / authentication. However, such systems may have large end-to-end delays.
[0036] As described above, the systems and techniques described herein can reduce end-to-end latency for voice-activated systems to support user verification or authentication (e.g., text-independent user verification or authentication). For example, instead of waiting until an audio sample including a keyword and an audio sample including a command are all completely completed by the user to start user authentication, the systems and techniques can use only the keyword audio sample from the keyword sample buffer to perform the first stage of the two-stage verification / authentication process, and if the confidence level of using the keyword audio sample is high (e.g., greater than a first threshold), the system can authenticate / authenticate the user (while also determining that the keyword is detected) without waiting until the command is completed to start the user verification / authentication process.
[0037] Additional details and aspects of the disclosure are described in more detail below with respect to the accompanying drawings.
[0038] Figure 1 100 is a conceptual diagram of inputting a voice command into a voice input device that incurs a delay due to delayed authentication. As used herein, the terms "voice" and "speech" are used interchangeably. The voice input device can be configured to receive voice input and perform various actions based on the input. An illustrative example of a voice input device is a smart speaker that can output audio and can be programmed with other functions or can be operated using voice input. Other illustrative examples of voice input devices include mobile devices (e.g., mobile phones), XR devices, systems or components of vehicles (e.g., media systems), or other devices. The voice input device can also be configured to be connected to another electronic device to be programmed or configured using a graphical user interface. Figure 5 An example of a smart speaker is shown in . Figure 6 and Figure 7 Illustrative examples of voice input devices are further illustrated in .
[0039] In order to operate the voice input device, a voice command is provided from the user and received by the voice input device. The voice command may include a keyword. In some cases, the keyword may be user defined, in which case the user may define the keyword during the registration or setup phase of the device (and it is not predefined by the manufacturer of the device). Text-independent authentication does not require a specific keyword to verify the user's authentication. For example, the user may be able to customize the keyword, such as selecting different possible keywords, or providing a customized keyword. In some examples, the keyword may be customized by the user through a user interface of a device connected to the voice input device, for example, by selecting a keyword or another method of inputting a keyword (e.g., by providing voice input defining the keyword).
[0040] At box 102, the voice input device may be configured to use an audio sensor (e.g., a microphone) to receive audio data (including speech input) and monitor the audio data to find keywords. In some cases, the voice input device may monitor keywords when in a low power state. In some examples, the voice input device may buffer audio data. The voice input device may analyze the audio data to determine whether keywords are detected. In some examples, keywords may be identified by comparing known patterns with voice commands to ascertain whether speech corresponds to keywords.
[0041] If a keyword is detected at box 102, the voice input device may obtain a second voice input (e.g., received as part of the same phrase that includes the keyword, or received after the voice input device prompts after the keyword is detected). At box 104, the voice input device may buffer the second voice input. In some cases, the voice input device may enter a higher power state after detecting the keyword. In some aspects, the second voice input may be a command, such as a function to be performed (e.g., start a timer, play music, etc.). On the other hand, the second voice input may be a query for information from the user (e.g., a request for the current time). The second voice input may be buffered so that the voice input device receives enough commands to perform user authentication or verification. In an illustrative example, the second voice input (e.g., a command) may include between two and four seconds. The second voice input may be longer due to complex queries and speech pauses. As Figure 1 As shown, the second voice input creates a first delay that varies over time based on the complexity of the second voice input.
[0042] After the second voice input is buffered, at box 106, the voice input device is configured to use the second voice input (and in some cases, the keyword and the first voice input) to perform text-independent user authentication or verification to determine whether the voice input corresponds to the user. In one aspect, the voice input device can store a voice model of the user's speech based on a registration or training process. The voice model may include characteristics of the user's speech. For example, the voice model may include pitch (e.g., pitch frequency), formants (e.g., formant frequency), and / or other characteristics of the user's speech based on the voice provided by the user during the registration or training process. The voice model can compare the voice input at box 104 with the voice model to authenticate that the user (e.g., the person providing the voice input) corresponds to a model of the user's speech.
[0043] Text-independent processing based on the second voice input (e.g., command) and in some cases based on keywords may consume a significant amount of time and may require comparing complex data objects to the voice model. The delay incurred by the text-independent processing at box 106 is also variable, depending on the length of the voice input, the noise quality of the voice input (e.g., signal-to-noise ratio (SRN)), and other factors. For example, a text-independent processing duration of 500 milliseconds (ms) may be incurred in some cases.
[0044] After authenticating the user, the voice input device is configured to provide at least the second voice input to the application (e.g., music application, timer, etc.) at box 108. In some cases, the voice input device can use a translation application to process the second voice input to convert speech into machine-readable content (e.g., word vectors, text, etc.) and eliminate the ambiguity of the meaning of the second voice input. In one aspect, the translation service can be an automatic speech recognition (ASR) or natural language processing (NLP) function that converts the input into machine-readable content. For example, NLP provides symbols (e.g., single words) that identify the relationship of the text within the second voice input. The second voice input can be processed by the voice input device and / or cloud services to eliminate the ambiguity of the meaning of the second voice input. In some aspects, the cloud service is used to eliminate the ambiguity of the meaning of the second voice input, because the translation service can use a dictionary with words represented by a multidimensional vector (e.g., 768 dimensions for the current NLP dictionary), which consumes a lot of storage space and changes continuously based on further training. In other cases, the dictionary can be stored locally on the device. The voice input device can pre-process the audio data, for example, performing local filtering and downsampling to reduce the size of the audio data. Depending on the complexity of the technology and language employed, the provision of the second audio input at block 108 is also variable. In some cases, the voice input device may provide a first audio input (eg, a keyword) and a second audio input (eg, a command) to the translation service.
[0045] At box 110, the voice input device is configured to receive a response associated with the second voice input, and can then act on the second voice input. For example, the response associated with the second voice input can be provided to an application executed in the voice input device, such as a multimedia application that is playing audio.
[0046] like Figure 1 As shown, voice input devices may have a significant amount of latency, and authentication of a user command may result in a duration in which the voice input device is processing the voice input but is unable to perform the intended function requested by the user. Second-person delays may exacerbate user frustration and may cause inconvenience to the user. For example, if a user requests information from a voice input device, and the voice input device consumes three seconds, and then informs the user that their voice input was not authenticated (e.g., as a result of a noisy environment with a low SNR), the delay may encourage the user to avoid using voice input.
[0047] Figure 2 is a block diagram of a voice input device 200 that reduces latency by improving authentication of voice input according to some aspects of the present disclosure. The voice input device 200 may perform a two-stage text-independent user authentication process (e.g., Figure 3 The voice input device includes an audio capture device 202, a processor 204, a memory 206, and a communication module 208.
[0048] The audio capture device 202 is configured to obtain the sound in the environment of the voice input device 200 and convert the sound into audio information (e.g., audio data). An example of the audio capture device 202 is a microphone, such as an audio transducer. In some cases, the voice input device 200 may include multiple audio capture devices 202 to improve the audio fidelity of the audio information. The processor is configured to receive instructions stored in the memory 206 and execute these instructions.
[0049] In one aspect, the memory 206 may store an audio processing engine 210, which is configured to process audio according to various aspects of the present disclosure. The memory may include a speech detection engine 212, which is configured to recognize audio information that may include the speech of the user. In some cases, the speech detection engine 212 can recognize the expressed words in the voice input and convert the words into text (e.g., speech to text synthesis). As described above, some devices may omit the speech detection engine because the high-fidelity model is more suitable for storage on the server for continued training and based on the size of the dictionary of the multidimensional vector. The voice input device 200 may also include a keyword detection engine 214, which is configured to detect or identify keywords based on a pattern. For example, the pattern can be represented by a spectral analysis within a time period.
[0050] The voice input device 200 may also store a voice model 216, which may be trained by a user of the voice input device 200 during an enrollment process or phase. For example, during an enrollment process, the voice input device 200 may request that the user provide voice input to the device, and the voice input device 200 or another device (e.g., a cloud computing device) may analyze the voice input to identify characteristics or patterns (e.g., pitch, formants, etc.) that are indicative of the user's speech patterns.
[0051] The voice input device 200 also includes a communication module 208, which is configured to transmit data across a physical interface (e.g., a wireless communication link) to perform various communication functions. The communication module may include short-range (e.g., Bluetooth low energy (BLE), Wi-Fi, etc.) communication circuits and long-range (e.g., cellular) communication circuits.
[0052] In some aspects, the audio processing engine 210 may include a logic function for controlling text-independent user authentication at the voice input device 200. In an exemplary aspect, the audio processing engine 210 may be configured to perform a two-stage text-independent user authentication process for voice input. For example, the audio processing engine 210 includes instructions for the processor 204 to perform text-independent user authentication based on comparing the voice input including the detected keywords (e.g., previously detected keywords) with the voice model 216 of the user associated with the voice input device 200 (corresponding to the first stage of the two-stage text-independent user authentication process). The comparison can use integers or floating point values to generate similarities, and the audio processing engine 210 includes instructions for the processor 204 to compare the similarities with a first threshold. If the similarity of the voice input including the detected keywords is higher than or equal to the first threshold, the audio processing engine 210 includes instructions for the processor 204 to determine that the voice input including the detected keywords corresponds to the user. Such similarity determination is separated from the detection keywords and involves authenticating that the authenticated user is an authenticated user (e.g., for text-independent user authentication).
[0053] In response to determining that the voice input including the detected keyword corresponds to the user, the audio processing engine 210 may include instructions for the processor 204 to authenticate the user and be configured to process the subsequent voice input before receiving the subsequent voice input (e.g., before receiving the command). For example, when the user speaks the detected keyword in a low noise environment, the processor 204 may identify that the similarity between the voice input and the voice model 216 is 0.95, which indicates a high correlation and is greater than the first threshold of 0.9. The value of the first threshold is an example and can be configured based on the device. For example, a smart speaker may have a lower first threshold than a mobile device.
[0054] In some aspects, the audio processing engine 210 may include instructions for the processor 204 to provide an indicator to identify the authentication of the user. For example, the processor 204 may authenticate the user, and may then provide a command to the executing application to indicate that the voice providing the speech is authenticated, and the application may provide instructions to change the visual indicator to indicate authentication. An example of a visual indicator may be a dot within a graphical user interface that may change color from red to green to indicate authentication, or may be a hardware component, such as an LED light that is changed to output green to indicate authentication. The output of the visual indicator provides visual feedback to the user that may be easily understood and informs the user that subsequent voice input will be processed.
[0055] In some other aspects, if the processor 204 does not authenticate the user based on voice input including the detected keyword, the audio processing engine 210 may include instructions for the processor 204 to capture a portion of the voice input including the query or command (e.g., for the second stage of a two-stage text-independent user authentication process). The entire query or command may take several seconds of voice input time, and the audio processing engine 210 may include instructions for the processor 204 to perform the authentication using a maximum amount of time (e.g., 3 seconds of voice input), which enables the processor 204 to perform the authentication in parallel with continuing to receive voice input. For example, the voice input device may be configured to buffer the voice input using a stream, and the data may be processed as the stream is received, rather than waiting for the entire voice input and then processing the entire voice input at once. This example allows the voice input device to continue to receive voice input, and a portion of the voice input may be provided for authentication based on a second comparison with the voice model 216.
[0056] The second comparison also generates a second similarity using an integer or floating point value, and the audio processing engine 210 includes instructions for the processor 204 to compare the second similarity with a second threshold. If the second similarity is greater than or equal to the second threshold, the audio processing engine 210 includes instructions for the processor 204 to determine that the voice input including the command or query corresponds to the user. In this case, the comparison of the voice input is more robust and can more accurately provide a comprehensive comparison with the voice model 216. The second threshold can therefore be lower to tolerate higher noise environments or conditions that can affect the audio quality of the voice input obtained by the audio capture device 202.
[0057] In this case, the voice input device 200 is configured to perform authentication during the input of the command or query, which can reduce the latency of the user's authentication. As described above, the voice input device 200 can also be configured to provide visual feedback in a graphical user interface or another visual indicator to notify the user that their identity has been authenticated based on the voice input. In some aspects, the voice input associated with the command or query can be less than the maximum amount of time, and the voice input device 200 can perform authentication of the entire voice input associated with the command or query.
[0058] Figure 3 900 is a flowchart illustrating an example of a method 300 for processing audio data according to certain aspects of the present disclosure. The method 300 may be performed by a computing device having an audio sensor, such as a mobile wireless communication device, a smart speaker, a camera, an XR device, a wireless-enabled vehicle, or another computing device. In an illustrative example, the computing system 900 may be configured to perform all or part of the method 300.
[0059] At block 302, a computing device (e.g., a smart speaker, a mobile communication device, etc.) is configured to obtain audio from a user. In one illustrative aspect, the computing device may be in a low power mode and configured to buffer the audio and then enter a higher power mode to determine whether the detected audio corresponds to a keyword. The computing device may include an analog-to-digital converter (ADC) to convert the received sound into first audio information. The computing device may also perform filtering to remove unnecessary information in the first audio information, such as noise and higher frequencies.
[0060] At block 304, the computing device detects a keyword in the audio provided by the user. For example, the computing device may include a predetermined model corresponding to the keyword and perform a comparison of the model with the audio to determine whether the keyword is detected within the first audio information. In some cases, the model may be at least partially trained during a training phase of the computing device.
[0061] At box 306, the computing device compares the first audio information to a speech model (e.g., speech model 216). As described above, the speech model is configured during training when the user reads content into the computing device and the computing device identifies speech patterns that are unique to the user. The comparison at box 306 produces a similarity or correlation that identifies the likelihood that the first audio information corresponds to the speech model.
[0062] At block 308, the computing device determines whether the similarity is greater than a first threshold. In some aspects, the first threshold is a value that has a high correlation with a relatively small amount of audio corresponding to the user's speech model. For example, a value may be empirically determined that indicates that even if additional audio information is obtained, the additional audio information may not significantly reduce the similarity. If the similarity is greater than or equal to the first threshold, the computing device may proceed to block 310.
[0063] At block 310, the computing device determines that the user (e.g., the speaker providing the voice input) corresponds to the user and authenticates the user. In some aspects, at block 310, a visual indication of user authentication may be output by the computing device. Referring back to block 308, if the similarity is less than the first threshold, the computing device may proceed to block 312.
[0064] At box 312, the computing device continues to obtain audio information, which is referred to as second audio information for clarity. In one aspect, at box 312, the computing device may detect the input of additional voice information and detect the start of a command or inquiry to the computing device. In addition, the computing device starts a timer in response to detecting a command or inquiry.
[0065] At box 314, the computing device identifies the second audio information from the audio information obtained based on the end of the command or query or the maximum duration of the timer. For example, if the maximum duration of the timer is 2 seconds, the computing device can extract the portion corresponding to the maximum duration of the timer from the obtained audio information. In another example, if the command or query ends before the maximum duration of the timer, the computing device can use the entire obtained portion of the audio information as the second audio information.
[0066] At block 316, the computing device compares the second audio information to a speech model (eg, speech model 216). In some aspects, the comparison at block 316 produces a second similarity or correlation that identifies a likelihood that the second audio information corresponds to the speech model.
[0067] At box 318, the computing device determines whether the second similarity is greater than a second threshold. In some aspects, the second threshold is less stringent than the first threshold because the duration of the second audio information is significantly longer than the first threshold, which can provide accurate determination at lower values. If the second similarity is greater than or equal to the second threshold, the computing device can proceed to box 310 to authenticate the user. However, if the second similarity is less than the second threshold, the computing device can proceed to box 320. At box 320, the computing device determines that the voice input does not correspond to the user and does not authenticate the user.
[0068] As described above, in response to authenticating the user, the computing device may enable authorized functions based on voice input. For example, in the illustrative example of a mobile communication device, the mobile communication device may authenticate the user to perform voice input, such as dialing a specific contact or sending a text message to the specific contact.
[0069] Although the foregoing aspects describe a single voice model for a single user, the foregoing aspects may include multiple voice models. For example, the smart speaker may be configured to authenticate different users, and different users may have different authorizations (e.g., access permissions).
[0070] Figure 4A and Figure 4B is a timing diagram illustrating different voice authentication scenarios according to aspects of the present disclosure;
[0071] Figure 4A A first example of user authentication performed by a computing device is illustrated. Figure 5 , Figure 6 and Figure 7 Various examples of computing devices are further described. Figure 4A , the computing device receives a first voice input 410 that begins at time 0 seconds and ends at time t1. In this illustrative aspect, the computing device compares the first voice input 410 with a keyword to determine that the first voice input 410 corresponds to the keyword. After identifying the keyword, the computing device then compares the first voice input 410 with a voice model associated with the user (e.g., voice model 216), and determines at time t2 that the similarity of the first voice input to the voice model (e.g., 95%) is less than a first threshold (e.g., 90%). Based on the comparison, the computing device then authenticates the user.
[0072] In this illustrative aspect, voice authentication using the first voice input significantly reduces latency.
[0073] Figure 4B A second example of user authentication performed by a computing device is illustrated. Figure 4B , the computing device receives a first voice input 420 that begins at time 0 seconds and ends at time t1. In this illustrative aspect, the computing device compares the first voice input 420 to the keyword to determine that the first voice input 420 corresponds to the keyword. After identifying the keyword, the computing device then compares the first voice input 420 to a voice model associated with the user (e.g., voice model 216), and determines that the similarity of the first voice input 420 to the voice model (e.g., 75%) is greater than a first threshold (e.g., 90%).
[0074] At time t2, the computing device begins receiving a second voice input corresponding to a command or query for the computing device (starting at time t2). In this aspect, the computing device may start a timer having a maximum duration configured to optimize the comparison of the voice input with the voice model. The computing device continues to obtain voice input, and at time t3, the computing device determines that the value of the timer corresponds to the maximum duration of the timer. The computing device may extract a portion 430 of the second voice input and compare the portion 430 of the second voice input with the voice model. At time t4, the computing device determines that the similarity (e.g., 83%) between the portion 430 of the second voice input and the voice model is greater than a second threshold value (e.g., 70%). In some aspects, the portion 430 of the second voice input may also include a first portion 420 for a second comparison.
[0075] exist Figure 4B In the example above, at time t4, the user continues to provide voice input. Figure 4B After the voice input is completed as shown, the computing device may then send a second portion 430 of the second voice input (or in some cases the entire second voice input) to the cloud service to perform speech recognition. If the user is not authenticated, the illustrative example will not perform any further processing of the voice input because the user is not authenticated and is therefore not authorized to perform any functions associated with the computing device.
[0076] Figure 5 is an illustration of a speaker device 500 including a voice input function for authenticating a speaker according to some aspects of the present disclosure. In some aspects, the speaker device 500 may include a voice assist function to enable voice input to provide convenient control of the speaker device 500. The speaker device 500 includes at least one audio capture device 502 disposed on a side of the speaker device 500. The speaker device 500 may also include an audio capture device 504 positioned on a top surface. In some aspects, the speaker device 500 may include a visual indicator 506 to provide a visual output to identify that the speaker device 500 is actively monitoring for voice input, such as a command or query. The visual indicator 506 may also provide a visual difference to indicate that the user is authenticated or not authenticated. For example, the visual indicator 506 may illuminate orange to indicate that the user is not authenticated, and may illuminate green to indicate that the user is authenticated. The speaker device 500 also includes at least one audio transducer 508 to output audio to the user, such as playing music or providing audio prompts. The speaker device may also include at least one port 510 for connecting to a power source or another computing device. Although Figure 6 The example of illustrates a USB Type-C port, but the port could be an analog stereo jack, or other analog or digital connector.
[0077] Figure 66 is an illustration of a mobile communication device 600 including voice input functionality for authenticating a speaker according to some aspects of the present disclosure. In some aspects, the mobile device includes a display 602 and a plurality of forward sensors 604. In one aspect, the forward sensors 604 may include an audio capture device. The mobile communication device 600 may also include audio capture devices at various locations, such as an audio capture device 606 located on a side and an audio capture device 608 located on a top surface. The mobile communication device 600 may include a plurality of audio capture devices to allow the mobile communication device 600 to be used in a hands-free mode and to allow a user to provide voice input. For example, the hands-free mode may be used when the user is driving and the graphical user interface should not be operated.
[0078] Figure 7 is an illustration of an unmanned ground vehicle 700 (such as an automated cleaning device) including voice input functionality for authenticating a speaker in accordance with some aspects of the present disclosure. In some aspects, the unmanned ground vehicle 700 performs visual simultaneous localization and mapping (VSLAM) to autonomously navigate an environment.
[0079] The unmanned ground vehicle 700 includes an image sensor 720 along the front surface of the ground vehicle 700. The unmanned ground vehicle 700 may also include a depth sensor 740. The ground vehicle 700 includes a plurality of wheels 715 along the bottom surface of the ground vehicle 700. The wheels 715 may be used as a means of transportation for the ground vehicle 700 and may be motorized using one or more motors. The motors, and therefore the wheels 715, may be actuated to move the unmanned ground vehicle 700 via a motion actuator. The ground vehicle may also include at least one audio capture device 740 for receiving voice input.
[0080] In some aspects, the unmanned ground vehicle 700 can be configured to use voice input for various purposes. In this illustrative aspect, the automated cleaning device can provide various information in response to the voice input, such as audibly outputting information related to the schedule. The systems and techniques disclosed herein can be used to improve the time when the unmanned ground vehicle 700 provides feedback based on authentication. For example, the unmanned ground vehicle 700 can include an LED indicator that provides a notification that the speaker is authenticated and is allowed to provide input to program various functions of the unmanned ground vehicle 700.
[0081] The various aspects disclosed herein are not limiting, and the system and technology can be applied to another mobile or fixed device. As an illustrative example, the system and technology can be applied to an automatic teller machine (ATM), an autonomous checkout system, an autonomous drone, etc.
[0082] Figure 8 800 is a flowchart illustrating an example of a method 800 for processing audio data according to certain aspects of the present disclosure. Methods 300 and 800 may be performed by a computing device (or a component of a computing device) having an audio capture device, such as a mobile wireless communication device, a smart speaker, a camera, an XR device, a wireless-enabled vehicle, or another computing device. In an illustrative example, computing system 900 may be configured to perform all or part of methods 300 and 800.
[0083] At box 802, the computing device may use an audio sensor of the user device to obtain first audio information from the user. The first audio information may include a keyword selected by the user. For example, during the setup of the voice function for the computing device, the computing device may receive a user input corresponding to a selection of text associated with the keyword in the user device. A person may change the keyword for various reasons, such as similarity that may confuse the name of the computing device. In other aspects, during the setup of the voice function for the computing device, the computing device may prompt the user to provide voice input to learn characteristics associated with the user's voice and train a model of the user's voice that can be used to uniquely identify the user. For example, the computing device may cause another computing device (e.g., a mobile phone) to display content for the user to read aloud. The computing device may also audibly prompt the user for various phrases.
[0084] At block 804, the computing device may determine whether the first audio information includes audio corresponding to a detected keyword (e.g., a previously detected keyword) that configures the user device to receive or process one or more commands from the user. As described above, the keyword may be selected by the user, and the training of the model of the user's speech may include repeating the keyword.
[0085] At block 806, the computing device may determine a similarity between the first audio information corresponding to the detected keyword and a model of the authenticated user based on the first audio information including the audio corresponding to the detected keyword. The similarity may be a correlation that identifies the likelihood that the first audio information is the user's voice.
[0086] At box 808, the computing device may determine whether to authenticate the user as the authenticated user based on a first comparison of the similarity between the first audio information and the model of the authenticated user and a first threshold. In some aspects, the first threshold is a high threshold and requires many features to match the model of the user's voice and provide a high confidence that is unlikely to degrade with additional input. In this case, the user can be authenticated using only the first audio information. The first threshold can be used in an environment without ambient noise that can affect the first comparison. For example, an environment with higher noise can prevent the first audio information from meeting the first threshold.
[0087] In some aspects, the computing device may not authenticate the user based on the first comparison to the first threshold. For example, if the user is watching a video and there is a significant amount of audio from the video, the noise floor may be increased and prevent the user from authenticating using a keyword. At this point, the computing device may output an audio indication or a visual indication that the keyword was detected. The user understands that the audio or visual indication indicates that the computing device is expecting audio information.
[0088] The computing device may then use the audio sensor of the user device to obtain the second audio information from the user. For example, the second audio information includes a command for the computing device to execute. In one exemplary aspect, the command may be a query that does not include keywords (e.g., what time, etc.). In one exemplary aspect, the computing device may start a timer for authenticating a portion of the second audio information, as further described below. In other aspects, the timer may be started based on the input of the first audio information.
[0089] In one aspect, upon obtaining the second audio information, the computing device may determine that the second audio information includes audio having the longest duration. For example, based on a timer, the computing device may determine that the amount of speech from the user (e.g., the second audio information) is equal to the longest duration. Based on this determination, the computing device may determine the similarity between the portion of the second audio information having the longest duration and the model of the authenticated user. After determining the similarity, the computing device may then determine whether to authenticate the user as an authenticated user based on a second comparison of the similarity between at least the portion of the second audio information and the model of the authenticated user and a second threshold different from the first threshold. In some cases, the similarity between at least a portion of the second audio information and the model of the authenticated user includes the similarity between the model of the authenticated user and a combination of at least a portion of the first audio information and the second audio information (e.g., a portion of the keywords and commands used in the second comparison). The first comparison and the second comparison may be part of the two-stage text-independent user verification process described above.
[0090] In some aspects, the second threshold is lower than the first threshold and provides sufficient determination based on analyzing a longer speech portion than the first portion of the speech. In an illustrative example, the maximum duration can be 3 seconds. By limiting the amount of speech (e.g., part of the second audio information), the computing device can attempt to authenticate the user in parallel with receiving the second audio information and reduce the delay of voice input. Although this example describes the first part of the second audio information for authentication, other examples are also possible. For example, the second part of the second audio information can be determined based on noise information to authenticate other parts of the second audio information.
[0091] In some other aspects, the computing device may determine that the second audio information is received before the maximum duration of the timer. For example, for a simple query (e.g., what time), the input of the second audio information may end before the maximum duration of the timer. The computing system may determine the similarity between the second audio information and the model of the authenticated user based on the fact that the user is not authenticated based on the first audio information. The computing system may then determine whether to authenticate the user as an authenticated user based on a third comparison of the similarity between at least the second audio information and the model of the authenticated user with a second threshold different from the first threshold. In some cases, the similarity between at least the second audio information and the model of the authenticated user includes the similarity between the model of the authenticated user and a combination of the first audio information and the second audio information (e.g., using keywords and commands in the third comparison). As described above, the first comparison and the second comparison can be part of the above-mentioned two-stage text-independent user verification process.
[0092] After completing the input of the second audio information, the computing system can provide the second audio information to the audio processing system for processing based on whether the user is authenticated as an authenticated user. For example, the audio processing system can convert the second audio information into machine-readable information such as text. Other forms of machine-readable information include extensible markup language (XML) or JavaScript object notation (JSON), which can identify metadata, such as uncertainty of words, tone information, etc. In some aspects, the audio processing system can perform functions such as natural language processing (NLP) functions to eliminate the ambiguity of commands and respond to the command. An example of NLP functions is named entity recognition (NER) and the identification entity is a word with a specific meaning, such as a name or a city name. Based on the further processing of the entity and machine-readable information in NLP, the audio processing system can try to understand the command in the second audio information and form a reply. The computing system can receive a reply from the audio processing system and then provide a response. For example, the computing system can provide local time, weather or other information related to the second audio information audibly.
[0093] In some examples, the processes described herein (e.g., methods 300 and 800 and / or other processes described herein) can be performed by a computing device or apparatus. In one example, methods 300 and 800 can be performed by a computer having Fig. 9 The computing architecture of the computing system 900 shown in FIG. Figure 2 , and the image capture and voice input device 200).
[0094] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, an AR glasses, a network-connected watch or smart watch or other wearable device), a server computer, an autonomous vehicle or a computing device of an autonomous vehicle, a robotic device, a television, and / or any other computing device having resource capabilities for performing the methods described herein, including methods 300 and 800. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the methods described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive IP-based data or other types of data.
[0095] The components of the computing device may be implemented in circuits. For example, each component may include and / or may be implemented using electronic circuits or other electronic hardware (which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits)), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.
[0096] Methods 300 and 800 are illustrated as logic flow diagrams, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally speaking, computer executable instructions include routines, programs, objects, components, and data structures that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the method.
[0097] Methods 300 and 800 and / or other methods or processes described herein may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors, implemented by hardware, or implemented by a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program that includes multiple instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0098] Fig. 9 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. Specifically, Fig. 9 An example of a computing system 900 is illustrated, which may be any computing device, for example, constituting an internal computing system, a remote computing system, a camera, or any components thereof, wherein the components of the system communicate with each other using a connection 905. The connection 905 may be a physical connection using a bus, or a direct connection into the processor 910, such as in a chipset architecture. The connection 905 may also be a virtual connection, a networked connection, or a logical connection.
[0099] In some aspects, computing system 900 is a distributed system, where the functionality described in the present disclosure can be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent a number of such components, each of which performs a portion or all of the functionality of the described component. In some aspects, each component can be a physical or virtual device.
[0100] The example computing system 900 includes at least one processing unit (CPU or processor) 910 and connections 905 that couple various system components including system memory 915 such as ROM 920 and RAM 925 to the processor 910. The computing system 900 may include a cache 912 of high-speed memory directly connected to the processor 910, close to the processor, or integrated as part of the processor.
[0101] Processor 910 may include any general purpose processor and hardware or software services, such as services 932, 934, and 936 stored in storage device 930, that are configured to control processor 910 as well as a dedicated processor where software instructions are incorporated into the actual processor design. Processor 910 may essentially be a completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0102] To enable user interaction, the computing system 900 includes an input device 945 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, and the like. The computing system 900 may also include an output device 935 that can be one or more of a plurality of output mechanisms. In some cases, a multimodal system may enable a user to provide multiple types of input / output to communicate with the computing system 900. The computing system 900 may include a communication interface 940, which may generally govern and manage user input and system output. The communication interface may perform or facilitate receiving and / or sending wired or wireless communications using a wired and / or wireless transceiver, including utilizing an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, Ports / plugs, Ethernet ports / plugs, Fiber optic ports / plugs, Dedicated wired ports / plugs, Wireless signal transmission, BLE wireless signal transmission, The communication interface 940 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of the computing system 900 based on receiving one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' GPS, Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There is no restriction to operating on any particular hardware arrangement, and thus the base features herein may be easily substituted for improved hardware or firmware arrangements as they are developed.
[0103] The storage device 930 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other type of computer-readable medium capable of storing data that can be accessed by a computer, such as a cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a cassette, a floppy disk, a floppy disk, a hard disk, a magnetic tape, a magnetic stripe / magnetic stripe, any other magnetic storage medium, flash memory, a memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) disc, a rewritable compact disc (CD) disc, a digital video disc (DVD) disc, a Blu-ray disc (BDD) disc, a holographic disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card chip, a Europay, an EMV chip, a Subscriber Identity Module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a RAM, a static RAM (SRAM), a dynamic RAM (DRAM), a ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASH EPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin transfer torque RAM (STT-RAM), another memory chip or box, and / or a combination thereof.
[0104] The storage device 930 may include software services, servers, services, etc., which, when the code defining such software is executed by the processor 910, causes the system to perform a function. In some aspects, a hardware service that performs a specific function may include a software component for performing a function stored in a computer-readable medium connected to necessary hardware components (such as the processor 910, connection 905, output device 935, etc.). The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. A computer-readable medium may include a non-transient medium that can store data and does not include carrier waves and / or transient electronic signals that are propagated wirelessly or on a wired connection. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media (such as CDs or DVDs), flash memory, memory, or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a process, function, subroutine, program, routine, subroutine, module, software package, category, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network sending, etc.
[0105] In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, one or more network interfaces configured to communicate and / or receive data, any combination thereof, and / or other components. The one or more network interfaces may be configured to communicate and / or receive wired and / or wireless data, including data in accordance with 3G, 4G, 5G, and / or other cellular standards, data in accordance with the Wi-Fi (802.11x) standard, data in accordance with the Bluetooth standard, and / or other components. TM Standard data, data according to IP standards and / or other types of data.
[0106] The components of the computing device may be implemented in circuits. For example, each component may include and / or may be implemented using electronic circuits or other electronic hardware (which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits)), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.
[0107] In some aspects, computer readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.
[0108] Specific details are provided in the above description to provide a detailed understanding of the various aspects and each example provided herein. However, it will be understood by those skilled in the art that these aspects can also be practiced without these specific details. For clarity, in some cases, the present technology may be presented as including separate functional blocks, including functional blocks comprising devices, device components, steps or routines in the method embodied in a combination of software or hardware and software. Additional components other than those components shown in the accompanying drawings and / or described herein may be used. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid confusing these aspects in unnecessary details. In other cases, known circuits, processes, algorithms, structures and techniques may be shown without unnecessary details to avoid confusing various aspects.
[0109] Various aspects may be described above as a process or method, which is depicted as a flow chart, a flow diagram, a data flow diagram, a structure diagram or a block diagram. Although a flow chart may describe an operation as a sequential process, many operations in the operation may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. The process is terminated when the operation of the process is completed, but the process may have additional steps not included in the accompanying drawings. The process may correspond to a method, a function, a process, a subroutine, a subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to a calling function or a main function.
[0110] The processes and methods according to the above examples can be implemented using stored computer executable instructions or computer executable instructions otherwise obtained from computer readable media. Such instructions may include, for example, instructions and data that enable or configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions in other ways. Parts of the computer resources used can be accessed through a network. Computer executable instructions can be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer readable media that can be used to store instructions, information used, and / or information created during the method according to the described examples include disks or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0111] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or plug-in cards. By way of additional examples, such functionality may also be implemented on circuit boards between different chips or different processes executed on a single device.
[0112] The instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0113] In the above description, various aspects of the present application are described with reference to their specific aspects, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the exemplary aspects of the present application have been described in detail herein, it is to be understood that each inventive concept can be implemented and adopted in various other ways, and the appended claims are not intended to be interpreted as including these variations, unless limited by the prior art. Various features and aspects of the above-mentioned applications can be used individually or in combination. In addition, without departing from the broader essence and scope of this specification, various aspects can be utilized in any number of environments and applications beyond those described herein. Therefore, the description and the accompanying drawings should be considered as illustrative rather than restrictive. For the purpose of illustration, each method is described in a specific order. It should be understood that, in alternative aspects, each method can be performed in a different order than described.
[0114] It should be understood by those of ordinary skill in the art that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of the present specification.
[0115] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0116] The phrase "coupled to" means that any component is directly or indirectly physically connected to another component, and / or any component is directly or indirectly in communication with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0117] Claim language or other language stating "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. For example, claim language stating "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0118] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the various aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. The technician may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be interpreted as departing from the scope of the present application.
[0119] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication device handsets, or integrated circuit devices with multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques may be implemented at least in part by a computer-readable data storage medium including a program code, the program code including instructions that, when executed, perform one or more of the above methods. A computer-readable data storage medium may form part of a computer program product, which may include packaging materials. A computer-readable medium may include a memory or data storage medium, such as a RAM (such as a synchronous dynamic random access memory (SDRAM)), a ROM, a non-volatile random access memory (NVRAM), an EEPROM, a flash memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0120] The program code may be executed by a processor, which may include one or more processors, such as one or more DSPs, general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the aforementioned structures, any combination of the aforementioned structures, or any other structure or device suitable for implementing the techniques described herein.
[0121] Illustrative aspects of the present disclosure include:
[0122] Aspect 1. A method for processing audio, the method comprising: using an audio sensor of a user device to obtain first audio information from a user; determining whether the first audio information includes audio corresponding to a detected keyword, the detected keyword configuring the user device to receive or process one or more commands from the user; based on the first audio information including the audio corresponding to the detected keyword, determining the similarity between the first audio information corresponding to the detected keyword and a model of an authenticated user; and determining whether to authenticate the user as the authenticated user based on a first comparison of the similarity between the first audio information and the model of the authenticated user with a first threshold.
[0123] Aspect 2. The method according to Aspect 1 further includes: using the audio sensor of the user device to obtain second audio information from the user; and providing the second audio information to an audio processing system for processing based on whether the user is authenticated as the authenticated user.
[0124] Aspect 3. The method according to aspect 2, wherein the second audio information includes a command.
[0125] Aspect 4. The method according to aspect 3, wherein the command does not include the keyword.
[0126] Aspect 5. According to the method described in any one of Aspects 2 to 4, the method further includes: determining the similarity between the second audio information and the model of the authenticated user based on the fact that the user has not been authenticated based on the first audio information; and determining whether to authenticate the user as the authenticated user based on a second comparison of the similarity between at least the second audio information and the model of the authenticated user with a second threshold different from the first threshold.
[0127] Aspect 6. A method according to Aspect 5, wherein the similarity between at least the second audio information and the model of the authenticated user includes the similarity between the model of the authenticated user and a combination of the first audio information and the second audio information.
[0128] Aspect 7. A method according to any one of Aspects 5 or 6, wherein the first comparison and the second comparison are part of a two-stage text-independent user verification process.
[0129] Aspect 8. According to the method described in any one of Aspects 2 to 4, the method further includes: when obtaining the second audio information, determining that the second audio information includes audio with the longest duration; determining the similarity between the part of the second audio information with the longest duration and the model of the authenticated user; and determining whether to authenticate the user as the authenticated user based on a second comparison between the similarity between at least the part of the second audio information and the model of the authenticated user and a second threshold different from the first threshold.
[0130] Aspect 9. The method according to aspect 8, further comprising determining, based on a timer, that the second audio information includes the audio having the longest duration.
[0131] Aspect 10. The method of any one of aspects 1 to 9, wherein the model of the authenticated user is based on speech including detected keywords from the authenticated user.
[0132] Aspect 11. The method according to any one of aspects 1 to 10, further comprising receiving, in the user device, a user input corresponding to a selection of text associated with the detected keyword.
[0133] Aspect 12. A device for processing audio. The device includes a memory (e.g., implemented in a circuit) and a processor (or multiple processors) coupled to the memory. The processor is configured to: use an audio sensor of a user device to obtain first audio information from a user; determine whether the first audio information includes audio corresponding to a detected keyword, and the detected keyword configures the user device to receive or process one or more commands from the user; based on the first audio information including the audio corresponding to the detected keyword, determine the similarity between the first audio information corresponding to the detected keyword and a model of an authenticated user; and determine whether to authenticate the user as the authenticated user based on a first comparison of the similarity between the first audio information and the model of the authenticated user with a first threshold.
[0134] Aspect 13. An apparatus according to Aspect 12, wherein the processor is configured to: use the audio sensor of the user device to obtain second audio information from the user; and provide the second audio information to an audio processing system for processing based on whether the user is authenticated as the authenticated user.
[0135] Clause 14. The apparatus according to clause 13, wherein the second audio information comprises a command.
[0136] Aspect 15. The apparatus according to aspect 14, wherein the command does not include the keyword.
[0137] Aspect 16. An apparatus according to any one of Aspects 13 to 15, wherein the at least one apparatus is configured to: determine the similarity between the second audio information and the model of the authenticated user based on the fact that the user is not authenticated based on the first audio information; and determine whether to authenticate the user as the authenticated user based on a second comparison of the similarity between at least the second audio information and the model of the authenticated user with a second threshold different from the first threshold.
[0138] Aspect 17. An apparatus according to Aspect 16, wherein the similarity between at least the second audio information and the model of the authenticated user includes a similarity between the model of the authenticated user and a combination of the first audio information and the second audio information.
[0139] Aspect 18. An apparatus according to any one of Aspects 16 or 17, wherein the first comparison and the second comparison are part of a two-stage text-independent user authentication process.
[0140] Aspect 19. An apparatus according to any one of Aspects 13 to 15, wherein the processor is configured to: when obtaining the second audio information, determine that the second audio information includes audio with a longest duration; determine the similarity between the portion of the second audio information with the longest duration and the model of the authenticated user; and determine whether to authenticate the user as the authenticated user based on a second comparison of the similarity between at least the portion of the second audio information and the model of the authenticated user with a second threshold different from the first threshold.
[0141] Aspect 20. The apparatus according to aspect 19, wherein the processor is configured to: determine that the second audio information includes the audio having the longest duration based on a timer.
[0142] Aspect 21. An apparatus according to any one of aspects 12 to 20, wherein the model of the authenticated user is based on speech including detected keywords from the authenticated user.
[0143] Aspect 22. An apparatus according to any one of aspects 12 to 21, wherein the processor is configured to: receive, in the user device, a user input corresponding to a selection of text associated with the detected keyword.
[0144] Aspect 23. An apparatus according to any one of aspects 12 to 22, wherein the apparatus is the user equipment.
[0145] Aspect 24: A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the operations of any one of aspects 1 to 11.
[0146] Aspect 25: An apparatus comprising means for performing the operations according to any one of aspects 1 to 11.
Claims
1. A method for processing audio, the method comprising: obtaining first audio information from a user using an audio sensor of a user device; determining whether the first audio information includes audio corresponding to a detected keyword, the detected keyword configuring the user device to receive or process one or more commands from the user; determining, based on the first audio information including the audio corresponding to the detected keyword, a similarity between the first audio information corresponding to the detected keyword and a model of the authenticated user; as well as Whether to authenticate the user as the authenticated user is determined based on a first comparison of the similarity between the first audio information and the model of the authenticated user and a first threshold.
2. The method according to claim 1, further comprising: obtaining second audio information from the user using the audio sensor of the user device; as well as The second audio information is provided to an audio processing system for processing based on whether the user is authenticated as the authenticated user. The method of claim 2 , wherein the second audio information comprises a command. The method according to claim 3 , wherein the command does not include the keyword.
5. The method according to any one of claims 2 to 4, further comprising: determining a similarity between at least the second audio information and the model of the authenticated user based on the user not being authenticated based on the first audio information; as well as Whether to authenticate the user as the authenticated user is determined based on a second comparison of the similarity between at least the second audio information and the model of the authenticated user with a second threshold different from the first threshold.
6. The method of claim 5, wherein the similarity between at least the second audio information and the model of the authenticated user comprises a similarity between the model of the authenticated user and a combination of the first audio information and the second audio information.
7. A method according to any one of claims 5 or 6, wherein the first comparison and the second comparison are part of a two-stage text-independent user authentication process.
8. The method according to any one of claims 2 to 4, further comprising: When obtaining the second audio information, determining that the second audio information includes audio having the longest duration; determining a similarity between the portion of the second audio information having the longest duration and the model of the authenticated user; as well as Whether to authenticate the user as the authenticated user is determined based on a second comparison of the similarity between at least the portion of the second audio information and the model of the authenticated user to a second threshold different from the first threshold. 9 . The method of claim 8 , further comprising determining, based on a timer, that the second audio information includes the audio having the longest duration.
10. The method of any one of claims 1 to 9, wherein the model of the authenticated user is based on speech including detected keywords from the authenticated user.
11. The method of any one of claims 1 to 10, further comprising receiving, in the user device, a user input corresponding to a selection of text associated with the detected keyword.
12. A device for processing audio, the device comprising: at least one memory; and at least one processor coupled to at least one memory and configured to: obtaining first audio information from a user using an audio sensor of a user device; determining whether the first audio information includes audio corresponding to a detected keyword, the detected keyword configuring the user device to receive or process one or more commands from the user; determining, based on the first audio information including the audio corresponding to the detected keyword, a similarity between the first audio information corresponding to the detected keyword and a model of the authenticated user; as well as Whether to authenticate the user as the authenticated user is determined based on a first comparison of the similarity between the first audio information and the model of the authenticated user and a first threshold.
13. The apparatus of claim 12, wherein the at least one processor is configured to: obtaining second audio information from the user using the audio sensor of the user device; and The second audio information is provided to an audio processing system for processing based on whether the user is authenticated as the authenticated user. The apparatus of claim 13 , wherein the second audio information comprises a command. The apparatus according to claim 14 , wherein the command does not include the keyword.
16. The apparatus according to any one of claims 13 to 15, wherein the at least one processor is configured to: determining a similarity between the at least second audio information and the model of the authenticated user based on the user not being authenticated based on the first audio information; and Whether to authenticate the user as the authenticated user is determined based on a second comparison of the similarity between at least the second audio information and the model of the authenticated user with a second threshold different from the first threshold.
17. The apparatus of claim 16, wherein the similarity between at least the second audio information and the model of the authenticated user comprises a similarity between the model of the authenticated user and a combination of the first audio information and the second audio information.
18. An apparatus according to any one of claims 16 or 17, wherein the first comparison and the second comparison are part of a two-stage text-independent user authentication process.
19. The apparatus according to any one of claims 13 to 15, wherein the at least one processor is configured to: When obtaining the second audio information, determining that the second audio information includes audio having the longest duration; determining a similarity between the portion of the second audio information having the longest duration and the model of the authenticated user; as well as Whether to authenticate the user as the authenticated user is determined based on a second comparison of the similarity between at least the portion of the second audio information and the model of the authenticated user to a second threshold different from the first threshold.
20. The apparatus of claim 19, wherein the at least one processor is configured to determine, based on a timer, that the second audio information includes the audio having the longest duration.
21. An apparatus according to any one of claims 12 to 20, wherein the model of the authenticated user is based on speech including detected keywords from the authenticated user.
22. An apparatus according to any one of claims 12 to 21, wherein the at least one processor is configured to: receive, in the user device, a user input corresponding to a selection of text associated with the detected keyword.
23. The apparatus according to any one of claims 12 to 22, wherein the apparatus is the user equipment.
24. A non-transitory computer readable medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to perform the following operations: obtaining first audio information from a user using an audio sensor of a user device; determining whether the first audio information includes audio corresponding to a detected keyword, the detected keyword configuring the user device to receive or process one or more commands from the user; determining, based on the first audio information including the audio corresponding to the detected keyword, a similarity between the first audio information corresponding to the detected keyword and a model of the authenticated user; as well as Whether to authenticate the user as the authenticated user is determined based on a first comparison of the similarity between at least the first audio information and the model of the authenticated user with a first threshold.
25. The non-transitory computer readable medium of claim 24, wherein the instructions, when executed by the one or more processors, cause the one or more processors to: obtaining second audio information from the user using the audio sensor of the user device; and The second audio information is provided to an audio processing system for processing based on whether the user is authenticated as the authenticated user.
26. The non-transitory computer-readable medium of claim 25, wherein the second audio information comprises a command.
27. The non-transitory computer readable medium of claim 26, wherein the command does not include the keyword.
28. The non-transitory computer readable medium of any one of claims 25 to 27, wherein the instructions, when executed by the one or more processors, cause the one or more processors to: Based on the user not being authenticated based on the first audio information, determining a similarity between the second audio information and the model of the authenticated user; and Whether to authenticate the user as the authenticated user is determined based on a second comparison of the similarity between at least the second audio information and the model of the authenticated user with a second threshold different from the first threshold.
29. The non-transitory computer-readable medium of claim 28, wherein the similarity between at least the second audio information and the model of the authenticated user comprises a similarity between the model of the authenticated user and a combination of the first audio information and the second audio information.
30. The non-transitory computer readable medium of any one of claims 28 or 29, wherein the first comparison and the second comparison are part of a two-stage text-independent user authentication process.