User voice profile management

By combining the segmenter and profile manager, a machine learning model is used to distinguish speaker audio segments without requiring active registration, solving the time-consuming problem in user voice profile management and achieving efficient speaker differentiation and profile management.

CN116583899BActive Publication Date: 2026-08-04QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2021-09-28
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing technologies, the active registration process for user voice profiles is time-consuming and inconvenient, and it is difficult to efficiently distinguish and manage multiple speakers.

Method used

By combining a segmenter and a profile manager, a machine learning segmentation model is used to distinguish speaker-homogeneous audio segments in an audio stream without requiring active user registration, and user voice profiles are generated or updated based on audio feature data.

Benefits of technology

It enables efficient differentiation and management of user voice profiles among multiple speakers, reducing training time and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116583899B_ABST
    Figure CN116583899B_ABST
Patent Text Reader

Abstract

An apparatus includes a processor configured to determine, in a first power mode, whether an audio stream corresponds to speech of at least two speakers. The processor is configured to analyze, in a second power mode, audio feature data of the audio stream to generate segmentation results based on determining that the audio stream corresponds to speech of at least two speakers. The processor is configured to perform a comparison of a plurality of user voice profiles to audio feature data sets of a plurality of audio feature data sets of a speaker homogenous audio segment to determine whether the audio feature data sets match any of the user voice profiles. The processor is configured to generate a user voice profile based on the plurality of audio feature data sets based on determining that the audio feature data sets do not match any of the plurality of user voice profiles.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to jointly owned U.S. nonprovisional patent application No. 17 / 115,158, filed on December 8, 2020, the contents of which are expressly incorporated herein by reference in their entirety. Technical Field

[0003] In summary, this disclosure relates to the management of user voice profiles. Background Technology

[0004] Technological advancements have led to smaller and more powerful computing devices. For example, a variety of portable personal computing devices exist today, including cordless phones (e.g., mobile phones and smartphones), small, lightweight tablets, and laptops that are easily carried by the user. These devices can transmit voice and data packets over wireless networks. Furthermore, many of these devices incorporate additional functionalities, such as digital cameras, digital camcorders, digital recorders, and audio file players. Moreover, these devices can process executable instructions, including software applications such as web browser applications that can be used to access the internet. Therefore, these devices can include significant computing power.

[0005] Such computing devices typically incorporate functionality for receiving audio signals from one or more microphones. For example, the audio signal could represent user speech captured by the microphone, external sounds captured by the microphone, or a combination thereof. Such devices can include applications that rely on user voice profiles, such as for user recognition. User voice profiles can be trained by having users speak pre-defined words or sentences using a script. This proactive user registration to generate user voice profiles is time-consuming and inconvenient. Summary of the Invention

[0006] According to one implementation of this disclosure, an apparatus for audio analysis includes a memory and one or more processors. The memory is configured to store multiple user voice profiles of multiple users. The one or more processors are configured to determine, in a first power mode, whether an audio stream corresponds to the speech of at least two different speakers. The one or more processors are further configured to, based on the determination that the audio stream corresponds to the speech of at least two different speakers, analyze audio feature data of the audio stream in a second power mode to generate segmentation results. The segmentation results indicate talker-homogeneous audio segments of the audio stream. The one or more processors are further configured to perform a comparison of a first set of audio feature data from the multiple user voice profiles with the first set of talker-homogeneous audio segments to determine whether the first set of audio feature data matches any of the multiple user voice profiles. The one or more processors are also configured to, based on the determination that the first set of audio feature data does not match any of the multiple user voice profiles, generate a first user voice profile based on the first set of audio feature data and add the first user voice profile to the multiple user voice profiles.

[0007] According to another implementation of this disclosure, an audio analysis method includes: determining, while the device is in a first power mode, whether an audio stream corresponds to the speech of at least two different speakers. The method further includes: based on the determination that the audio stream corresponds to the speech of at least two different speakers, analyzing audio feature data of the audio stream in a second power mode to generate segmentation results. The segmentation results indicate speaker-homogeneous audio segments of the audio stream. The method further includes: performing a comparison at the device of a first set of audio feature data from a first plurality of audio feature data sets of a plurality of user voice profiles with the first speaker-homogeneous audio segments to determine whether the first set of audio feature data matches any of the plurality of user voice profiles. The method further includes: based on the determination that the first set of audio feature data does not match any of the plurality of user voice profiles: generating a first user voice profile at the device based on the first plurality of audio feature data sets, and adding the first user voice profile to the plurality of user voice profiles at the device.

[0008] According to another implementation of this disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to determine in a first power mode whether an audio stream corresponds to the speech of at least two different speakers. The instructions, when executed by the one or more processors, further cause the processors to analyze audio feature data of the audio stream in a second power mode to generate segmentation results based on the determination that the audio stream corresponds to the speech of at least two different speakers. The segmentation results indicate speaker-homogeneous audio segments of the audio stream. The instructions, when executed by the one or more processors, further cause the one or more processors to perform a comparison of a first set of audio feature data from a first plurality of audio feature data sets of a plurality of user speech profiles with the first speaker-homogeneous audio segments to determine whether the first set of audio feature data matches any of the plurality of user speech profiles. The instructions, when executed by the one or more processors, further cause the one or more processors to generate a first user speech profile based on the first plurality of audio feature data sets and add the first user speech profile to the plurality of user speech profiles based on the determination that the first set of audio feature data does not match any of the plurality of user speech profiles.

[0009] According to another implementation of this disclosure, an apparatus includes units for storing multiple user voice profiles of multiple users. The apparatus further includes units for determining, in a first power mode, whether an audio stream corresponds to the speech of at least two different speakers. The apparatus also includes units for analyzing audio feature data of the audio stream in a second power mode to generate segmentation results. The audio feature data is analyzed in the second power mode based on the determination that the audio stream corresponds to the speech of at least two different speakers. The segmentation results indicate speaker-homogeneous audio segments of the audio stream. The apparatus further includes units for performing a comparison of a first set of audio feature data from the multiple user voice profiles with a first set of speaker-homogeneous audio segments to determine whether the first set of audio feature data matches any of the multiple user voice profiles. The apparatus also includes units for generating a first user voice profile based on the first set of audio feature data. The first user voice profile is generated based on the determination that the first set of audio feature data does not match any of the multiple user voice profiles. The apparatus further includes units for adding the first user voice profile to the multiple user voice profiles.

[0010] Other aspects, advantages, and features of this disclosure will become apparent upon examination of the entire application, which includes the following parts: description of the drawings, detailed description, and claims. Attached Figure Description

[0011] Figure 1This is a block diagram illustrating a specific example of user voice profile management based on some examples of this disclosure.

[0012] Figure 2A These are schematic diagrams illustrating specific aspects of a system operable to perform user voice profile management, based on some examples of the contents of this disclosure.

[0013] Figure 2B These are some examples based on the content of this disclosure. Figure 2A A schematic diagram illustrating the components of the system.

[0014] Figure 3 These are illustrative diagrams illustrating aspects of operations associated with user voice profile management, based on some examples of this disclosure.

[0015] Figure 4 These are illustrative diagrams illustrating aspects of operations associated with user voice profile management, based on some examples of this disclosure.

[0016] Figure 5 These are illustrative diagrams illustrating aspects of operations associated with user voice profile management, based on some examples of this disclosure.

[0017] Figure 6 These are illustrative diagrams illustrating aspects of operations associated with user voice profile management, based on some examples of this disclosure.

[0018] Figure 7 These are illustrative diagrams illustrating aspects of operations associated with user voice profile management, based on some examples of this disclosure.

[0019] Figure 8 These are illustrative diagrams illustrating aspects of operations associated with user voice profile management, based on some examples of this disclosure.

[0020] Figure 9 These are illustrative diagrams illustrating aspects of operations associated with user voice profile management, based on some examples of this disclosure.

[0021] Figure 10 These are some examples that can be derived from this disclosure. Figure 2A A schematic diagram illustrating a specific implementation of the user voice profile management method executed by the system.

[0022] Figure 11 Examples of integrated circuits operable to perform user voice profile management are shown, based on some examples of this disclosure.

[0023] Figure 12These are schematic diagrams of mobile devices operable to perform user voice profile management, based on some examples of this disclosure.

[0024] Figure 13 This is a schematic diagram of an earphone operable to perform user voice profile management, based on some examples of this disclosure.

[0025] Figure 14 These are schematic diagrams illustrating, based on some examples of this disclosure, a wearable electronic device operable to perform user voice profile management.

[0026] Figure 15 This is a schematic diagram of a voice-controlled speaker system operable to perform user voice profile management, based on some examples of the contents of this disclosure.

[0027] Figure 16 These are schematic diagrams illustrating, based on some examples of this disclosure, a virtual reality or augmented reality headset operable to perform user voice profile management.

[0028] Figure 17 This is a schematic diagram of a first example of a vehicle operable to perform user voice profile management, based on some examples of the present disclosure.

[0029] Figure 18 This is a schematic diagram of a second example of a vehicle operable to perform user voice profile management, based on some examples of the present disclosure.

[0030] Figure 19 This is a block diagram of a specific illustrative example of a device operable to perform user voice profile management, based on some examples of the contents of this disclosure. Detailed Implementation

[0031] Training user voice profiles using active user registration (where users speak a predetermined set of words or sentences) can be time-consuming and inconvenient. For example, users must plan ahead and spend time training their voice profiles. The user voice profile management system and method disclosed herein enable differentiation among multiple speakers without using active user registration. For example, an audio stream corresponding to the speech of one or more users is received by a segmenter. The segmenter generates segmentation results that indicate speaker-homogeneous audio segments of the audio stream. As used herein, a “speaker-homogeneous audio segment” includes audio portions (e.g., audio frames) representing the speech of the same speaker. For example, the segmentation results identify a set of audio frames representing the speech of the same speaker. A profile manager compares the audio features of the audio frames in the set of audio frames to determine whether the audio features match any of a plurality of stored user voice profiles. In response to determining that the audio features do not match any of the stored user voice profiles, the profile manager generates a user voice profile based at least in part on the audio features. Alternatively, in response to determining that audio features match a stored user voice profile, the profile manager updates the stored user voice profile at least in part based on the audio features. Thus, passive registration can be used to generate or update user voice profiles, for example, during a telephone call or conference. The profile manager can also generate or update multiple user voice profiles during a conversation between multiple speakers. In a particular example, the profile manager provides profile identifiers of the generated or updated voice profiles to one or more additional audio applications. For example, an audio application could perform speech-to-text conversion on an audio stream to generate a transcript with tags indicating the speaker for the corresponding text.

[0032] Specific aspects of this disclosure are described below with reference to the accompanying drawings. In the description, common features are designated throughout the drawings by common reference numerals. In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically distinct, the same reference numerals are used for each feature, and different instances are distinguished by adding letters to the reference numerals. When features are referred to herein as a group or a type (e.g., when a specific one of the features is not referenced), reference numerals are used without distinguishing letters. However, when a specific feature of the same type is referred to herein, reference numerals are used with distinguishing letters. For example, referencing… Figure 1 Multiple frames are shown and associated with reference numerals 102A, 102B, and 102C. When referring to a specific one of these frames (such as frame 102A), the distinguishing letter "A" is used. However, when referring to any one of these frames or to these frames as a group, reference numeral 102 is used without the distinguishing letter.

[0033] As used herein, various terms are used only for the purpose of describing a particular implementation and are not intended to restrict the implementation. For example, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context explicitly indicates otherwise. Furthermore, some features described herein are singular in some implementations and plural in others. For example, Figure 2A It describes a system that includes one or more processors ( Figure 2A The device 202 (referred to as "processor 220") indicates that in some implementations, device 202 includes a single processor 220, while in other implementations, device 202 includes multiple processors 220. For ease of reference, such features are generally described as "one or more" features and are referred to thereafter in the singular unless an aspect relating to multiple features is being described.

[0034] As used herein, the terms “comprise,” “comprises,” and “comprising” are used interchangeably with “include,” “includes,” or “including.” Furthermore, the term “wherein” is used interchangeably with “where.” As used herein, “exemplary” indicates an example, implementation, and / or aspect, and should not be construed as restrictive or indicating a preference or preferred implementation. As used herein, ordinal terms used to modify elements (e.g., structure, component, operation, etc.) (e.g., “first,” “second,” “third,” etc.) do not themselves indicate any priority or order of that element relative to another element, but merely distinguish that element from another element with the same name (but using ordinal terms). As used herein, the term “set” refers to one or more specific elements, while the term “multiple” refers to multiple (e.g., two or more) specific elements.

[0035] As used herein, “coupling” can include “communicationally coupled,” “electrically coupled,” or “physically coupled,” and may (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof). As an illustrative, non-limiting example, two electrically coupled devices (or components) may be included in the same device or different devices and may be connected via electronics, one or more connectors, or inductive coupling. In some implementations, two communicationally coupled (e.g., electrically communicating) devices (or components) may directly or indirectly transmit and receive signals (e.g., digital or analog signals) via one or more wires, buses, networks, etc. As used herein, “directly coupled” can include two devices coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) without an intervening component.

[0036] In this disclosure, terms such as “determine,” “calculate,” “estimate,” “shift,” and “adjust” can be used to describe how one or more operations are performed. It should be noted that these terms should not be construed as restrictive, and similar operations can be performed using other techniques. Furthermore, as mentioned herein, “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” are used interchangeably. For example, “generating,” “calculating,” “estimate,” or “determining” a parameter (or signal) can refer to actively generating, estimating, calculating, or determining the parameter (or signal), or it can refer to using, selecting, or accessing a parameter (or signal) that has already been generated (e.g., by another component or device).

[0037] Figure 1 Example 100 of user voice profile management is shown. In example 100, segmenter 124 and profile manager 126 cooperate to process audio stream 141 to distinguish voices from multiple speakers without using active user registration by the speakers.

[0038] Audio stream 141 comprises multiple discrete parts, which in Figure 1 The frames are represented as frames 102A, 102B, and 102C. In this example, each frame 102 represents or encodes a portion of the audio from audio stream 141. For example, each frame 102 could represent half a second of audio from the audio stream. In other examples, frames of different sizes or durations could be used.

[0039] Audio stream 141 is provided as input to segmenter 124. Segmenter 124 is configured to segment audio stream 141 and identify each segment as containing speech from a single speaker, speech from multiple speakers, or silence. For example, in Figure 1 In this process, segmenter 124 has identified a first set of audio portions 151A, which together form a speaker-homogeneous audio segment 111A. Similarly, segmenter 124 has identified a second set of audio portions 151C, which together form a second speaker-homogeneous audio segment 111B. Segmenter 124 also identifies a set of audio portions 151B, which together form a silent or mixed speaker audio segment 113. The silent or mixed speaker audio segment 113 represents speech from multiple speakers or sound without speech (e.g., silence or non-speech noise).

[0040] In a specific example, as described in more detail below, segmenter 124 segments audio stream 141 using one or more machine learning segmentation models (e.g., neural networks) trained to perform speaker segmentation. In this example, pre-registration of speakers is not required. Instead, segmenter 124 is trained to distinguish between two or more previously unknown speakers by comparing speaker characteristics between different audio frames of audio stream 141. The specific number of speakers that segmenter 124 can distinguish depends on the configuration and training of the machine learning segmentation models. For example, in a particular aspect, segmenter 124 can be configured to distinguish between three speakers, in which case the machine learning segmentation model may include five output layer nodes corresponding to speaker 1 output node, speaker 2 output node, speaker 3 output node, silence output node, and mixed output node. In this aspect, each output node is trained to generate a segmentation score as output, which indicates the probability that a set of analyzed audio portions 151 is associated with the corresponding output node. For example, the output node of speaker 1 generates a set of indicators for audio segment 151 representing the segment scores of the first speaker's speech, the output node of speaker 2 generates a set of indicators for audio segment 151 representing the segment scores of the second speaker's speech, and so on.

[0041] In a particular implementation, when segmenter 124 is configured to differentiate between three speakers, the machine learning segmentation model may include four output layer nodes. For example, the four output layer nodes may include a speaker 1 output node, a speaker 2 output node, a speaker 3 output node, and a silence output node, but not a mixed output node. In this implementation, the mixed speech is indicated by a set of indicator audio portions 151 from multiple speaker output nodes, representing the segmentation scores of the corresponding speaker's speech.

[0042] In a particular implementation, when segmenter 124 is configured to differentiate between three speakers, the machine learning segmentation model may include three output layer nodes. For example, the three output layer nodes may include a speaker 1 output node, a speaker 2 output node, and a speaker 3 output node, but not a silence output node. In this implementation, silence is indicated by the set of indicator audio portions 151 generated by each of the speaker output nodes not representing the segment score of the corresponding speaker's speech. For example, silence is indicated when speaker 1 output node generates a set of indicator audio portions 151 not representing the segment score of the first speaker's speech, speaker 2 output node generates a set of indicator audio portions 151 not representing the segment score of the second speaker's speech, and speaker 3 output node generates a set of indicator audio portions 151 not representing the segment score of the third speaker's speech. In some aspects, as used herein, "silence" may refer to "lack of speech," such as "non-speech noise."

[0043] Each of the audio portions 151 of a speaker-homogeneous audio segment 111 comprises multiple frames 102 of an audio stream 141. For example, each of the audio portions 151A may comprise ten (10) audio frames 102 representing five (5) seconds of sound. In other examples, different numbers of frames are included in each audio portion, or the frames may have different sizes, such that each audio portion 151A represents more or less than ten seconds of sound. Furthermore, each speaker-homogeneous audio segment 111 comprises multiple audio portions 151. The number of audio portions 151 in each speaker-homogeneous audio segment 111 is variable. For example, a speaker-homogeneous audio segment 111 may continue until the speaker's speech is interrupted, for example by a period of silence (e.g., a silence of threshold duration) or by the speech of another speaker.

[0044] Segmenter 124 provides segmentation results, identifying speaker homogeneous audio segments 111, to profile manager 126. Profile manager maintains user voice profiles (USPs) 150 in memory. Each user voice profile 150 is associated with a profile identifier (ID) 155. In certain aspects, profile ID 155 and user voice profile 150 are generated by profile manager 126 (e.g., profile ID 155 and user voice profile 150 are not based on user pre-registration).

[0045] In response to the segmentation results, profile manager 126 compares the audio portion 151 from speaker homogeneous audio segment 111 with user voice profile 150. If the audio portion 151 matches one of the user voice profiles 150 (e.g., is sufficiently similar), profile manager 126 updates user voice profile 150 based on the audio portion 151. For example, if the audio portion 151A of speaker homogeneous audio segment 111A is sufficiently similar to user voice profile 150A, profile manager 126 uses audio portion 151A to update user voice profile 150A.

[0046] If audio portion 151 does not match any of the user voice profiles 150, the profile manager 126 adds the user voice profile 150 based on audio portion 151. For example, in Figure 1 In the process, profile manager 126 generates user voice profile 150C based on the audio portion 151C of speaker homogeneous audio segment 111C, and assigns profile ID 155C to user voice profile 150C.

[0047] Profile manager 126 also generates output indicating the speaker or speaker variation in audio stream 141. For example, this output may include profile ID 155 of a user speech profile 150 that matches a speaker-homogeneous audio segment 111. One or more audio analysis applications 180 generate results based on the speaker or speaker variation. For example, audio analysis application 180 may transcribe detected speech to generate text and may indicate in the text when a speaker variation occurred.

[0048] refer to Figure 2A This document discloses specific illustrative aspects of a system configured to perform user voice profile management, and designates it generally as 200. System 200 includes a device 202 coupled to a microphone 246. Device 202 is configured to use... Figure 1 The segmenter 124 and profile manager 126 are used to perform user voice profile management. In a particular aspect, the device 202 includes one or more processors 220, which include a feature extractor 222, a segmenter 124, a profile manager 126, a speaker detector 278, one or more audio analysis applications 180, or a combination thereof.

[0049] Feature extractor 222 is configured to generate a set of audio feature data representing features of audio portions (e.g., audio frames) of an audio stream. Segmenter 124 is configured to indicate audio portions (or sets of audio feature data) representing the speech of the same speaker. Profile manager 126 is configured to generate (or update) user speech profiles based on audio portions (or sets of audio feature data) representing the speech of the same speaker. Speaker detector 278 is configured to determine a count of speakers detected in the audio stream. In a particular implementation, speaker detector 278 is configured to activate segmenter 124 in response to the detection of multiple speakers in the audio stream. In this implementation, when speaker detector 278 detects a single speaker in the audio stream, segmenter 124 is bypassed, and profile manager 126 generates (or updates) a user speech profile corresponding to that single speaker. In a particular implementation, one or more audio analysis applications 180 are configured to perform audio analysis based on user speech profiles.

[0050] In one aspect, device 202 includes memory 232 coupled to one or more processors 220. In another aspect, memory 232 includes one or more buffers, such as buffer 268. Memory 232 is configured to store one or more thresholds, such as segmentation threshold 257. Figure 2A (as in "Seg.Threshold"). In a particular context, this one or more thresholds are based on user input, configuration settings, default data, or a combination thereof.

[0051] In a particular aspect, memory 232 is configured to store data generated by feature extractor 222, speaker detector 278, segmenter 124, profile manager 126, one or more audio analysis applications 180, or combinations thereof. For example, memory 232 is configured to store multiple user speech profiles 150 of multiple users 242, segmentation results 236 (… Figure 2A The “segmentation results” in the data set 252, audio feature data set 151, and segmentation scores 254 are included. Figure 2A "Segmented scores" in the data set, segmented results 256 ( Figure 2A The "data set segmentation result" in the document, profile ID 155, or a combination thereof. Memory 232 is configured to store profile update data 272, user interaction data 274, etc. Figure 2A "User interaction data" in the context of data processing (or a combination thereof).

[0052] Device 202 is configured to receive audio stream 141 via a modem, network interface, input interface, or from microphone 246. In one aspect, audio stream 141 includes one or more audio segments 151. For example, audio stream 141 may be divided into a set of audio frames corresponding to audio segments 151, where each audio frame represents a time window portion of audio stream 141. In other examples, audio stream 141 may be divided in another manner to generate audio segments 151. Each audio segment 151 of audio stream 141 includes or represents silence, speech from one or more of users 242, or other sounds. The set of audio segments 151 representing speech from a single user is called speaker-homogeneous audio segments 111. Each speaker-homogeneous audio segment 111 includes multiple audio segments 151 (e.g., multiple audio frames). In one aspect, speaker-homogeneous audio segment 111 includes at least a threshold-counted number of audio frames (e.g., 5 audio frames). In one aspect, speaker-homogeneous audio segment 111 includes a consecutive set of audio segments 151 corresponding to the speech of the same user. In a particular aspect, a continuous set of audio portions 151 may include one or more subsets of audio portions 151, wherein each subset corresponds to a less-than-threshold silence that indicates a natural short pause in speech.

[0053] Audio stream 141 may include various combinations of the following: speaker-homogeneous audio segments, audio segments corresponding to silence, audio segments corresponding to multiple speakers, or combinations thereof. As an example, in Figure 2A In this audio stream 114, audio stream 114 includes an audio portion 151A of speaker-homogeneous audio segment 111A corresponding to the speech of user 242A, an audio portion 151B of audio segment 113 corresponding to silence (or non-speech noise), and an audio portion 151C of speaker-homogeneous audio segment 111B corresponding to the speech of user 242B. In other examples, audio stream 114 includes different sets or arrangements of audio segments. Although an audio portion is described as referring to an audio frame, in other implementations, an audio portion refers to a portion of an audio frame, multiple audio frames, audio data corresponding to a specific speech or playback duration, or a combination thereof.

[0054] Feature extractor 222 is configured to extract (e.g., determine) audio features from audio stream 141 to generate an audio feature data set 252. For example, feature extractor 222 is configured to extract audio features from audio portion 151 of audio stream 141 to generate an audio feature data set (AFDS) 252. In a particular aspect, the audio feature data set 252 includes audio feature vectors, such as embedding vectors. In a particular aspect, the audio feature data set 252 indicates the Mel-frequency cepstral coefficients (MFCCs) of audio portion 151. In a particular example, feature extractor 222 generates one or more audio feature data sets 252A by extracting audio features from audio portion 151A. Feature extractor 222 generates one or more audio feature data sets 252B by extracting audio features from audio portion 151B. Feature extractor 222 generates one or more audio feature data sets 252C by extracting audio features from audio portion 151C. The audio feature data set 252 includes one or more audio feature data sets 252A, one or more audio feature data sets 252B, one or more audio feature data sets 252C, or a combination thereof.

[0055] In an illustrative example, feature extractor 222 extracts audio features from each frame of audio stream 141 and provides the audio features of each frame to segmenter 124. In a particular aspect, segmenter 124 is configured to generate a set of segmentation scores (e.g., segmentation scores 254) for the audio features of a specific number of audio frames (e.g., 10 audio frames). For example, audio segment 151 comprises a specific number of audio frames (e.g., 10 audio frames). The audio features of the specific number of audio frames (e.g., used by segmenter 124 to generate the specific set of segmentation scores) correspond to audio feature data set 252. For example, feature extractor 222 extracts first audio features from a first audio frame, second audio features from a second audio frame, and so on (including a tenth audio feature from a tenth audio frame). Segmenter 124 generates a first segmentation score 254 based on the first audio features, the second audio features, and so on (including a tenth audio feature). For example, the first audio features, the second audio features, and up to the tenth audio features correspond to the first audio feature data set 252. Similarly, feature extractor 222 extracts the eleventh audio feature of the eleventh audio frame, the twelfth audio feature of the twelfth audio frame, and so on (including the twentieth audio feature of the twentieth audio frame). Segmenter 124 generates a second segmentation score 254 based on the eleventh audio feature, the twelfth audio feature, and so on (including the twentieth audio feature). For example, the eleventh audio feature, the twelfth audio feature, and so on up to the twentieth audio feature correspond to the second audio feature data set 252. It should be understood that generating the segmentation score set based on ten audio frames is provided as an illustrative example. In other examples, segmenter 124 generates the segmentation score set based on fewer or more than ten audio frames. For example, audio segment 151 includes fewer or more than ten audio frames.

[0056] Segmenter 124 is configured to generate a set of segmentation scores (e.g., segmentation scores 254) for each set of audio feature data. For example, in response to inputting audio feature data set 252 into segmenter 124, segmenter 124 generates multiple segmentation scores 254. The number of segmentation scores 254 generated in response to audio feature data set 252 depends on the number of speakers that segmenter 124 is trained to distinguish. As an example, segmenter 124 is configured to distinguish the speech of K different speakers by generating a set of K segmentation scores 254. In this example, each segmentation score 254 indicates the probability that the set of audio feature data input to segmenter 124 represents the speech of the corresponding speaker. For example, when segmenter 124 is configured to distinguish the speech of three (3) different speakers (e.g., speaker 292A, speaker 292B, and speaker 292C), K equals three (3). In this illustrative example, segmenter 124 is configured to output three (3) segment scores 254 for each audio feature data set 252 input to segmenter 124, such as segment score 254A, segment score 254B, and segment score 254C. In this illustrative example, segment score 254A indicates the probability that audio feature data set 252 represents the speech of speaker 292A, segment score 254B indicates the probability that audio feature data set 252 represents the speech of speaker 292B, and segment score 254C indicates the probability that audio feature data set 252 represents the speech of speaker 292C. In other examples, segmenter 124 is configured to distinguish (K in the above example) speakers whose counts are greater than three or less than three.

[0057] Speaker 292 corresponds to the set of speakers most recently (e.g., during the segmentation window) detected by segmenter 124. In one aspect, speaker 292 does not need to be pre-registered for differentiation by segmenter 124. Segmenter 124 achieves passive registration of multiple users by differentiating between the voices of multiple users who are not pre-registered. The segmentation window includes up to a specific count of audio segments (e.g., 20 audio frames), audio segments processed by segmenter 124 during a specific time window (e.g., 20 milliseconds), or audio segments corresponding to a specific speech duration or playback duration.

[0058] exist Figure 2AIn the example shown, an audio feature data set 252 representing the characteristics of the audio portion 151 of audio stream 141 can be provided as input to segmenter 124. In this example, audio feature data set 252 represents the speech of two or more of the users 242; for example, audio feature data set 252A represents the speech of user 242A, audio feature data set 252B represents silence, and audio feature data set 252C represents the speech of user 242B. In a particular implementation, segmenter 124 has no prior information about user 242. For example, user 242 has not pre-registered with device 202. In response to the input of audio feature data set 252, segmenter 124 outputs segment scores 254A, 254B, and 254C. Each segment score 254 indicates the probability that audio feature data set 252 represents the speech of the corresponding speaker 292, and each of the segment scores 254 is compared with a segmentation threshold 257. If a segmentation threshold 257 is satisfied for one of the segmentation scores 254 of the audio feature data set 252, it indicates that the speech of the corresponding speaker 292 has been detected in the audio feature data set 252. For example, if the segmentation threshold 257 is satisfied for segmentation score 254A of the audio feature data set 252, it indicates that the speech of speaker 292A has been detected in the audio feature data set 252 (and the audio portion 151 represented by the audio feature data set 252). A similar operation is performed for each of the audio feature data sets 252A, 252B, and 252C.

[0059] Segmenter 124 uses speaker 292 as a placeholder for an unknown user during the segmentation window (e.g., for user 242, unknown to segmenter 124 and associated with speech represented by audio feature data set 252). For example, audio feature data set 252A corresponds to the speech of user 242A. Segmenter 124 generates a segmentation score 254A for each of the audio feature data sets 252A that satisfies segmentation threshold 257, indicating that audio feature data set 252A corresponds to the speech of speaker 292A (e.g., a placeholder for user 242A). As another example, audio feature data set 252C corresponds to the speech of user 242B. Segmenter 124 generates a segmentation score 254B for each of the audio feature data sets 252C that satisfies segmentation threshold 257, indicating that audio feature data set 252C corresponds to the speech of speaker 292B (e.g., a placeholder for user 242B).

[0060] In a particular implementation, when the speech of speaker 292A (e.g., user 242A) has not been detected within the duration of the segmentation window (e.g., the threshold duration has expired since the previous speech associated with speaker 292A was detected), segmenter 124 may reuse speaker 292A (e.g., segmentation score 254A) as a placeholder for another user (e.g., user 242C). Segmenter 124 can distinguish speech associated with more than a predetermined number of speakers (e.g., more than K speakers) in audio stream 141 by reusing the speaker placeholder for another user when the previous user associated with the speaker placeholder has not spoken during the segmentation window. In a particular implementation, in response to determining that the speech of each of speaker 292A (e.g., user 242A), speaker 292B (e.g., user 242B), and speaker 292C (e.g., user 242C) is detected within the segmentation window, and determining that the speech associated with another user (e.g., user 242D) is detected, segmenter 124 reuses a speaker placeholder (e.g., speaker 292A) based on determining that the speech of speaker 292A (e.g., user 242A) was least recently detected.

[0061] In certain aspects, segmenter 124 includes or corresponds to a trained machine learning system, such as a neural network. For example, analyzing audio feature data set 252 includes applying a speaker segmentation neural network (or another machine learning-based system) to the audio feature data set 252.

[0062] In a particular aspect, segmenter 124 generates a dataset segmentation result 256 based on a segmentation score 254. The dataset segmentation result 256 indicates the presence of a speaker 292 (if any) in the audio portion 151. For example, the dataset segmentation result 256 output by segmenter 124 indicates that the speech of speaker 292 was detected in response to determining that the segmentation score 254 for speaker 292 satisfies (e.g., greater than) a segmentation threshold 257. For instance, when the segmentation score 254A of the audio feature dataset 252 satisfies the segmentation threshold 257, segmenter 124 generates a dataset segmentation result 256 (e.g., "1") for the audio feature dataset 252, indicating that the speech of speaker 292A was detected in the audio portion 151. In another example, when each of the segmentation scores 254A and 254B of the audio feature data set 252 satisfies the segmentation threshold 257, the segmenter 124 generates a data set segmentation result 256 (e.g., "1, 2") for the audio feature data set 252 to indicate that speech from speakers 292A and 292B (e.g., multiple speakers) has been detected in audio segment 151. In a particular example, when each of the segmentation scores 254A, 254B, and 254C for the audio feature data set 252 fails to satisfy the segmentation threshold 257, the segmenter 124 generates a data set segmentation result 256 (e.g., "0") for the audio feature data set 252 to indicate that silence (or non-speech audio) has been detected in audio segment 151. The segmentation result 236 for the audio part 151 (or audio feature data set 252) includes the segmentation score 254 for the audio part 151 (or audio feature data set 252), the data set segmentation result 256, or both.

[0063] Segmenter 124 is configured to provide segmentation results 236 for audio portion 151 (e.g., audio feature data set 252) to profile manager 126. Profile manager 126 is configured to generate user voice profile 150 at least in part based on audio feature data set 252 in response to determining that audio feature data set 252 does not match any of the plurality of user voice profiles 150. In a particular aspect, profile manager 126 is configured to generate user voice profile 150 based on speaker homogeneous audio segments 111. For example, profile manager 126 is configured to generate user voice profile 150A for speaker 292A (e.g., a placeholder for user 242A) based on audio feature data segment 152A of speaker homogeneous audio segment 111A. User voice profile 150A represents (e.g., models) the speech of user 242A. Alternatively, profile manager 126 is configured to update user voice profile 150 based on audio feature data set 252 in response to determining that audio feature data set 252 matches user voice profile 150. For example, profile manager 126 is configured to update user voice profile 150A representing the speech of user 242A based on subsequent audio portions that match user voice profile 150A, regardless of which speaker 292 is used as a placeholder for user 242A for the subsequent audio portions. In a particular aspect, profile manager 126 outputs profile ID 155 of user voice profile 150 in response to generating or updating user voice profile 150.

[0064] In one implementation, speaker detector 278 is configured to determine the count of speakers detected in audio stream 141 based on audio features extracted from audio stream 141. In one aspect, speaker detector 278 determines the speaker count based on an audio feature data set 252 extracted by feature extractor 222. For example, the audio features used by speaker detector 278 to determine the speaker count may be the same as the audio features used by segmenter 124 to generate segmentation results 236 and by profile manager 126 to generate or update user voice profile 150. In an alternative aspect, speaker detector 278 determines the speaker count based on audio features extracted by a second feature extractor different from feature extractor 222. In this aspect, the audio features used by speaker detector 278 to determine the speaker count may be different from the audio features used by segmenter 124 to generate segmentation results 236 and by profile manager 126 to generate or update user voice profile 150. In one aspect, the speaker detector 278 activates the segmenter 124 in response to detecting at least two distinct speakers in the audio stream 141. For example, when multiple speakers are detected in the audio stream 141, the segmenter 124 processes the audio feature data set 252. Alternatively, when the speaker detector 278 detects the speech of a single speaker in the audio stream 141, the segmenter 124 is bypassed, and the profile manager 126 processes the audio feature data set 252 to generate or update the user speech profile 150.

[0065] In some implementations, device 202 corresponds to or is included in one or more types of devices. In an illustrative example, one or more processors 220 are integrated into a headphone device including microphone 246, for example, see reference 1. Figure 13 Further described. In other examples, one or more processors 220 are integrated in at least one of: mobile phone or tablet computer devices (as described in reference). Figure 12 As described), wearable electronic devices (as referenced) Figure 14 As described), voice-controlled speaker system (as referenced) Figure 15 (as described), or virtual reality headsets or augmented reality headsets (as referenced) Figure 16 (As described). In another illustrative example, one or more processors 220 are integrated into a vehicle that also includes a microphone 246, for example, see reference 1. Figure 17 and Figure 18 Further description.

[0066] During operation, one or more processors 220 receive an audio stream 141 corresponding to the speech of one or more users 242 (e.g., user 242A, user 242B, user 242C, user 242D, or combinations thereof). In a particular example, one or more processors 220 receive the audio stream 141 from a microphone 246 that captures the speech of one or more users. In another example, the audio stream 141 corresponds to an audio playback file stored in memory 232, and one or more processors 220 receive the audio stream 141 from memory 232. In a particular aspect, one or more processors 220 receive the audio stream 141 from another device via an input interface or network interface (e.g., the network interface of a modem).

[0067] During the feature extraction phase, feature extractor 222 generates an audio feature data set 252 for audio stream 141. For example, feature extractor 222 generates audio feature data set 252 by determining features of audio portions 151 of audio stream 141. In a particular example, audio stream 141 includes audio portions 151A, 151B, 151C, or combinations thereof. Feature extractor 222 generates an audio feature data set 252A representing features of audio portion 151A, an audio feature data set 252B representing features of audio portion 151B, and an audio feature data set 252C representing features of audio portion 151C, or combinations thereof. For example, feature extractor 222 generates an audio feature data set 252 (e.g., a feature vector) for audio portion 151 (e.g., an audio frame) by extracting audio features from audio portion 151.

[0068] During the segmentation phase, segmenter 124 analyzes audio feature data set 252 to generate segmentation results 236. For example, segmenter 124 analyzes audio feature data set 252 (e.g., feature vectors) of audio segment 151 (e.g., audio frame) to generate a segmentation score 254 for audio segment 151. For example, segmentation score 254 includes segmentation score 254A (e.g., 0.6), which indicates the likelihood that audio segment 151 corresponds to the speech of speaker 292A. Segmentation score 254 also includes segmentation score 254B (e.g., 0) and segmentation score 254C (e.g., 0), which indicate the likelihood that audio segment 151 corresponds to the speech of speaker 292B and speaker 292C, respectively. In a particular aspect, in response to determining that segmentation score 254A satisfies segmentation threshold 257 and that each of segmentation scores 254B and 254C fails to satisfy segmentation threshold 257, segmenter 124 generates a data set segmentation result 256 indicating that audio segment 151 corresponds to the speech of speaker 292A and does not correspond to the speech of speaker 292B or speaker 292C. Segmenter 124 generates a segmentation result 236 indicating the segmentation score 254, data set segmentation result 256, or both for audio segment 151.

[0069] In a particular example, during the segmentation phase, in response to determining that each of a plurality of segmentation scores (e.g., segmentation score 254A and segmentation score 254B) satisfies segmentation threshold 257, segmenter 124 generates segmentation result 236, which indicates that audio portion 151 corresponds to the speech of a plurality of speakers (e.g., speaker 292A and speaker 292B).

[0070] The profile manager 126 processes the audio portion 151 (e.g., audio feature data set 252) based on the segmentation results 236, as referenced. Figure 2B Further description. In Figure 2B In this memory, memory 232 includes a registration buffer 234, a probe buffer 240, or a combination thereof. For example, memory 232 includes a registration buffer 234 and a probe buffer 240 assigned to each of the speakers 292. For instance, memory 232 includes a registration buffer 234A and a probe buffer 240A assigned to speaker 292A, a registration buffer 234B and a probe buffer 240B assigned to speaker 292B, and a registration buffer 234C and a probe buffer 240C assigned to speaker 292C. Memory 232 is configured to store a registration threshold 264, a profile threshold 258, a silence threshold 294, or a combination thereof. Memory 232 is configured to store data indicating the following: a stop condition 270, a voice profile result 238, and a silence count 262. Figure 2B (Sil Count) or combinations thereof.

[0071] Profile manager 126 is configured to determine, during the profile check phase, whether audio feature data set 252 matches an existing user voice profile 150. In one aspect, profile manager 126 uses the same audio features as those used by segmenter 124 to generate segmentation result 236 for comparison or updating with user voice profile 150. In another aspect, profile manager 126 uses a second audio feature different from the first audio feature used by segmenter 124 to generate segmentation result 236 for comparison or updating with user voice profile 150.

[0072] In a particular implementation, profile manager 126 is configured to collect an audio feature data set 252 corresponding to the same speaker in probe buffer 240 before comparing it with user voice profile 150, to improve the accuracy of the comparison. If the audio feature data set 252 matches an existing user voice profile, profile manager 126 is configured to update the existing user voice profile based on the audio feature data set 252 during the update phase. If the audio feature data set 252 does not match an existing user voice profile, profile manager 126 is configured to add the audio feature data set 252 to registration buffer 234 during the registration phase, and generate user voice profile 150 based on the audio feature data set 252 stored in registration buffer 234 in response to determining that the audio feature data set 252 stored in registration buffer 234 meets registration threshold 264.

[0073] During the profile check phase, in response to determining that no user voice profile is available and segmentation result 236 indicates that audio segment 151 corresponds to the speech of a speaker (e.g., speaker 292A), profile manager 126 adds audio feature data set 252 to the registration buffer 234 (e.g., registration buffer 234A) specified for speaker 292 and proceeds to the registration phase.

[0074] In one aspect, in response to determining that at least one user voice profile 150 is available, profile manager 126 performs a comparison of audio feature data set 252 with at least one user voice profile 150 to determine whether audio feature data set 252 matches any of the at least one user voice profile 150. In response to determining that at least one user voice profile 150 is available and segmentation result 236 indicates that audio segment 151 corresponds to the speech of speaker 292 (e.g., speaker 292A), profile manager 126 adds audio feature data set 252 to a probe buffer 240 (e.g., probe buffer 240A) specified for speaker 292.

[0075] Profile manager 126 determines whether the set of audio feature data (e.g., including audio feature data set 252) stored in probe buffer 240 matches any of at least one user voice profile 150. For example, profile manager 126 generates voice profile result 238 based on a comparison of the set of audio feature data (e.g., including audio feature data set 252) of probe buffer 240 (e.g., probe buffer 240A) with each of at least one user voice profile 150. For example, profile manager 126 generates voice profile result 238A based on a comparison of the set of audio feature data (e.g., including audio feature data set 252) of probe buffer 240 (e.g., probe buffer 240A) with user voice profile 150A.

[0076] In one aspect, in response to determining that a single set of audio feature data (e.g., audio feature data set 252) is available in probe buffer 240 (e.g., probe buffer 240A), profile manager 126 generates a voice profile result 238A based on a comparison of the single audio feature data set and the user voice profile 150A. Alternatively, in response to determining that multiple sets of audio feature data (e.g., including audio feature data set 252) are available in probe buffer 240 (e.g., probe buffer 240A), profile manager 126 generates a voice profile result 238A based on a comparison of the multiple audio feature data sets and the user voice profile 150A. For example, profile manager 126 generates a first data set result based on a comparison of audio feature data set 252 and user voice profile 150A, generates a second data set result based on a comparison of a second audio feature data set of probe buffer 240 and user voice profile 150A, generates an additional data set result based on a comparison of additional audio feature data set of probe buffer 240 and user voice profile 150A, or a combination thereof. Profile manager 126 generates voice profile result 238A based on the first data set result, the second data set result, the additional data set result (e.g., their weighted average), or a combination thereof. In a particular aspect, higher weights are assigned to the data set result of the audio feature data set most recently added to probe buffer 240.

[0077] Voice profile result 238A indicates the likelihood that the set of audio feature data matches the user voice profile 150A. Similarly, profile manager 126 generates voice profile result 238B based on a comparison of the set of audio feature data (e.g., including audio feature data set 252) of probe buffer 240 (e.g., probe buffer 240A) with the user voice profile 150B.

[0078] In a particular aspect, profile manager 126 selects the voice profile result 238 that indicates the highest probability of matching the audio feature data set 252 with the corresponding user voice profile 150. For example, in response to determining that voice profile result 238A indicates a higher matching probability compared to voice profile result 238B (e.g., greater than or equal to), profile manager 126 selects voice profile result 238A. In response to determining that voice profile result 238A (e.g., the voice profile result 238A indicating the highest probability of matching) satisfies (e.g., greater than or equal to) profile threshold 258, profile manager 126 determines that the audio feature data set stored in probe buffer 240 (e.g., probe buffer 240A) matches the user voice profile 150A and proceeds to the update phase. Alternatively, in response to determining that the voice profile result 238A (e.g., the voice profile result 238A indicating the highest probability of a match) fails to meet (e.g., less than) the profile threshold 258, the profile manager 126 determines that the set of audio feature data stored in the probe buffer 240 (e.g., probe buffer 240A) does not match any of the user's voice profiles 150 and proceeds to the registration phase.

[0079] During the update phase, in response to determining that the audio feature data set 252 matches the user voice profile 150 (e.g., user voice profile 150A), the profile manager 126 updates the user voice profile 150 and outputs the profile ID 155 of the user voice profile 150. The profile manager 126 updates the user voice profile 150 based on the audio feature data set stored in the probe buffer 240 (matching the audio feature data set stored in the probe buffer 240). Therefore, the user voice profile 150A evolves over time to match changes in the user's voice.

[0080] During the registration phase, in response to determining that segmentation result 236 indicates that audio feature data set 252 represents the speech of speaker 292 (e.g., speaker 292A), profile manager 126 adds audio feature data set 252 to the registration buffer 234 (e.g., registration buffer 234A) corresponding to speaker 292. Profile manager 126 determines whether the audio feature data set stored in registration buffer 234 satisfies registration threshold 264. In one aspect, in response to determining that the count of audio feature data sets is greater than or equal to registration threshold 264 (e.g., 48 audio feature data sets), profile manager 126 determines that the audio feature data set stored in registration buffer 234 satisfies registration threshold 264. In another aspect, in response to determining that the speech duration (e.g., playback duration) of the audio feature data set is greater than or equal to registration threshold 264 (e.g., 2 seconds), profile manager 126 determines that the audio feature data set stored in registration buffer 234 satisfies registration threshold 264.

[0081] In response to determining that the set of audio feature data stored in registration buffer 234 fails to meet registration threshold 264, profile manager 126 avoids generating user speech profile 150 based on the set of audio feature data stored in registration buffer 234 and continues processing subsequent audio portions of audio stream 141. In one aspect, profile manager 126 continues to add subsequent sets of audio feature data representing the speech of speaker 292 (e.g., speaker 292A) to registration buffer 234 (e.g., registration buffer 234A) until stop condition 270 is met. For example, in response to determining that the count of the set of audio feature data stored in registration buffer 234 (e.g., including audio feature data set 252) meets registration threshold 264, a silence longer than the threshold is detected in audio stream 141, or both, profile manager 126 determines that stop condition 270 is met, as described herein. For example, stopping condition 270 is met when there is enough audio feature data set in registration buffer 234 to generate a user voice profile, or when speaker 292 appears to have stopped speaking.

[0082] In a particular aspect, in response to determining that the set of audio feature data stored in the registration buffer 234 (e.g., including audio feature data set 252) satisfies the registration threshold 264, the profile manager 126 generates a user voice profile 150C based on the set of audio feature data stored in the registration buffer 234, resets the registration buffer 234, adds the user voice profile 150C to multiple user voice profiles 150, outputs the profile ID 155 of the user voice profile 150C, and continues processing subsequent audio portions of the audio stream 141. The profile manager 126 thus generates the user voice profile 150C based on the set of audio feature data corresponding to the audio portions of the same speaker 292 (e.g., speaker 292A), which is stored in the registration buffer 234 (e.g., registration buffer 234A) specified for speaker 292 (e.g., speaker 292A). Using multiple sets of audio feature data to generate user voice profiles 150C improves the accuracy of user voice profiles 150A in representing the speech of speaker 292A (e.g., user 242A). Therefore, by generating user voice profiles for users (who do not need to pre-register and do not need to speak predetermined words or sentences for the purpose of generating user voice profiles), segmenter 124 and profile manager 126 enable the passive registration of multiple users.

[0083] In certain aspects, audio segments corresponding to multiple speakers are skipped or ignored when generating or updating user voice profile 150. For example, in response to a segmentation result 236 indicating that audio segment 151 corresponds to the speech of multiple speakers, profile manager 126 ignores audio feature data set 252 of audio segment 151 and continues processing subsequent audio segments of audio stream 141. Ignoring audio feature data set 252 includes avoiding comparison of audio feature data set 252 with multiple user voice profiles 150, avoiding updating user voice profile 150 based on audio feature data set 252, avoiding generating user voice profile 150 based on audio feature data set 252, or combinations thereof.

[0084] In certain aspects, audio segments corresponding to silences shorter than a threshold (e.g., natural short pauses in the same user's speech) are not used to generate or update the user speech profile 150, but are tracked to detect silences longer than the threshold. For example, during the segmentation phase, segmenter 124 generates segmentation results 236 for audio feature data set 252, indicating that audio segment 151 corresponds to silence. In response to determining that audio segment 151 corresponds to silence, profile manager 126 increments the silence count 262 (e.g., by 1). In one aspect, in response to determining that the silence count 262 is greater than or equal to the silence threshold 294 (e.g., indicating a longer pause after the user has finished speaking), profile manager 126 resets registration buffer 234 (e.g., registration buffer 234A, registration buffer 234B, and registration buffer 234C) (e.g., marking it as empty), resets probe buffer 240 (e.g., probe buffer 240A, probe buffer 240B, and probe buffer 240C) (e.g., marking it as empty), resets the silence count 262 (e.g., resetting it to 0), or a combination thereof, and continues processing subsequent audio portions of audio stream 141. In another aspect, in response to determining that the silence count 262 is greater than or equal to the silence threshold 294, profile manager 126 determines that stop condition 270 is met. In response to determining that stop condition 270 is met, profile manager 126 resets registration buffer 234 (e.g., registration buffer 234A, registration buffer 234B, and registration buffer 234C).

[0085] In one aspect, profile manager 126 provides a notification to a display device coupled to device 202. This notification indicates that user speech analysis is in progress. In another aspect, profile manager 126 selectively processes audio stream 141 based on user input indicating whether user speech analysis should be performed.

[0086] Return to Figure 2A In a particular aspect, profile manager 126 maintains profile update data 272 to track how many user voice profiles 150 are generated or updated during audio stream 141 processing. For example, in response to updating (or generating) a user voice profile 150, profile manager 126 updates profile update data 272. In a particular example, in response to updating user voice profile 150A, profile manager 126 updates profile update data 272 to indicate that user voice profile 150A has been updated. As another example, in response to generating user voice profile 150C, profile manager 126 updates profile update data 272 to indicate that user voice profile 150C has been updated. In response to determining that a first count of multiple user voice profiles 150 indicated by profile update data 272 has been updated during audio stream 141 processing, profile manager 126 outputs this first count as a count of speakers detected in audio stream 141.

[0087] In a particular aspect, profile manager 126 maintains user interaction data 274 to track the duration of detected speech that matches each of a plurality of user voice profiles 150. Profile manager 126 updates user interaction data 274 based on updating (or generating) user voice profiles 150. For example, in response to updating user voice profile 150A based on audio portion 151, profile manager 126 updates user interaction data 274 to instruct the user associated with user voice profile 150A to interact during the duration of speech in audio portion 151. As another example, in response to generating user voice profile 150C based on audio portion 151, profile manager 126 updates user interaction data 274 to instruct the user associated with user voice profile 150C to interact during the duration of speech in audio portion 151. For example, after a user voice profile 150 is generated or updated based on the audio portion of speaker homogeneous audio segment 111, user interaction data 274 instructs the user associated with the user voice profile 150 to interact during the duration of the speech in speaker homogeneous audio segment 111. In a particular aspect, profile manager 126 outputs user interaction data 274.

[0088] In certain aspects, profile manager 126 provides profile ID 155, profile update data 272, user interaction data 274, additional information, or a combination thereof, to one or more audio analysis applications 180. For example, audio analysis application 180 performs speech-to-text conversion on audio feature data set 252 to generate a transcript of audio stream 141. Audio analysis application 180 tags the text corresponding to audio feature data set 252 in the transcript based on profile ID 155 for audio feature data set 252 received from profile manager 126.

[0089] In one aspect, one or more processors 220 are configured to operate in one of a plurality of power modes. For example, one or more processors 220 are configured to operate in power mode 282 (e.g., an always-on power mode) or power mode 284 (e.g., an on-demand power mode). In one aspect, power mode 282 is a lower power mode compared to power mode 284. For example, one or more processors 220 conserve energy by operating in power mode 282 (compared to power mode 284) and switch to power mode 284 when needed to activate components that are not operational in power mode 282.

[0090] In a specific example, some functions of device 202 are active in power mode 284 but inactive in power mode 282. For example, speaker detector 278 can be activated in both power mode 282 and power mode 284. In this example, feature extractor 222, segmenter 124, profile manager 126, one or more audio analysis applications 180, or combinations thereof, can be activated in power mode 284 but not in power mode 282. When audio stream 141 corresponds to the speech of a single speaker, segmenter 124 is not required to distinguish between audio portions corresponding to different speakers. Maintaining (or switching to) power mode 282 reduces overall resource consumption when segmenter 124 is not required. Speaker detector 278 is configured to determine in power mode 282 whether audio stream 141 corresponds to the speech of at least two different speakers. In response to the output of the speaker detector 278 indicating that the audio stream 141 corresponds to the speech of at least two different speakers, one or more processors 220 are configured to switch from power mode 282 to power mode 284 and activate segmenter 124. For example, segmenter 124 analyzes audio feature data set 252 in power mode 284 to generate segmentation results 236.

[0091] In a specific example, the speaker detector 278 and profile manager 126 can be activated in power mode 282 and power mode 284. In this example, feature extractor 222, segmenter 124, one or more audio analysis applications 180, or a combination thereof, can be activated in power mode 284 but not in power mode 282. For example, in response to the output of speaker detector 278 indicating that a single speaker has been detected, one or more processors 220 maintain or switch to power mode 282. In power mode 282, profile manager 126 generates or updates a user speech profile 150 for a single speaker based on audio feature data set 252. Alternatively, in response to the output of speaker detector 278 indicating that audio stream 141 corresponds to the speech of at least two different speakers, one or more processors 220 switch from power mode 282 to power mode 284 and activate segmenter 124. For example, segmenter 124 analyzes audio feature data set 252 in power mode 284 to generate segmentation results 236.

[0092] In a particular example, feature extractor 222, speaker detector 278, segmenter 124, or combinations thereof, may be activated in power mode 282 and power mode 284. In this example, profile manager 126, one or more audio analysis applications 180, or combinations thereof, may be activated in power mode 284 but not in power mode 282. In a particular aspect, one or more processors 220 are configured to switch from power mode 282 to power mode 284 and activate profile manager 126, one or more audio analysis applications 180, or combinations thereof, in response to determining that segmentation result 236 indicates that audio stream 141 corresponds to the speech of at least two different speakers. For example, profile manager 126 performs a comparison of audio feature data set 252 with multiple user speech profiles 150 in power mode 284.

[0093] In one aspect, in response to determining that segmentation result 236 indicates that audio stream 141 corresponds to the speech of at least two different speakers, one or more processors 220 process subsequent audio portions of audio stream 141 in power mode 284. For example, feature extractor 222, segmenter 124, or both operate in power mode 284 to process subsequent audio portions. In another aspect, feature extractor 222, speaker detector 278, segmenter 124, or a combination thereof determine audio information of audio stream 141 in power mode 282 and provide the audio information to one or more audio analysis applications 180 in power mode 284. This audio information includes speaker counts, voice activity detection (VAD) information, or both, indicated in audio stream 141.

[0094] In a particular implementation, one or more portions of the audio stream 141, the audio feature data set 252, or a combination thereof are stored in a buffer 268, and one or more processors 220 access one or more portions of the audio stream 141, the audio feature data set 252, or a combination thereof from the buffer 268. For example, one or more processors 220 store an audio portion 151 in the buffer 268. A feature extractor 222 retrieves the audio portion 151 from the buffer 268 and stores the audio feature data set 252 in the buffer 268. A segmenter 124 retrieves the audio feature data set 252 from the buffer 268 and stores a segmentation score 254, a data set segmentation result 256, or a combination thereof, of the audio feature data set 252 in the buffer 268. A profile manager 126 retrieves the audio feature data set 252, the segmentation score 254, the data set segmentation result 256, or a combination thereof from the buffer 268. In one aspect, profile manager 126 stores profile ID 155, profile update data 272, user interaction data 274, or a combination thereof, in buffer 268. In another aspect, one or more audio analysis applications 180 retrieve profile ID 155, profile update data 272, user interaction data 274, or a combination thereof, from buffer 268.

[0095] Therefore, system 200 enables the registration and updating of passive user voice profiles for multiple speakers. For example, multiple user voice profiles 150 can be generated and updated in the background during normal operation of device 202 without requiring user 242 to speak predetermined words or sentences from a script.

[0096] Although microphone 246 is shown as coupled to device 202, in other implementations, microphone 246 may be integrated into device 202. While a single microphone 246 is shown, in other implementations, one or more additional microphones 146 configured to capture user speech may be included.

[0097] Although system 200 is shown as including a single device 202, in other implementations, the implementation operations described as being performed at device 202 can be distributed across multiple devices. For example, operations described as being performed by one or more of the feature extractor 222, speaker detector 278, segmenter 124, profile manager 126, or one or more audio analysis applications 180 can be performed at device 202, and operations described as being performed by the other of the feature extractor 222, speaker detector 278, segmenter 124, profile manager 126, or one or more audio analysis applications 180 can be performed at a second device.

[0098] refer to Figure 3This illustrates an illustrative aspect of operation 300 associated with user voice profile management. In a particular aspect, one or more operations in operation 300 are performed by… Figure 1 Segmenter 124, Profile Manager 126 Figure 2A The feature extractor 222, one or more processors 220, devices 202, systems 200, or combinations thereof are used to perform the operation.

[0099] During speaker segmentation 302 Figure 2A The feature extractor 222 generates an audio feature data set 252 based on the audio stream 141, as shown in the reference. Figure 2A As described. Segmenter 124 analyzes audio feature data set 252 to generate segmentation results 236, as referenced. Figure 2A As described.

[0100] During the 304 period of voice profile management Figure 1 The profile manager 126 determines at 306 whether the audio feature data set 252 corresponds to a registered speaker. For example, the profile manager 126 determines whether the audio feature data set 252 matches any user voice profile 150, as referenced. Figure 2B As described. In response to determining at 306 that the audio feature data set 252 matches a user voice profile 150A with profile ID 155, the profile manager 126 updates the user voice profile 150A at 308 based at least in part on the audio feature data set 252. Alternatively, in response to determining at 306 that the audio feature data set 252 does not match any of the plurality of user voice profiles 150, and segmentation result 236 indicating that the audio feature data set 252 represents the speech of speaker 292A, the profile manager 126 adds the audio feature data set 252 to the registration buffer 234A specified for speaker 292A at 310.

[0101] In response to determining at 312 that the count of the audio feature data set of the registration buffer 234A (or the speech duration of the audio feature data set of the registration buffer 234A) is greater than the registration threshold 264, the profile manager 126 registers the speaker at 314. For example, the profile manager 126 generates a user speech profile 150C based on the audio feature data set of the registration buffer 234A and adds the user speech profile 150C to multiple user speech profiles 150, as referenced. Figure 2B As described, the profile manager 126 continues to process subsequent audio portions of the audio stream 141.

[0102] The segmentation result 236 generated during speaker segmentation 302 thus ensures that the set of audio feature data corresponding to the speech of the same speaker is collected in the same registration buffer during voice profile management 304 for speaker registration. Generating user voice profiles 150C based on multiple audio feature data sets improves the accuracy of user voice profiles 150C in representing speaker speech.

[0103] refer to Figure 4 This illustrates an illustrative aspect of operation 400 associated with user voice profile management. In a particular aspect, one or more operations in operation 400 are performed by… Figure 1 Segmenter 124, Profile Manager 126 Figure 2A The feature extractor 222, one or more processors 220, devices 202, systems 200, or combinations thereof are used to perform the operation.

[0104] Audio stream 141 includes audio segments 151A to 151I. During speaker segmentation 302, Figure 1 Segmenter 124 generates segmentation scores 254A, 254B, and 254C for each of the audio sections 151A-I, as shown in the reference. Figure 2A As described.

[0105] Segment score 254 indicates that audio segment 151A corresponds to the speech of the same single speaker (e.g., designated speaker 292A). For example, segment score 254A of each speaker in audio segment 151A satisfies segment threshold 257. Segment scores 254B and 254C of each speaker in audio segment 151A do not satisfy segment threshold 257.

[0106] During voice profile management 304, profile manager 126 adds audio portion 151A (e.g., the corresponding set of audio feature data) to the registration buffer 234A associated with speaker 292A. Profile manager 126 generates user voice profile 150A based on audio portion 151A (e.g., the corresponding set of audio feature data).

[0107] In a particular aspect, segmentation 254 indicates that audio segment 151B corresponds to the speech of multiple speakers, such as speaker 292A and another speaker (e.g., designated as speaker 292B). Figure 4In this process, profile manager 126 updates user voice profile 150A based on audio portion 151B (e.g., the corresponding set of audio feature data). In one aspect, profile manager 126 also adds audio portion 151B to a registration buffer 234B associated with speaker 292B. Alternatively, profile manager 126 ignores audio portions 151B corresponding to multiple speakers. For example, profile manager 126 avoids using audio portion 151B to update or generate user voice profile 150.

[0108] Segment score 254 indicates that audio segment 151C corresponds to the speech of speaker 292B (e.g., a single speaker). Profile manager 126 adds audio segment 151C to registration buffer 234B. In response to determining that the audio segment (e.g., the corresponding audio feature data set) stored in registration buffer 234B fails to meet registration threshold 264, profile manager 126 avoids generating user speech profile 150 based on the audio segment (e.g., the corresponding audio feature data set) stored in registration buffer 234B. In one aspect, the audio segment (e.g., the corresponding audio feature data set) stored in registration buffer 234B includes audio segment 151B (e.g., the corresponding audio feature data set) and audio segment 151C (e.g., the corresponding audio feature data set). In an alternative aspect, the audio segment (e.g., the corresponding audio feature data set) stored in registration buffer 234B includes audio segment 151C (e.g., the corresponding audio feature data set) and does not include audio segment 151B (e.g., the corresponding audio feature data set).

[0109] Segment score 254 indicates that audio segment 151D corresponds to the speech of another single speaker (e.g., designated as speaker 292C). Profile manager 126 adds a first subset of audio segment 151D (e.g., the corresponding set of audio feature data) to registration buffer 234C. In response to determining that the first subset of audio segment 151D (e.g., the corresponding set of audio feature data) stored in registration buffer 234C satisfies registration threshold 264, profile manager 126 generates user speech profile 150B based on the first subset of audio segment 151D (e.g., the corresponding set of audio feature data) stored in registration buffer 234C. Profile manager 126 updates user speech profile 150B based on a second subset of audio segment 151D.

[0110] Segment score 254 indicates that audio segment 151E corresponds to a silence greater than the threshold. For example, the count of audio segment 151E is greater than or equal to the silence threshold 294. In response to determining that audio segment 151E corresponds to a silence greater than the threshold, profile manager 126 resets registration buffer 234.

[0111] Segment score 254 indicates that audio segment 151F corresponds to the speech of a single speaker (e.g., designated as speaker 292A). In response to determining that each of the audio segments 151F matches a user voice profile 150B, profile manager 126 updates the user voice profile 150B based on the audio segment 151F. Because speaker designations (e.g., speaker 292A) are reused, audio segments 151D and 151F are associated with different designated speakers (e.g., speaker 292C and speaker 292A), even though audio segments 151D and 151F correspond to the same speaker (e.g., speaker 292A). Figure 2A The voice of user 242C is matched with the voice profile of the same user (e.g., user voice profile 150B).

[0112] Segment score 254 indicates that audio segment 151G corresponds to the speech of a single speaker (e.g., designated as speaker 292B). In response to determining that a first subset of audio segment 151G does not match any of the user voice profiles 150, profile manager 126 adds the first subset of audio segment 151G to the registration buffer 234B associated with speaker 292B. Profile manager 126 generates user voice profile 150C based on the first subset of audio segment 151G and updates user voice profile 150C based on a second subset of audio segment 151G. Because the speaker designation (e.g., speaker 292B) is reused, audio segments 151C and 151G, associated with the same designated speaker (e.g., speaker 292B), can correspond to the speech of the same user or different users.

[0113] Segment score 254 indicates that audio segment 151H corresponds to a silence greater than the threshold. In response to determining that audio segment 151H corresponds to a silence greater than the threshold, profile manager 126 resets registration buffer 234.

[0114] Segment score 254 indicates that audio segment 151I corresponds to the speech of a single speaker (e.g., designated as speaker 292C). In response to determining that each of the audio segments 151I matches a user voice profile 150A, profile manager 126 updates the user voice profile 150A based on the audio segment 151I. Because the speaker designation (e.g., speaker 292C) is reused, audio segments 151A and 151I are associated with different designated speakers (e.g., speaker 292A and speaker 292C), even though audio segments 151A and 151I correspond to the same user (e.g., [missing information]). Figure 2AThe profile manager 126 takes the voice of user 242A and matches it with the same user voice profile (e.g., user voice profile 150A). Alternatively, in response to determining that audio portion 151I does not match any of the multiple user voice profiles 150, the profile manager 126 adds a first subset of audio portion 151I to the registration buffer 234C associated with speaker 292C and generates user voice profile 150D based on the first subset of audio portion 151I. By reusing speaker designations (e.g., speaker 292C), the profile manager 126 can generate (or update) user profiles with a larger count than a predetermined count (e.g., K) of speakers 292 that can be distinguished by the segmenter 124.

[0115] refer to Figure 5 This illustrates an illustrative aspect of operation 500 associated with user voice profile management. In a particular aspect, one or more operations in operation 500 are performed by… Figure 1 Segmenter 124, Profile Manager 126 Figure 2A The feature extractor 222, one or more processors 220, devices 202, systems 200, or combinations thereof are used to perform the operation.

[0116] Audio stream 141 includes audio portion 151A, audio portion 151B, and audio portion 151C. For example, audio portion 151A includes audio portion 151D (e.g., an audio frame), one or more additional audio portions, and audio portion 151E. Audio portion 151B includes audio portion 151F, one or more additional audio portions, and audio portion 151G. Audio portion 151C includes audio portion 151H, one or more additional audio portions, and audio portion 151I.

[0117] In a particular aspect, the data set segmentation result 256A of each of the audio portions 151A indicates that the audio portion 151A corresponds to the speech of speaker 292A. For example, the data set segmentation result 256D of the audio portion 151D (e.g., "1") indicates that the audio portion 151D represents the speech of speaker 292A. As another example, the data set segmentation result 256E of the audio portion 151E (e.g., "1") indicates that the audio portion 151E represents the speech of speaker 292A.

[0118] The data set segmentation result 256B of each audio portion 151B indicates that audio portion 151B corresponds to silence (or non-speech noise). For example, the data set segmentation result 256F (e.g., "0") of audio portion 151F indicates that audio portion 151F represents silence (or non-speech noise). As another example, the data set segmentation result 256G (e.g., "0") of audio portion 151G indicates that audio portion 151G represents silence (or non-speech noise).

[0119] The data set segmentation result 256C for each of the audio portions 151C indicates that the audio portion 151C corresponds to the speech of speaker 292B. For example, the data set segmentation result 256H (e.g., "2") of the audio portion 151H indicates that the audio portion 151H represents the speech of speaker 292B. As another example, the data set segmentation result 256I (e.g., "2") of the audio portion 151I indicates that the audio portion 151I represents the speech of speaker 292B.

[0120] Figure 590 is a visual depiction of an example of segmentation result 236. For example, audio segment 151A represents the speech of speaker 292A (e.g., a single speaker), and therefore audio segment 151A corresponds to speaker-homogeneous audio segment 111A of audio stream 141. Audio segment 151B represents silence, and therefore audio segment 151B corresponds to audio segment 113A of audio stream 141 (e.g., not a speaker-homogeneous audio segment). Audio segment 151C represents the speech of speaker 292B (e.g., a single speaker), and therefore audio segment 151C corresponds to speaker-homogeneous audio segment 111B of audio stream 141.

[0121] Figure 592 is a visual depiction of an example of voice profile result 238. Profile manager 126 generates a user voice profile 150A based on a first subset of audio portion 151A. After generating user voice profile 150A, profile manager 126 determines voice profile result 238A by comparing subsequent audio portions (e.g., a subsequent set of audio feature data) with user voice profile 150A. Voice profile result 238A of audio portion 151 indicates the probability that audio portion 151 matches user voice profile 150A. Profile manager 126 determines voice profile result 238A of the first subset of audio portion 151C by comparing it with user voice profile 150A. In response to determining that voice profile result 238A of the first subset of audio portion 151C is less than profile threshold 258, profile manager 126 determines that the first subset of audio portion 151C does not match user voice profile 150A.

[0122] In response to determining that a first subset of audio portion 151C does not match user voice profile 150A, profile manager 126 generates user voice profile 150B based on the first subset of audio portion 151C. After generating user voice profile 150B, profile manager 126 determines voice profile result 238B by comparing subsequent audio portions with user voice profile 150B. Voice profile result 238B indicates the probability that an audio portion matches user voice profile 150B. For example, voice profile result 238B for a second subset of audio portion 151C indicates that a second subset of audio portion 151C matches user voice profile 150B. In a particular aspect, profile manager 126 generates a graphical user interface (GUI) including graphics 590, graphics 592, or both, and provides the GUI to a display device.

[0123] refer to Figure 6 This illustrates an illustrative aspect of operation 600 associated with user voice profile management. In a particular aspect, one or more operations in operation 600 are performed by… Figure 1 Segmenter 124, Profile Manager 126 Figure 2A The feature extractor 222, one or more processors 220, devices 202, systems 200, or combinations thereof are used to perform the operation.

[0124] Audio stream 141 includes audio segments 151J corresponding to the speech of multiple speakers. For example, audio segment 151J includes audio segment 151K (e.g., an audio frame), one or more additional audio segments, and audio segment 151L. In a particular aspect, the data set segmentation result 256D of each of the audio segments 151J indicates that audio segment 151J corresponds to the speech of speakers 292A and 292B. For example, the data set segmentation result 256K of audio segment 151K (e.g., "1, 2") indicates that audio segment 151K represents the speech of speakers 292A and 292B. As another example, the data set segmentation result 256L of audio segment 151L (e.g., "1, 2") indicates that audio segment 151L represents the speech of speakers 292A and 292B. Since audio segment 151J represents the speech of multiple speakers, audio segment 151J corresponds to audio segment 113B (e.g., a non-speaker homogeneous audio segment).

[0125] After generating the user voice profile 150A, the profile manager 126 determines a voice profile result 238A by comparing subsequent audio portions (e.g., a set of subsequent audio feature data) with the user voice profile 150A. The profile manager 126 determines a voice profile result 238A for audio portion 151J by comparing it with the user voice profile 150A. In a particular aspect, the voice profile result 238A for audio portion 151J is lower than the voice profile result 238A for audio portion 151A because audio portion 151J includes the voice of speaker 292B in addition to the voice of speaker 292A.

[0126] refer to Figure 7 This illustrates an illustrative aspect of operation 700 associated with user voice profile management. In a specific aspect, one or more operations in operation 700 are performed by feature extractor 222, segmenter 124, profile manager 126, ... Figure 2A It may be executed by one or more processors 220, devices 202, systems 200 or combinations thereof.

[0127] Audio stream 141 includes audio portion 151J and audio portion 151K. For example, audio portion 151J includes audio portion 151L (e.g., an audio frame), one or more additional audio portions, and audio portion 151M. Audio portion 151K includes audio portion 151N (e.g., an audio frame), one or more additional audio portions, and audio portion 151O.

[0128] In a particular aspect, the data set segmentation result 256J of each of the audio portions 151J indicates that the audio portion 151J represents the speech of speaker 292C (e.g., a single speaker), and therefore the audio portion 151J corresponds to speaker homogeneous audio segment 111C. The data set segmentation result 256K of each of the audio portions 151K indicates that the audio portion 151K represents silence (or non-speech noise), and therefore the audio portion 151K corresponds to audio segment 113C.

[0129] After generating the user voice profile 150A, the profile manager 126 determines the voice profile result 238A of the audio portion 151J by comparing it with the user voice profile 150A. In response to determining that the voice profile result 238A is less than the profile threshold 258, the profile manager 126 determines that the audio portion 151J does not match the user voice profile 150A.

[0130] In response to determining that audio portion 151J does not match user voice profile 150A, profile manager 126 stores audio portion 151J in registration buffer 234C associated with speaker 292C. In response to determining that audio portion 151J stored in registration buffer 234C fails to meet registration threshold 264, profile manager 126 avoids generating user voice profile 150 based on audio portion 151J stored in registration buffer 234C. In response to determining that audio portion 151K indicates greater than threshold silence, profile manager 126 resets registration buffer 234 (e.g., marks it as empty). Therefore, when speaker 292C appears to have stopped speaking, audio portion 151J is removed from registration buffer 234C.

[0131] refer to Figure 8 This illustrates an illustrative aspect of operation 800 associated with user voice profile management. In a particular aspect, one or more operations in operation 800 are performed by… Figure 1 Segmenter 124, Profile Manager 126 Figure 2A The feature extractor 222, one or more processors 220, devices 202, systems 200, or combinations thereof are used to perform the operation.

[0132] Figure 1 The segmenter 124 performs speaker segmentation 302 at 804. For example, the segmenter 124 receives an audio feature data set 252 from the feature extractor 222 at time T and generates a segmentation score 254 for the audio feature data set 252 for the audio portion 151, as shown in the reference. Figure 2A As described.

[0133] Figure 1 At 806, the profile manager 126 determines whether any of the segment scores 254 satisfy the segment threshold 257. For example, in response to determining that none of the segment scores 254 satisfy the segment threshold 257, the profile manager 126 determines that the audio feature data set 252 represents silence (or non-speech noise) and increments the silence count 262 (e.g., by 1). After incrementing the silence count 262, the profile manager 126 determines at 808 whether the silence count 262 is greater than the silence threshold 294.

[0134] In response to determining at 808 that the silence count 262 is greater than the silence threshold 294, profile manager 126 performs a reset at 810. For example, profile manager 126 performs the reset by resetting registration buffer 234 (e.g., marking it as empty), probe buffer 240 (e.g., marking it as empty), silence count 262 (e.g., resetting it to 0), or a combination thereof, and returns to 804 to process subsequent sets of audio feature data for audio stream 141. Alternatively, in response to determining at 808 that the silence count 262 is less than or equal to the silence threshold 294, profile manager 126 returns to 804 to process subsequent sets of audio feature data for audio stream 141.

[0135] In response to determining at 806 that at least one of the segment scores 254 satisfies the segment threshold 257, the profile manager 126 adds an audio feature data set 252 to at least one of the probe buffers 240 at 812. For example, in response to determining that the segment score 254A associated with speaker 292A satisfies the segment threshold 257, the profile manager 126 determines that the audio feature data set 252 represents the speech of speaker 292A and adds the audio feature data set 252 to the probe buffer 240A associated with speaker 292A. In a particular implementation, the audio feature data set 252 representing the speech of multiple speakers 292 is added to multiple probe buffers 240 corresponding to the multiple speakers 292. For example, in response to determining that each of segment scores 254A and 254B satisfies the segment threshold 257, the profile manager 126 adds the audio feature data set 252 to probe buffers 140A and 140B. In an alternative implementation, the audio feature data set 252 representing the speech of multiple speakers 292 is ignored and not added to the probe buffer 240.

[0136] At 816, profile manager 126 determines whether the corresponding speaker (e.g., speaker 292A) is registered. For example, profile manager 126 determines whether speaker 292 (e.g., speaker 292A) is registered by comparing the audio feature data set (e.g., including audio feature data set 252) of the corresponding probe buffer 240 (e.g., probe buffer 240A) with multiple user voice profiles 150.

[0137] In response to determining at 816 that speaker 292 (e.g., speaker 292A) has not been registered, profile manager 126 determines at 818 whether audio feature data set 252 has passed a quality check. For example, in response to determining that audio feature data set 252 corresponds to multiple speakers 292, profile manager 126 determines that audio feature data set 252 has failed the quality check. Alternatively, in response to determining that audio feature data set 252 corresponds to a single speaker, profile manager 126 determines that audio feature data set 252 has passed the quality check.

[0138] In response to determining at 818 that audio feature data set 252 failed the quality check, profile manager 126 returns to 804 to process subsequent audio feature data sets for audio stream 141. Alternatively, in response to determining at 818 that audio feature data set 252 passed the quality check, profile manager 126 adds at 820 the audio feature data set 252 representing the speech of speaker 292 (e.g., speaker 292A) to the registration buffer 234 (e.g., registration buffer 234A) associated with speaker 292.

[0139] Profile manager 126 determines at 822 whether the count of the audio feature data set stored in registration buffer 234 (e.g., registration buffer 234A) is greater than registration threshold 264. In response to determining at 822 that the count of the audio feature data set in each of the registration buffers 234 (e.g., registration buffer 234A) is less than or equal to registration threshold 264, profile manager 126 returns to 804 to process subsequent audio feature data sets of audio stream 141. Alternatively, in response to determining that the count of the audio feature data set in registration buffer 234 (e.g., registration buffer 234A) is greater than registration threshold 264, profile manager 126 generates user voice profile 150A at 824, adds user voice profile 150A to multiple user voice profiles 150, and returns to 804 to process subsequent audio feature data sets of audio stream 141.

[0140] In response to determining at 816 that speaker 292A has been registered, profile manager 126 determines at 826 whether audio feature data set 252 (or the audio feature data set of probe buffer 240 associated with speaker 292 whose speech is represented by audio feature data set 252) passes a quality check. In response to determining at 826 that audio feature data set 252 (or the audio feature data set of probe buffer 240) fails the quality check, profile manager 126 returns to 804 to process subsequent audio feature data sets of audio stream 141. In response to determining at 826 that audio feature data set 252 (or the audio feature data set of probe buffer 240) passes the quality check, profile manager 126 updates user speech profile 150A (matching audio feature data set 252) based on audio feature data set 252 (or the audio feature data set of probe buffer 240) and returns to 804 to process subsequent audio feature data sets of audio stream 141. In an alternative approach, a quality check is performed at 826 before adding the audio feature data set 252 to the probe buffer 240. For example, in response to determining that the audio feature data set 252 fails the quality check, the profile manager 126 avoids adding the audio feature data set 252 to the probe buffer 240 and returns to 804 to process subsequent audio feature data sets for the audio stream 141.

[0141] refer to Figure 9 This illustrates an illustrative aspect of operation 900 associated with user voice profile management. In a particular aspect, one or more operations in operation 900 are performed by… Figure 1 Segmenter 124, Profile Manager 126 Figure 2A The feature extractor 222, the speaker detector 278, one or more processors 220, devices 202, systems 200, or combinations thereof are used to perform this.

[0142] One or more processors 220 add audio features (e.g., audio feature data set 252) to buffer 268 at time T in power mode 282. Figure 2A Speaker detector 278 determines at 904 whether multiple speakers are detected in audio stream 141. For example, speaker detector 278 determines that multiple speakers are detected in response to determining that audio features (e.g., audio feature data set 252) represent the speech of multiple speakers. In another example, speaker detector 278 determines that multiple speakers are detected in response to determining that audio features (e.g., audio feature data set 252) represent the speech of a second speaker after the speech of a first speaker has been detected in previous audio features (e.g., previous audio feature data set).

[0143] In response to determining at 904 that multiple speakers have not yet been detected in audio stream 141, speaker detector 278 continues processing subsequent audio features of audio stream 141. Alternatively, in response to determining at 904 that multiple speakers have been detected in audio stream 141, speaker detector 278 switches one or more processors 220 from power mode 282 to power mode 284 and activates one or more applications 920 at 906. In a particular aspect, the one or more applications 920 include feature extractor 222, segmenter 124, profile manager 126, one or more audio analysis applications 180, or a combination thereof. In a particular aspect, speaker detector 278 generates at least one of a wake-up signal or an interrupt to switch one or more processors 220 from power mode 282 to power mode 284 to activate one or more applications 920.

[0144] Speaker detector 278 determines at 910 in power mode 284 whether multiple speakers have been detected. For example, speaker detector 278 determines whether multiple speakers have been detected after a threshold time has expired since a previous determination that multiple speakers have been detected. In response to determining that multiple speakers have been detected, speaker detector 278 avoids transitioning to power mode 282. Alternatively, in response to determining that no multiple speakers have been detected within a threshold count of the audio feature data set, speaker detector 278 switches one or more processors 220 from power mode 284 to power mode 282.

[0145] Therefore, one or more processors 220 conserve energy by operating in power mode 282 (compared to power mode 284) and switch to power mode 284 when needed to activate components that are not operational in power mode 282. Selectively switching to power mode 284 reduces the overall power consumption of device 202.

[0146] refer to Figure 10 This illustrates a specific implementation of a method 1000 for managing user voice profiles. In a particular aspect, one or more operations of method 1000 are performed by... Figure 1 Segmenter 124, Profile Manager 126 Figure 2A The speaker detector 278, one or more processors 220, device 202, system 200, or a combination thereof shall be used to perform the operation.

[0147] Method 1000 includes: at 1002, determining, in a first power mode, whether the audio stream corresponds to the speech of at least two different speakers. For example, Figure 2A The speaker detector 278 in power mode 282 determines whether audio stream 141 corresponds to the speech of at least two different speakers, as shown in the reference. Figure 2A As described.

[0148] Method 1000 includes: at 1004, based on determining that the audio stream corresponds to the speech of at least two different speakers, analyzing audio feature data of the audio stream in a second power mode to generate segmentation results. For example, Figure 2A One or more processors 220, based on determining that the audio stream 141 corresponds to the speech of at least two different speakers, switch to power mode 284 and activate segmenter 124, as referenced. Figure 2A As described. Segmenter 124 analyzes the audio feature data set 252 of audio stream 141 in power mode 284 to generate segmentation results 236, as referenced. Figure 2A As described. Segmentation result 236 indicates speaker-homogeneous audio segments of audio stream 141 (e.g., speaker-homogeneous audio segment 111A and speaker-homogeneous audio segment 111B), as referenced. Figure 2A As described.

[0149] Method 1000 further includes, at 1006, performing a comparison of a first set of audio feature data from a first plurality of audio feature data sets of homogeneous audio segments of a first speaker with multiple user voice profiles to determine whether the first set of audio feature data matches any of the multiple user voice profiles. For example, Figure 1 The profile manager 126 performs a comparison of audio feature data set 252 in one or more audio feature data sets 252A of multiple user voice profiles 150 and speaker homogeneous audio segments 111A to determine whether audio feature data set 252 matches any of the multiple user voice profiles 150, as shown in the reference. Figure 2B As described.

[0150] Method 1000 further includes: at 1008, based on determining that a first set of audio feature data does not match any of the plurality of user voice profiles: generating a first user voice profile based on the first plurality of audio feature data sets and adding the first user voice profile to the plurality of user voice profiles. For example, Figure 1 The profile manager 126 determines that the audio feature data set 252 does not match any of the multiple user voice profiles 150, generates a user voice profile 150C based on at least one subset of one or more audio feature data sets 252A, and adds the user voice profile 150C to the multiple user voice profiles 150, as shown in the reference. Figure 2B As described.

[0151] Method 1000 enables the generation of user speech profiles based on a set of audio feature data of homogeneous audio segments from the same speaker. Compared to generating user speech profiles based on a single set of audio feature data, using multiple sets of audio feature data corresponding to the speech of the same speaker improves the accuracy of the user speech profile in representing the speaker's speech. Passive registration can be used to generate user speech profiles without requiring the user to pre-register or speak predetermined words or sentences.

[0152] Figure 10 Method 1000 can be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit (e.g., a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. For example, Figure 10 Method 1000 can be executed by the processor that executes the instructions, for example, refer to Figure 19 As described.

[0153] Figure 11 The device 202 is depicted as an implementation 1100 of an integrated circuit 1102 including one or more processors 220. The one or more processors 220 include multiple applications 1122. Applications 1122 include a feature extractor 222, a speaker detector 278, a segmenter 124, a profile manager 126, one or more audio analysis applications 180, or combinations thereof. The integrated circuit 1102 also includes an audio input 1104 (e.g., one or more bus interfaces) to enable receiving an audio stream 141 for processing. The integrated circuit 1102 also includes a signal output 1106 (e.g., a bus interface) to enable sending an output signal 1143, such as a profile ID 155. The integrated circuit 1102 enables user voice profile management to be implemented as a system including a microphone (e.g., as in...). Figure 12 The mobile phone or tablet device depicted in the image, such as... Figure 13 The headphones depicted in the text, as in Figure 14 The wearable electronic devices depicted in the text, such as Figure 15 The voice-controlled speaker system depicted in the text, such as in Figure 16 The virtual reality headset or augmented reality headset depicted in the text, or as in Figure 17 or Figure 18 Components of the vehicle depicted in the text.

[0154] Figure 12Implementation 1200 is depicted, wherein device 202 includes mobile device 1202, such as a mobile phone or tablet device, as an illustrative and non-limiting example. Mobile device 1202 includes microphone 246 and display screen 1204. Components of one or more processors 220 (including application 1122) are integrated into mobile device 1202 and are shown using dashed lines to indicate internal components that are generally not visible to the user of mobile device 1202. In a particular example, feature extractor 222, segmenter 124, and profile manager 126 of application 1122 operate to manage a user voice profile, which is then used to perform one or more operations at mobile device 1202, such as launching a graphical user interface or otherwise displaying other information associated with the user's voice (e.g., a conversation transcript) at display screen 1204 (e.g., via an integrated "smart assistant" application).

[0155] Figure 13 An implementation 1300 is depicted, wherein device 202 includes headset device 1302. Headset device 1302 includes microphone 246. Components of one or more processors 220 (including application 1122) are integrated into headset device 1302. In a particular example, feature extractor 222, segmenter 124, and profile manager 126 of application 1122 operate to manage user voice profiles, which can cause headset device 1302 to perform one or more operations at headset device 1302, such as sending information corresponding to the user's voice to a second device (not shown) (e.g., Figure 2B (e.g., profile update data 272, user interaction data 274, or both) for further processing, or a combination thereof.

[0156] Figure 14An implementation 1400 is depicted, in which device 202 includes wearable electronic device 1402, shown as a "smartwatch". Application 1122 and microphone 246 are integrated into wearable electronic device 1402. In a particular example, feature extractor 222, segmenter 124, and profile manager 126 of application 1122 operate to manage a user voice profile, which is then used to perform one or more operations at wearable electronic device 1402, such as launching a graphical user interface or otherwise displaying other information associated with the user's voice on display screen 1404 of wearable electronic device 1402. For example, wearable electronic device 1402 may include display screen 1404 configured to display notifications (e.g., options for adding calendar events) based on user voice detected by wearable electronic device 1402. In a particular example, wearable electronic device 1402 includes a haptic device that provides haptic notifications (e.g., vibration) in response to detection of user voice. For example, a haptic notification could allow a user to view the wearable electronic device 1402 to see a displayed notification indicating that a keyword spoken by the user has been detected. Therefore, the wearable electronic device 1402 could alert a user with a hearing impairment or a user wearing headphones to the detection of their speech. In a specific example, the wearable electronic device 1402 could display a transcript of a conversation in response to the detection of speech.

[0157] Figure 15 Implementation 1500 includes a wireless speaker and a voice activation device 1502. The wireless speaker and voice activation device 1502 may have a wireless network connection and are configured to perform assistant operations. One or more processors 220, a microphone 246, or a combination thereof, including application 1122, are included in the wireless speaker and voice activation device 1502. The wireless speaker and voice activation device 1502 also includes a speaker 1504. During operation, in response to receiving a spoken command from a user's voice identified as associated with a user's voice profile 150A via the operation of the feature extractor 222, segmenter 124, and profile manager 126 of application 1122, the wireless speaker and voice activation device 1502 can perform assistant operations, such as via a voice activation system (e.g., an integrated assistant application). Assistant operations may include adjusting the temperature, playing music, turning on lights, etc. For example, an assistant operation may be performed in response to receiving a command following a keyword or key phrase (e.g., "Hello, assistant"). In certain aspects, assistant operations include executing user-specific commands (e.g., “Set an appointment for 2 p.m. tomorrow in my calendar” or “Enhance the heating in my room”) for a user associated with the user’s voice profile 150A.

[0158] Figure 16An implementation 1600 is depicted, wherein device 202 includes a portable electronic device corresponding to a virtual reality, augmented reality, or mixed reality headset 1602. Application 1122, microphone 246, or a combination thereof are integrated into the headset 1602. A visual interface device 1620 is positioned in front of the user's eyes to enable the display of augmented reality or virtual reality images or scenes to the user when the headset 1602 is worn. In a particular example, the visual interface device is configured to display a notification indicating user speech detected in an audio signal received from microphone 246. In a particular aspect, the visual interface device is configured to display a dialogue transcript picked up by microphone 246.

[0159] Figure 17 An implementation 1700 is depicted, wherein device 202 corresponds to or is integrated within vehicle 1702 (shown as a manned or unmanned aerial device (e.g., a package delivery drone)). Application 1122, microphone 246, or a combination thereof are integrated into vehicle 1702. Speech analysis can be performed based on audio signals received from microphone 246 of vehicle 1702, for example, to generate a transcript of the conversation captured by microphone 246.

[0160] Figure 18Another implementation 1800 is depicted, in which device 202 corresponds to or is integrated within vehicle 1802 (shown as a car). Vehicle 1802 includes one or more processors 220 containing application 1122. Vehicle 1802 also includes microphone 246. Microphone 246 is positioned to capture the speech of one or more passengers of vehicle 1802. User speech analysis can be performed based on audio signals received from microphone 246 of vehicle 1802. In some implementations, user speech analysis can be performed based on audio signals received from an internal microphone (e.g., microphone 246) (e.g., conversations between passengers of vehicle 1802). For example, user speech analysis can be used to set calendar events for users associated with a particular user voice profile based on conversations detected in vehicle 1802 (e.g., “We’re going on a picnic Saturday afternoon” and “Sure, that would be great”). In some implementations, user speech analysis may be performed based on audio signals received from an external microphone (e.g., microphone 246) (e.g., a user speaking outside vehicle 1802). In a particular implementation, in response to detecting a specific dialogue between users associated with a specific voice profile, application 1122 initiates one or more operations on vehicle 1802 based on the detected dialogue, the detected user, or both, for example, by providing feedback or information via display 1820 or one or more speakers (e.g., speaker 1830) (e.g., “User 1 has a prior commitment to arrange a picnic at 4 pm on Saturday until 3 pm?”).

[0161] refer to Figure 19 A block diagram depicts a specific illustrative implementation of the device, which is generally designated as 1900. In various implementations, device 1900 may have more than... Figure 19 The components shown may be more or fewer. In an illustrative implementation, device 1900 may correspond to device 202. In an illustrative implementation, device 1900 may perform the reference... Figure 1-18 One or more operations described.

[0162] In a particular implementation, device 1900 includes a processor 1906 (e.g., a central processing unit (CPU)). Device 1900 may include one or more additional processors 1910 (e.g., one or more DSPs). In a particular aspect, Figure 2A One or more processors 220 in the processor correspond to processor 1906, processor 1910, or a combination thereof. Processor 1910 may include feature extractor 222, speaker detector 278, segmenter 124, profile manager 126, one or more audio analysis applications 180, or a combination thereof.

[0163] Device 1900 may include memory 1986 and CODEC 1934. In a particular aspect, memory 1986 corresponds to Figure 2A The memory 232. The memory 1986 may include instructions 1956, which may be executed by one or more additional processors 1910 (or processor 1906) to implement the functions described by the reference feature extractor 222, speaker detector 278, segmenter 124, profile manager 126, one or more audio analysis applications 180, or combinations thereof. The device 1900 may include a wireless controller 2841940 coupled to antenna 1952 via transceiver 1950. In a particular aspect, the device 1900 includes a modem coupled to transceiver 1950.

[0164] Device 1900 may include a display 1928 coupled to display controller 1926. One or more speakers 1992, microphone 246, or a combination thereof may be coupled to CODEC 1934. CODEC 1934 may include digital-to-analog converter (DAC) 1902, analog-to-digital converter (ADC) 1904, or both. In a particular implementation, CODEC 1934 may receive analog signals from microphone 246, convert the analog signals to digital signals using ADC 1904, and provide the digital signals to one or more processors 1910. One or more processors 1910 may process the digital signals. In a particular implementation, one or more processors 1910 may provide digital signals to CODEC 1934. CODEC 1934 may use ADC 1902 to convert the digital signals to analog signals and may provide the analog signals to speaker 1992.

[0165] In a particular implementation, device 1900 may be included in a system-in-package (SoC) or a system-on-a-chip (SoC) 1922. In a particular implementation, memory 1986, processor 1906, processor 1910, display controller 1926, CODEC 1934, wireless controller 284 1940, and transceiver 1950 are included in a SoC or SoC 1922. In a particular implementation, input device 1930 and power supply 1944 are coupled to SoC 1922. Furthermore, in a particular implementation, such as... Figure 19 As shown, the display 1928, input device 1930, speaker 1992, microphone 246, antenna 1952, and power supply 1944 are external to the system-on-chip device 1922. In a particular implementation, each of the display 1928, input device 1930, speaker 1992, microphone 246, antenna 1952, and power supply 1944 may be coupled to a component of the system-on-chip device 1922, such as an interface or controller.

[0166] Device 1900 may include smart speakers, soundbars, mobile communication devices, smartphones, cellular phones, laptops, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radio units, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, headphones, augmented reality headphones, virtual reality headphones, aircraft, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computing devices, communication devices, Internet of Things (IoT) devices, virtual reality (VR) devices, base stations, mobile devices, or any combination thereof.

[0167] In conjunction with the described implementation, an apparatus includes a unit for storing multiple user voice profiles of multiple users. For example, the storage unit includes... Figure 2A The memory 232, device 202, system 200, memory 1986, device 1900, one or more other circuits or components, or any combination thereof, configured to store multiple user voice profiles.

[0168] The device also includes a unit for determining whether the audio stream corresponds to the speech of at least two different speakers in a first power mode. For example, the unit for determination includes... Figure 2A The speaker detector 278, one or more processors 220, device 202, system 200, processor 1906, one or more processors 1910, device 1900, are configured in a first power mode to determine whether an audio stream corresponds to the speech of at least two different speakers, or one or more other circuits or components, or any combination thereof.

[0169] The device also includes a unit for analyzing audio feature data of the audio stream to generate segmentation results. For example, the unit for analysis includes... Figure 2A The segmenter 124, one or more processors 220, device 202, system 200, processor 1906, one or more processors 1910, device 1900, one or more other circuits or components, or any combination thereof, configured to analyze audio feature data. Segmentation result 236 indicates speaker-homogeneous audio segments of audio stream 141.

[0170] The apparatus further includes a unit for performing a comparison of a first audio feature data set from a plurality of user voice profiles with a homogeneous audio segment of a first speaker, to determine whether the first audio feature data set matches any of the plurality of user voice profiles. For example, the unit for performing the comparison includes... Figure 2AThe profile manager 126, one or more processors 220, device 202, system 200, processor 1906, one or more processors 1910, device 1900, one or more other circuits or components configured to perform the comparison, or any combination thereof.

[0171] The apparatus also includes a unit for generating a first user voice profile based on a first plurality of audio feature data sets. For example, the unit for generating the first user voice profile includes... Figure 2A The profile manager 126, one or more processors 220, device 202, system 200, processor 1906, one or more processors 1910, device 1900, one or more other circuits or components, or any combination thereof, configured to generate a first user voice profile. The user voice profile 150A is generated based on the determination that the audio feature data set 252 does not match any of the plurality of user voice profiles 150.

[0172] The device also includes a unit for adding a first user voice profile to multiple user voice profiles. For example, the unit for adding the first user voice profile includes... Figure 2A The profile manager 126, one or more processors 220, device 202, system 200, processor 1906, one or more processors 1910, device 1900, one or more other circuits or components, or any combination thereof, configured to add a first user voice profile.

[0173] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 1986) includes instructions (e.g., instruction 1956) that, when executed by one or more processors (e.g., one or more processors 1910 or 1906), cause the one or more processors to determine, in a first power mode (e.g., power mode 282), whether an audio stream (e.g., audio stream 141) corresponds to the speech of at least two different speakers. The instructions, when executed by the one or more processors, also cause the processors to analyze audio feature data of the audio stream (e.g., audio feature data set 252) to generate segmentation results (e.g., segmentation results 236). The segmentation results indicate speaker-homogeneous audio segments of the audio stream (e.g., speaker-homogeneous audio segment 111A and speaker-homogeneous audio segment 111B). When executed by one or more processors, the instructions further cause the processors to perform a comparison of a first audio feature data set (e.g., audio feature data set 252) within a first plurality of audio feature data sets (e.g., audio feature data set 252A) of a plurality of user voice profiles (e.g., plurality of user voice profiles 150) and a first speaker homogeneous audio segment (e.g., speaker homogeneous audio segment 111A), to determine whether the first audio feature data set matches any of the plurality of user voice profiles. When executed by one or more processors, the instructions further cause the processors, based on the determination that the first audio feature data set does not match any of the plurality of user voice profiles, to: generate a first user voice profile (e.g., user voice profile 150A) based on the first plurality of audio feature data sets, and add the first user voice profile to the plurality of user voice profiles.

[0174] Specific aspects of this disclosure are described below in the first set of interconnected clauses:

[0175] According to Clause 1, an apparatus for audio analysis includes: a memory configured to store multiple user voice profiles of multiple users; and one or more processors configured to: in a first power mode, determine whether an audio stream corresponds to the speech of at least two different speakers; based on the determination that the audio stream corresponds to the speech of at least two different speakers, analyze audio feature data of the audio stream in a second power mode to generate segmentation results indicating speaker-homogeneous audio segments of the audio stream; perform a comparison of a first set of audio feature data from a first plurality of audio feature data sets of the multiple user voice profiles with a first set of speaker-homogeneous audio segments to determine whether the first set of audio feature data matches any of the multiple user voice profiles; and based on the determination that the first set of audio feature data does not match any of the multiple user voice profiles: generate a first user voice profile based on the first plurality of audio feature data sets; and add the first user voice profile to the multiple user voice profiles.

[0176] Clause 2 includes the device described in Clause 1, wherein the first audio feature data set includes a first audio feature vector.

[0177] Clause 3 includes the device described in Clause 1 or Clause 2, wherein one or more processors are configured to analyze audio feature data by applying a speaker segmentation neural network to the audio feature data.

[0178] Clause 4 includes a device according to any one of Clauses 1 to 3, wherein one or more processors are configured to: indicate, based on a determined segmentation result, that a first audio feature data set corresponds to the speech of a first speaker, and that the first audio feature data set does not match any of a plurality of user speech profiles; store the first audio feature data set in a first registration buffer associated with the first speaker; and store a subsequent audio feature data set corresponding to the speech of the first speaker in the first registration buffer until a stopping condition is met, wherein a first plurality of audio feature data sets of homogeneous audio segments of the first speaker include the first audio feature data set and the subsequent audio feature data sets.

[0179] Clause 5 includes the device described in Clause 4, wherein one or more processors are configured to determine that a stop condition is met in response to determining that a silence longer than a threshold is detected in the audio stream.

[0180] Clause 6 includes a device pursuant to any one of Clauses 4 to 5, wherein one or more processors are configured to add a particular set of audio feature data to a first registration buffer based at least in part on determining that a particular set of audio feature data corresponds to the speech of a single speaker, wherein the single speaker includes a first speaker.

[0181] Clause 7 includes a device according to any one of Clauses 1 to 6, wherein one or more processors are configured to: generate a first user voice profile based on a first plurality of audio feature data sets stored in a first registration buffer, based on a count of a first plurality of audio feature data sets that determine a first speaker homogeneous audio segment being greater than a registration threshold.

[0182] Clause 8 includes a device pursuant to any one of Clauses 1 to 7, wherein one or more processors are configured to: update a particular user's voice profile based on the first audio feature data set, based on determining that the first audio feature data set matches a particular user's voice profile.

[0183] Clause 9 includes the device described in Clause 8, wherein one or more processors are configured to: update a specific user voice profile based on the first audio feature data set, at least in part, based on determining that the first audio feature data set corresponds to the speech of a single speaker.

[0184] Clause 10 includes a device according to any one of Clauses 1 to 9, wherein one or more processors are configured to: determine whether a second set of audio feature data in a second plurality of audio feature data sets of homogeneous audio segments of a second speaker matches any one of a plurality of user voice profiles.

[0185] Clause 11 includes the device described in Clause 10, wherein one or more processors are configured to: generate a second user voice profile based on a second plurality of audio feature data sets that do not match any of the plurality of user voice profiles; and add the second user voice profile to the plurality of user voice profiles.

[0186] Clause 12 includes the device described in Clause 10, wherein one or more processors are configured to: update a specific user voice profile based on the second audio feature data set, based on determining that a second audio feature data set matches a specific user voice profile among a plurality of user voice profiles.

[0187] Clause 13 includes an apparatus according to any one of Clauses 1 to 12, wherein the memory is configured to store profile update data, and wherein one or more processors are configured to: update the profile update data in response to generating a first user voice profile to indicate that the first user voice profile has been updated; and output the first count as a count of speakers detected in the audio stream based on determining that a first count of a plurality of user voice profiles has been updated.

[0188] Clause 14 includes an apparatus according to any one of Clauses 1 to 13, wherein the memory is configured to store user interaction data, and wherein one or more processors are configured to: update the user interaction data based on the duration of a homogeneous audio segment of a first speaker in response to generating a first user voice profile, to instruct a first user associated with the first user voice profile to interact during the duration of the voice; and at least output the user interaction data.

[0189] Clause 15 includes the device pursuant to any one of Clauses 1 to 14, wherein the first power mode is a lower power mode compared to the second power mode.

[0190] Clause 16 includes the device described in Clause 1, wherein one or more processors are configured to: determine audio information of an audio stream in a first power mode, the audio information including speaker counts, voice activity detection (VAD) information, or both, detected in the audio stream; activate one or more audio analysis applications in a second power mode; and provide the audio information to one or more audio analysis applications.

[0191] Clause 17 includes a device pursuant to any one of Clauses 1 to 16, wherein one or more processors are configured to: in response to determining that one or more second audio segments of an audio stream correspond to multiple speakers, avoid updating multiple user voice profiles based on one or more second audio segments.

[0192] Specific aspects of this disclosure are described below in the second set of interconnected clauses:

[0193] According to Clause 18, an audio analysis method includes: determining, while the device is in a first power mode, whether an audio stream corresponds to the speech of at least two different speakers; based on the determination that the audio stream corresponds to the speech of at least two different speakers, analyzing audio feature data of the audio stream in a second power mode to generate segmentation results indicating speaker-homogeneous audio segments of the audio stream; performing a comparison at the device of a first set of audio feature data from a first plurality of audio feature data sets of a plurality of user voice profiles with the first speaker-homogeneous audio segments to determine whether the first set of audio feature data matches any of the plurality of user voice profiles; and based on the determination that the first set of audio feature data does not match any of the plurality of user voice profiles: generating a first user voice profile at the device based on the first plurality of audio feature data sets; and adding the first user voice profile to the plurality of user voice profiles at the device.

[0194] Clause 19 includes the method described in accordance with Clause 18, and further includes applying a speaker segmentation neural network to audio feature data.

[0195] Clause 20 includes the method described in accordance with Clause 18 or Clause 19, and further includes: determining, based on a segmentation result, that a first audio feature data set corresponds to the speech of a first speaker, and that the first audio feature data set does not match any of a plurality of user speech profiles; storing the first audio feature data set in a first registration buffer associated with the first speaker; and storing subsequent audio feature data sets corresponding to the speech of the first speaker in the first registration buffer until a stopping condition is met, wherein a first plurality of audio feature data sets of homogeneous audio segments of the first speaker include the first audio feature data set and the subsequent audio feature data sets.

[0196] Clause 21 includes the method described in accordance with Clause 20, and further includes: at the device, determining that a stop condition is met in response to determining that a silence longer than a threshold is detected in the audio stream.

[0197] Clause 22 includes the method described in accordance with Clause 20 or Clause 21, and further includes: at the device, adding a specific audio feature data set to a first registration buffer, at least in part based on determining that a specific audio feature data set corresponds to the speech of a single speaker, wherein the single speaker includes a first speaker.

[0198] Clause 23 includes the method according to any one of Clauses 18 to 22, and further includes: generating a first user voice profile based on the first plurality of audio feature data sets stored in a first registration buffer, based on the determination that the count of homogeneous audio segments of the first speaker is greater than a registration threshold.

[0199] Clause 24 includes the method pursuant to any one of Clauses 18 to 23, and further includes: updating the specific user voice profile based on the first audio feature data set based on determining that a first audio feature data set matches a specific user voice profile.

[0200] Clause 25 includes the method described in accordance with Clause 24, and further includes: updating a specific user's voice profile based on the first audio feature data set, at least in part, based on determining that the first audio feature data set corresponds to the voice of a single speaker.

[0201] Clause 26 includes the method pursuant to any one of Clauses 18 to 25, and further includes: updating the specific user voice profile based on the second audio feature data set in a second plurality of audio feature data sets that determines a second audio feature data set of homogeneous audio segments of a second speaker and matching it with a specific user voice profile in a plurality of user voice profiles.

[0202] Specific aspects of this disclosure are described below in the third set of interconnected clauses:

[0203] According to Clause 27, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors, cause the processor to: determine, in a first power mode, whether an audio stream corresponds to the speech of at least two different speakers; based on the determination that the audio stream corresponds to the speech of at least two different speakers, analyze audio feature data of the audio stream in a second power mode to generate segmentation results indicating speaker-homogeneous audio segments of the audio stream; perform a comparison of a first set of audio feature data from a plurality of user voice profiles with a first set of audio feature data from a plurality of first speaker-homogeneous audio segments to determine whether the first set of audio feature data matches any of the plurality of user voice profiles; and based on the determination that the first set of audio feature data does not match any of the plurality of user voice profiles: generate a first user voice profile based on the first plurality of audio feature data sets; and add the first user voice profile to the plurality of user voice profiles.

[0204] Clause 28 includes a non-transitory computer-readable storage medium as described in Clause 27, wherein the instructions, when executed by one or more processors, cause the processor to: generate a first user voice profile based on a first plurality of audio feature data sets stored in a first registration buffer, based on the determination that a count of homogeneous audio segments of a first speaker is greater than a registration threshold.

[0205] Specific aspects of this disclosure are described below in the fourth set of interconnected clauses:

[0206] According to Clause 29, an apparatus includes: a unit for storing multiple user voice profiles of multiple users; a unit for determining, in a first power mode, whether an audio stream corresponds to the speech of at least two different speakers; a unit for analyzing audio feature data of the audio stream in a second power mode to generate segmentation results, the audio feature data being analyzed in the second power mode based on the determination that the audio stream corresponds to the speech of at least two different speakers, wherein the segmentation results indicate speaker-homogeneous audio segments of the audio stream; a unit for performing a comparison of a first set of audio feature data from the multiple user voice profiles with a first set of first speaker-homogeneous audio segments to determine whether the first set of audio feature data matches any of the multiple user voice profiles; a unit for generating a first user voice profile based on the first set of first audio feature data, the first user voice profile being generated based on the determination that the first set of audio feature data does not match any of the multiple user voice profiles; and a unit for adding the first user voice profile to the multiple user voice profiles.

[0207] Clause 30 includes the apparatus described in Clause 29, wherein the unit for storage, the unit for determination, the unit for analysis, the unit for execution, the unit for generation, and the unit for addition are integrated into at least one of: mobile communication devices, smartphones, cellular phones, smart speakers, soundbars, laptops, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radio units, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, headphones, augmented reality headphones, virtual reality headphones, aircraft, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computing devices, communication devices, Internet of Things (IoT) devices, virtual reality (VR) devices, base stations, mobile devices, or any combination thereof.

[0208] Those skilled in the art will further understand that the various illustrative logic blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various illustrative components, blocks, configurations, modules, circuits, and steps have been generally described above regarding their functionality. Whether this functionality is implemented as hardware or processor-executable instructions depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, and such implementation decisions should not be construed as departing from the scope of this disclosure.

[0209] The steps of the methods or algorithms described in conjunction with the implementations disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compressed optical disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium can be a component of the processor. The processor and storage medium can reside in an application-specific integrated circuit (ASIC). The ASIC can reside in a computing device or user terminal. Alternatively, the processor and storage medium can reside as discrete components in a computing device or user terminal.

[0210] The prior description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the broadest possible scope consistent with the principles and novel features defined by the following claims.

Claims

1. A device for audio analysis, comprising: The memory is configured to store multiple user voice profiles for multiple users; as well as One or more processors, which are configured as follows: In a first power mode, it is determined whether the audio stream corresponds to the speech of at least two different speakers, wherein the first power mode is a lower power mode of the device; Based on determining that the audio stream corresponds to the speech of at least two different speakers, the system switches from a first power mode to a second power mode and analyzes the audio feature data of the audio stream in the second power mode to generate segmentation results, wherein the second power mode is a higher power mode compared to the first power mode, and the segmentation results indicate speaker-homogeneous audio segments of the audio stream. Perform a comparison of a first audio feature data set from a first plurality of audio feature data sets of homogeneous audio segments of a first speaker with the plurality of user voice profiles to determine whether the first audio feature data set matches any user voice profile among the plurality of user voice profiles; and Based on the determination that the segmentation result indicates that the first audio feature data set corresponds to the speech of the first speaker, and that the first audio feature data set does not match any of the plurality of user speech profiles: The first audio feature data set is stored in a first registration buffer associated with the first speaker; The set of subsequent audio feature data corresponding to the speech of the first speaker is stored in the first registration buffer until the stopping condition is met, wherein the first plurality of audio feature data sets of the homogeneous audio segments of the first speaker include the first audio feature data set and the set of subsequent audio feature data. A first user voice profile is generated based on the first plurality of audio feature data sets; and Add the first user voice profile to the plurality of user voice profiles.

2. The apparatus of claim 1, wherein, The first audio feature data set includes a first audio feature vector.

3. The apparatus of claim 1, wherein, The one or more processors are configured to analyze the audio feature data by applying a speaker segmentation neural network to the audio feature data.

4. The apparatus of claim 1, wherein, The one or more processors are configured to: determine that the stopping condition is met in response to determining that a silence longer than a threshold is detected in the audio stream.

5. The device according to claim 1, wherein, The one or more processors are configured to add the third audio feature data set to the first registration buffer, at least in part, based on determining that the third audio feature data set corresponds to the speech of a single speaker, wherein the single speaker includes the first speaker.

6. The device according to claim 1, wherein, The one or more processors are configured to generate the first user voice profile based on the first plurality of audio feature data sets stored in the first registration buffer, based on the determination that the count of the first speaker homogeneous audio segments is greater than a registration threshold.

7. The device according to claim 1, wherein, The one or more processors are configured to update the third user voice profile based on the first audio feature data set, based on determining that the first audio feature data set matches the third user voice profile.

8. The device according to claim 7, wherein, The one or more processors are configured to update the third user voice profile based on the first audio feature data set, at least in part, based on determining that the first audio feature data set corresponds to the speech of a single speaker.

9. The device according to claim 1, wherein, The one or more processors are configured to: determine whether a second set of audio feature data in a second plurality of audio feature data sets of homogeneous audio segments of a second speaker matches any user voice profile in the plurality of user voice profiles.

10. The device according to claim 9, wherein, The one or more processors are configured to: determine that the second audio feature data set does not match any of the plurality of user voice profiles; A second user voice profile is generated based on the second set of multiple audio feature data. as well as Add the second user voice profile to the plurality of user voice profiles.

11. The device according to claim 9, wherein, The one or more processors are configured to update the fourth user voice profile based on the second audio feature data set, based on determining that the second audio feature data set matches the fourth user voice profile among the plurality of user voice profiles.

12. The device according to claim 1, wherein, The memory is configured to store profile update data, and wherein the one or more processors are configured to: In response to generating the first user voice profile, the profile update data is updated to indicate that the first user voice profile has been updated; and Based on the determination that the profile update data indicates that a first count of the plurality of user voice profiles has been updated, the first count is output as the count of the speaker detected in the audio stream.

13. The device according to claim 1, wherein, The memory is configured to store user interaction data, and wherein the one or more processors are configured to: In response to generating the first user voice profile, the user interaction data is updated based on the duration of a homogeneous audio segment of the first speaker to indicate that the first user associated with the first user voice profile interacts within the duration of the audio segment; and At least the user interaction data should be output.

14. The device according to claim 1, wherein, The one or more processors are configured to: In the first power mode, audio information of the audio stream is determined, including speaker counts detected in the audio stream, voice activity detection (VAD) information, or both. Activate one or more audio analysis applications in the second power mode; as well as The audio information is provided to one or more audio analysis applications.

15. The device according to claim 1, wherein, The one or more processors are configured to: in response to determining that the segmentation result indicates that one or more second audio segments of the audio stream correspond to multiple speakers, avoid updating the multiple user voice profiles based on the one or more second audio segments.

16. An audio analysis method, comprising: In a first power mode, it is determined whether the audio stream corresponds to the speech of at least two different speakers, wherein the first power mode is a lower power mode of the device; Based on determining that the audio stream corresponds to the speech of at least two different speakers, the system switches from a first power mode to a second power mode and analyzes the audio feature data of the audio stream in the second power mode to generate segmentation results, wherein the second power mode is a higher power mode compared to the first power mode, and the segmentation results indicate speaker-homogeneous audio segments of the audio stream. At the device, a comparison is performed on a first set of audio feature data from a first plurality of audio feature data sets of multiple user voice profiles and homogeneous audio segments of a first speaker to determine whether the first set of audio feature data matches any of the multiple user voice profiles; and Based on the determination that the segmentation result indicates that the first audio feature data set corresponds to the speech of the first speaker, and that the first audio feature data set does not match any of the plurality of user speech profiles: The first audio feature data set is stored in a first registration buffer associated with the first speaker; The set of subsequent audio feature data corresponding to the speech of the first speaker is stored in the first registration buffer until the stopping condition is met, wherein the first plurality of audio feature data sets of the homogeneous audio segments of the first speaker include the first audio feature data set and the set of subsequent audio feature data. At the device, a first user voice profile is generated based on the first plurality of audio feature data sets; and The first user voice profile is added to the plurality of user voice profiles at the device.

17. The method of claim 16, further comprising: A speaker segmentation neural network is applied to the audio feature data.

18. The method of claim 16, further comprising: At the device, in response to determining that a silence longer than a threshold is detected in the audio stream, the stopping condition is determined to be met.

19. The method of claim 16, further comprising: At the device, the third audio feature data set is added to the first registration buffer based at least in part on determining that the third audio feature data set corresponds to the speech of a single speaker, wherein the single speaker includes the first speaker.

20. The method of claim 16, further comprising: Based on the determination that the count of the first plurality of audio feature data sets stored in the first registration buffer for the first speaker's homogeneous audio segments is greater than the registration threshold, the first user voice profile is generated based on the first plurality of audio feature data sets.

21. The method of claim 16, further comprising: Based on the determination that the first audio feature data set matches the second user voice profile, the second user voice profile is updated based on the first audio feature data set.

22. The method of claim 21, further comprising: The second user voice profile is updated based at least in part on the determination that the first audio feature data set corresponds to the speech of a single speaker.

23. The method of claim 16, further comprising: Based on the second audio feature data set in the second plurality of audio feature data sets that determine the homogeneous audio segments of the second speaker, and the third user voice profile in the plurality of user voice profiles, the third user voice profile is updated based on the second audio feature data set.

24. A non-transitory computer-readable storage medium storing instructions, said instructions, when executed by one or more processors, causing said one or more processors to perform the following operations: In the first power mode, it is determined whether the audio stream corresponds to the speech of at least two different speakers, where, The first power mode is a lower power mode; Based on determining that the audio stream corresponds to the speech of at least two different speakers, the system switches from a first power mode to a second power mode and analyzes the audio feature data of the audio stream in the second power mode to generate segmentation results, wherein the second power mode is a higher power mode compared to the first power mode, and the segmentation results indicate speaker-homogeneous audio segments of the audio stream. Perform a comparison of a first audio feature data set from a first plurality of audio feature data sets of multiple user voice profiles with homogeneous audio segments of a first speaker to determine whether the first audio feature data set matches any user voice profile among the plurality of user voice profiles; and Based on the determination that the segmentation result indicates that the first audio feature data set corresponds to the speech of the first speaker, and that the first audio feature data set does not match any of the plurality of user speech profiles: The first audio feature data set is stored in a first registration buffer associated with the first speaker; The set of subsequent audio feature data corresponding to the speech of the first speaker is stored in the first registration buffer until the stopping condition is met, wherein the first plurality of audio feature data sets of the homogeneous audio segments of the first speaker include the first audio feature data set and the set of subsequent audio feature data. A first user voice profile is generated based on the first plurality of audio feature data sets; and Add the first user voice profile to the plurality of user voice profiles.

25. The non-transitory computer-readable storage medium according to claim 24, wherein, When executed by the one or more processors, the instructions cause the one or more processors to perform the following operations: based on determining that the count of the first plurality of audio feature data sets stored in the first registration buffer for the first speaker's homogeneous audio segments is greater than a registration threshold, generate the first user voice profile based on the first plurality of audio feature data sets.

26. An apparatus comprising: A unit used to store multiple user voice profiles; A unit for determining whether an audio stream corresponds to the speech of at least two different speakers in a first power mode, wherein the first power mode is a lower power mode of the device; A unit for converting from a first power mode to a second power mode and analyzing audio feature data of the audio stream in the second power mode to generate segmentation results based on determining that the audio stream corresponds to the speech of at least two different speakers, wherein the second power mode is a higher power mode compared to the first power mode, and the segmentation results indicate speaker-homogeneous audio segments of the audio stream. A unit for performing a comparison of a first audio feature data set in a first plurality of audio feature data sets of the plurality of user voice profiles with a first speaker homogeneous audio segment, to determine whether the first audio feature data set matches any user voice profile in the plurality of user voice profiles; A unit for storing the first audio feature data set in a first registration buffer associated with the first speaker and storing subsequent audio feature data sets corresponding to the first speaker's speech in the first registration buffer based on determining that the segmentation result indicates that the first audio feature data set corresponds to the speech of the first speaker and that the first audio feature data set does not match any of the plurality of user speech profiles, until a stopping condition is met, wherein the first plurality of audio feature data sets of the homogeneous audio segment of the first speaker include the first audio feature data set and the subsequent audio feature data sets; A unit for generating a first user voice profile based on the first plurality of audio feature data sets, the first user voice profile being generated based on determining that the first audio feature data set does not match any of the plurality of user voice profiles; and A unit for adding the first user voice profile to the plurality of user voice profiles.

27. The apparatus according to claim 26, wherein, The storage unit, the determination unit, the analysis unit, the execution unit, the generation unit, and the addition unit are integrated into at least one of the following: mobile communication devices, smartphones, cellular phones, smart speakers, soundbars, laptops, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radio units, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, headphones, augmented reality headphones, virtual reality headphones, aircraft, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computing devices, communication devices, Internet of Things (IoT) devices, virtual reality (VR) devices, base stations, mobile devices, or any combination thereof.