Method, device, and system for performing pinned-state connectionist sequence classification

PS-CSC addresses inefficiencies in neural network training for sequence-based tasks by combining sequence and cluster boundary information, leading to faster and more efficient training processes.

JP2025528850AActive Publication Date: 2025-09-02DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025508841
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-12
Filing Date
2023-08-22
Publication Date
2025-09-02
Estimated Expiration
2043-08-22

AI Technical Summary

Technical Problem

Existing machine learning methods for tasks like speaker diarization and handwriting recognition face inefficiencies due to the need for extensive computational resources and time in training neural networks, particularly when dealing with sequence information and cluster boundary identification.

Method used

The implementation of Pinned-State Connectionist Sequential Classification (PS-CSC) methods, which utilize a loss function combining sequence and cluster boundary information to train neural networks, reducing the number of valid paths and computational overhead by pinning known states, thereby accelerating convergence.

Benefits of technology

PS-CSC enables faster and more efficient training of neural networks for tasks such as speaker diarization and handwriting recognition by minimizing computational resources and time required for convergence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528850000001_ABST
    Figure 2025528850000001_ABST
Patent Text Reader

Abstract

Some disclosed methods include receiving an observation sequence including a plurality of extracted features, each of the plurality of extracted features corresponding to a sequential signal in a series of sequential signals; determining a lattice of posterior probabilities, the lattice including a probability of each observation sequence corresponding to one label class among a plurality of label classes; and applying a loss function to the lattice of posterior probabilities according to ground truth values, the applying the loss function including applying both sequential information and cluster boundary information. Some methods include updating parameters for determining the lattice according to a loss determined by the loss function; and performing the above operations until one or more convergence criteria are met.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to machine learning. [Background technology]

[0002] Several methods, devices, and systems for machine learning are known, and while existing devices, systems, and methods for machine learning may provide benefits in some situations, improvements to the devices, systems, and methods are desirable. Summary of the Invention

[0003] [Notation and Nomenclature] Throughout this disclosure, including the claims, the terms "speaker," "loudspeaker," and "audio reproduction transducer" are used interchangeably to refer to any acoustically emitting transducer (or set of transducers) driven by a single speaker feed. A speaker may be implemented to include multiple transducers (e.g., woofers and tweeters) that may be driven by a single common speaker feed or by multiple speaker feeds. In some examples, the speaker signal may undergo different processing in different circuit branches that are coupled to different transducers.

[0004] Throughout this disclosure, including the claims, the expression "performing an operation on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing the operation on the signal or data directly or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing before performing the operation on the signal).

[0005] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to a plurality of inputs, M of which are generated by the subsystem and the remaining XM inputs are from external sources) may also be referred to as a decoder system.

[0006] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., by firmware or software) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing on audio or other acoustic data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0007] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used to mean a direct or indirect connection. Thus, when a first device is coupled to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.

[0008] As used herein, a "smart device" is an electronic device that can operate interactively and / or autonomously to some degree and is generally configured to communicate with one or more other devices (or networks) via various wireless protocols, such as Bluetooth®, Zigbee®, near-field communications, Wi-Fi, Light Fidelity (Li-Fi), 3G, 4G, 5G, etc. Some notable types of smart devices are smartphones, smart cards, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smart watches, smart bands, smart keychains, and smart audio devices. The term "smart device" can also refer to devices that exhibit some characteristics of ubiquitous computing, such as artificial intelligence.

[0009] Here, we use the expression "smart audio device" to refer to a smart device that is either a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of virtual assistant functionality). A single-purpose audio device is a device (e.g., a television (TV)) that includes or is coupled to at least one microphone (optionally also including or coupled to at least one speaker and / or at least one camera) and is designed primarily or primarily to achieve one purpose. For example, a TV is typically capable of (and is considered capable of) playing audio from program material, but in most cases, modern TVs run some kind of operating system on which applications, including television viewing applications, run locally. In this sense, single-purpose audio devices equipped with speakers and microphones are often configured to run local applications and / or services to directly use the speakers and microphones. Several single-purpose audio devices may be configured to be grouped to achieve audio playback over a zone or user-defined area.

[0010] One common type of multipurpose audio device is a smart audio device, such as a “smart speaker,” which may be configured to implement at least some aspects of virtual assistant functionality, although other aspects of the virtual assistant functionality may be implemented by one or more other devices, such as one or more servers with which the multipurpose audio device is configured to communicate. Such multipurpose audio devices may be referred to herein as “virtual assistants.” A virtual assistant is a device (e.g., a smart speaker or voice-assistant-integrated device) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera). In some examples, a virtual assistant may be cloud-enabled in some sense or otherwise provide the ability to utilize multiple devices (separate from the virtual assistant) for applications that are not fully implemented in the virtual assistant itself. In other words, at least some aspects of the virtual assistant functionality, such as voice recognition functionality, may be implemented (at least in part) by one or more servers or other devices with which the virtual assistant can communicate over a network, such as the Internet. Virtual assistants may sometimes work together, for example, in a conditionally defined, individualized manner. For example, two or more virtual assistants may work together in the sense that one of them, e.g., the one that is most confident that it heard the wake word, responds to the wake word. Connected virtual assistants may, in some implementations, form a kind of constellation that may be managed by one main application that may be (or may implement) a virtual assistant.

[0011] Herein, "wake word" is used broadly to refer to any sound (human-uttered word or other sound) that the smart audio device is configured to wake up in response to detecting ("listening") for the sound (using at least one microphone included in or coupled to the smart audio device or at least one other microphone). In this context, "waking up" refers to the device entering a state in which it awaits (i.e., listens for) an acoustic command. In some cases, what may be referred to herein as a "wake word" may include more than one word, e.g., a phrase.

[0012] Here, the term "wake word detector" refers to a device (or software including instructions for configuring a device) configured to continuously search for alignment between real-time acoustic (e.g., speech) features and a trained model. Typically, a wake word event is triggered when the wake word detector determines that the probability that the wake word has been detected exceeds a predefined threshold. For example, the threshold may be a predetermined threshold that is adjusted to provide a reasonable compromise between false acceptance and false rejection rates. Following a wake word event, the device may enter a state (which may be referred to as an "awake" or "attentive" state) in which it listens for commands and passes received commands to a larger and more computationally intensive recognition engine.

[0013] As used herein, the terms "program stream" and "content stream" refer to a collection of one or more audio signals, and in some cases video signals, where at least some of the signals are intended to be listened to together. For example, a music selection, a movie soundtrack, a movie, a television program, an audio portion of a television program, a podcast, a live voice call, a synthesized voice response from a smart assistant, etc. In some cases, a content stream may include multiple versions of at least a portion of an audio signal, e.g., the same dialogue in more than one language. In such cases, only one version of the audio data or portion thereof (e.g., a version corresponding to a single language) is intended to be played at a time.

[0014] [overview] At least some aspects of the present disclosure may be implemented by one or more methods. In some cases, the methods may be implemented, at least in part, by a control system and / or via instructions (e.g., software) stored on one or more non-transitory media. Some methods may include receiving, by the control system, an observation sequence including a plurality of extracted features. Each extracted feature of the plurality of extracted features may correspond to a sequential signal in a series of sequential signals. Some methods may include determining, by the control system, a lattice of posterior probabilities. The lattice may include a probability of each observation sequence corresponding to one label class of a plurality of label classes.

[0015] Some methods may include applying, by a control system, a loss function to a lattice of posterior probabilities according to ground truth values. In some examples, applying the loss function may include applying both sequential information and cluster boundary information. According to some examples, applying the loss function may include determining one or more valid paths between observations in the lattice.

[0016] Some methods may include updating, by the control system, parameters for determining the lattice of posterior probabilities according to a loss determined by the loss function. Some methods may include continuing to perform the receiving, determining, applying, and updating processes until the control system determines that one or more convergence criteria are met.

[0017] According to some examples, performing the above process may provide a trained neural network that is implemented by a control system. Some methods may include performing one or more downstream tasks with the trained neural network.

[0018] In some examples, the series of sequential signals may be or include a series of audio signals, and in some such examples, the one or more downstream tasks may include acoustic event detection, speaker diarization, automatic speech recognition, or a combination thereof.

[0019] According to some examples, the series of sequential signals may be or include a series of handwritten images, and in some such examples, one or more downstream tasks may include handwriting recognition.

[0020] In some examples, the cluster boundary information may be incomplete cluster boundary information, but in some alternative examples, the cluster boundary information may be complete cluster boundary information.

[0021] Some methods may include determining, by a control system, a plurality of extracted features. In some such examples, the method may include updating, by the control system, parameters for determining the plurality of extracted features according to a loss determined by the loss function.

[0022] In some examples, the sequence information of the loss function, the cluster boundary information of the loss function, or both can be variable. According to some examples, the loss function may combine elements of connectionist time-series classification and cross-entropy loss.

[0023] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure may be embodied by one or more non-transitory media having software stored thereon.

[0024] At least some aspects of the present disclosure may be implemented by an apparatus. For example, one or more devices (e.g., a system including one or more devices) may be capable of at least partially performing the methods disclosed herein. In some implementations, the apparatus is or includes an audio processing system with an interface system and a control system. The control system may include one or more general-purpose single- or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof. The control system may be configured to perform some or all of the methods disclosed herein.

[0025] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. It should be noted that the relative dimensions of the following figures may not be to scale.

[0026] Like reference numbers and designations in the various drawings indicate like elements. [Brief explanation of the drawings]

[0027] [Figure 1A] FIG. 1 is a block diagram illustrating example components of an apparatus capable of implementing various aspects of the disclosure. [Figure 1B] An example of an audio environment is shown below. [Figure 2] 1 illustrates examples of blocks that may be included in some disclosed implementations of pinned-state connectionist sequence classification (PS-CSC). [Figure 3] 1 illustrates examples of blocks that may be included in some disclosed implementations of pinned-state connectionist sequence classification (PS-CSC). [Figure 4] 1 shows an example of the blocks involved in training a neural network for a speaker diarization use case using PS-CSC. [Figure 5] Figures A to C show the results of experiments using the conventional method in the speaker diarization use case. [Figure 6A] We present experimental results using the PS-CSC method in the speaker diarization use case. [Figure 6B] We present experimental results using the PS-CSC method in the speaker diarization use case. [Figure 7] The results of two additional experiments are shown. [Figure 8] FIG. 1 is a flow diagram illustrating an example of the disclosed method. DETAILED DESCRIPTION OF THE INVENTION

[0028] Machine learning methods based on cross-entropy loss and its variants and machine learning methods based on connectionist temporal classification (CTC) use different loss objectives and are generally used for different purposes in the machine learning community. Machine learning methods based on cross-entropy loss are typically used to train neural network models for classification problems, such as image-based face classification and speech signal-based speaker identification. Machine learning methods based on cross-entropy loss may involve obtaining and / or learning cluster boundary information. In contrast, machine learning methods based on CTC are typically applied when sequence information needs to be explored, such as in automatic speech recognition or handwriting recognition.

[0029] This disclosure provides examples of machine learning methods based on cluster boundary information and sequence information. Such methods may be referred to herein as Pinned-State Connectionist Sequential Classification (PS-CSC) methods. In accordance with some PS-CSC examples, the loss function used to train the neural network may be based on sequence information as well as cluster boundary information. In some PS-CSC examples, the cluster boundary information may be or may include incomplete cluster boundary information. In some PS-CSC examples, the loss function used to train the neural network may combine elements of CTC and cross-entropy loss.

[0030] In some PS-CSC examples, the disclosed neural network training process may include receiving an observation sequence including multiple extracted features. According to some examples, the observation sequence may be a time sequence, while in other examples, the observation sequence may be a spatial sequence. Each extracted feature may correspond to a sequential signal in a series of sequential signals. In some cases, the training process may include determining a lattice of posterior possibilities. The lattice may include a probability of each observation sequence corresponding to one label class among multiple label classes. According to some examples, the labels may be monotonically aligned with the observations.

[0031] According to some PS-CSC examples, applying the loss function may include determining one or more valid paths between the observations in the lattice. In some such examples, pinning the observations in the lattice reduces the number of valid paths, thereby reducing the time and computational overhead required for convergence during the neural network training process.

[0032] Neural networks trained according to some of the disclosed methods can be configured to perform one or more downstream tasks. The downstream tasks can vary depending on the type of signals in the series of sequential signals. In some cases, the series of sequential signals can be or include a series of handwritten images, and the downstream task can be handwriting recognition. In other examples, the series of sequential signals can be or include a series of audio signals, and the downstream task can be acoustic event detection, talker diarization, and / or automatic speech recognition.

[0033] FIG. 1A is a block diagram illustrating example components of a device capable of implementing various aspects of the present disclosure. As with other figures provided herein, the types, number, and arrangement of elements shown in FIG. 1A are provided by way of example only. Other implementations may include more, fewer, and / or different types, numbers, and arrangements of elements. According to some examples, device 100 may be configured to perform at least some of the methods disclosed herein. In some implementations, device 100 may be or include one or more components of an office workstation, one or more components of a home entertainment system, or the like. For example, device 100 may be a laptop computer, a tablet device, a mobile device (e.g., a cellular phone), a smart home hub, a television, or other type of device.

[0034] According to some alternative implementations, apparatus 100 may be or include a server. In some such examples, apparatus 100 may be or include an encoder. In some examples, apparatus 100 may be or include a decoder. Thus, in some cases, apparatus 100 may be a device configured for use in an environment, such as a home environment, while in other cases, apparatus 100 may be a device configured for use in the "cloud," such as a server.

[0035] In this example, device 100 includes interface system 105 and control system 110. Interface system 105, in some implementations, may be configured for communication with one or more other devices in an environment. The environment, in some examples, may be a home environment. In other examples, the environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. Interface system 105, in some implementations, may be configured to exchange control information and associated data with other devices in the environment. The control information and associated data, in some examples, may relate to one or more software applications that device 100 is executing.

[0036] The interface system 105, in some implementations, may be configured to receive or provide a content stream. In some examples, the content stream may include video data and audio data corresponding to the video data. The audio data may include, but may not be limited to, an audio signal. In some cases, the audio data may include spatial data, such as channel data and / or spatial metadata. The metadata may be provided, for example, by what may be referred to herein as an "encoder."

[0037] Interface system 105 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more universal serial bus (USB) interfaces). According to some implementations, interface system 105 may include one or more wireless interfaces. Interface system 105 may include one or more devices that implement a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, a gesture sensor system, or combinations thereof. Thus, although some such devices are represented individually in FIG. 1A , such devices may, in some examples, correspond to aspects of interface system 105.

[0038] In some examples, the interface system 105 may include one or more interfaces between a memory system, such as any of the memory systems 115 shown in FIG. 1A, and the control system 110. Alternatively, or additionally, the control system 110 may include a memory system in some cases. The interface system 105 may be configured to receive input from one or more microphones in the environment in some implementations.

[0039] The control system 110 may include, for example, a general purpose single or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or combinations thereof.

[0040] In some implementations, control system 110 may reside in more than one device. For example, in some implementations, a portion of control system 110 may reside in a device within one of the environments referenced herein, while other portions of control system 110 may reside in a device external to the environment, such as a server, a mobile device (e.g., a smartphone, or a tablet computer), or the like. In other examples, a portion of control system 110 may reside in a device within one of the environments represented herein, while other portions of control system 110 may reside in one or more other devices in the environment. For example, the functionality of the control system may be shared by an orchestrating device (e.g., what is referred to herein as a smart home hub) and one or more other devices in the environment. In other examples, a portion of control system 110 may reside in a device implementing a cloud-based service, such as a server, while other portions of control system 110 may reside in other devices implementing the cloud-based service, such as other servers, memory devices, etc. Interface system 105 may also reside in more than one device, in some examples.

[0041] In some implementations, control system 110 may be configured to at least partially perform the methods disclosed herein. According to some examples, control system 110 may be configured to receive an observation sequence including a plurality of extracted features. Each extracted feature of the plurality of extracted features may correspond to a sequential signal in a series of sequential signals. In some examples, control system 110 may be configured to determine the plurality of extracted features.

[0042] According to some examples, control system 110 may be configured to determine a lattice of posterior probabilities. The lattice may include the probability of each observation sequence corresponding to one label class of a plurality of label classes.

[0043] In some examples, the control system 110 may be configured to apply a loss function to the lattice of posterior probabilities according to the ground truth values, hi some such examples, applying the loss function may include applying both sequence information and cluster boundary information.

[0044] According to some examples, control system 110 may be configured to update parameters for determining the lattice of posterior probabilities according to the loss determined by the loss function. In some examples, control system 110 may be configured to train the neural network by performing the above process until control system 110 determines that one or more convergence criteria are satisfied.

[0045] As described elsewhere herein, control system 110 may reside on a single device or on multiple devices, depending on the particular implementation. In some examples, all of the above processes may be performed by the same device. In some alternative examples, the above processes may be performed by more than one device. For example, the embedding may be generated by one device and the lattice of posterior probabilities may be performed by one or more other devices. In some such examples, one or more processes may be performed by one or more services configured to implement cloud-based services.

[0046] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may reside, for example, in any memory system 115 shown in FIG. 1A and / or in control system 110. Thus, various innovative aspects of the subject matter described in this disclosure may be embodied in one or more non-transitory media having software stored thereon. The software may include, for example, instructions for controlling at least one device to perform some or all of the methods disclosed herein. The software may be executable by one or more components of a control system, such as, for example, control system 110 of FIG. 1A.

[0047] In some examples, device 100 may include any of microphone systems 120 shown in FIG. 1A . Any of microphone systems 120 may include one or more microphones. According to some examples, any of microphone systems 120 may include an array of microphones. In some examples, the microphone array may be configured to determine information such as direction of arrival (DOA) and / or time of arrival (TOA), for example, according to instructions from control system 110. The microphone array may in some cases be configured for receive-side beamforming, for example, according to instructions from control system 110. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker of a speaker system, a smart audio device, or the like. In some examples, device 100 may not include microphone system 120. However, in some such implementations, device 100 may still be configured to receive microphone data corresponding to one or more microphones in the environment or corresponding to one or more microphones in other environments via interface system 105. In some such implementations, a cloud-based implementation of device 100 may be configured to receive, via interface system 105, microphone data or metadata corresponding to microphone data acquired by one or more microphones in the environment.

[0048] According to some implementations, device 100 may include any loudspeaker system 125 shown in FIG. 1A. Any loudspeaker system 125 may include one or more loudspeakers, sometimes referred to herein as "speakers" or more generally as "audio reproduction transducers." In some examples (e.g., cloud-based implementations), device 100 may not include loudspeaker system 125.

[0049] In some implementations, device 100 may include any of sensor systems 130 shown in FIG. 1A . Any of sensor systems 130 may include one or more touch sensors, gesture sensors, motion detectors, cameras, eye-tracking devices, or combinations thereof. In some implementations, one or more cameras may include one or more free-standing cameras. In some examples, one or more cameras, eye-trackers, etc. of any of sensor systems 130 may be present in a television, a mobile phone, a smart speaker, or combinations thereof. In some examples, device 100 may not include sensor system 130. However, in some such implementations, device 100 may still be configured to receive, via interface system 105, data from one or more sensors (e.g., cameras, eye-trackers, etc.) present in or on other devices in the environment.

[0050] 1A . Optional display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some cases, optional display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples, optional display system 135 may include one or more displays of a television, laptop, mobile device, smart audio device, or other type of device. In some examples in which device 100 includes display system 135, sensor system 130 may include a touch sensor system and / or a gesture sensor system near one or more displays of display system 135. According to some such implementations, control system 110 may be configured to control display system 135 to present one or more graphical user interfaces (GUIs).

[0051] According to some such examples, apparatus 100 may be or include a smart audio device, such as a smart speaker. In some such implementations, apparatus 100 may be or include a keyword detector. For example, apparatus 100 may be configured to implement (at least in part) a virtual assistant.

[0052] Figure 1B illustrates an example audio environment. As with other figures provided herein, the types, numbers, and arrangements of elements shown in Figure 1B are provided by way of example only. Other implementations of audio environment 150 may include more, fewer, and / or different types, numbers, and arrangements of elements.

[0053] In this example, person 102 and person 104 are seated at table 101, which has audio device 111 positioned thereon. According to this example, audio device 111 is configured to capture audio signals, including but not limited to, audio signals corresponding to the speech of person 102 and person 104, via microphones 103A, 103B, and 103C. The speech of person 102 and person 104 generates not only direct sound waves 107 and 109, but also sound waves 106 and 108 reflected from ceiling 112 and other parts of audio environment 150, which are captured by audio device 111.

[0054] According to some examples, control system 110a of audio device 111 may be configured to provide acoustic event detection functionality, automatic speech recognition functionality, and / or speaker diarization functionality. Alternatively, or additionally, control system 110b of server 116, with which audio device 111 is configured to communicate via network 1171, may be configured to provide acoustic event detection functionality, automatic speech recognition functionality, and / or speaker diarization functionality. This disclosure discloses methods, devices, and systems that may be used to train neural networks to provide such functionality and / or other functionality, such as handwriting recognition functionality. This disclosure also discloses trained neural networks configured to provide such functionality.

[0055] Figures 2 and 3 show examples of blocks that may be included in some disclosed implementations of pinned-state connectionist sequence classification (PS-CSC). As with other figures provided herein, the types, numbers, and arrangements of elements shown in Figures 2 and 3 are given by way of example only. Other implementations may include more, fewer, and / or different types, numbers, and arrangements of elements.

[0056] According to this example, Figure 2 illustrates blocks implemented by an instance of control system 110 described with reference to Figure 1A. As noted elsewhere herein, control system 110 may, in some cases, reside in more than one device. In this example, control system 110 implements at least a posterior probability generator 204 and a loss calculator 206. According to this example, loss calculator 206 is configured to apply a loss function that implements PS-CSC. In some alternative examples, control system 110 may also implement feature extractor 202.

[0057] In this example, the posterior probability generator 204 is configured to generate a posterior probability lattice 205. The posterior probability lattice 305 shown in Figure 3 is an example of the posterior probability lattice 205 shown in Figure 2. Similarly, the ground truth label sequence 303[L1, L2, L3] shown in Figure 3 is an example of the ground truth label sequence 208 shown in Figure 2. The ground truth label sequence 208 may reside in a data structure stored in a memory of the control system 110 or a memory accessible by the control system 110, for example.

[0058] According to this example, posterior probability generator 204 is configured to generate posterior probability lattice 205 from observation sequence 203. In this example, observation sequence 203 includes a plurality of extracted features extracted by feature extractor 202. In this example, each extracted feature generated by feature extractor 202 corresponds to a sequential signal in series of sequential signals 201 and was generated therefrom by feature extractor 202.

[0059] The extracted features in the observation sequence 203 can vary depending on the particular implementation. In some examples, the series of sequential signals may be a time sequence of an audio signal. In some such examples, the extracted features in the observation sequence 203 may simply be time samples of the audio signal, while in other examples, the extracted features in the observation sequence 203 may be a vector representation of the time samples of the audio signal. In some examples, the extracted features may be or include frequency bands or bin energies. According to some examples, the extracted features may be or include transform bins stacked across multiple channels.

[0060] However, in other examples, the series of sequential signals may be a spatial sequence, such as a spatial sequence of handwritten images. In some such examples, the extracted features in the observation sequence 203 may be a vector representation of the segmented input image. In some examples, the extracted features in the observation sequence 203 may be a vector representation of the segmented input image, including each particular field defining the length and width of the image segment.

[0061] If the series of sequential signals is a time sequence of audio signals containing human speech, in some examples, the ground truth label sequence 208 may be or include the ID of each speaker in a speaker sequence. The ground truth label sequence 208 may be expressed as [L1, L2, L3], [L1, L2, L3, L4], etc., [L1, L2, ], where each "L" value corresponds to a label, such as a speaker ID.

[0062] In some examples where the series of sequential signals is a time sequence of audio signals, including an audio signal corresponding to human speech, the posterior probability lattice 205 may be configured to generate the posterior probability lattice 205 by inferring the probability of each extracted feature in the observation sequence 203 to be of a class indicated by one of the ground truth labels in the ground truth label sequence 208. An example of a class is the ID of each speaker in a speaker sequence. In some such examples where the ground truth label sequence 208 is [L1, L2, L3, L4], the posterior probability generator 204 may be configured to generate the posterior probability lattice 205 by inferring the probability of each extracted feature in the observation sequence 203 to be of the class [L1, L2, L3, L4].

[0063] In the example shown in FIG. 2 , the loss calculator 206 is configured to apply a loss function that implements PS-CSC. Various examples and details are described below. Here, the loss calculator 206 is configured to apply the loss function to values ​​of the posterior probability lattice 205 according to the ground truth label sequence 208. According to this example, the loss calculator 206 is configured to provide updated parameters, represented in FIG. 2 as a parameter update block 207, according to the loss determined by the loss function. The parameter update block 207 can be, for example, a data structure stored in a memory of the control system 110 or a memory accessible by the control system 110. In this example, the loss calculator 206 is configured to provide the updated parameters to the posterior probability generator 204 to generate the posterior probability lattice 205. According to some examples, the loss calculator 206 can be configured to provide the updated parameters to the feature extractor 202 to determine a plurality of extracted features in the observation sequence 203.

[0064] A group of extracted features with the same label, such as a group of extracted consecutive features corresponding to the same speaker, is an example of what may be referred to herein as a "cluster." A transition from one class of extracted features to another class of extracted features, in this example, from extracted features corresponding to one speaker to extracted features corresponding to another speaker, indicates or corresponds to what may be referred to herein as a "cluster boundary." For example, a cluster of extracted consecutive features corresponding to the same speaker (also referred to as a class of consecutive observations) can be said to have a first cluster boundary immediately preceding the first such extracted feature and a second cluster boundary immediately preceding the last extracted feature in the series corresponding to the same speaker. For example, referring to FIG. 3, observations 302A and 302B can be seen to correspond to label / speaker L1, and observation 302C can be seen to correspond to label / speaker L2. Thus, a cluster boundary 310 exists between observations 302B and 302C.

[0065] As mentioned above, the ground truth sequence [L1, L2, L3] shown in Figure 3 is an example of the ground truth label sequence 208 shown in Figure 2. Following this example, the posterior probability lattice 305 shows four circles for each of the observations 302A-302N, each representing the probability that the corresponding observation is in class L1, L2, L3, or L4.

[0066] In this example, the underlying labels are monotonically aligned with the observations. For example, an observation that immediately follows another observation aligned with L2 will not be aligned with L1. Instead, the observation can only be aligned with L2 or L3.

[0067] According to the example shown in FIG. 3 , the filled circles correspond to “pinned states,” i.e., the state or label of each corresponding observation is known in advance. In contrast, the unfilled circles correspond to “normal states,” i.e., the state or label of each corresponding observation is not known in advance. Because the states of consecutive observations 302B and 302C are known, and because the underlying labels are known to be monotonically aligned with the observations, there is only one valid path in the posterior probability lattice 305 between the states of consecutive observations 302B and 302C. Because the states of observations 302C and 302E are also known, there are only two valid paths in the posterior probability lattice 305 between the states of observations 302C and 302E. Without pinned state information or cluster boundary information, there would be many more potentially valid paths between the states of the observations in the posterior probability lattice 305 for different classes.

[0068] Therefore, a PS-CSC implementation can result in faster convergence and lower computational overhead. Furthermore, the posterior probability generator 204 does not need to determine the probability that an observation has a pinned state. Therefore, pinning states in some PS-CSC implementations can result in even lower computational overhead. Thus, Figure 3 illustrates, or at least suggests, the potential advantages of a PS-CSC implementation.

[0069] As mentioned above, in the examples described with reference to Figures 2 and 3, the loss calculation unit 206 is configured to apply a loss function that implements PS-CSC. Some examples are described below.

[0070] [Example of PS-CSC loss function] One observed sequence X=[x1, x2, x3, x] generated by a feature extractor such as feature extractor 202 from the corresponding sequential signal. N ], for example, assume that there is only one instance of the observation sequence 203 shown in FIG. 2. In some examples, there may be multiple observation sequences. In some such examples, the loss function may provide a sum of the individual loss values. However, without loss of generality, the following example includes one observation sequence for brevity. The corresponding ground truth label in one such example is [L1, L2, L3], which is the ground truth label sequence 303 shown in FIG. 3. Following the following example, a path through the posterior probability lattice is represented as π. One possible path is π1 = [L1, L1, L2, L2, L3]. By applying the independence assumption, which can be the same as that of connectionist time series classification (CTC), the probability of the entire path can be expressed as:

number

[0071] In one example, the loss function can represent the sum of all possible paths, as follows:

number

[0072] The pinned states shown in Figure 3 mean that a specific state must be passed by a valid path. The pinned states are given as prior knowledge in this example. The number of valid paths is reduced due to the pinned states. CTC loss can be considered an extreme example of PS-CSC where no pinned states exist.

[0073] At the other extreme, if the alignment of all observations in a given observation sequence is known, then there is only one possible path. This single path is denoted by π for simplicity. ce Then, the PS-CSC objective function can be expressed as:

number

[0074] [Speaker diarization example] Traditionally, in speaker recognition systems (including speaker identification and verification), an embedding extractor is first trained. This is done using techniques such as the defactor process described in Kinnunen, Tomi, and Haizhou Li, "An overview of text-independent speaker recognition: From features to supervectors," Speech communication 52, no. 1 (2010), pages 12-40, and the i-vector system described in Dehak, Najim, et al., "Front-end factor analysis for speaker verification," IEEE Transactions on Audio, Speech, and Language Processing 19, no. 4 (2010): pages 788-798, as well as those using neural networks as extractors (e.g., Snyder, David, et al.). al., "X-vectors: Robust dnn embeddings for speaker recognition," 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 5329-5333. (IEEE, 2018). Subsequent cosine distance similarity or probabilistic linear discriminant analysis (PLDA) is typically then applied or trained to measure the similarity between the two embeddings.

[0075] The embedding extractor is trained using the cross-entropy by SoftMax or its variants (angular SoftMax, additive marginal SoftMax, etc.) objective function, and uses the cross-entropy loss to map aggregated utterance-level embeddings to corresponding class labels. In other words, the extractor learns to accumulate identity-related information in an utterance.

[0076] In some use cases, for example, in online speaker diarization, the goal is to identify who is speaking and when. This process is usually performed within a limited time (which may be a few seconds or less), which means that the amount of information to be gathered is also limited. The duration is constrained by the tolerance interval between two decision boundaries, which is typically 250 to 300 milliseconds. Furthermore, in the above example, the embedding extractor is not trained sequentially in the diarization scenario.

[0077] Some disclosed examples include using PS-CSC to learn both speaker ID and temporal alignment. Losses applied by a loss function can enable the neural network to learn how to map sequential embeddings to sequence labels, which can be considered a sequence-to-sequence problem. Unlike speech recognition use cases, if two segments are from the same speaker, no blanking is required because the same speaker label is shared by both of them. Since the boundaries of some labels can be known in advance, the boundary information can be provided to PS-CSC prior to the training process. This is not possible with CTC or cross-entropy loss.

[0078] Figure 4 shows an example of the blocks involved in using PS-CSC to train a neural network for the speaker diarization use case. As with other figures provided herein, the types, numbers, and arrangements of elements shown in Figure 4 are given by way of example only. Other implementations may include more, fewer, and / or different types, numbers, and arrangements of elements.

[0079] According to this example, FIG. 4 illustrates blocks performed by an instance of control system 110 described with reference to FIG. 1A. As noted elsewhere herein, control system 110 may, in some cases, reside in more than one device. In this example, control system 110 is configured to implement at least a posterior probability generator 404, a loss calculator 406, and a parameter update block 407, which are instances of posterior probability generator 204, loss calculator 206, and parameter update block 207, respectively, of FIG. 2. According to this example, loss calculator 406 is configured to apply a loss function that implements PS-CSC. In some alternative examples, control system 110 may also implement an embedding extractor 402, which is an instance of feature extractor 202 of FIG. 2.

[0080] In this example, input sequential signal 401 is audio data samples containing audio signals corresponding to speaker utterances or speakers spk_684, spk_379, spk_167, spk_1195, and spk_527, in that order. Accordingly, ground truth label sequence 408 includes labels L1, L2, L3, L4, and L5, which correspond to speakers spk_684, spk_379, spk_167, spk_1195, and spk_527, respectively. Following this example, embedding extractor 402 generates observation sequence 403 from input sequential signal 401, which are speaker embeddings corresponding to speakers spk_684, spk_379, spk_167, spk_1195, and spk_527.

[0081] According to this example, the posterior probability generator 404 is configured to generate the posterior probability lattice 405. In this example, the posterior probability generator 404 is configured to generate the posterior probability lattice 405 by inferring the probability of each of the extracted features in the observation sequence 403 being of a class indicated by one of the ground truth labels in the ground truth label sequence 408. In this example, the ground truth label sequence 408 is [L1, L2, L3, L4, L5], and the posterior probability generator 404 is configured to generate the posterior probability lattice 405 by inferring the probability of each of the extracted features in the observation sequence 403 being of class L1, L2, L3, L4, or L5.

[0082] In the example shown in FIG. 4 , the loss calculator 406 is configured to apply a loss function that implements PS-CSC. Here, the loss calculator 406 is configured to apply the loss function to values ​​of the posterior probability lattice 405 according to a ground truth label sequence 408. According to this example, the loss calculator 406 is configured to provide updated parameters, represented in FIG. 4 as a parameter update block 407, according to the loss determined by the loss function. In this example, the loss calculator 406 is configured to provide the updated parameters to the posterior probability generator 404 to generate the posterior probability lattice 405. According to some examples, the loss calculator 406 may be configured to provide the updated parameters to the embedding extractor 402 to determine a plurality of extracted features in the observation sequence 403.

[0083] [Experimental Results] The inventors conducted experiments to demonstrate the potential of PS-CSC in speaker diarization scenarios. Figures 5A, 5B, and 5C show the results of experiments using conventional methods in the speaker diarization use case. In the examples represented by Figures 5A-5C, the embedding extractor was trained using cross-entropy loss. In these examples, no sequence information was introduced during the training process. Figures 6A and 6B show the results of experiments using PS-CSC in the speaker diarization use case. In the examples represented by Figures 6A and 6B, sequence information was provided during the training process, and the embedding extractor was trained using CTC. The methods represented by Figures 5A-5C and Figures 6A-6B shared the same dataset and model. In each case, the SoftMax score was then used to indicate the quality of the extracted speaker embeddings for diarization.

[0084] In the experiments depicted in Figures 5A-5C, no sequence boundary information was provided during training. Using PS-CSC, separability improved and training time decreased, both of which were obtained, at least in part, from the sequence boundary information provided during the training process.

[0085] In the experiment represented by Figures 5A-5C, the ground truth label sequence is [spk_684, spk_379, spk_167, spk_1195, spk_527]. Figure 5A shows an example input audio segment 503 in the time domain. Figure 5B shows a graph 502 including curves 505A, 505B, 505C, 505D, and 505E representing the SoftMax scores for utterances of speakers spk_684, spk_379, spk_167, spk_1195, and spk_527, respectively, when the neural network was trained using the CE cost function but without sequence information. Figure 5B shows that the neural network trained without sequence information performed poorly in predicting speaker alignment.

[0086] Figure 5C shows graph 501, including curves 504A, 504B, 504C, 504D, and 504E, representing the SoftMax scores for utterances of speakers spk_684, spk_379, spk_167, spk_1195, and spk_527, respectively, when the neural network is trained with a sequence-loss cost function (CTC). By comparing Figure 5C with Figure 5B, it can be seen that the neural network performs much better when provided with sequence information during the training process. For example, as is evident from graph 501, frames from time units 0 to 5 align with spk_684 with high probability. However, no clear trend is evident in graph 502.

[0087] The experiments represented by Figures 5A-5C show that using sequence information when training neural networks is highly useful for the speaker diarization use case.

[0088] 6A and 6B illustrate the benefits of using sequence and alignment information in a PS-CSC implementation. With some prior alignment information, also referred to herein as pinning information, the number of valid paths through the lattice of posterior probabilities that need to be computed during training is smaller.

[0089] FIG. 6A shows an example of an input audio segment 603A in the time domain, as well as alignment information 610A and 610B. FIG. 6A also shows an example of the posterior probability lattice 205 of FIG. 2, which in this case is the posterior probability lattice 605A. Following this example, as also seen in FIG. 3, the black circles in the posterior probability lattice 605A indicate pinning states, with each pinning information corresponding to one instance of alignment information 610A and 610B. In this example, only the alignments corresponding to spk_684 and spk_379 are known prior to the neural network training process. FIG. 6A also shows regions 604A and 604B within the posterior probability lattice 605A. In this example, regions 604A and 604B are regions where there may be a valid path through the posterior probability lattice 605A, given the known alignment information 610A and 610B.

[0090] FIG. 6B also shows an example of an input audio segment 603B in the time domain, but without alignment information. FIG. 6B also shows a posterior probability lattice 605B. According to this example, no alignment information is provided, so there are no black circles in the posterior probability lattice 605A indicating a pinned state. In this example, because no prior alignment information was provided prior to training, valid paths through the posterior probability lattice 605B may exist within region 602. Given only alignment information 610A and 610B, it can be seen that regions 604A and 604B are each much smaller than region 602. Thus, because no prior alignment information was provided prior to training, the number of valid paths through the posterior probability lattice 605B is significantly greater than the number of valid paths through the posterior probability lattice 605A.

[0091] 7 shows the results of two additional experiments. In this example, graph 700 shows the loss determined by the loss function on the vertical axis and the number of iterations of the neural network training process on the horizontal axis. Following this example, curve 705 shows the loss values ​​corresponding to the CTC-based neural network training process, while curve 710 shows the loss values ​​corresponding to the PS-CSC-based neural network training process. It can be seen that after approximately 45,000 iterations, the loss corresponding to the PS-CSC-based neural network training process was consistently less than the loss corresponding to the CTC-based neural network training process.

[0092] 8 is a flow diagram illustrating one example of the disclosed method. The blocks of method 800, as well as other methods described herein, are not necessarily performed in the order shown. According to some examples, one or more blocks may be performed in parallel. Furthermore, some similar methods may include more or fewer blocks than shown and / or described.

[0093] Method 800 may be performed by a device or system, such as device 100 shown in FIG. 1 and described above. In some examples, device 100 includes at least control system 110 as disclosed herein. In some examples, at least some aspects of method 800 may be performed by one or more devices in an audio environment, such as an audio system controller (e.g., what may be referred to herein as a smart home hub) or other components of an audio system, such as a television, a television control module, a laptop computer, a mobile device (e.g., a cellular phone), etc. However, in some implementations, at least some blocks of method 800 may be performed by one or more devices, such as one or more systems, that may be configured to implement cloud-based services.

[0094] In this example, block 805 includes receiving, by the control system, an observation sequence including a plurality of extracted features. According to some examples, each extracted feature of the plurality of extracted features may correspond to a sequential signal in a series of sequential signals. In some examples, the series of sequential signals may be a time sequence including audio signals, and in some cases may include audio signals corresponding to speech. In some alternative examples, the series of sequential signals may be a time sequence including other types of audio signals, signals corresponding to stock market data, signals corresponding to biological data, etc. According to some alternative examples, the series of sequential signals may be a spatial sequence, such as a spatial sequence including handwriting images, DNA sequence images, or other images.

[0095] According to this example, block 810 includes determining, by the control system, a lattice of posterior probabilities. In some examples, the lattice may include a probability of each observed sequence corresponding to one label class of a plurality of label classes. In some examples, the label classes may correspond to the identities of each of a plurality of speakers.

[0096] In this example, block 815 includes applying, by the control system, a loss function to the lattice of posterior probabilities according to the ground truth values. According to this example, applying the loss function includes applying both sequence information and cluster boundary information. According to some examples, the sequence information and / or cluster boundary information of the cost function can be variable. In some examples, the cluster boundary information may be incomplete cluster boundary information. However, in some alternative examples, the cluster boundary information may be complete cluster boundary information. In some examples, the loss function may combine elements of connectionist time-series classification and cross-entropy loss. According to some examples, applying the loss function may include determining one or more valid paths between the observations in the lattice.

[0097] According to this example, block 820 includes updating, by the control system, parameters for determining the lattice of posterior probabilities according to the losses determined by the loss function. In some examples, block 820 may include updating parameters used by posterior probability generator 204 of FIG. 2 according to parameters of parameter update block 207.

[0098] In this example, block 825 includes continuing to execute blocks 805 through 820 until the control system determines that one or more convergence criteria have been satisfied. In some examples, the control system 110 may be configured to determine that convergence has been achieved when the training process reaches a state where the loss determined by the loss function has settled within an error range around a final value, or where the loss determined by the loss function is no longer decreasing, or where the loss determined by the loss function has not decreased for a predetermined number of steps or epochs.

[0099] In some examples, performing blocks 805 through 820 may provide a trained neural network that is implemented by control system 110. Method 800 may, in some examples, include performing one or more downstream tasks with the trained neural network.

[0100] According to some examples, the series of sequential signals may be or include a series of audio signals, and in some such examples, the one or more downstream tasks may include acoustic event detection, speaker diarization, automatic speech recognition, or a combination thereof.

[0101] Method 800, in some examples, may include determining, by a control system, the plurality of extracted features. In some such examples, the method may include updating, by the control system, parameters for determining the plurality of extracted features according to a loss determined by the loss function.

[0102] In some examples, the series of sequential signals may be or include a series of handwritten images, and in some such examples, one or more downstream tasks may include handwriting recognition.

[0103] Some aspects of the present disclosure include systems or devices configured (e.g., programmed) to perform one or more examples of the disclosed methods, as well as tangible computer-readable media (e.g., disks) storing code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of various operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor can be or include a computer system including input devices, memory, and a processing subsystem, where the processing system is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0104] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed or otherwise configured) to perform necessary processing on an audio signal, including performing one or more examples of the disclosed methods. Alternatively, the disclosed systems (or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) that is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system may be implemented as a general-purpose processor or DSP configured (or programmed) to perform one or more examples of the disclosed methods, and the system may also include other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device.

[0105] Another aspect of the present disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) storing code (e.g., executable code) for performing one or more examples of the disclosed methods or steps thereof.

[0106] While specific embodiments of and applications of the present disclosure have been described herein, it will be apparent to those skilled in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the present disclosure as described and claimed herein. While particular forms of the present disclosure have been illustrated and described, it should be understood that the disclosure should not be limited to the specific embodiments described and illustrated, or to the specific methods described.

[0107] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 513,294, filed July 12, 2023, and U.S. Provisional Patent Application No. 63 / 401,042, filed August 25, 2022, which are incorporated herein by reference.

Claims

1. (a) receiving, by a control system, an observation sequence including a plurality of extracted features, each extracted feature of the plurality of extracted features corresponding to a sequential signal in a series of sequential signals; (b) determining, by the control system, a lattice of posterior probabilities, the lattice including a probability of each observation sequence corresponding to one label class of a plurality of label classes; (c) applying, by the control system, a loss function to the lattice of posterior probabilities according to ground truth values, wherein applying the loss function includes applying both sequence information and cluster boundary information; and (d) updating, by the control system, parameters for determining the lattice of posterior probabilities according to the loss determined by the loss function; (e) continuing to perform (a) through (d) until the control system determines that one or more convergence criteria are met; and A method having the following.

2. performing (a) through (e) provides a trained neural network implemented by the control system; The method of claim 1.

3. performing one or more downstream tasks with the trained neural network. The method of claim 2.

4. the series of sequential signals comprises a series of audio signals; the one or more downstream tasks include at least one of acoustic event detection, speaker diarization, and automatic speech recognition; The method of claim 3.

5. the series of sequential signals includes a series of handwritten images; the one or more downstream tasks include handwriting recognition; The method of claim 3.

6. the cluster boundary information includes incomplete cluster boundary information; 6. The method according to any one of claims 1 to 5.

7. the cluster boundary information includes complete cluster boundary information; 7. The method according to any one of claims 1 to 6.

8. applying the loss function includes determining one or more valid paths between observations in the lattice.

8. The method according to any one of claims 1 to 7.

9. determining, by the control system, the plurality of extracted features.

9. The method according to any one of claims 1 to 8.

10. and updating, by the control system, parameters for determining the plurality of extracted features according to the loss determined by the loss function.

10. The method of claim 9.

11. The loss function combines elements of connectionist time-series classification and cross-entropy loss.

11. The method according to any one of claims 1 to 10.

12. the sequence information of the loss function, the cluster boundary information of the loss function, or both are variable; 12. The method according to any one of claims 1 to 11.

13. Apparatus configured to perform the method of any one of claims 1 to 12.

14. A system configured to perform the method of any one of claims 1 to 12.

15. One or more non-legal media storing instructions for controlling one or more devices to perform the method of any one of claims 1 to 12.

Citation Information

Patent Citations

  • Efficient connectionist temporal classification for binary classification

    US20180232632A1

  • Handwriting recognition with language modeling

    US20220138453A1

  • Learning device, speech recognition device, learning method, speech recognition method, learning program, and speech recognition program

    WO2022024202A1