Method, device, and system for performing pinned state connector series classification.

PS-CSC methods address inefficiencies in neural network training for speaker diarization and speech recognition by combining cluster boundary and sequence information, achieving faster convergence and reduced computational overhead.

JP7836463B2Active Publication Date: 2026-03-26DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-08-22
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing machine learning methods for tasks like speaker diarization and automatic speech recognition face inefficiencies due to the need for extensive computational resources and time in training neural networks, particularly when dealing with sequential information.

Method used

The implementation of Pinned-State Connectionist Sequential Classification (PS-CSC) methods, which utilize a loss function combining cluster boundary and sequence information to train neural networks, reducing the number of valid paths and computational overhead during training.

Benefits of technology

PS-CSC methods enable faster convergence and lower computational overhead by pinning known states, allowing for efficient training of neural networks for tasks such as speaker diarization and automatic speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007836463000004
    Figure 0007836463000004
  • Figure 0007836463000005
    Figure 0007836463000005
  • Figure 0007836463000006
    Figure 0007836463000006
Patent Text Reader

Abstract

Some disclosed methods include receiving an observation sequence including a plurality of extracted features, each of the plurality of extracted features corresponding to a sequential signal in a series of sequential signals; determining a lattice of posterior probabilities, the lattice including a probability of each observation sequence corresponding to one label class among a plurality of label classes; and applying a loss function to the lattice of posterior probabilities according to ground truth values, the applying the loss function including applying both sequential information and cluster boundary information. Some methods include updating parameters for determining the lattice according to a loss determined by the loss function; and performing the above operations until one or more convergence criteria are met.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to machine learning. [Background technology]

[0002] Several methods, devices, and systems are known for machine learning. While existing devices, systems, and methods for machine learning may offer advantages in some situations, improvements are desired. [Overview of the Initiative]

[0003] [Notation and Nomenclature] Throughout this disclosure, including the claims, the terms “speaker,” “loudspeaker,” and “audio reproduction transducer” are used synonymously to represent any acoustic emission transducer (or set of transducers) driven by a single speaker feed. A speaker is implemented to include multiple transducers (e.g., a woofer and a tweeter) which may be driven by a single common speaker feed or multiple speaker feeds. In some examples, speaker signals may undergo different processing in different circuit branches coupled to different transducers.

[0004] Throughout this disclosure, including the claims, the expression "operating on" a signal or data (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used in a broad sense to mean operating on a signal or data directly, or on a processed version of a signal or data (e.g., on a version of a signal that has undergone preliminary filtering or preprocessing before the operation on the signal is performed).

[0005] Throughout this disclosure, including the claims, the term “system” is used in a broad sense to describe a device, system, or subsystem. For example, a subsystem that implements a decoder may be called a decoder system, and a system that includes such a subsystem (for example, a system that generates X output signals in response to multiple inputs, where M inputs are generated by the subsystem and the remaining XM inputs are from an external source) may also be called a decoder system.

[0006] Throughout this disclosure, including the claims, the term “processor” is used in a broad sense to describe a system or device that is programmable or otherwise configurable (e.g., by firmware or software) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing on audio or other acoustic data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0007] Throughout this disclosure, including the claims, the terms “combine” or “become” are used to mean direct or indirect connection. Thus, when the first device is coupled to the second device, the connection may be through a direct connection or through an indirect connection via other devices and connections.

[0008] As used herein, “smart device” is an electronic device that can operate to some extent interactively and / or autonomously and is generally configured to communicate with one or more other devices (or networks) via various radio protocols such as Bluetooth®, Zigbee®, Near Field Communication, Wi-Fi, Light Fidelity (Li-Fi), 3G, 4G, and 5G. Some notable types of smart devices include smartphones, smart cards, smart thermostats, smart doorbells, smart locks, smart refrigerators, phablets and tablets, smartwatches, smart bands, smart keychains, and smart audio devices. The term “smart device” may also refer to devices that exhibit some characteristics of ubiquitous computing, such as artificial intelligence.

[0009] Here, we use the term “smart audio device” to describe a smart device that is either a single-purpose audio device or a multi-purpose audio device (e.g., an audio device that implements at least some aspects of virtual assistant functionality). A single-purpose audio device is a device that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera) and is designed to serve a single purpose (e.g., a television (TV)). For example, a TV can (or is considered capable of) playing audio from program material, but in most cases, modern TVs run some operating system on which applications, including a television viewing application, run locally. In this sense, a single-purpose audio device with a speaker and microphone is often configured to run local applications and / or services for direct use of the speaker and microphone. Some single-purpose audio devices may be configured to group together to enable audio playback across zones or user-defined areas.

[0010] One common type of multipurpose audio device is a smart audio device, such as a “smart speaker,” which may be configured to implement at least some aspects of virtual assistant functionality, while other aspects of virtual assistant functionality may be implemented by one or more other devices, such as one or more servers, with which the multipurpose audio device is configured to communicate. Such a multipurpose audio device may be referred to herein as a “virtual assistant.” A virtual assistant is a device (e.g., a smart speaker or voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally also includes or is coupled to at least one speaker and / or at least one camera). In some examples, a virtual assistant may be cloud-enabled in some sense, or in other sense, provide the ability to utilize multiple devices (separate from the virtual assistant) for applications that are not fully implemented in the virtual assistant itself. In other words, at least some aspects of virtual assistant functionality, such as speech recognition, may be implemented (at least partially) by one or more servers or other devices with which the virtual assistant can communicate over a network, such as the Internet. Virtual assistants may sometimes work together in specific, conditionally defined ways. For example, two or more virtual assistants may work together in the sense that one of them, for instance, the one that is most confident it heard the wake word, will respond to the wake word. Connected virtual assistants may, in some implementations, form a kind of constellation managed by a single main application that can be (or implements) a virtual assistant.

[0011] Here, “wake word” is used in a broad sense to represent any sound (a word spoken by a person, or any other sound), and the smart audio device is configured to wake up in response to the detection of sound (“hearing”) (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, “wake up” means that the device enters a state of waiting for (i.e., listening to) an acoustic command. In some cases, what may be called a “wake word” here may include more than one word, such as a phrase.

[0012] Here, the term “wake word detector” refers to a device (or software containing instructions for configuring the device) configured to continuously search for alignments between real-time acoustic (e.g., speech) features and a trained model. Typically, a wake word event is triggered when the wake word detector determines that the probability of a wake word being detected exceeds a predefined threshold. For example, the threshold may be a predetermined threshold adjusted to give a reasonable compromise between the false acceptance rate and the false rejection rate. Following a wake word event, the device can enter a state (which may be called an “awake” or “attention” state) where it listens for commands and passes received commands to a larger, more computationally intensive recognition engine.

[0013] As used herein, the terms “program stream” and “content stream” refer to a collection of one or more audio signals, and in some cases video signals, where at least some of the signals are intended to be heard together. Examples include music selections, movie soundtracks, movies, television programs, audio portions of television programs, podcasts, live voice calls, synthesized voice responses from smart assistants, etc. In some cases, a content stream may contain multiple versions of at least some of the audio signals, for example, the same dialogue in more than one language. In such cases, only one version of the audio data or a portion thereof (for example, a version corresponding to a single language) is intended to be played at a time.

[0014] [overview] At least some aspects of this disclosure may be implemented by one or more methods. In some cases, the methods may be implemented at least in part by a control system and / or via instructions (e.g., software) stored in one or more non-temporary media. Some methods may include the control system receiving an observation sequence containing a plurality of extracted features. Each extracted feature of the plurality of extracted features may correspond to a sequential signal in a sequence of sequential signals. Some methods may include the control system determining a grid of posterior probabilities. The grid may contain the probabilities of each observation sequence corresponding to one of a plurality of label classes.

[0015] Some methods may involve a control system applying a loss function to a posterior probability grid according to ground truth values. In some examples, applying the loss function may involve applying both sequential information and cluster boundary information. In some examples, applying the loss function may involve determining one or more valid paths between observations in the grid.

[0016] Some methods may involve a control system updating parameters for determining the posterior probability grid according to a loss determined by a loss function. Other methods may involve continuing to perform the receiving, determining, applying, and updating processes until the control system determines that one or more convergence criteria are met.

[0017] Following several examples, performing the above process may provide a trained neural network implemented by a control system. Some methods may involve having the trained neural network perform one or more downstream tasks.

[0018] In some examples, the sequence of signals may also include a sequence of audio signals. In some such examples, one or more downstream tasks may include acoustic event detection, speaker dialization, automatic speech recognition, or a combination thereof.

[0019] As in some examples, a sequence of signals may be, or include, a sequence of handwritten images. In some such examples, one or more downstream tasks may include handwriting recognition.

[0020] In some examples, the cluster boundary information may be incomplete. However, in some alternative examples, the cluster boundary information may be complete.

[0021] Some methods may involve a control system determining multiple extracted features. In some such examples, the method may involve the control system updating parameters for determining the multiple extracted features according to a loss determined by a loss function.

[0022] In some examples, the series information of the loss function, the cluster boundary information of the loss function, or both can be variable. According to some examples, the loss function can combine elements of connectionist time series classification and cross-entropy loss.

[0023] Some or all of the operations, functions, and / or methods described herein can be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media can include memory devices such as, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc., as described herein. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented by one or more non-transitory media storing software.

[0024] At least some aspects of the present disclosure can be implemented by an apparatus. For example, one or more devices (e.g., a system including one or more devices) can be capable of at least partially performing the methods disclosed herein. In some implementations, the apparatus is or includes an audio processing system with an interface system and a control system. The control system can include one or more general-purpose single or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof. The control system can be configured to implement some or all of the methods disclosed herein.

[0025] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be to scale.

[0026] Like reference numerals and characters in the various drawings indicate like elements.

Brief Description of the Drawings

[0027] [Figure 1A] FIG. 1 is a block diagram illustrating examples of components of an apparatus that can implement various aspects of the present disclosure. [Figure 1B] FIG. 2 shows an example of an audio environment. [Figure 2] FIG. 3 shows examples of blocks that may be included in the implementation of some disclosed pin-stopping state connector series classifications (PS-CSCs). [Figure 3] FIG. 4 shows examples of blocks that may be included in the implementation of some disclosed pin-stopping state connector series classifications (PS-CSCs). [Figure 4] FIG. 5 shows an example of blocks involved in training a neural network for use cases of speaker diarization using PS-CSC. [Figure 5] A - C show the results of experiments using conventional methods in use cases of speaker diarization. [Figure 6A] FIG. 6 shows the results of experiments using the PS-CSC method in use cases of speaker diarization. [Figure 6B] FIG. 7 shows the results of experiments using the PS-CSC method in use cases of speaker diarization. [Figure 7] FIG. 8 shows the results of two additional experiments. [Figure 8] FIG. 9 is a flowchart illustrating an example of the disclosed method.

Best Mode for Carrying Out the Invention

[0028] Machine learning methods based on cross-entropy loss and its variations, and machine learning methods based on Connectionist Temporal Classification (CTC), use different loss objectives and are generally used for different purposes in the machine learning community. Machine learning methods based on cross-entropy loss are typically used to train neural network models for classification problems in use cases such as image-based face classification and utterance-based speaker identification. Machine learning methods based on cross-entropy loss may include acquiring and / or learning cluster boundary information. In contrast, machine learning methods based on CTC are typically applied when sequential information needs to be explored, in use cases such as automatic speech recognition and handwriting recognition.

[0029] This disclosure provides examples of machine learning methods based on cluster boundary information and sequence information. Such methods may be referred to herein as Pinned-State Connectionist Sequential Classification (PS-CSC) methods. According to some examples of PS-CSC, the loss function used to train the neural network may be based not only on cluster boundary information but also on sequence information. In some examples of PS-CSC, the cluster boundary information may be incomplete, or may include incomplete cluster boundary information. In some examples of PS-CSC, the loss function used to train the neural network may be a combination of CTC and cross-entropy loss elements.

[0030] In some PS-CSC examples, the disclosed neural network training process may involve receiving an observation sequence containing multiple extracted features. According to some examples, the observation sequence may be a time sequence, while in others, it may be a spatial sequence. Each extracted feature may correspond to a sequential signal in a sequence of sequential signals. In some cases, the training process may involve determining a lattice of posterior possibilities. The lattice may contain the probability of each observation sequence corresponding to one of several label classes. According to some examples, the labels may be monotonically aligned with the observations.

[0031] As in some PS-CSC examples, applying a loss function may involve determining one or more valid paths between observations in the grid. In some such examples, pinning the observations in the grid reduces the number of valid paths, thus reducing the time and computational overhead required for convergence during the neural network training process.

[0032] A neural network trained according to several disclosed methods may be configured to perform one or more downstream tasks. The downstream tasks can vary depending on the type of signals in the sequence of sequential signals. In some cases, the sequence of sequential signals may be or include a sequence of handwritten images, and the downstream task may be handwriting recognition. In other examples, the sequence of sequential signals may be or include a sequence of audio signals, and the downstream tasks may be acoustic event detection, speaker diarization, and / or automatic speech recognition.

[0033] Figure 1A is a block diagram showing an example of components of an apparatus that can implement various aspects of the present disclosure. As with other figures provided herein, the type, number, and arrangement of elements shown in Figure 1A are given only as an example. Other implementations may include more, fewer, and / or different types, numbers, and arrangements of elements. According to some examples, apparatus 100 may be configured to perform at least some of the methods disclosed herein. In some implementations, apparatus 100 may be one or more components of an office workstation, one or more components of a home entertainment system, etc., or include such. For example, apparatus 100 may be a laptop computer, a tablet device, a mobile device (e.g., a cellular phone), a smart home hub, a television, or other type of device.

[0034] In some alternative implementations, device 100 may be a server or include a server. In some such examples, device 100 may be an encoder or include an encoder. In some examples, device 100 may be a decoder or include a decoder. Thus, in some cases, device 100 may be a device configured for use in an environment such as a home environment, while in other cases, device 100 may be a device configured for use in a “cloud,” such as a server.

[0035] In this example, the device 100 includes an interface system 105 and a control system 110. In some implementations, the interface system 105 may be configured for communication with one or more other devices in an environment. In some examples, the environment may be a home environment. In other examples, the environment may be other types of environments, such as an office environment, a car environment, a train environment, a street or sidewalk environment, a park environment, etc. In some implementations, the interface system 105 may be configured to exchange control information and related data with other devices in the environment. In some examples, the control information and related data may relate to one or more software applications that the device 100 is running.

[0036] In some implementations, the interface system 105 may be configured to receive or supply a content stream. In some examples, the content stream may include video data and audio data corresponding to the video data. The audio data may include, but is not limited to, audio signals. In some cases, the audio data may include spatial data, such as channel data and / or spatial metadata. The metadata may be supplied, for example, by what may be referred to herein as an “encoder”.

[0037] The interface system 105 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more Universal Serial Bus (USB) interfaces). In some implementations, the interface system 105 may include one or more wireless interfaces. The interface system 105 may include one or more devices that implement a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, a gesture sensor system, or a combination thereof. Thus, although some such devices are shown individually in Figure 1A, such devices may, in some examples, correspond to aspects of the interface system 105.

[0038] In some examples, the interface system 105 may include one or more interfaces between the control system 110 and a memory system, such as any memory system 115 shown in Figure 1A. Alternatively, or additionally, the control system 110 may include a memory system in some cases. In some implementations, the interface system 105 may be configured to receive input from one or more microphones in the environment.

[0039] The control system 110 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gates or transistor logic, discrete hardware components, or a combination thereof.

[0040] In some implementations, the control system 110 may reside on more than one device. For example, in some implementations, part of the control system 110 may reside on a device in one of the environments referenced herein, while other parts of the control system 110 may reside on devices outside the environment, such as a server or a mobile device (e.g., a smartphone or tablet computer). In other examples, part of the control system 110 may reside on a device in one of the environments represented herein, while other parts of the control system 110 may reside on one or more other devices in the environment. For example, the functions of the control system may be shared between an orchestrating device (e.g., what is here called a smart home hub) and one or more other devices in the environment. In other examples, part of the control system 110 may reside on a device implementing cloud-based services, such as a server, while other parts of the control system 110 may reside on other devices implementing cloud-based services, such as other servers or memory devices. The interface system 105 may also reside on more than one device in some examples.

[0041] In some implementations, the control system 110 may be configured to perform at least partially the methods disclosed herein. According to some examples, the control system 110 may be configured to receive an observation sequence containing a plurality of extracted features. Each of the plurality of extracted features may correspond to a sequential signal in a sequence of sequential signals. In some examples, the control system 110 may be configured to determine the plurality of extracted features.

[0042] According to several examples, the control system 110 may be configured to determine a grid of posterior probabilities. The grid may include the probabilities of each observation sequence corresponding to one of several label classes.

[0043] In some examples, the control system 110 may be configured to apply a loss function to a grid of posterior probabilities according to ground truth values. In some such examples, applying the loss function may include applying both sequence information and cluster boundary information.

[0044] In some examples, the control system 110 may be configured to update the parameters for determining the posterior probability grid according to the loss determined by the loss function. In some examples, the control system 110 may be configured to train the neural network by performing the above process until the control system 110 determines that one or more convergence criteria are satisfied.

[0045] As described elsewhere in this specification, the control system 110 may reside in a single device or in multiple devices, depending on the particular implementation. In some examples, all of the above processes may be performed by the same device. In some alternative examples, the above processes may be performed by two or more devices. For example, the embedding may be generated by one device, and the posterior probability grid may be performed by one or more other devices. In some such examples, one or more processes may be performed by one or more services configured to implement cloud-based services.

[0046] Some or all of the methods described herein may be executed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-temporary media. Such non-temporary media may include, but are not limited to, memory devices such as random-access memory (RAM) devices and read-only memory (ROM) devices, as described herein. One or more non-temporary media may be located, for example, in any memory system 115 and / or control system 110 shown in Figure 1A. Thus, various innovative aspects of the subject matter described herein can be realized on one or more non-temporary media on which software is stored. The software may include, for example, instructions that control at least one device to execute some or all of the methods disclosed herein. The software may be executable by one or more components of a control system, such as the control system 110 in Figure 1A.

[0047] In some examples, the device 100 may include an arbitrary microphone system 120 shown in Figure 1A. An arbitrary microphone system 120 may include one or more microphones. According to some examples, an arbitrary microphone system 120 may include an array of microphones. In some examples, the array of microphones may be configured to determine information such as direction of arrival (DOA) and / or time of arrival (TOA) in accordance with commands from, for example, a control system 110. In some cases, the array of microphones may be configured for receiving beamforming in accordance with commands from, for example, a control system 110. In some implementations, one or more microphones may be part of, or associated with, other devices such as speakers in a speaker system, smart audio devices, etc. In some examples, the device 100 may not include a microphone system 120. However, in some such implementations, the device 100 may still be configured to receive microphone data via the interface system 105 corresponding to one or more microphones in an environment or one or more microphones in other environments. In some such implementations, the cloud-based implementation of device 100 may be configured to receive microphone data or metadata corresponding to microphone data acquired by one or more microphones in the environment via the interface system 105.

[0048] In some implementations, the apparatus 100 may include any loudspeaker system 125 shown in Figure 1A. Any loudspeaker system 125 may include one or more loudspeakers, which may be referred to herein as “speakers” or more generally as “audio playback transducers.” In some examples (e.g., cloud-based implementations), the apparatus 100 may not include the loudspeaker system 125.

[0049] In some implementations, device 100 may include any sensor system 130 shown in Figure 1A. Any sensor system 130 may include one or more touch sensors, gesture sensors, motion detectors, cameras, eye-tracking devices, or a combination thereof. In some implementations, one or more cameras may include one or more free-standing cameras. In some examples, one or more cameras, eye-trackers, etc. of any sensor system 130 may be present in a television, mobile phone, smart speaker, or a combination thereof. In some examples, device 100 may not include sensor system 130. However, in some such implementations, device 100 may still be configured to receive data from one or more sensors (e.g., cameras, eye-trackers, etc.) present in or on other devices in the environment via interface system 105.

[0050] In some embodiments, the apparatus 100 may include any display system 135 shown in Figure 1A. Any display system 135 may include one or more displays, such as one or more light-emitting diode (LED) displays. In some cases, any display system 135 may include one or more organic light-emitting diode (OLED) displays. In some examples, any display system 135 may include one or more displays of a television, laptop, mobile device, smart audio device, or other type of device. In some examples where the apparatus 100 includes a display system 135, the sensor system 130 may include a touch sensor system and / or a gesture sensor system near one or more displays of the display system 135. According to some such embodiments, the control system 110 may be configured to control the display system 135 to present one or more graphical user interfaces (GUIs).

[0051] In some such examples, the device 100 may be or include a smart audio device such as a smart speaker. In some such implementations, the device 100 may be or include a keyword detector. For example, the device 100 may be configured to implement (at least partially) a virtual assistant.

[0052] Figure 1B shows an example of an audio environment. As with other figures provided herein, the type, number, and arrangement of elements shown in Figure 1B are given only as an example. Other implementations of the audio environment 150 may include more, fewer, and / or different types, numbers, and arrangements of elements.

[0053] In this example, persons 102 and 104 are seated at table 101, on which an audio device 111 is placed. According to this example, the audio device 111 is configured to capture audio signals, including but not limited to audio signals corresponding to the speech of persons 102 and 104, via microphones 103A, 103B, and 103C. The speech of persons 102 and 104 generates not only direct sound waves 107 and 109, but also sound waves 106 and 108 reflected from the ceiling 112 and other parts of the audio environment 150, and these signals are captured by the audio device 111.

[0054] According to some examples, the control system 110a of the audio device 111 may be configured to provide acoustic event detection, automatic speech recognition, and / or speaker dialization functions. Alternatively or additionally, the control system 110b of the server 116, which is configured to communicate with the audio device 111 via the network 1171, may be configured to provide acoustic event detection, automatic speech recognition, and / or speaker dialization functions. The disclosure also discloses methods, devices, and systems that may be used to train a neural network to provide such functions and / or other functions such as handwriting recognition. The disclosure also discloses a trained neural network configured to provide such functions.

[0055] Figures 2 and 3 show examples of blocks that may be included in some disclosed implementations of the Pinned State Connectionist Series Classification (PS-CSC). As with other figures provided herein, the types, number, and arrangement of elements shown in Figures 2 and 3 are given only as examples. Other implementations may include more, fewer, and / or different types, numbers, and arrangements of elements.

[0056] Following this example, Figure 2 shows a block implemented by an instance of the control system 110 described with reference to Figure 1A. As described elsewhere in this specification, the control system 110 may, in some cases, be present in more than one device. In this example, the control system 110 implements at least a posterior probability generator 204 and a loss calculation unit 206. Following this example, the loss calculation unit 206 is configured to apply a loss function that realizes PS-CSC. In some alternative examples, the control system 110 may also implement a feature extraction unit 202.

[0057] In this example, the posterior probability generation unit 204 is configured to generate a posterior probability grid 205. The posterior probability grid 305 shown in Figure 3 is a specific example of the posterior probability grid 205 shown in Figure 2. Similarly, the ground truth label sequence 303 [L1, L2, L3] shown in Figure 3 is an example of the ground truth label sequence 208 shown in Figure 2. The ground truth label sequence 208 may reside, for example, in a data structure stored in the memory of the control system 110 or in a memory accessible by the control system 110.

[0058] In this example, the posterior probability generation unit 204 is configured to generate a posterior probability grid 205 from the observation sequence 203. In this example, the observation sequence 203 includes a number of extracted features extracted by the feature extraction unit 202. In this example, each extracted feature generated by the feature extraction unit 202 corresponds to a sequential signal in a sequence 201 of sequential signals, which was then generated by the feature extraction unit 202.

[0059] The extracted features in observation sequence 203 can vary depending on the specific implementation. In some examples, the sequence of sequential signals may be a time sequence of audio signals. In some such examples, the extracted features in observation sequence 203 may simply be time samples of the audio signals, while in other examples, the extracted features in observation sequence 203 may be a vector representation of the time samples of the audio signals. In some examples, the extracted features may be or include frequency bands or bin energies. In some examples, the extracted features may be or include stacked transformation bins across multiple channels.

[0060] However, in other examples, the sequence of sequential signals may be a spatial sequence, such as a spatial sequence of handwritten images. In some such examples, the extracted features in observation sequence 203 may be a vector representation of a segmented input image. In some examples, the extracted features in observation sequence 203 may be a vector representation of a segmented input image containing specific fields defining the length and width of each image segment.

[0061] If the sequence of sequential signals is a time sequence of audio signals including human speech, then in some examples the ground truth label sequence 208 may be the ID of each speaker in the speaker sequence, or it may contain the ID of each speaker. The ground truth label sequence 208 may be expressed as [L1,L2,L3], [L1,L2,L3,L4], etc., [L1,L2,...], where each "L" value corresponds to a label such as a speaker ID.

[0062] In some examples where the sequence of sequential signals is a time sequence of audio signals containing audio signals corresponding to human speech, the posterior stochastic grid 205 may be configured to generate the posterior stochastic grid 205 by inferring the probability of each extracted feature in the observation sequence 203 to belong to a class indicated by one of the ground truth labels in the ground truth label sequence 208. An example of a class is the ID of each speaker in a sequence of speakers. In some such examples where the ground truth label sequence 208 is [L1,L2,L3,L4], the posterior stochastic generator 204 may be configured to generate the posterior stochastic grid 205 by inferring the probability of each extracted feature in the observation sequence 203 to belong to class [L1mL2,L3,L4].

[0063] In the example shown in Figure 2, the loss calculation unit 206 is configured to apply a loss function that implements PS-CSC. Various examples and details are described below. Here, the loss calculation unit 206 is configured to apply the loss function to the values ​​of the posterior stochastic grid 205 according to the ground truth label sequence 208. According to this example, the loss calculation unit 206 is configured to supply updated parameters, shown in Figure 2 as a parameter update block 207, according to the loss determined by the loss function. The parameter update block 207 can be, for example, a data structure stored in the memory of the control system 110 or in memory accessible by the control system 110. In this example, the loss calculation unit 206 is configured to supply the updated parameters to the posterior stochastic generation unit 204 to generate the posterior stochastic grid 205. According to some examples, the loss calculation unit 206 may be configured to supply the updated parameters to the feature extraction unit 202 to determine a plurality of extracted features in the observation sequence 203.

[0064] A group of extracted features with the same label, such as a group of extracted continuous features corresponding to the same speaker, is an example of what may be referred to herein as a “cluster.” The transition from one class of extracted features to another, in this example, from an extracted feature corresponding to one speaker to an extracted feature corresponding to another speaker, represents or corresponds to what may be referred to herein as a “cluster boundary.” For example, a cluster of extracted continuous features corresponding to the same speaker (also called a class of continuous observations) can be said to have a first cluster boundary immediately preceding the first such extracted feature and a second cluster boundary immediately preceding the last extracted feature in the continuum corresponding to the same speaker. For example, referring to Figure 3, we can see that observations 302A and 302B correspond to label / speaker L1, and observation 302C corresponds to label / speaker L2. Therefore, there is a cluster boundary 310 between observation 302B and observation 302C.

[0065] As mentioned above, the ground truth sequence [L1, L2, L3] shown in Figure 3 is an example of the ground truth label sequence 208 shown in Figure 2. Following this example, the posterior stochastic grid 305 shows four circles for each of the observations 302A to 302N. Each circle represents the probability that the corresponding observation belongs to class L1, L2, L3, or L4.

[0066] In this example, the underlying label is monotonically aligned with the observation. For example, an observation immediately following another observation aligned with L2 will not be aligned with L1. Instead, that observation may only be aligned with L2 or L3.

[0067] As shown in the example in Figure 3, the black circles correspond to "pinned states," meaning that the state or label of each corresponding observation is known beforehand. In contrast, the unfilled circles correspond to "normal states," meaning that the state or label of each corresponding observation is not known beforehand. Since the states of the continuous observations 302B and 302C are known, and the underlying labels are known to be monotonically aligned with the observations, there is only one valid path within the posterior stochastic grid 305 between the states of the continuous observations 302B and 302C. Since the states of observations 302C and 302E are also known, there are only two valid paths within the posterior stochastic grid 305 between the states of observations 302C and 302E. If pinned state information or cluster boundary information were not available, there would be more potentially valid paths between the states of observations within the posterior stochastic grid 305 for different classes.

[0068] Therefore, the PS-CSC implementation can result in faster convergence and lower computational overhead. Furthermore, the posterior probability generator 204 does not need to determine the probability that an observation has a pinned state. Therefore, pinned states in some PS-CSC implementations can result in even lower computational overhead. Thus, Figure 3 represents, or at least suggests, the potential advantages of the PS-CSC implementation.

[0069] As described above, in the example described with reference to FIGS. 2 and 3, the loss calculation unit 206 is configured to apply a loss function that realizes PS-CSC. Some examples are described below.

[0070] [Examples of PS-CSC loss functions] One observation sequence X = [x1, x2, x3, ··· x N , for example, assuming that only one instance of the observation sequence 203 shown in FIG. 2 exists. In some examples, there may be multiple observation sequences. In some such examples, the loss function may provide the sum of individual loss values. However, without loss of generality, the following example includes one observation sequence for simplicity. The corresponding ground truth label in one such example is the ground truth label sequence 303 [L1, L2, L3] shown in FIG. 3. According to the following example, the path through the posterior probability lattice is represented as π. One possible path is π1 = [L1, L1, L2, L2, ··· L3]. By applying the independence assumption, which can be the same independence assumption as that of Connectionist Temporal Classification (CTC), the probability of the entire path can be expressed as follows:

Equation

[0071] For example, the loss function can represent the sum of all possible paths as follows:

number

[0072] The pinned states shown in Figure 3 mean that a particular state must be passed through a valid path. The pinned states are given as prior knowledge in this example. Due to the pinned states, the number of valid paths is reduced. The CTC loss can be considered an extreme example of PS-CSC where no pinned states exist.

[0073] In another extreme case, if the alignment of all observations in a given sequence of observations is known, there is only one possible path. This single path is π for simplicity. ce It can be called this. In that case, the PS-CSC objective function can be expressed as follows:

number

[0074] [Example of speaker diarization] Traditionally, in speaker recognition systems (including speaker identification and matching), an embedded extractor is trained first. This includes the Defactor process disclosed in Kinnunen, Tomi, and Haizhou Li, "An overview of text-independent speaker recognition: From features to supervectors", Speech Communication 52, no. 1 (2010), pages 12-40, and the i-vector system disclosed in Dehak, Najim, et al., "Front-end factor analysis for speaker verification", IEEE Transactions on Audio, Speech, and Language Processing 19, no. 4 (2010): pages 788-798, as well as those using neural networks as extractors (e.g., Snyder, David, et al.). This is the x-vector system disclosed in al., "X-vectors: Robust dnn embeddings for speaker recognition", 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp. 5329-5333. (IEEE, 2018). Subsequently, cosine distance similarity or probabilistic linear discriminant analysis (PLDA) is typically then applied or trained to measure the similarity between the two embeddings.

[0075] The embedding extractor is trained using cross-entropy with an objective function based on SoftMax or its variations (angular SoftMax, additive marginal SoftMax, etc.), and uses the cross-entropy loss to map aggregated utterance level embeddings to corresponding class labels. In other words, the extractor learns to accumulate ID-related information from a single utterance.

[0076] In some use cases, such as online speaker dialization, the goal is to identify who is speaking and when. This process typically takes place within a limited timeframe (which may be a few seconds or less). Therefore, the amount of information to be collected is also limited. The duration is constrained by the allowable interval between the boundaries of two decisions, which is typically 250 to 300 milliseconds. Furthermore, in the above example, the embedded extractor is not trained sequentially in the dialization scenario.

[0077] Some disclosed examples involve using PS-CSC to learn both speaker ID and time alignment. The loss applied by the loss function can allow the neural network to learn how to map sequential embeddings to sequence labels, which can be viewed as a sequence-to-sequence problem. Unlike speech recognition use cases, blanks are unnecessary when two segments are from the same speaker, as the same speaker label is shared by both. Since the boundaries of some labels may be known in advance, that boundary information can be fed into PS-CSC before the training process, which is not possible with CTC or cross-entropy loss.

[0078] Figure 4 shows an example of the blocks included in using PS-CSC to train a neural network in the speaker diarization use case. As with other figures provided herein, the type, number, and arrangement of elements shown in Figure 4 are given only as an example. Other implementations may include more, fewer, and / or different types, numbers, and arrangements of elements.

[0079] Following this example, Figure 4 shows the blocks implemented by an instance of the control system 110 described with reference to Figure 1A. As described elsewhere in this specification, the control system 110 may in some cases be present in more than one device. In this example, the control system 110 is configured to implement at least a posterior probability generator 404, a loss calculation unit 406, and a parameter update block 407, which are instances of the posterior probability generator 204, loss calculation unit 206, and parameter update block 207 in Figure 2, respectively. Following this example, the loss calculation unit 406 is configured to apply a loss function that realizes PS-CSC. In some alternative examples, the control system 110 may also implement an embedding extraction unit 402, which is an instance of the feature extraction unit 202 in Figure 2.

[0080] In this example, the input sequential signal 401 is an audio data sample containing, in that order, audio signals corresponding to the speaker's utterance or speakers spk_684, spk_379, spk_167, spk_1195, and spk_527. Therefore, the ground truth label sequence 408 contains labels L1, L2, L3, L4, and L5 corresponding to speakers spk_684, spk_379, spk_167, spk_1195, and spk_527, respectively. Following this example, the embedding extraction unit 402 generates an observation sequence 403 from the input sequential signal 401, which is a speaker embedding corresponding to speakers spk_684, spk_379, spk_167, spk_1195, and spk_527.

[0081] In this example, the posterior probability generator 404 is configured to generate a posterior probability grid 405. In this example, the posterior probability generator 404 is configured to generate the posterior probability grid 405 by inferring that the probability of each extracted feature in the observation sequence 403 belongs to a class indicated by one of the ground truth labels in the ground truth label sequence 408. In this example, the ground truth label sequence 408 is [L1, L2, L3, L4, L5], and the posterior probability generator 404 is configured to generate the posterior probability grid 405 by inferring that the probability of each extracted feature in the observation sequence 403 belongs to class L1, L2, L3, L4, or L5.

[0082] In the example shown in Figure 4, the loss calculation unit 406 is configured to apply a loss function that realizes PS-CSC. Here, the loss calculation unit 406 is configured to apply the loss function to the values ​​of the posterior stochastic grid 405 according to the ground truth label sequence 408. According to this example, the loss calculation unit 406 is configured to supply updated parameters, which are represented as a parameter update block 407 in Figure 4, according to the loss determined by the loss function. In this example, the loss calculation unit 406 is configured to supply the updated parameters to the posterior stochastic generation unit 404 in order to generate the posterior stochastic grid 405. According to some examples, the loss calculation unit 406 may be configured to supply the updated parameters to the embedding extraction unit 402 in order to determine several extracted features in the observation sequence 403.

[0083] [Experimental Results] The inventors conducted experiments to demonstrate the potential of PS-CSC in speaker diarization scenarios. Figures 5A, 5B, and 5C show the results of experiments using conventional methods in speaker diarization use cases. In the examples shown in Figures 5A-5C, the embedding extractor was trained using cross-entropy loss. In these examples, sequence information was not introduced during the training process. Figures 6A and 6B show the results of experiments using PS-CSC in speaker diarization use cases. In the examples shown in Figures 6A and 6B, sequence information was supplied during the training process, and the embedding extractor was trained using CTC. The methods shown in Figures 5A-5C and 6A-6B shared the same dataset and model. In all cases, the SoftMax score was then used to indicate the quality of the speaker embeddings extracted for diarization.

[0084] In the experiments shown in Figures 5A-5C, sequence boundary information was not provided during training. Using PS-CSC improved separability and reduced training time. Both of these improvements were obtained, at least partially, from sequence boundary information provided during the training process.

[0085] In the experiment shown in Figures 5A-5C, the ground truth label sequence is [spk_684,spk_379,spk_167,spk_1195,spk_527]. Figure 5A shows an example of the input audio segment 503 in the time domain. Figure 5B shows graph 502, which includes curves 505A, 505B, 505C, 505D, and 505E, respectively, representing the SoftMax scores of the utterances of speakers spk_684, spk_379, spk_167, spk_1195, and spk_527 when the neural network was trained using the CE cost function but without sequence information. Figure 5B shows that the neural network trained without sequence information did not perform well in predicting speaker alignment.

[0086] Figure 5C shows Graph 501, which includes curves 504A, 504B, 504C, 504D, and 504E, representing the SoftMax scores of utterances by speakers spk_684, spk_379, spk_167, spk_1195, and spk_527, respectively, when the neural network is trained using the Sequence Loss Cost Function (CTC). Comparing Figure 5C with Figure 5B, it can be seen that the neural network performs much better when sequence information is provided during the training process. For example, as is clear from Graph 501, frames from time unit 0 to 5 are aligned to spk_684 with a high probability. However, no clear trend is observed in Graph 502.

[0087] The experiments shown in Figures 5A-5C demonstrate that using sequence information when training neural networks is highly useful in speaker dialization use cases.

[0088] Figures 6A and 6B illustrate the advantages of using sequence and alignment information in the PS-CSC implementation. Using several pre-alignment information pieces, also referred to herein as pinning information, reduces the number of valid paths through the posterior probability grid that need to be computed during training.

[0089] Figure 6A shows not only an example of the input audio segment 603A in the time domain, but also the alignment information 610A and 610B. Figure 6A also shows an example of the posterior probability grid 205 in Figure 2, which in this case is the posterior probability grid 605A. Following this example, as can also be seen in Figure 3, the black circles in the posterior probability grid 605A indicate pinning states, and each pinning information corresponds to one of the instances of the alignment information 610A and 610B. In this example, only the alignments corresponding to spk_684 and spk_379 are known before the neural network training process. Figure 6A shows regions 604A and 604B within the posterior probability grid 605A. In this example, given the known alignment information 610A and 610B, regions 604A and 604B are regions where a valid path through the posterior probability grid 605A may exist.

[0090] Figure 6B also shows an example of the input audio segment 603B in the time domain, but without alignment information. Figure 6B also shows the posterior probability grid 605B. As in this example, since alignment information is not provided, there are no black circles in the posterior probability grid 605A indicating the pinned state. In this example, since no prior alignment information was provided before training, there may be valid paths through the posterior probability grid 605B within region 602. Even with only alignment information 610A and 610B provided, it can be seen that regions 604A and 604B are much smaller than region 602. Therefore, since no prior alignment information was provided before training, the number of valid paths through the posterior probability grid 605B is significantly greater than the number of valid paths through the posterior probability grid 605A.

[0091] Figure 7 shows the results of two additional experiments. In this example, graph 700 shows the loss determined by the loss function on the vertical axis and the number of iterations of the neural network training process on the horizontal axis. Following this example, curve 705 shows the loss value corresponding to the CTC-based neural network training process, while curve 710 shows the loss value corresponding to the PS-CSC-based neural network training process. After approximately 45,000 iterations, it can be seen that the loss corresponding to the PS-CSC-based neural network training process was consistently lower than the loss corresponding to the CTC-based neural network training process.

[0092] Figure 8 is a flowchart illustrating an example of the disclosed method. The blocks of Method 800, as with other methods described herein, are not necessarily performed in the order shown. In some examples, one or more blocks may be performed in parallel. Furthermore, some similar methods may include more or fewer blocks than those illustrated and / or described.

[0093] Method 800 may be carried out by an apparatus or system such as the apparatus 100 shown in Figure 1 and described above. In some examples, apparatus 100 includes at least the control system 110 disclosed herein. In some examples, at least some aspects of Method 800 may be carried out by one or more devices in the audio environment, for example, by an audio system controller (for example, which may be called a smart home hub herein), or by other components of the audio system such as a television, a television control module, a laptop computer, or a mobile device (for example, a cellular phone). However, in some implementations, at least some blocks of Method 800 may be carried out by one or more devices that can be configured to provide cloud-based services, such as one or more systems.

[0094] In this example, block 805 includes receiving an observation sequence containing multiple extracted features by the control system. According to some examples, each extracted feature of the multiple extracted features may correspond to a sequential signal in a sequence of sequential signals. In some examples, the sequence of sequential signals may be a time sequence containing audio signals, and in some cases may contain audio signals corresponding to speech. In some alternative examples, the sequence of sequential signals may be a time sequence containing other types of audio signals, signals corresponding to stock market data, signals corresponding to biological data, etc. According to some alternative examples, the sequence of sequential signals may be a spatial sequence, such as a spatial sequence containing a handwritten image, a DNA sequence image, or other images.

[0095] Following this example, block 810 includes the control system determining a grid of posterior probabilities. In some examples, the grid may contain the probabilities of each observation sequence corresponding to one of several label classes. In some examples, the label classes may correspond to the respective IDs of several speakers.

[0096] In this example, block 815 involves the control system applying a loss function to a posterior probability grid according to the ground truth value. Applying the loss function according to this example involves applying both sequence information and cluster boundary information. According to some examples, the sequence information and / or cluster boundary information of the cost function can be variable. In some examples, the cluster boundary information may be incomplete; however, in some alternative examples, the cluster boundary information may be complete. In some examples, the loss function may combine elements of connectionist time series classification and cross-entropy loss. According to some examples, applying the loss function may involve determining one or more valid paths between observations in the grid.

[0097] In this example, block 820 includes updating the parameters for determining the posterior probability grid according to the loss determined by the loss function, as performed by the control system. In some examples, block 820 may include updating the parameters used by the posterior probability generator 204 in Figure 2 according to the parameters of parameter update block 207.

[0098] In this example, block 825 includes continuing to execute blocks 805 through 820 until the control system determines that one or more convergence criteria have been met. In some examples, the control system 110 may be configured to determine that convergence has been achieved when the training process reaches a state in which the loss determined by the loss function has settled within an error range around the final value, or when the loss determined by the loss function is no longer decreasing, or when the loss determined by the loss function does not decrease for a predetermined number of steps or epochs.

[0099] In some examples, performing blocks 805 to 820 may provide a trained neural network to be implemented by the control system 110. Method 800 may, in some examples, involve the trained neural network performing one or more downstream tasks.

[0100] As in some examples, a sequence of signals may be, or include, a sequence of audio signals. In some such examples, one or more downstream tasks may include acoustic event detection, speaker dialization, automatic speech recognition, or a combination thereof.

[0101] Method 800 may, in some examples, include determining a plurality of extracted features by a control system. In some such examples, the method may include updating parameters for determining the plurality of extracted features by a control system according to a loss determined by a loss function.

[0102] In some examples, a sequence of signals may be, or include, a sequence of handwritten images. In some such examples, one or more downstream tasks may include handwriting recognition.

[0103] Some aspects of this disclosure include systems or devices configured (e.g., programmed) to perform one or more examples of the disclosed methods, or tangible computer-readable media (e.g., disks) storing code for performing one or more examples of the disclosed methods or steps thereof. For example, some of the disclosed systems may be, or may include, programmable general-purpose processors, digital signal processors, or microprocessors programmed in software or firmware and / or otherwise configured to perform any of a variety of operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be, or may include, a computer system including input devices, memory, and processing subsystems, the processing system being programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0104] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on an audio signal, including the execution of one or more examples of the disclosed methods. Alternatively, the disclosed system (or its elements) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) programmed in software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the system of the invention may be implemented as a general-purpose processor or DSP configured (or programmed) to execute one or more examples of the disclosed methods, and the system may also include other elements (e.g., one or more loudspeakers and / or one or more microphones). The general-purpose processor configured to execute one or more examples of the disclosed methods may be coupled with input devices (e.g., a mouse and / or keyboard), memory, and display devices.

[0105] Another aspect of this disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) storing code (e.g., an executable coder) for performing one or more examples of the disclosed method or steps thereof.

[0106] While specific embodiments and applications of the Disclosure have been described herein, as will be apparent to those skilled in the art, numerous modifications to the embodiments and applications described herein are possible without departing from the scope of the Disclosure as described and claimed herein. Although specific forms of the Disclosure have been illustrated and described herein, it should be understood that the Disclosure should not be limited to the specific embodiments or methods described herein.

[0107] [Cross-reference of related applications] This application claims priority rights to U.S. Provisional Patent Application No. 63 / 513294, filed on 12 July 2023, and U.S. Provisional Patent Application No. 63 / 401042, filed on 25 August 2022, which are incorporated herein by reference.

Claims

1. (a) The control system extracts a sequence of observed features from a sequence of sequential signals, wherein each extracted feature corresponds to one of the sequential signals, and the sequential signals are either a sequence of audio signals or a sequence of handwritten images. (b) The control system determines a grid of posterior probabilities, the grid including the probability that each of the extracted features belongs to a class indicated by one of a plurality of ground truth labels in the ground truth label sequence, (c) The control system applies a loss function to the posterior probability grid according to the ground truth labels, the application of the loss function includes applying both sequence information and cluster boundary information from a plurality of the extracted features to connect the labels of the extracted features within the grid and determine one or more valid paths including one or more pinning labels known in advance for one or more corresponding extracted features, wherein a cluster is a group of extracted features having the same label, and the cluster boundary information indicates a transition from one cluster of the extracted features to another class of the extracted features, (d) The control system updates the parameters for determining the posterior probability grid according to the loss determined by the loss function, (e) Continue performing (a) to (d) until the control system determines that one or more convergence criteria are met. It has, Performing (a) to (e) provides a trained neural network implemented by the control system, The aforementioned trained neural network further performs one or more downstream tasks, If the sequential signal is a sequence of audio signals, the one or more downstream tasks include acoustic event detection, speaker dialization, and automatic speech recognition; if the sequential signal is a sequence of handwritten images, the one or more downstream tasks include handwriting recognition. method.

2. The aforementioned cluster boundary information includes incomplete cluster boundary information. The method according to claim 1.

3. The aforementioned cluster boundary information includes complete cluster boundary information. The method according to claim 1.

4. The control system further comprises determining the plurality of extracted features. The method according to claim 1.

5. The control system further includes updating parameters for determining the plurality of extracted features according to the loss determined by the loss function. The method according to claim 4.

6. The loss function combines elements of connectionist time series classification and cross-entropy loss. The method according to claim 1.

7. At least one of the sequence information of the loss function and the cluster boundary information of the loss function is variable. The method according to claim 1.

8. A device comprising one or more devices and one or more computer-readable media storing computer executable instructions, When the computer executable instruction is executed by one or more devices, it causes the one or more devices to perform the method according to any one of claims 1 to 7. system.

9. One or more non-temporary computer-readable media storing computer-executable instructions that, when executed by one or more devices, cause one or more devices to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Efficient connectionist temporal classification for binary classification

    US20180232632A1

  • Handwriting recognition with language modeling

    US20220138453A1

  • Learning device, speech recognition device, learning method, speech recognition method, learning program, and speech recognition program

    WO2022024202A1