Speech recognition and interlaced audio input enhancement using data analysis

Convolutional neural networks with augmentation parameters enhance speech recognition by refining interlaced audio input, reducing word error rates through diarization and voiceprint attribution.

JP7776236B2Active Publication Date: 2025-11-26INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023513524
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-09
Filing Date
2021-08-24
Publication Date
2025-11-26
Estimated Expiration
2041-08-24

AI Technical Summary

Technical Problem

Existing speech recognition technologies struggle with converting interlaced audio input containing altered or atypical sounds, background noise, and multiple speakers, leading to increased word error rates.

Method used

The use of convolutional neural networks with augmentation parameters to enhance audio signals by increasing spacing between samples, refining audio input through diarization, and attributing speech content to individual speakers based on voiceprints.

Benefits of technology

Reduces word error rates in speech recognition by accurately converting interlaced audio input to text, even in noisy environments and with multiple speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007776236000001
    Figure 0007776236000001
  • Figure 0007776236000002
    Figure 0007776236000002
  • Figure 0007776236000003
    Figure 0007776236000003
Patent Text Reader

Abstract

The present disclosure comprises using augmentation of audio content from interlaced audio input for speech recognition. A training model is initiated to determine augmentation parameters for each of a plurality of audible sounds of audio content from a plurality of speakers received at a computer as audio input. As part of the training model, variations in each of a plurality of independent sounds are determined in response to audio stimuli, and the independent sounds are derived from the audio input. The present disclosure applies the augmentation parameters based on the variations in each of the independent sounds. A voiceprint is constructed for each of the speakers based on each of the independent sounds and the augmentation parameters. The audio content is attributed to each of the plurality of speakers based at least in part on the voiceprint and the independent sounds.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to computer-implemented techniques for speech recognition of speech content from audio input, more particularly, audio input including interlaced speech content and conversion or transcription of speech content to text.

[0002] Computer-based techniques can be used to convert human speech into text. Human speech may include, for example, speaking alone or in a group, singing, etc. During human speech, converting the audio output or audio output signal to be converted to text can be difficult. For example, if the sounds are altered or less typical than the typical phonetics of the words, speech recognition and conversion can be difficult. For example, sounds may be stretched or mixed with one or more other noises. In one example, background noise may be present when a speaker is speaking. In another example, a group of speakers may be speaking, and there may be speaker overlap. In another example, background noise may occur when one or more speakers are speaking. In another example, a speaker may unintentionally or intentionally alter the typical phonetics of one or more words for emphasis, or as part of an unorthodox or atypical speech pattern, or as part of an accent. Such altered sounds or atypical sounds, or both, result in speech that is difficult for speech identification and speech-to-text conversion when the speaker speaks. Summary of the Invention

[0003] The present disclosure recognizes the shortcomings and problems associated with current techniques for speech recognition using enhancement of speech content from interlaced audio input.

[0004] The present invention can analyze speech content from interlaced audio input for speech recognition of each of multiple speakers and provide conversion of the speech content to text. For example, the challenges of speech recognition and conversion can be overcome using the present invention when the sounds are altered or less typical than the typical phonetics of words, when the speech content contains altered or atypical sounds, or both, from speakers for identification and speech-to-text conversion.

[0005] One problem can arise, for example, when an artist is singing, and some of the words may change or be altered in a way that follows harmonics rather than general phonetics. In another example, in noisy environments, the mixing of sound waves and tones can increase the error rate of words when converting. For example, at large events, the roar of a loud crowd or the sounds of a sporting event can interfere with the audio signal.

[0006] The present invention provides speech recognition using augmentation of an audio signal, or audio input, to increase the spacing between samples or audio samplings before attempting to recognize words or analyzing audio content to recognize a word or words. In one example, according to the present invention, convolutional neural networks (CNNs) with various augmentation parameters are trained and can be applied to these problems. Furthermore, expected environmental noise and voice type may indicate which augmentation to use. In yet another example, each speaker may be assigned an augmentation parameter through machine learning. In a group setting of speech or singing, the augmentation parameters may be weighted together by the group based on the amplitude of each speaker.

[0007] In an aspect according to the present invention, a computer-implemented method for speech recognition uses augmentation of speech content from interlaced audio input. The method includes initiating a training model to determine augmentation parameters for each of a plurality of audible sounds of speech content from a plurality of speakers received at the computer as audio input. The method includes determining, as part of the training model, a variation of each of a plurality of independent sounds in response to an audio stimulus. The plurality of independent sounds are derived from the audio input. The method includes applying the augmentation parameters based on the variation of each of the plurality of independent sounds, respectively. A voiceprint is constructed for each of the plurality of speakers based on the plurality of independent sounds and the augmentation parameters, respectively. The method includes attributing speech content to each of the plurality of speakers, respectively, based at least in part on the voiceprint and the plurality of independent sounds.

[0008] One advantage of the present invention includes using a method according to the present invention to reduce word error rates when converting speech content from interlaced audio input using multi-speaker speech recognition to text.

[0009] In a related aspect, the method further comprises generating text from the attributed audio content.

[0010] In a related aspect, the method further comprises using the computer to display the text on a screen or monitor in communication with the computer or the device, or both.

[0011] In a related aspect, the method further comprises transmitting the text to a computer or device, or both, via an electronic communication system for display on a screen or monitor in communication with the computer or device, or both.

[0012] In a related aspect, the method includes displaying the text on a screen or monitor in communication with the computer or device, or both.

[0013] In a related aspect, the audio input may comprise a plurality of audible tones, the audio input being received at a computer, the plurality of audible tones including speech content from a plurality of speakers.

[0014] In a related aspect, the method further comprises using the computer to enhance the audio input, the enhancing comprising separating sounds in the audio input.

[0015] In a related aspect, the method further comprises refining the audio input for each of a plurality of speakers using diarization.

[0016] In a related aspect, the learning model includes a convolutional neural network (CNN) for receiving a plurality of independent sounds and determining a change in each of the plurality of independent sounds in response to an audio stimulus using diarization.

[0017] In a related aspect, the method further comprises layering sounds in the refined audio input into multiple separate sounds using a diarization of the audio input.

[0018] In a related aspect, the method further comprises receiving, at a computer, an audio input having a plurality of audible sounds, the plurality of audible sounds including speech content from a plurality of speakers. The method further comprises augmenting the audio input using the computer, the augmenting comprising separating sounds in the audio input. The method comprises refining the audio input for each of the plurality of speakers using diarization, and layering the sounds in the refined audio input into a plurality of separate sounds using the diarization of the audio input.

[0019] In a related aspect, separating sounds in the audio input includes distinguishing environmental or background sounds from speech from one speaker of multiple speakers.

[0020] In a related aspect, refining the audio input for each of a plurality of speakers using diarization includes partitioning the audio input into homogeneous segments for speaker identification.

[0021] In another aspect according to the invention, a system for speech recognition uses enhancement of audio content from interlaced audio input and includes a computer system having a computer processor, a computer-readable storage medium, and program instructions stored on the computer-readable storage medium that are executable by the processor, causing the computer system to perform the following functions: initiate a learning model to determine enhancement parameters for each of a plurality of audible sounds of audio content from a plurality of speakers received at the computer as audio input; determine, as part of the learning model, variations in each of a plurality of independent sounds in response to audio stimuli, where the plurality of independent sounds are derived from the audio input; apply the enhancement parameters based on the variations in each of the plurality of independent sounds, respectively; construct a voiceprint for each of the plurality of speakers based on the plurality of independent sounds and the enhancement parameters, respectively; and attribute audio content to each of the plurality of speakers based at least in part on the voiceprint and the plurality of independent sounds, respectively.

[0022] One advantage of the present invention includes reducing word error rates when using a system according to the present invention to convert speech content from interlaced audio input to text using multi-speaker speech recognition.

[0023] In a related aspect, the system further comprises generating text from the attributed audio content.

[0024] In a related aspect, the system further comprises using the computer to display the text on a screen or monitor in communication with the computer or the device, or both.

[0025] In a related aspect, the system further comprises transmitting the text to a computer or device, or both, via an electronic communication system for display on a screen or monitor in communication with the computer or device, or both.

[0026] In a related aspect, the audio input may comprise a plurality of audible tones, the audio input being received at a computer, the plurality of audible tones including speech content from a plurality of speakers.

[0027] In a related aspect, the system further comprises augmenting the audio input using the computer, where augmenting comprises separating sounds in the audio input.

[0028] In a related aspect, the system further comprises refining the audio input for each of a plurality of speakers using diarization.

[0029] In a related aspect, the learning model includes a convolutional neural network (CNN) for receiving a plurality of independent sounds and determining a change in each of the plurality of independent sounds in response to an audio stimulus using diarization.

[0030] In a related aspect, the system further comprises layering sounds in the refined audio input into multiple separate sounds using diarization of the audio input.

[0031] In another aspect according to the invention, a computer program product for speech recognition uses augmentation of audio content from interlaced audio input and includes a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a computer to cause the computer to perform the following functions: initiate a learning model to determine augmentation parameters for each of a plurality of audible sounds of audio content from a plurality of speakers received at the computer as audio input; determine, as part of the learning model, variations in each of a plurality of independent sounds in response to audio stimuli, where the plurality of independent sounds are derived from the audio input; apply the augmentation parameters based on the variations in each of the plurality of independent sounds, respectively; construct a voiceprint for each of the plurality of speakers based on the plurality of independent sounds and the augmentation parameters, respectively; and attribute audio content to each of the plurality of speakers based at least in part on the voiceprint and the plurality of independent sounds, respectively.

[0032] One advantage of the present invention includes a procedure for reducing word error rates when using a computer program product according to the present invention to convert speech content from interlaced audio input to text using multi-speaker speech recognition.

[0033] In a related aspect, the computer program product further comprises generating text from the attributed audio content.

[0034] In a related aspect, the computer program product further comprises instructions for using the computer to display the text on a screen or monitor in communication with the computer or a device, or both.

[0035] In a related aspect, the computer program product further comprises transmitting the text via an electronic communication system to a computer or device, or both, for display on a screen or monitor in communication with the computer or device, or both. [Brief explanation of the drawings]

[0036] These and other objects, features, and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the invention, which should be read in conjunction with the accompanying drawings, in which various features of the drawings are not to scale, as the illustrations are for clarity in facilitating those skilled in the art to understand the invention in relation to the detailed description, and the drawings are discussed below.

[0037] [Figure 1] FIG. 1 is a schematic block diagram outlining a system, system functions or components, and methodology for speech recognition using speech content enhancement from interlaced audio input according to one embodiment of the present disclosure.

[0038] [Figure 2] 2 is a flowchart illustrating a method according to an embodiment of the present disclosure, implemented using the system shown in FIG. 1, for speech recognition using speech content enhancement from interlaced audio input according to an embodiment of the present disclosure.

[0039] [Figure 3] 1 is a series of tables illustrating one embodiment of an extension according to the present disclosure.

[0040] [Figure 4] 2 is a flowchart illustrating another embodiment of a method according to the present disclosure for speech recognition using speech content enhancement from interlaced audio input, implemented using the system shown in FIG. 1 .

[0041] [Figure 5] 5 is a flowchart continuing from the flowchart shown in FIG. 4, depicting a continuation of the method shown in FIG. 4, according to one embodiment of the present invention.

[0042] [Figure 6] FIG. 6 is a functional schematic block diagram illustrating, for educational purposes, the sequence of operations and functional methodology illustrating the functional features of the present disclosure associated with the embodiments shown in FIGS. 1, 2, 3, 4, and 5 for speech recognition using audio content enhancement from interlaced audio input.

[0043] [Figure 7] FIG. 6 is a functional schematic block diagram illustrating, for educational purposes, a sequence of operations and a functional methodology illustrating functional features of the present disclosure associated with the embodiments shown in FIGS. 1, 2, 3, 4, and 5 for speech recognition using audio content enhancement from interlaced audio input.

[0044] [Figure 8] 1 is a schematic block diagram illustrating a computer system according to one embodiment of the present disclosure, which may be incorporated in whole or in part into one or more computers or devices shown in FIG. 1 and which cooperates with the systems and methods shown in FIGS. 1, 2, 3, 4, 5, 6, and 7.

[0045] [Figure 9] 1 is a schematic block diagram of a system depicting system components interconnected using a bus, in whole or in part, in accordance with one or more embodiments of the present disclosure, for use with embodiments of the present disclosure.

[0046] [Figure 10] FIG. 1 is a block diagram illustrating a cloud computing environment according to one embodiment of the present invention.

[0047] [Figure 11] FIG. 2 is a block diagram illustrating an abstract model layer according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0048] The following description, which refers to the accompanying drawings, is provided to facilitate a comprehensive understanding of exemplary embodiments of the present invention, as defined by the claims and their equivalents. While various specific details are included to facilitate understanding, these are considered to be merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications to the embodiments described herein can be made without departing from the scope of the present invention. Furthermore, descriptions of well-known functions and configurations may be omitted for the sake of clarity and conciseness.

[0049] The terms and phrases used in the following description and claims are not limited to their bibliographical meanings, but are merely used to enable a clear and consistent understanding of the present invention. Therefore, it should be apparent to those skilled in the art that the following description of exemplary embodiments of the present invention is provided for illustrative purposes only, and is not intended to limit the present invention as defined by the appended claims and their equivalents.

[0050] The singular forms "a," "an," and "the" should be understood to include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "component surface" includes a reference to the presence of one or more of such surfaces unless the context clearly dictates otherwise.

[0051] Embodiments according to the present disclosure analyze speech content from an interlaced audio input to provide speech recognition for each of multiple speakers, thereby providing word recognition and identification and conversion of the speech content to text. The present disclosure enables speech recognition and speech-to-text conversion when sounds are altered or less typical than the typical phonetics of words, for example, when the speech content includes altered or atypical sounds from a speaker, or both.

[0052] Embodiments of the present disclosure provide speech recognition using augmentation of the audio signal or audio input to increase the spacing between samples or audio samplings before attempting to recognize words or analyzing audio content to recognize a word or words. In one example according to the present disclosure, convolutional neural networks (CNNs) with various augmentation parameters can be trained and applied to these problems. In another example, expected environmental noise and voice type may indicate which augmentation to use. In another example, each speaker may be assigned an augmentation parameter through machine learning. In another example, in a group setting of speech or singing, the augmentation parameters may be weighted together by the group based on the amplitude of each speaker.

[0053] Embodiments of the present disclosure may provide predictions of environmental noise and voice type to set the expansion parameters. In another example, embodiments of the present disclosure may classify voice types (e.g., singing, talking) that contribute to the expansion parameters. In another example, embodiments of the present disclosure may fit the voice spread to other independent models. In another example, embodiments of the present disclosure may comprise averaging the expansion parameters based on speaker diarization and group models. Furthermore, in another example, embodiments of the present disclosure may comprise social expansion transfer of knowledge.

[0054] Thereby, embodiments of the present disclosure comprise modeling augmentation from environmental noise and predicting augmentation parameters. The extent of augmentation can be mapped, and further, social aspects can be combined with augmentation metrics for each person in the conversation.

[0055] 1 and 2, a method 100 (FIG. 2) in accordance with an embodiment of the present disclosure is provided for speech recognition using enhancement of speech content from interlaced audio input. Referring to FIG. 2, the method comprises a series of operational blocks for implementing one embodiment of the present disclosure. Referring to FIG. 2, the method 100 comprises, as seen in block 104, initiating a learning model 320 (see FIG. 6) to determine enhancement parameters 324 for each of a plurality of audible sounds 62 of speech content 64 from a plurality of human speakers 52 received at the computer 22 as audio input 60.

[0056] Referring to FIG. 6, a functional system 300 has components and operations related to an embodiment according to the present disclosure and is used herein with reference to the methods and systems shown in FIGS. 1, 2, 3, 4 and 5.

[0057] In one example, a group of speakers may be speaking together, and audio output from the group of speakers may be received as audio input using a computer or device, such as using a microphone on the device or in communication with the device or computer.

[0058] In one example, a spectrogram may be generated and used as a visual representation of the spectrum of frequencies of a signal as it varies over time, such as in an audio signal. A spectrogram may also be referred to as a sonograph, voiceprint, or voicegram.

[0059] A spectrogram can be formed and a DFT (Discrete Fourier Transform) applied to identify possible distinctive speakers. The DFT can convert a finite sequence of equally spaced sample functions into an equal length sequence of equally spaced samples of the Discrete Time Fourier Transform (DTFT), which is a complex-valued function of frequency. Initial expansion variables can be initialized to estimates for each DFT.

[0060] In one example, if it is known who is speaking or singing in a group, the augmentation parameters can be adjusted or specified based on the information. Such identifying information can be gleaned from social media input, observations, for example.

[0061] The plurality of audible sounds may include, for example, one or more human speakers 14 or users in the vicinity 50 speaking as a plurality of human speakers 52 or users and producing audible sounds 62. The audible sounds may include, for example, human speech in a conversation, solo speech, singing, a group of speakers singing, etc. Thus, the audible sounds 62 produce and have audio content 64.

[0062] The audible sound may be received at the computer via a device 20, such as a microphone at computer 22 or a mobile device, as audio input 60 for processing by the techniques of this disclosure, and the computer, alone or in combination with a control device of control system 70, may transmit the audio file to another computer 72 or server, such as a remote computer or server (via communications network 45, e.g., the Internet). In another example, the audible sound in the audio file may be processed by the techniques of this disclosure locally on the computer, or in combination with a remote computer or server, or both.

[0063] The learning model 320 may include machine learning using the parameters, for example, the extended parameters may be assigned to each of multiple speakers or users using machine learning.

[0064] Sound enhancement in the context of speech recognition can be defined as increasing the space between sounds or sound samples. In this disclosure, enhancement is performed before attempting to recognize words from the sounds.

[0065] The expansion parameters may include a specified amount of space between sound samples, or specify a range of space between sound samples.

[0066] Extension variables can be assigned to each potential speaker and used in the training model.

[0067] 3, representative image 150 of tables 154, 158, and 162 showing the expansion of image 166 may be, for example, a sound image and may be layered into portions or sound samples 168. Image 150 shows variations in expansion parameter D. In table 154, expansion parameter 172 is equal to 1, and image 166 has no spacing. In second table 158, expansion parameter 174 is equal to 2, and the image has sound samples 168 with space 180 between samples. In third table 162, expansion parameter 176 is equal to 3, and the image has sound samples 168 with more space 180 between samples.

[0068] The method includes, as part of the learning model, determining a change in each of a plurality of independent sounds in response to an audio stimulus, where the independent sounds are derived from the audio input, as seen in block 108. For example, the audio stimulus may include an environmental stimulus. A change in the sound, or an independent sound change 322, may be determined in response to the environmental stimulus.

[0069] In one embodiment, the audio input can be refined for each of multiple speakers using diarization. For example, diarization can include partitioning the input audio stream into segments corresponding to speaker identities; in one example, the segments can be homogeneous. The diarized signal can be used to layer the audio input into individual sounds.

[0070] Diarization can be used in deep learning to refine the audio input from each of multiple speakers to attribute to each speaker. In one example, if there is error from the DFT or deep learning, or both approaches to speaker identification, the diarization parameters can be averaged together.

[0071] In one example, a voiceprint can be constructed. In one example, environmental stimuli can be performed and a determination can be made as to how the layered data changes. The expansion parameters can be modified based on changes in audio input data (e.g., speech data). For example, the expansion parameters can increase as the speech becomes more extended. Furthermore, the expansion parameters for each speaker can be based on the correlation coefficient of the independent signals. The correlation coefficient (R value) is a value given in a tabulation in the regression output. R-squared is called the coefficient of determination, i.e., R x R to obtain the R-squared value. The coefficient of determination is the square of the correlation coefficient.

[0072] In one example, the R-squared correlation metric determines how to combine pairs of most correlated speakers together. For example, the rank of R can be shifted from 0 to 0.5, so that at most, paired speakers contribute 50% of the coordinated expansion.

[0073] The method comprises applying an expansion parameter based on each variation of each independent sound, as seen in block 112 .

[0074] The method comprises constructing a voiceprint for each of the speakers based on the independent sound and extension parameters, respectively, as seen in block 116 .

[0075] The method, as seen in block 120, comprises attributing audio content to each of a plurality of speakers based at least in part on the respective voiceprints and individual sounds.

[0076] The method may include generating text from the attributed audio content, as seen in block 124 .

[0077] If, as determined in block 126, the generated text is to be displayed locally, for example on a local computer, the method continues to block 130. If, as determined in block 126, the generated text is not to be displayed locally but rather on a device computer display or monitor, the method continues to block 128.

[0078] The method includes, in response to displaying the text location as determined in block 126, displaying the text on a screen or monitor in communication with the computer or device, as seen in block 130.

[0079] In response to not displaying the text locally as determined in block 126, the method may include transmitting the text to the computer or device, or both, via an electronic communications system for display on a screen or monitor in communication with the computer or device, or both, as seen in block 128. The method may continue with displaying the text on a screen or monitor in communication with the computer or device, as seen in block 130.

[0080] The method may include a training model 320 having a CNN 326 (convolutional neural network) for receiving the individual sounds and determining the variation of each of the individual sounds in response to the audio stimulus using diarization.

[0081] A CNN (convolutional neural network) may be at least part of deep learning, and a CNN is a type of deep neural network. A CNN involves a mathematical operation generally defined as two functions generating a third function, called a convolution. A convolution is a special type of linear operation. Thus, a convolutional network is a neural network that uses convolution instead of general matrix multiplication in at least one of its multiple layers.

[0082] The method may include receiving an audio input at a computer, the audio input may have a plurality of audible sounds. Further, the audible sounds may include speech content from multiple speakers.

[0083] The method may comprise using a computer to enhance the audio input, where the enhancement may include separating sounds in the audio input.

[0084] The method may include refining the augmented audio input 302 for each of a plurality of speakers using diarization of the sounds 308 .

[0085] The method may further comprise layering sounds 306 in the refined audio input 304 into separate sounds 310 using diarization of the audio input.

[0086] The method may comprise separating sounds in an audio input, comprising distinguishing environmental or background sounds from speech from one speaker of a plurality of speakers.

[0087] The method may comprise refining the audio input for each of a plurality of speakers using diarization, which may include partitioning the audio input into homogeneous segments for speaker identification.

[0088] In another embodiment according to the present disclosure, and referring to Figure 4, as seen in block 204, a computer-implemented method 200 for speech recognition using speech content enhancement from interlaced audio input comprises receiving, at a computer, an audio input having a plurality of audible sounds, the audible sounds including speech content from a plurality of speakers. The operational blocks of method 200 shown in Figures 4 and 5 may be similar to the operational blocks shown in Figure 2. The method shown in Figures 4 and 5 is intended as another exemplary embodiment that may include aspects / operations already shown and discussed in this disclosure.

[0089] The method 200 includes using a computer to enhance the audio input, as seen in block 208, where enhancing comprises separating sounds in the audio input.

[0090] The method 200 includes refining the audio input for each of a plurality of speakers using diarization, as seen in block 212. The method may include using the diarization of the audio input to layer sounds in the audio input into separate sounds, as seen in block 216.

[0091] The method 200 includes initiating a learning model to determine enhancement parameters for each of the audible sounds, as seen in block 220 .

[0092] The method 200 may include a learning model having a convolutional neural network (CNN) for receiving the individual sounds and determining the variation of each of the individual sounds in response to the audio stimulus using diarization, as seen in block 222.

[0093] The method 200 includes determining, as part of the learning model, a change in each of a plurality of independent sounds in response to an audio stimulus, as seen at block 224 .

[0094] The method 200 includes applying an expansion parameter based on each variation of each independent sound, as seen in block 228 .

[0095] The method 200 includes constructing a voiceprint 330 for each of the speakers 52 based on the individual sounds 310 and the extension parameters 324, respectively, as seen in block 232.

[0096] The method 200 includes attributing audio content to each of a plurality of speakers based at least in part on the voiceprints and individual sounds, respectively, as seen in block 236. The attributed audio content 332 can be used to generate text.

[0097] Referring to FIG. 5, the method 200 includes generating text 334 from the attributed audio content 332, as seen in block 240.

[0098] The method 200 further includes transmitting the text to a computer or device, or both, via an electronic communication system for display on a screen or monitor in communication with the computer or device, or both, as seen in block 244. In another example, the communication may be performed from the group consisting of SMS, email, instant message, navigation software. Such examples are intended to be illustrative and non-exhaustive.

[0099] The method 200 may further comprise displaying the text on a screen or monitor in communication with the computer or device, as seen at block 248 .

[0100] 7, a functional system 400 according to one embodiment of the present disclosure, and illustrating and supporting embodiments described herein, includes components and operations for speech recognition using enhancements of speech content from interlaced audio input. The system 400 includes a group of human speakers 402 that output audio output. The audio output is received for training each distinct signal using enhancements, as seen in block 404. The system can train enhancements of the audio input signal based on diarization, as seen in block 406.

[0101] The system comprises layering an audio input signal using diarization, as seen in block 410. The system comprises performing sounds, e.g., environmental stimuli, and classifying the diarized audio input signal, as seen in block 412. The system comprises setting individual and group augmentations based on the environmental stimuli, as seen in block 414. The system comprises generating an audio output, as seen in block 416. The system comprises generating a text output using the audio output 416 based on the augmentations and the environmental stimuli, as seen in block 418.

[0102] In one example, the system may predict a speaker's future signal using a predictive speaker signal technique or method / system to predict the speaker's signal, in one example by predicting how the speaker signal will change based on external noise, as seen in block 450. Such prediction is not the focus of this disclosure.

[0103] 1 and 2, the computer may be part of a remote computer or server, such as remote server 1100 (FIG. 8). In another example, computer 72 may be part of control system 70 and provide performance of the functionality of the present disclosure. In another embodiment, computer 22 may be part of mobile device 20 and provide performance of the functionality of the present disclosure. In yet another embodiment, performance of some of the functionality of the present disclosure may be shared between the control system computer and the mobile device computer, e.g., the control system serves as a back end for one or more programs embodying the present disclosure, and the mobile device computer serves as a front end for the one or more programs.

[0104] The computer may be part of the mobile device or a remote computer in communication with the mobile device. In another example, the mobile device and remote computer may work in cooperation to implement the methods of the present disclosure using stored program code or instructions for performing the method features described herein. In one example, the mobile device 20 may include a computer 22 having a processor 15 and a storage medium 34 storing an application 40. The application may incorporate program instructions for performing the features of the present disclosure using the processor 15. In another example, the mobile device 20 application 40 may have executable program instructions for a front end of a software application that incorporates the method features of the present disclosure into the program instructions, while one or more back end programs 74 of the software application stored on the computer 72 of the control system 70 communicate with the mobile device computer to perform other method features. The control system 70 and the mobile device 20 may communicate using a communications network 45, such as the Internet.

[0105] Accordingly, the method 100 according to one embodiment of the present disclosure may be embodied in one or more computer programs or applications 40 stored on an electronic storage medium 34 and executable by a processor 15 as part of a computer on a mobile device 20. For example, a human speaker or user 14 has a device 20, which can communicate with a control system 70. Other users (not shown) may have similar devices and similarly communicate with the control system. The application may be stored, in whole or in part, on a computer or computers on the mobile device and in a control system communicating with the device using a communications network 45, such as the Internet. It is contemplated that the application may access all or a portion of the program instructions to implement the method of the present disclosure. The program or application may communicate with a remote computer system via a communications network 45 (e.g., the Internet) to access data and collaborate with programs stored on the remote computer system. Such interactions and mechanisms are described in further detail herein and reference is made to components of a computer system, such as computer-readable storage media, which are shown in one embodiment in FIG. 8 and, in this regard, are described in more detail with reference to one or more computer systems 1010.

[0106] Thus, in one example, control system 70 is in communication with device 20, which may have application 40. Device 20 communicates with control system 70 using communication network 45.

[0107] In another example, control system 70 may have a front-end computer belonging to one or more users, such as device 20, and a back-end computer embodied as a control system.

[0108] 1, device 20 may include computer 22, computer-readable storage medium 34, and operating system, and / or program, and / or software application 40, which may include program instructions executable using processor 15. These features are illustrated herein in FIG. 1 and with reference to one or more computer systems 1010, which may include one or more general computer components, and also in one embodiment of a computer system illustrated in FIG. 8.

[0109] The methods according to the present disclosure may include a computer as part of a control system for implementing features of the methods according to the present disclosure. In another example, a computer as part of a control system may function in conjunction with a mobile device computer for implementing features of the methods according to the present disclosure. In another example, a computer for implementing features of the methods may be part of a mobile device, thereby implementing the methods locally.

[0110] It will be understood that the features illustrated in Figures 6 and 7 are functional representations of features of the present disclosure, and such features are shown in embodiments of the systems and methods of the present disclosure for illustrative purposes to clarify the functionality of the features of the present disclosure.

[0111] Specifically, with respect to control system 70, devices 20 of one or more users 14 may be in communication with control system 70 via communications network 50. In the embodiment of the control system shown in FIG. 1 , control system 70 includes a computer 72 having a database 76 and one or more programs 74 stored on a computer-readable storage medium 73. In the embodiment of the present disclosure shown in FIG. 1 , devices 20 communicate with control system 70 and one or more programs 74 stored on computer-readable storage medium 73. Control system includes computer 72 having a processor 75, which also has access to database 76.

[0112] The control system 70 may include a storage medium 80 for holding a user registration 82 and a user's device for analysis of audio input. Such registration may include a user profile 83, which may include user data provided by the user for account registration and setup. In one embodiment, methods and systems incorporating the present disclosure include a control system (commonly referred to as a backend) in combination with and working with a front end of the method and system, which may be an application 40. In one example, the application 40 is stored on a device, e.g., device 20, and can access data and additional programs in the application's backend, e.g., the control system 70.

[0113] A control system may also be part of a software application implementation, or may represent a software application having a front-end user portion and a back-end portion that provide functionality, or both. In one embodiment, a method and system incorporating the present disclosure includes a control system (which may be generally referred to as a back-end of a software application incorporating a portion of a method and system of an embodiment of the present disclosure) in combination with and working in cooperation with a front-end of a software application incorporating another portion of the method and system of the present disclosure on a device, such as in the example shown in FIG. 1 of device 20 having application 40. Application 40 is stored on device 20 and can access data and additional programs in the back-end of the application, for example, in program 74 stored in control system 70.

[0114] The program 74 may include, in whole or in part, a series of executable steps for implementing the method of the present disclosure. A program incorporating the method may be stored, in whole or in part, on a computer-readable storage medium on the control system, or, in whole or in part, on the device 20. It is contemplated that the control system 70 may not only store user profiles, but may also, in one embodiment, interact with a website for viewing on the device's display, or in another example, the Internet, to receive user input related to the method and system of the present disclosure. While FIG. 1 shows one or more profiles 83, it is understood that the method may involve multiple profiles, users, registrations, etc. It is contemplated that multiple users or groups of users may use the control system to register and provide profiles for use by the method and system of the present disclosure.

[0115] With respect to the collection of data pertaining to the present disclosure, such uploading or creation of a profile is voluntary by one or more users and is therefore initiated by and with user authorization, thereby allowing users to opt in to establishing an account with a profile in accordance with the present disclosure. Similarly, data received by the system or entered or received as input is voluntary by one or more users and is therefore initiated by and with user authorization, thereby allowing users to opt in to entering data in accordance with the present disclosure. Such user authorization also includes the user's option to cancel such profile or account, and / or data entry, thereby opting out of communication and data capture at the user's discretion. It is further understood that any stored or collected data is intended to be securely stored and unavailable without user authorization, and not available to the public or unauthorized users, or both. It is understood that such stored data is deleted upon the user's request, and in a secure manner. It is also understood that any use of such stored data, in accordance with this disclosure, is only with the authorization and consent of the user.

[0116] In one or more embodiments of the present invention, a user may opt in or register with a control system, voluntarily providing data or information, or both, with the user's consent and authorization, in a process in which the data is stored and used in one or more methods of the present disclosure. The user may also register one or more user electronic devices for use with one or more methods and systems according to the present disclosure. As part of registration, the user may also identify and authorize access to one or more activities or other systems (e.g., audio and / or video systems). Such opt-in to registration and authorization to data collection and / or storage is voluntary, and the user may request deletion of data (including profile and / or profile data), unsubscription, or opt-out of any registration, or a combination thereof. Such opt-out is understood to include disposal of the data in its entirety in a secure manner.

[0117] In one example, artificial intelligence (AI) may be used, in whole or in part, to learn models to determine the expansion parameters.

[0118] In another example, the control system 70 may be, in whole or in part, an artificial intelligence (AI) system. For example, the control system may be one or more components of an AI system.

[0119] It is also understood that the method 100 according to an embodiment of the present disclosure can be incorporated into an (artificial intelligence) AI device that can communicate with the respective AI system and the respective AI system platform. Accordingly, such a program or application incorporating the method of the present disclosure can be part of the AI ​​system, as described above. In one embodiment according to the present invention, it is contemplated that the control system can communicate with the AI ​​system, or in another example, can be part of the AI ​​system. The control system can also represent a software application having a front-end user portion and a back-end portion that provides functionality, which in one or more examples can interact with, encompass, or be part of a larger system, such as an AI system. In one example, an AI device can be associated with an AI system that can be, in whole or in part, a control system or a content distribution system, or both, and that can be remote from the AI ​​device. Such an AI system can be represented by one or more servers storing a program on a computer-readable medium that can communicate with one or more AI devices. The AI ​​system can communicate with the control system, and in one or more embodiments, the control system can be, in whole or in part, an AI system, or vice versa.

[0120] As discussed herein, it is understood that downloads or downloadable data may be initiated using voice commands or using a mouse, touch screen, etc. In such examples, a mobile device may be user-initiated, or an AI device may be enabled for use with user consent and permission. Other examples of AI devices include devices that include a microphone, a speaker, and have access to a cellular or mobile network, a communications network, or the Internet, such as a vehicle having a computer and cellular or satellite communications, or, in another example, an IoT (Internet of Things) device such as a home appliance with cellular network or Internet access.

[0121] As used herein, a set is understood to be a collection of distinct objects or elements. The objects or elements that make up a set can be anything, such as numbers, letters of the alphabet, other sets, etc. It is further understood that a set can be a single element, such as a single thing or number, in other words, a set of a single element.

[0122] Referring to FIG. 8, one embodiment of a system or computer environment 1000 according to the present disclosure includes a computer system 1010, shown in the form of a general-purpose computing device. The method 100 may be embodied, for example, in a program 1060 including program instructions embodied on a computer-readable storage device or computer-readable storage medium, e.g., generally referred to as computer memory 1030 and more specifically as computer-readable storage medium 1050. Such memory or computer-readable storage medium, or both, may include non-volatile memory or non-volatile storage, also known and referred to as non-transitory computer-readable storage medium or non-transitory computer-readable storage medium. For example, such non-volatile memory may also be a disk storage device, including one or more hard drives. For example, the memory 1030 may include a storage medium 1034, such as RAM (random access memory) or ROM (read-only memory), and a cache memory 1038. The program 1060 is executable by the processor 1020 of the computer system 1010 (to execute program steps, code, or program code). Additional data storage may be embodied as a database 1110 containing data 1114. The computer system 1010 and the program 1060 are general depictions of a computer and program that may be local to a user or provided as a remote service (e.g., as a cloud-based service), and in further examples may be provided using a website accessible using the communications network 1200 (e.g., interacting with a network, the Internet, or a cloud service). It is understood that, herein, the computer system 1010 also generally represents a computing device or computer included in a device, such as a laptop computer or desktop computer, or one or more servers, either alone or as part of a data center.The computer system may include a network adapter / interface 1026 and an input / output (I / O) interface 1022. The I / O interface 1022 allows for the input and output of data with external devices 1074 that may be connected to the computer system. The network adapter / interface 1026 may provide for communication between the computer system and a network, generally depicted as communications network 1200.

[0123] The computer 1010 may be described in the general context of computer system-executable instructions, such as program modules, being executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Method steps and system components and techniques may be embodied in modules of the program 1060 for performing the respective method steps and system tasks. Modules are generally represented in the diagram as program modules 1064. The program 1060 and program modules 1064 may execute specific program procedures, routines, subroutines, instructions, or code.

[0124] The methods of the present disclosure may be executed locally on a device, such as a mobile device, or may be executed as a service on a server 1100, which may be remote, for example, and accessible using a communications network 1200. The program or executable instructions may be provided as a service by a provider. The computer 1010 may be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices linked through the communications network 1200. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.

[0125] More specifically, system or computer environment 1000 includes a computer system 1010 shown in the form of a general-purpose computing device with exemplary peripheral devices. Components of computer system 1010 may include, but are not limited to, one or more processors or processing units 1020, a system memory 1030, and a bus 1014 that couples various system components including the system memory 1030 to the processor 1020.

[0126] Bus 1014 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures, including, by way of example only, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0127] The computer 1010 may include a variety of computer-readable media. Such media may be any available media accessible by the computer 1010 (e.g., a computer system or server) and may include both volatile and nonvolatile media, as well as removable and non-removable media. The computer memory 1030 may include additional computer-readable media in the form of volatile memory, such as random access memory (RAM) 1034 and / or cache memory 1038. The computer 1010 may further include other removable / non-removable, volatile / non-volatile computer storage media, such as a portable computer-readable storage medium 1072. In one embodiment, the computer-readable storage medium 1050 may be provided for reading from and writing to non-removable, non-volatile magnetic media. The computer-readable storage medium 1050 may be embodied as, for example, a hard drive. Additional memory and data storage may be provided, for example, as a storage system 1110 (e.g., a database) for storing data 1114 and communicating with the processing unit 1020. The database may be stored on or part of server 1100. Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical media, may be provided, each of which may be connected to bus 1014 by one or more data media interfaces. As further depicted and described below, memory 1030 may include at least one program product, which may include one or more program modules configured to perform the functions of embodiments of the present invention.

[0128] The methods described in this disclosure may be embodied in one or more computer programs, generally referred to as programs 1060, for example, and stored in memory 1030 within computer-readable storage medium 1050. Program 1060 may include program modules 1064. Program modules 1064 generally may perform the functionality and / or methodology of embodiments of the invention as described herein. The one or more programs 1060 are stored in memory 1030 and executable on processing unit 1020. By way of example, memory 1030 may store an operating system 1052, one or more application programs 1054, other program modules, and program data on computer-readable storage medium 1050. It will be understood that program 1060, and the operating system 1052 and application programs 1054 stored on computer-readable storage medium 1050, are similarly executable on processing unit 1020. It is also understood that applications 1054 and programs 1060 are shown generally and may include or be part of one or more applications and programs described in this disclosure, or vice versa, i.e., applications 1054 and programs 1060 may be part of or be part of one or more applications or programs described in this disclosure. It is also understood that control system 70 in communication with a computer system may include all or part of computer system 1010 and its components, or the control system may communicate with all or part of computer system 1010 and its components as a remote computer system, or both, to achieve the control system functionality described in this disclosure. Control system functionality may include, for example, storing, processing, and executing software instructions to perform the functions of the present disclosure.It is also understood that one or more computers or computer systems shown in FIG. 1 may similarly include all or some of the computer system 1010 and its components, or one or more computers may be in communication with all or some of the computer system 1010 and its components as remote computer systems, or both, to accomplish the computer functions described in this disclosure.

[0129] In one embodiment according to the present disclosure, one or more programs may be stored on one or more computer-readable storage media such that the programs are embodied and / or encoded on the computer-readable storage media. In one example, the stored program may include program instructions for execution by a processor or a computer system having a processor to perform a method or to cause the computer system to perform one or more functions. For example, in one embodiment according to the present disclosure, a program embodying a method is embodied or encoded on a computer-readable storage medium, which includes a non-transitory or non-transitory computer-readable storage medium and is defined as a non-transitory or non-transitory computer-readable storage medium. Thus, an embodiment or example according to the present disclosure of a computer-readable storage medium does not include a signal, and the embodiment may include one or more non-transitory or non-transitory computer-readable storage media. Thus, in one example, the program may be recorded on a computer-readable storage medium and be structurally and functionally interrelated with the medium.

[0130] The computer 1010 may communicate with one or more external devices 1074, such as a keyboard, pointing device, display 1080, one or more devices that allow a user to interact with the computer 1010, or any device that allows the computer 1010 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.), or a combination thereof. Such communication may occur via an input / output (I / O) interface 1022. Still further, the computer 1010 may communicate with one or more networks 1200, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter / interface 1026. As shown, the network adapter 1026 is in communication with other components of the computer 1010 via a bus 1014. Although not shown, it should be understood that other hardware and / or software components may be used in conjunction with the computer 1010. Examples include, but are not limited to, microcode, device drivers 1024, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.

[0131] It is understood that a computer, or a program running on computer 1010, may communicate with a server embodied as server 1100 via one or more communications networks embodied as communications network 1200. Communications network 1200 may include transmission media and network links, including, for example, wireless, wired, or fiber optic, as well as routers, firewalls, switches, and gateway computers. A communications network may include connections such as wires, wireless communications links, or fiber optic cables. A communications network may represent a worldwide collection of networks and gateways, such as the Internet, communicating with each other using various protocols, such as Lightweight Directory Access Protocol (LDAP), Transport Control Protocol / Internet Protocol (TCP / IP), Hypertext Transport Protocol (HTTP), Wireless Application Protocol (WAP), etc. A network may include several different types of networks, such as, for example, an intranet, a local area network (LAN), or a wide area network (WAN).

[0132] In one example, a computer may use a network that may access websites on the web (World Wide Web) using the Internet. In one embodiment, a computer 1010, including a mobile device, may use a communication system or network 1200, which may include the Internet, or a public switched telephone network (PSTN), e.g., a cellular network. The PSTN may include telephone lines, fiber optic cables, microwave transmission links, cellular networks, and communications satellites. The Internet may facilitate numerous searching and texting techniques, such as submitting a query to a search engine via text message (SMS), multimedia messaging service (MMS) (related to SMS), email, or a web browser using, for example, a mobile phone or laptop computer. The search engine may obtain search results, i.e., links to websites, documents, or other downloadable data corresponding to the query, and may also provide the search results to a user via the device, e.g., as a search result web page.

[0133] 9, an exemplary system 1500 for use with embodiments of the present disclosure is illustrated. The system 1500 includes multiple components and elements connected via a system bus 1504 (also referred to as a bus). At least one processor (CPU) 1510 is connected to the other components via the system bus 1504. A cache 1570, a read-only memory (ROM) 1512, a random access memory (RAM) 1514, an input / output (I / O) adapter 1520, an audio adapter 1530, a network adapter 1540, a user interface adapter 1552, a display adapter 1560, and a display device 1562 are also operatively coupled to the system bus 1504 of the system 1500.

[0134] One or more storage devices 1522 are operatively coupled to the system bus 1504 by an I / O adapter 1520. The storage device 1522 may be, for example, a disk storage device (e.g., a magnetic or optical disk storage device), a solid-state magnetic device, or the like. The storage device 1522 may be the same type of storage device or a different type of storage device. The storage device may include, for example, but is not limited to, a hard drive or flash memory and may be used to store one or more programs 1524 or applications 1526. The programs and applications are shown as general components executable using the processor 1510. The programs 1524 or applications 1526, or both, may comprise all or part of a program or application discussed in this disclosure, and vice versa, i.e., the programs 1524 and applications 1526 may be part of other applications or programs discussed in this disclosure. The storage devices may communicate with the control system 70, which has various functions described in this disclosure.

[0135] Speakers 1532 are operatively coupled to the system bus 1504 by an audio adapter 1530. Transceiver 1542 is operatively coupled to the system bus 1504 by a network adapter 1540. Display 1562 is operatively coupled to the system bus 1504 by a display adapter 1560.

[0136] One or more user input devices 1550 are operatively coupled to the system bus 1504 by a user interface adapter 1552. The user input device 1550 may be, for example, a keyboard, a mouse, a keypad, an image capture device, a motion sensing device, a microphone, a device incorporating the functionality of at least two of the foregoing devices, or the like. Other types of input devices may be used while maintaining the spirit of the invention. The user input devices 1550 may be the same type of user input device or different types of user input devices. The user input devices 1550 are used to input and output information to and from the system 1500.

[0137] The present invention may be a system, method, or computer program product, or combination thereof, integrated at any possible level of technical detail. The computer program product may have a computer-readable storage medium (or multiple computer-readable storage media) containing computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0138] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as ridge structures in grooves on which instructions are recorded, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over wires.

[0139] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device, or may be downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may comprise copper cables, optical fibers, wireless networks, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to a computer-readable storage medium within the computing / processing device for storage.

[0140] The computer-readable program instructions for carrying out the operations of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk® or C++, and procedural programming languages ​​such as the “C” programming language or similar programming languages. The computer-readable program instructions may run entirely on the user's computer, as a standalone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions to personalize the electronic circuitry by utilizing state information of the computer readable program instructions to perform aspects of the present invention.

[0141] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations or block diagrams or combinations thereof, and combinations of blocks in the flowchart illustrations or block diagrams or combinations thereof, can be implemented by computer-readable program instructions.

[0142] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the processor of the computer or other programmable data processing apparatus, form means for implementing the functions / acts specified in one or more blocks of the flowcharts or block diagrams, or both. These computer-readable program instructions may also be stored on a computer-readable storage medium capable of instructing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, whereby the computer-readable storage medium having stored thereon instructions comprises an article of manufacture including instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts or block diagrams, or both.

[0143] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device and cause the computer, other programmable apparatus, or other device to execute a series of operating procedures to create a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the functions / operations specified in one or more blocks of the flowcharts or block diagrams, or both.

[0144] The flowcharts and block diagrams in the figures of this disclosure illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions, that implements a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be implemented as a single procedure, executed concurrently, substantially concurrently, partially, or fully in a time-overlapping manner, or the blocks may even be executed in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks within the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.

[0145] Although this disclosure includes detailed descriptions of cloud computing, it should be understood that practice of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0146] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0147] The features are as follows: On-Demand Self-Service: Cloud consumers can unilaterally provision computing capacity, such as server time and network storage, automatically as needed, without requiring human interaction with the service provider. Wide network access: The capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (eg, cell phones, laptops, and PDAs). Resource Pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated according to demand. Consumers generally have no control or knowledge over the exact location of the provided resources, although there is a sense of location independence in that it may be possible to specify location at a higher level of abstraction (e.g., country, state, or data center). Rapid Elasticity: This capacity can be rapidly and elastically provisioned, sometimes automatically, to quickly scale out and rapidly released to quickly scale in. To the consumer, the capacity available for provisioning often appears unlimited, and can be purchased at any time and in any quantity. Measured Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services utilized.

[0148] The service model is as follows: Software as a Service (SaaS): The consumer is offered the ability to use a provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings. Platform as a Service (PaaS): The capability created or offered by a consumer to deploy consumer-created or acquired applications, written using programming languages ​​and tools supported by the provider, on a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does control the deployed applications and, in some cases, the application-hosting environment configuration. Infrastructure as a Service (IaaS): The ability offered to consumers is to provision processing, storage, network, and other underlying computing resources onto which they can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but rather controls the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0149] The deployment model is as follows: Private Cloud: This cloud infrastructure operates solely for an organization. It may be managed by the organization or a third party and may exist on-premise or off-premise. Community Cloud: This cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies and compliance considerations). It can be managed by the organization or a third party and can reside on-premises or off-premises. Public Cloud: This cloud infrastructure is made available to the general public or large industry groups and is owned by an organization that sells cloud services. Hybrid Cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain distinct entities but are tied together by standardized or proprietary technologies that allow data and application portability (e.g., cloud bursting for load balancing between clouds).

[0150] Cloud computing environments are service-oriented, emphasizing statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0151] 10 , an exemplary cloud computing environment 2050 is shown. As shown, the cloud computing environment 2050 includes one or more cloud computing nodes 2010 with which a local computing device used by a cloud consumer may communicate, such as, for example, a personal digital assistant (PDA) or cellular phone 2054A, a desktop computer 2054B, a laptop computer 2054C, or an automobile computer system 2054N, or a combination thereof. The nodes 2010 may communicate with each other. The nodes 2010 may be physically or virtually grouped in one or more networks (not shown), such as a private cloud, a community cloud, a public cloud, or a hybrid cloud, or a combination thereof, as described hereinabove. This enables the cloud computing environment 2050 to provide infrastructure, a platform, or software, or a combination thereof, as a service without the cloud consumer having to maintain resources on the local computing device. It should be understood that the types of computing devices 2054A-N shown in FIG. 10 are intended as examples only, and that computing node 2010 and cloud computing environment 2050 can communicate with any type of computerized device through any type of network or network-addressable connection (e.g., using a web browser) or combination thereof.

[0152] Referring now to Figure 11, a set of functional abstraction layers provided by cloud computing environment 2050 (Figure 10) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 11 are intended to be examples only, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0153] Hardware and software layer 2060 includes hardware and software components. Examples of hardware components include mainframe 2061, RISC (reduced instruction set computer) architecture-based servers 2062, servers 2063, blade servers 2064, storage devices 2065, and networks and networking components 2066. In some embodiments, software components include network application server software 2067 and database software 2068.

[0154] The virtualization layer 2070 provides an abstraction layer from which the following examples of virtual entities may be provided: virtual servers 2071, virtual storage 2072, virtual networks including virtual private networks 2073, virtual applications and operating systems 2074, and virtual clients 2075.

[0155] In one example, management layer 2080 may provide the functionality described below. Resource provisioning 2081 provides dynamic procurement of computing and other resources used to execute tasks within the cloud computing environment. Metering and pricing 2082 tracks costs as resources are utilized within the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 2083 provides access to the cloud computing environment to consumers and system administrators. Service level management 2084 provides cloud computing resource allocation and management to ensure required service levels are met. Service level agreement (SLA) planning and fulfillment 2085 provides proactive coordination and procurement of cloud computing resources in anticipation of future requirements according to SLAs.

[0156] Workload tier 2090 provides examples of functionality for which a cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this tier include mapping and navigation 2091, software development and lifecycle management 2092, virtual classroom instructional delivery 2093, data analytics processing 2094, transaction processing 2095, and specifically speech recognition from audio input using augmentation of speech content from one or more humans (or human speakers) from interlaced audio input 2096.

[0157] The description of various embodiments of the present invention has been presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Similarly, examples of features or functionality of embodiments of the present disclosure described herein, whether used to describe a particular embodiment or listed as an example, are not intended to limit the embodiments of the present disclosure described herein or to limit the disclosure to the examples described herein. Such examples are intended to be examples or illustrative and not exhaustive. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best explain the principles, practical applications, or technical improvements to technology found in the marketplace of the embodiments, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A computer-implemented method for speech recognition using speech content enhancement from audio input, comprising: initiating a training model to determine extension parameters for each of a plurality of audible sounds of speech content from a plurality of speakers received at the computer as audio input; using the learned model to determine a change in each of a plurality of independent sounds in response to an audio stimulus, wherein the plurality of independent sounds are derived from the audio input; applying the expansion parameters based on the variation of each of the plurality of independent sounds, respectively; constructing a voiceprint for each of the plurality of speakers based on the plurality of independent sounds and the extension parameters, respectively; and attributing the audio content to each of the plurality of speakers based at least in part on the voiceprint and the plurality of independent sounds, respectively. A computer-implemented method comprising:

2. generating text from the attributed audio content; The computer-implemented method of claim 1 further comprising:

3. using said computer to display said text on a screen or monitor in communication with said computer or device, or both; The computer-implemented method of claim 2 further comprising:

4. transmitting said text via an electronic communications system to another computer or device, or both, for display on a screen or monitor in communication with said other computer or device, or both. The computer-implemented method of claim 2 further comprising:

5. 5. The computer-implemented method of claim 1, wherein the audio input comprises the plurality of audible tones, the audio input being received at the computer, and the plurality of audible tones including speech content from the plurality of speakers.

6. augmenting the audio input using the computer, wherein the augmenting comprises separating the sounds in the audio input; The computer-implemented method of claim 1 , further comprising:

7. refining the audio input for each of the plurality of speakers using diarization. The computer-implemented method of claim 1 , further comprising:

8. 8. The computer-implemented method of claim 7, wherein the learning model comprises a convolutional neural network (CNN) for receiving the plurality of independent sounds and using the diarization to determine the change in each of the plurality of independent sounds in response to the audio stimulus.

9. layering the sounds in the refined audio input into a plurality of separate sounds using the diarization of the audio input. The computer-implemented method of claim 7 further comprising:

10. receiving, at the computer, the audio input having the plurality of audible sounds, wherein the plurality of audible sounds includes speech content from the plurality of speakers; augmenting the audio input using the computer, wherein the augmenting comprises separating the sounds in the audio input; refining the audio input for each of the plurality of speakers using diarization; and layering the sounds in the refined audio input into a plurality of separate sounds using the diarization of the audio input. The computer-implemented method of claim 1 , further comprising:

11. The computer-implemented method of claim 10 , wherein separating the sounds in the audio input comprises distinguishing environmental or background sounds from speech from one speaker of the multiple speakers.

12. 11. The computer-implemented method of claim 10, wherein refining the audio input for each of the plurality of speakers using the diarization comprises partitioning the audio input into homogeneous segments for speaker identification.

13. A system for speech recognition using speech content enhancement from audio input, comprising:

1. A computer system having a computer processor, a computer readable storage medium, and program instructions stored on the computer readable storage medium that are executable by the computer processor, the program instructions causing the computer system to perform the following functions: Initiating a training model to determine enhancement parameters for each of a plurality of audible sounds of speech content from a plurality of speakers received at the computer as audio input; using the learned model to determine a change in each of a plurality of distinct sounds in response to an audio stimulus, wherein the plurality of distinct sounds are derived from the audio input; applying the expansion parameters based on the variation of each of the plurality of independent sounds, respectively; constructing a voiceprint for each of the plurality of speakers based on the plurality of independent sounds and the extension parameters, respectively; and attributing the audio content to each of the plurality of speakers based at least in part on the voiceprint and the plurality of independent sounds, respectively. A system that executes the following.

14. generating text from said attributed audio content; The system of claim 13 further comprising:

15. Using said computer to display said text on a screen or monitor in communication with said computer or device, or both. The system of claim 14 further comprising:

16. Transmitting said text via an electronic communications system to another computer or device, or both, for display on a screen or monitor in communication with said other computer or device, or both. The system of claim 14 further comprising:

17. 17. The system of claim 13, wherein the audio input comprises the plurality of audible tones, the audio input being received at the computer, and the plurality of audible tones including speech content from the plurality of speakers.

18. augmenting the audio input using the computer, wherein the augmenting comprises separating the sounds in the audio input.

18. The system of claim 13, further comprising:

19. Refining the audio input for each of the plurality of speakers using diarization.

19. The system of claim 13, further comprising:

20. 20. The system of claim 19, wherein the learning model comprises a convolutional neural network (CNN) for receiving the plurality of independent sounds and using the diarization to determine a change in each of the plurality of independent sounds in response to an audio stimulus.

21. using the diarization of the audio input to layer the sounds in the refined audio input into multiple separate sounds; 20. The system of claim 19, further comprising:

22. A computer program for speech recognition using augmentation of speech content from audio input, the computer program comprising: initiating a training model to determine enhancement parameters for each of a plurality of audible sounds of speech content from a plurality of speakers received at the computer as audio input; using the learned model to determine a change in each of a plurality of distinct sounds in response to an audio stimulus, wherein the plurality of distinct sounds are derived from the audio input; applying the expansion parameters based on the variation of each of the plurality of independent sounds, respectively; constructing a voiceprint for each of the plurality of speakers based on the plurality of independent sounds and the extension parameters, respectively; and attributing the audio content to each of the plurality of speakers based at least in part on the voiceprint and the plurality of independent sounds, respectively. A computer program to be executed.

23. The computer, generating text from the attributed audio content; 23. The computer program of claim 22, further comprising:

24. The computer, using said computer to display said text on a screen or monitor in communication with said computer or device, or both; 24. The computer program of claim 23, further comprising:

25. The computer, transmitting said text via an electronic communications system to another computer or device, or both, for display on a screen or monitor in communication with said other computer or device, or both; 24. The computer program of claim 23, further comprising:

Citation Information

Patent Citations

  • Association device, association method, and computer program

    JP2009237353A

  • Full-band scalable audio codec

    JP2012032803A

  • The number of speakers estimation device, the number of speakers estimation method, and program

    JP2018063313A

  • Voice activity detection

    JP2018517928A

  • Speaker diarization using speaker embedding(s) and trained generative model

    WO2020068056A1