Speech recognition using data analysis and dilation of speech content from separated audio input

By employing CNNs with dilation parameters and machine learning to analyze and adjust for external noise, the method enhances speech recognition accuracy in noisy or multi-speaker environments, addressing the challenges of altered and atypical sounds.

JP7776237B2Active Publication Date: 2025-11-26INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023515727
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-09
Filing Date
2021-08-24
Publication Date
2025-11-26
Estimated Expiration
2041-08-24

AI Technical Summary

Technical Problem

Current speech recognition technologies struggle with accurately converting speech content from audio inputs that include altered or atypical sounds, such as prolonged sounds, background noise, overlapping speakers, or intentional modifications, leading to increased word error rates.

Method used

The use of convolutional neural networks (CNNs) with dilation parameters to separate and analyze speech content from multiple speakers, predicting modifications based on external noise, and adjusting dilation to reduce word error rates through machine learning and environmental noise consideration.

Benefits of technology

Reduces word error rates by effectively recognizing and converting speech content from audio inputs with modified or atypical sounds, improving the accuracy of speech-to-text conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007776237000001
    Figure 0007776237000001
  • Figure 0007776237000002
    Figure 0007776237000002
  • Figure 0007776237000003
    Figure 0007776237000003
Patent Text Reader

Abstract

Speech recognition using data analysis and dilation of speech content from a separated audio input. The present disclosure includes using dilation of speech content from a separated audio input for speech recognition. An audio input from a speaker and a predicted modification of the audio input based on external noise are received at a CNN (Convolutional Neural Network). In the CNN, diarization is applied to the audio input to predict how the dilation of the speech content from the speaker will modify the audio input and generate a CNN output. A resulting dilation is determined from the CNN output. A word error rate is determined for the dilated CNN output to determine the accuracy of the speech-to-text output. Tuning parameters are set to modify the range of dilation based on the word error rate, and the resulting dilation of the CNN output is adjusted based on the tuned parameters to reduce the word error rate.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to techniques for computer-based speech recognition of speech content from an audio input. More specifically, the audio input includes speech content separated from the audio input. [Background technology]

[0002] The technique can be used to convert human speech into text using a computer. Human speech can include, for example, spoken words, songs, alone or in groups. During human speech, converting the speech output or converting the speech output signal into text can be difficult. For example, speech recognition and conversion can be difficult when sounds are altered or less typical than the typical pronunciation of a word. For example, sounds may be prolonged or mixed with one or more other noises. In one example, there may be background noise when a speaker is speaking. In another example, a group of speakers may speak and there may be overlapping speakers. In another example, background noise may occur when one or more speakers are speaking. In another example, a speaker may unintentionally or intentionally alter the typical pronunciation of one or more words for emphasis, as part of an illegal or atypical speech pattern, or as part of an accent. Such altered or atypical sounds, or a combination thereof, when a speaker speaks can result in speech that is difficult for speech identification and speech-to-text conversion. Summary of the Invention

[0003] This disclosure recognizes shortcomings and problems associated with current techniques for speech recognition that use dilation of speech content from interlaced audio input.

[0004] The present invention can analyze speech content from interlaced audio input for speech recognition of each of multiple speakers and provide speech content to text conversion. For example, the challenges of speech recognition and conversion when sounds are modified or less typical than typical word pronunciations can be overcome using the present invention when the speech content includes modified sounds or atypical sounds, or a combination thereof, from speakers for speech-to-text conversion.

[0005] For example, one problem may arise when an artist sings a song and some of the words are altered or modified in a manner that follows harmony rather than common pronunciation. In another example, in a noisy environment, sound waves and sound mixing can increase the word error rate during conversion. For example, in a loud event, a loud crowd, or a sporting event, sounds may occlude the speech code. The present invention includes speech recognition using dilation of the audio signal, audio input, to increase the space between samples or audio samples before attempting to recognize words or analyzing the audio content to recognize one or more words. In one example according to the present invention, convolutional neural networks (CNNs) with different dilation parameters can be trained and applied to these problems. Furthermore, expected environmental noise and voice type may dictate which dilation to use. Additionally, in another example, each speaker can be assigned a dilation parameter through machine learning. In a group setting of a conversation or song, dilation parameters can be weighted together for each group based on the amplitude of each speaker.

[0006] The present invention involves predicting the audio signal or speech content of each human speaker into the future. In one example, the present invention can predict how the speaker's audio signal or speech content will change based on external noise. The output can be input to a CNN without dilation, and based on future trends, it can be predicted how the dilation of the speech content will change. The resulting dilation can be determined and applied to speech-to-text conversion.

[0007] In one aspect, according to the present invention, a computer-implemented method for speech recognition uses dilation of speech content from a separated audio input, comprising receiving an audio input in a convolutional neural network (CNN) and receiving predicted modifications to the audio input based on external noise, the audio input including speech content from a speaker. The method comprises applying diarization to the audio input in the CNN to predict how the dilation of the speech content from the speaker will modify the audio input and generating a CNN output. A resulting dilation from the CNN output is determined, the resulting dilation of the CNN output including separating the sounds of the audio input. A word error rate for the dilated CNN output is determined to determine the accuracy of the speech-to-text output. Tuning parameters are set to modify the range of dilation based on the word error rate. The method comprises adjusting the resulting dilation of the CNN output based on the tuning parameters to reduce the word error rate.

[0008] One advantage of the present invention includes reducing word error rates when converting speech content from an audio input separated using speech recognition to convert speech from the audio input to text using a method according to the present invention.

[0009] In a related aspect, the method further comprises identifying speech content from the speaker from the audio input based on the training corrections applied to the adjusted resulting dilation for the speaker.

[0010] In a related aspect, the method includes generating text from the identified audio content.

[0011] In a related aspect, the audio input and the predicted modifications are received without dilation of the audio content.

[0012] In a related aspect, a method comprises adjusting a resulting dilation of an audio input based on a word error rate using a grid search to reduce the word error rate.

[0013] In a related aspect, a method includes, in a computer, receiving predicted audio input for a speaker, the predicted audio input including speech content for the speaker; generating environmental stimulus audio input for the predicted audio input; and predicting changes in the audio input for the speaker based on the environmental stimulus audio input.

[0014] In a related aspect, the method further comprises sharing the adjusted resulting dilation of the expected audio input with a social network; generating a learning correction from the sharing of the adjusted resulting dilation of the expected audio input; applying the learning correction to the adjusted resulting dilation for the speaker; identifying speech content from the speaker from the audio input based on the learning correction applied to the adjusted resulting dilation for the speaker; and generating text from the identified speech content.

[0015] In a related aspect, the method further comprises sharing the adjusted resulting dilation of the expected audio input with a social network; generating a learning modification from the sharing of the adjusted resulting dilation of the expected audio input; and applying the learning modification to the adjusted resulting dilation for the speaker.

[0016] In a related aspect, the method further comprises identifying speech content from the speaker from the audio input based on the training corrections applied to the adjusted resulting dilation for the speaker.

[0017] In a related aspect, the method further comprises generating text from the identified audio content.

[0018] In a related aspect, the method further comprises receiving, at the CNN, dilation parameters for speech content of one of the multiple speakers, the dilation parameters being derived from audio input from the multiple speakers.

[0019] In a related aspect, the audio input from multiple speakers is an interlaced audio input.

[0020] In another aspect according to the present invention, a system for speech recognition employs dilation of speech content from a separated audio input, including a computer system, the computer system including a computer processor, a computer-readable storage medium, and program instructions stored on the computer-readable storage medium executable by the processor, causing the computer system to perform the following functions: receiving an audio input at a convolutional neural network (CNN) and receiving predicted modifications to the audio input based on external noise, the audio input having speech content from a speaker; applying diarization to the audio input at the CNN to predict how dilation of the speech content from the speaker will modify the audio input to generate a CNN output; determining a resulting dilation from the CNN output, the resulting dilation of the CNN output including separating sounds of the audio input; determining a word error rate for the dilated CNN output to determine accuracy for the speech-to-text output; setting tuning parameters to modify the extent of the dilation based on the word error rate; and adjusting the resulting dilation of the CNN output based on the tuning parameters to reduce the word error rate.

[0021] One advantage of the present invention includes reducing word error rates when converting speech content from an audio input separated using speech recognition to convert speech from the audio input to text using a method according to the present invention.

[0022] In a related aspect, the system further comprises identifying speech content from the speaker from the audio input based on the learning corrections applied to the adjusted resulting dilation for the speaker.

[0023] In a related aspect, the system further comprises generating text from the identified audio content.

[0024] In a related aspect, the audio input and the predicted modifications are received without dilation of the audio content.

[0025] In a related aspect, the system further comprises adjusting a resulting dilation of the audio input based on the word error rate using a grid search to reduce the word error rate.

[0026] In a related aspect, the system further comprises receiving, at the computer, an expected audio input for the speaker, the expected audio input including speech content for the speaker; generating an environmental stimulus audio input for the expected audio input; and predicting changes in the audio input for the speaker based on the environmental stimulus audio input.

[0027] In a related aspect, the system further comprises sharing the adjusted resulting dilation of the expected audio input with a social network; generating a learning correction from the sharing of the adjusted resulting dilation of the expected audio input; applying the learning correction to the adjusted resulting dilation for the speaker; identifying speech content from the speaker from the audio input based on the learning correction applied to the adjusted resulting dilation for the speaker; and generating text from the identified speech content.

[0028] In a related aspect, the system further comprises sharing the adjusted resulting dilation of the expected audio input with a social network; generating a learned modification from the sharing of the adjusted resulting dilation of the expected audio input; and applying the learned modification to the adjusted resulting dilation for the speaker.

[0029] In a related aspect, the system includes identifying speech content from a speaker from the audio input based on a training correction applied to a dilation of the adjusted result for the speaker.

[0030] In a related aspect, the system further comprises generating text from the identified audio content.

[0031] In another aspect, according to the present invention, a computer program product for speech recognition using dilation of speech content from a separated audio input includes a computer-readable storage medium having program instructions embodied thereon. The program instructions are executable by a computer to cause the computer to perform computer-implemented functions, including receiving an audio input at a convolutional neural network (CNN) and receiving predicted modifications to the audio input based on external noise, the audio input having speech content from a speaker; applying diarization to the audio input at the CNN to predict how dilation of the speech content from the speaker will modify the audio input to generate a CNN output; determining a resulting dilation from the CNN output, the resulting dilation of the CNN output including separating sounds from the audio input; determining a word error rate for the dilated CNN output to determine accuracy for the speech-to-text output; setting tuning parameters to modify the extent of the dilation based on the word error rate; and adjusting the resulting dilation of the CNN output based on the tuning parameters to reduce the word error rate.

[0032] One advantage of the present invention includes reducing word error rates when converting speech content from an audio input separated using speech recognition to convert speech from the audio input to text using a method according to the present invention.

[0033] In a related aspect, the computer program product further comprises identifying speech content from a speaker from the audio input based on the training corrections applied to the adjusted resulting dilation for the speaker.

[0034] In a related aspect, the computer program product further includes generating text from the identified audio content. [Brief explanation of the drawings]

[0035] These and other objects, features, and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the invention, which should be read in conjunction with the accompanying drawings, in which various features of the drawings are not to scale, as the illustrations are for clarity in facilitating understanding of the invention by those skilled in the art with respect to the detailed description, and the drawings are discussed below.

[0036] [Figure 1] 1 is a schematic block diagram illustrating a system overview, system features or components, and methods for speech recognition using dilation of speech content from interlaced audio input, according to an embodiment of the present invention;

[0037] [Figure 2] 2 is a flowchart illustrating a method implemented using the system shown in FIG. 1 for speech recognition using dilation of speech content from interlaced audio input, according to an embodiment of the present invention.

[0038] [Figure 3] 1 is a series of tables illustrating dilation embodiments according to the present disclosure.

[0039] [Figure 4] 2 is a flowchart illustrating another embodiment of a method according to the present disclosure for speech recognition using dilation of speech content from interlaced audio input, implemented using the system shown in FIG. 1 .

[0040] [Figure 5] 5 is a flowchart continuing from the flowchart shown in FIG. 4, depicting a continuation of the method shown in FIG. 4, according to an embodiment of the present invention.

[0041] [Figure 6] FIG. 6 is a functional schematic block diagram showing, for illustrative purposes, a sequence of operations and functional methods illustrating functional features of the present disclosure in relation to the embodiments shown in FIGS. 1, 2, 3, 4 and 5 for speech recognition using dilation of speech content from interlaced audio input.

[0042] [Figure 7] FIG. 6 is a functional schematic block diagram showing, for illustrative purposes, a sequence of operations and functional methods illustrating functional features of the present disclosure related to the embodiments shown in FIGS. 1, 2, 3, 4 and 5 for speech recognition using dilation of speech content from interlaced audio input.

[0043] [Figure 8] 2 is a flowchart illustrating a method according to an embodiment of the present disclosure, implemented using the system shown in FIG. 1, for speech recognition using dilation of speech content from a separated audio input, in accordance with an embodiment of the present invention.

[0044] [Figure 9]2 is a flowchart illustrating another embodiment of a method according to the present disclosure for speech recognition using dilation of speech content from a separated audio input, implemented using the system shown in FIG. 1 .

[0045] [Figure 10] 10 is a flowchart continuing from the flowchart shown in FIG. 9, depicting a continuation of the method shown in FIG. 9, according to an embodiment of the present disclosure.

[0046] [Figure 11] FIG. 11 is a functional schematic block diagram showing, for illustrative purposes, a sequence of operations and functional methods illustrating functional features of the present disclosure related to the embodiments shown in FIGS. 8, 9 and 10 for speech recognition using dilation of speech content from a separated audio input.

[0047] [Figure 12] FIG. 11 is a functional schematic block diagram showing, for illustrative purposes, a sequence of operations and functional methods illustrating functional features of the present disclosure relating to the embodiments shown in FIGS. 8, 9 and 10 for speech recognition using dilation of speech content from a separated audio input.

[0048] [Figure 13] 1 is a schematic block diagram depicting a computer system according to an embodiment of the present disclosure that may be incorporated in whole or in part into one or more computers or devices shown in FIG. 1 and that cooperates with the systems and methods shown in FIGS. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12.

[0049] [Figure 14] 1 is a schematic block diagram of a system depicting system components interconnected using a bus, in whole or in part, in accordance with and for use with one or more embodiments of the present disclosure.

[0050] [Figure 15] FIG. 1 is a block diagram depicting a cloud computing environment according to an embodiment of the present invention.

[0051] [Figure 16] FIG. 2 is a block diagram illustrating abstract model layers according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0052] The following description, which refers to the accompanying drawings, is provided to facilitate a comprehensive understanding of exemplary embodiments of the present invention, as defined by the claims and their equivalents. It includes various specific details to aid in understanding, but these should be considered merely as examples. Therefore, those skilled in the art will recognize that various changes and modifications to the embodiments described herein can be made without departing from the scope of the present invention. Additionally, for clarity and conciseness, descriptions of well-known functions and structures may be omitted.

[0053] The terms and phrases used in the following description and claims are not limited to their bibliographical meanings, but are merely used to enable a clear and consistent understanding of the present invention. Therefore, it should be apparent to those skilled in the art that the following description of exemplary embodiments of the present invention is provided for illustrative purposes only, and not for the purpose of limiting the present invention, which is defined by the appended claims and their equivalents.

[0054] The singular forms "a," "an," and "the" should be understood to include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a surface of a component" includes reference to the presence of one or more such surfaces unless the context clearly dictates otherwise.

[0055] Embodiments according to the present disclosure analyze speech content from interlaced audio input to provide speech recognition for each of multiple speakers, thereby providing word recognition and identification and speech content to text conversion. The present disclosure enables speech recognition and speech to text conversion when sounds are modified or less common than the typical pronunciation of a word, for example, when the speech content includes modified sounds or atypical sounds from a speaker, or a combination thereof.

[0056] Embodiments of the present disclosure include speech recognition using dilation of an audio signal or audio input to increase the spacing between samples or audio samplings before attempting to recognize words or analyzing audio content to recognize one or more words. In one example according to the present disclosure, convolutional neural networks (CNNs) with different dilation parameters can be trained and applied to these problems. In another example, expected environmental noise and audio type can dictate which dilation to use. In another example, each speaker can be assigned dilation parameters through machine learning. In another example, in a group setting of a conversation or song, dilation parameters can be weighted together for each group based on the amplitude of each speaker.

[0057] An embodiment of the present disclosure may thereby provide expected environmental noise to set dilation parameters. In another example, an embodiment of the present disclosure classifies audio type (e.g., singing, speaking) to contribute to dilation parameters. In another example, an embodiment of the present disclosure fits audio spread to other independent models. In another example, an embodiment of the present disclosure may include speaker diarization and average dilation parameters based on group models. And in another example, an embodiment of the present disclosure may include social dilation transfer of knowledge.

[0058] Accordingly, embodiments of the present disclosure include modeling dilation from environmental noise and predicting dilation parameters. Dilation spread can be mapped, and further, social aspects can be combined with dilation metrics for each person in the conversation.

[0059] 1 and 2, a method 100 (FIG. 2) for system 10 (FIG. 1) according to an embodiment of the present disclosure is provided for speech recognition using dilation of speech content from interlaced audio input. Referring to FIG. 2, the method includes a series of operational blocks for implementing an embodiment according to the present disclosure. Referring to FIG. 2, method 100 includes, as in block 104, initiating a learning model 320 (see FIG. 6) to determine dilation parameters 324 for each of a plurality of audible sounds 62 of speech content 64 from a plurality of human speakers 52 received as audio input 60 by computer 22.

[0060] Referring to FIG. 6, a functional system 300 includes components and operations for embodiments according to the present disclosure and is used herein with reference to the methods and systems shown in FIGS. 1, 2, 3, 4 and 5.

[0061] In one example, a group of speakers may speak together, and audio output from the group of speakers may be received as audio input using a computer or device, for example, using a microphone on the device or in communication with the device or computer.

[0062] In one example, a spectrogram may be generated and used as a visual representation of the spectrum of frequencies of a signal, which varies over time, such as an audio signal. A spectrogram may also be referred to as a sonograph, voiceprint, or voicegram.

[0063] A spectrogram may be created and a DFT (Discrete Fourier Transform) may be applied to determine potential unique speakers. The DFT may convert a finite series of equally spaced samples of a function into an identical length series of equally spaced samples of a Discrete Time Fourier Transform (DTFT), which is a complex-valued function of frequency. An initial dilation variable may be initialized to each DFT estimate.

[0064] In one example, when it is known who is speaking or singing in a group, dilation parameters can be adjusted or specified based on the information. Such identifying information can be gathered, for example, from social media entries, observations.

[0065] The plurality of audible sounds may include one or more human speakers 14 or users, for example, as a plurality of human speakers 52 or users in a neighborhood 50 speaking and generating audible sounds 62. The audible sounds may include, for example, humans speaking in a conversation, a single voice, a song, a group of speakers singing, etc. And the audible sounds 62 thus generate and include audio content 64.

[0066] The audible sound may be received at the computer as audio input 60 via a microphone on computer 22 or device 20, such as a mobile device, and the computer, alone or in combination with a control device of control system 70, may transmit the audio file to another computer 72 or server, such as a remote computer or server, via communications network 45, e.g., the Internet, for processing in accordance with the techniques of this disclosure. In another example, the audible sound in the audio file may be processed in accordance with the techniques of this disclosure locally on the computer or in combination with a remote computer or server.

[0067] The learning model 320 may include machine learning using parameters. For example, dilation parameters may be assigned to each of multiple speakers or users using machine learning.

[0068] Dilation of sounds for speech recognition may be defined as increasing the space between sounds or sound samples. In this disclosure, dilation is performed before attempting to recognize words from sounds.

[0069] The dilation parameters may include a specified amount of space between sound samples or may specify a range of space between sound samples. A dilation variable may be assigned to each potential speaker and used in the training model.

[0070] In one example, referring to FIG. 3 , representative images 150 of tables 154, 158, and 162 depict the dilation of an image 166, which may be, for example, a fragment or sound image layered into sound samples 168. The images 150 depict variations in the dilation parameter D. In table 154, the dilation parameter 172 is equal to 1, and the image 166 has no space. In second table 158, the dilation parameter 174 is equal to 2, and the image has sound samples 168 with space 180 between samples. In third table 162, the dilation parameter 176 is equal to 3, and the image has sound samples 168 with more space 180 between samples.

[0071] The method includes determining, as part of the learning model, a variation of each of a plurality of independent sounds in response to an audio stimulus, the independent sounds being derived from the audio input, as in block 108. For example, the audio stimulus may include an environmental stimulus. The sound variation, or independent sound variation 322, may be determined in response to the environmental stimulus.

[0072] In one embodiment, the audio input may be refined for each of multiple speakers using diarization. For example, diarization may involve the process of partitioning the input audio stream into segments corresponding to speaker identities, and in one example, the segments may be homogeneous. The diarized signal may be used to layer the audio input into independent sounds.

[0073] Diarization can be used in deep learning to refine the audio input from each of multiple speakers that can be attributed to each speaker. In one example, if there is an error from a DFT or deep learning or combination approach to speaker identification, the diarization parameters can be averaged together.

[0074] In one example, a voiceprint can be constructed. In one example, environmental stimuli can be played and a determination can be made as to how the layered data changes. The dilation parameters can be modified based on changes in the audio input data (e.g., speech data). For example, if the voice is lengthened, the dilation parameters can be increased. Additionally, the dilation parameters for each speaker can be based on the correlation coefficient of the independent signals. The coefficient of correlation (R value) is the value given in the summary table in the regression output. R-squared is called the coefficient of determination, i.e., R x R to obtain the value of R-squared. The coefficient of determination is the square of the coefficient of correlation.

[0075] In one example, the R-squared correlation metric determines how pairwise most correlated speakers are combined together. For example, the rank of R can be shifted between 0 and 0.5, so that, at most, a pair of speakers contributes 50% of the adjusted dilation.

[0076] The method comprises applying dilation parameters based on the variation of each of the independent sounds, respectively, as in block 112 .

[0077] The method comprises constructing a voiceprint for each of the speakers based on each of the independent sound and dilation parameters, as per block 116 .

[0078] The method includes attributing speech content to each of a plurality of speakers based at least in part on the respective voiceprints and individual sounds, as per block 120 .

[0079] The method may include generating text from the attributed audio content, as in block 124 .

[0080] If the generated text is to be displayed locally, e.g., on a local computer, as determined in block 126, the method continues to block 130. If the generated text is not to be displayed locally, as determined in block 126, the method continues to block 128 for display on a device or computer display or monitor.

[0081] In response to displaying the text location as determined in block 126, the method comprises displaying the text on a screen or monitor in communication with the computer or device, as in block 130.

[0082] In response to not displaying the text locally as determined in block 126, the method may include transmitting the text via an electronic communication system to a computer or device, or combination thereof, for display on a screen or monitor in communication with the computer or device, as in block 128. The method may continue with displaying the text on a screen or monitor in communication with the computer or device, as in block 130.

[0083] The method may include a learning model 320 including a CNN 326 (convolutional neural network) for receiving the independent sounds and determining a modification of each of the independent sounds in response to the audio stimulus using diarization.

[0084] CNN (Convolutional Neural Network) can be at least part of deep learning, and CNN is a class of deep neural networks. CNNs involve a mathematical operation commonly defined as two functions that produce a third function, called a convolution. A convolution is a specialized type of linear operation. Thus, a convolutional network is a neural network that uses convolution instead of general-purpose matrix multiplication in at least one of its layers.

[0085] The method may include receiving an audio input that may include a plurality of audible sounds, the audio input may include speech content from a plurality of speakers.

[0086] The method may further comprise dilating the audio input using the computer, which may include separating sounds in the audio input.

[0087] The method may further comprise refining the dilated audio input 302 for each of the multiple speakers using diarization of the sounds 308 .

[0088] The method may further comprise layering the sounds 306 in the refined audio input 304 into independent sounds 310 using a diarization of the audio input.

[0089] The method may include separating sounds in an audio input and may comprise distinguishing environmental or background sounds from sounds from a speaker of a plurality of speakers.

[0090] The method may comprise refining the audio input for each of a plurality of speakers using diarization, which may include partitioning the audio input into homogeneous segments with respect to speaker identification.

[0091] In another embodiment according to the present disclosure, and referring to Figure 4, a computer-implemented method 200 for speech recognition using dilation of speech content from an interlaced audio input includes receiving, at a computer, an audio input including multiple audible sounds, the audible sounds including speech content from multiple speakers, as in block 204. The operational blocks of method 200 shown in Figures 4 and 5 may be similar to the operational blocks shown in Figure 2. The method shown in Figures 4 and 5 is intended as another exemplary embodiment that may include aspects / operations already shown and discussed in this disclosure.

[0092] The method 200 includes, as in block 208, dilating the audio input using a computer, where the dilation includes separating sounds in the audio input.

[0093] The method 200 includes refining the audio input for each of a plurality of speakers using diarization, as in block 212. The method includes layering the sounds in the audio input into independent sounds using the diarization of the audio input, as in block 216.

[0094] The method 200 includes initiating a learning model to determine dilation parameters for each of the audible sounds, as in block 220 .

[0095] The method 200 includes, as in block 222, a learning model including a convolutional neural network (CNN) for receiving the independent sounds and determining a modification of each of the independent sounds in response to the audio stimulus using diarization.

[0096] The method 200 includes determining, as part of a learning model, a modification of each of a plurality of independent sounds in response to an audio stimulus, as per block 224 .

[0097] The method 200 includes applying dilation parameters based on the variation of each of the independent sounds, as in block 228 .

[0098] The method 200 includes constructing a voiceprint 330 for each of the speakers 52 based on each of the independent sounds 310 and dilation parameters 324 , as per block 232 .

[0099] The method 200 includes attributing speech content to each of a plurality of speakers based at least in part on the respective voiceprints and individual sounds, as in block 236. The attributed speech content 332 may be used to generate text.

[0100] Referring to FIG. 5, the method 200 includes generating text 334 from the attributed audio content 332, as per block 240.

[0101] Method 200 further includes transmitting the text via an electronic communication system to a computer or device, or a combination thereof, for display on a screen or monitor in communication with the computer or device, or a combination thereof, as in block 244. In another example, the communication may be implemented from the group consisting of SMS, email, instant messaging, and navigation software. Such examples are intended to be illustrative and non-exhaustive.

[0102] Method 200 may further comprise displaying the text on a screen or monitor in communication with the computer or device, as per block 248 .

[0103] 7, a functional system 400 illustrating and supporting embodiments described herein, according to an embodiment of the present disclosure, includes components and operations for speech recognition using dilation of speech content from interlaced audio input. The system 400 includes a group of human speakers 402 that output audio output. The audio output is received to train each distinct signal using dilation, as in block 404. The system may train the dilation of the audio input signal based on the diarization, as in block 406.

[0104] The system includes layering the audio input signals using diarization, as in block 410. The system includes playing sounds, e.g., environmental stimuli, and grouping the diarized audio input signals, as in block 412. The system includes setting individual and group dilations based on the environmental stimuli, as in block 414. The system includes generating an audio output, as in block 416. The system includes generating a text output, as in block 418, using the audio output 416 based on the dilation and the environmental stimuli.

[0105] In one example, the system may predict the speaker's signal using a predictive speaker signal technique or method / system by predicting how the speaker's signal will change based on external noise, as in block 450, to predict the speaker's signal for the future. Such prediction is not the focus of this disclosure.

[0106] 1 and 2, the computer may be a remote computer or part of a remote server, such as remote server 1100 (FIG. 8). In another example, computer 72 may be part of control system 70 and provide performance of the functionality of the present disclosure. In another embodiment, computer 22 may be part of mobile device 20 and provide performance of the functionality of the present disclosure. In yet another embodiment, some performance of the functionality of the present disclosure may be shared between a control system computer and a mobile device computer, e.g., the control system acts as a back end for one or more programs embodying the present disclosure, and the mobile device computer acts as a front end for the one or more programs.

[0107] The computer may be part of the mobile device or a remote computer that communicates with the mobile device. In another example, the mobile device and the remote computer may work in cooperation to implement the methods of the present disclosure using stored program code or instructions for performing the method features described herein. In one example, the mobile device 20 may include a computer 22 having a processor 15 and a storage medium 34 that stores an application 40. The application may incorporate program instructions for performing the features of the present disclosure using the processor 15. In another example, the mobile device 20 application 40 may have executable program instructions for a software application front end that incorporates the method features of the present disclosure into the program instructions, while one or more back-end programs 74 of the software application stored on the computer 72 of the control system 70 communicate with the mobile device computer to perform other method features. The control system 70 and the mobile device 20 may communicate using a communications network 45, such as the Internet.

[0108] Accordingly, a method 100 according to one embodiment of the present disclosure may be embodied in one or more computer programs or applications 40 stored on an electronic storage medium 34 and executable by a processor 15 as part of a computer on a mobile device 20. For example, a human speaker or user 14 may have a device 20, which may communicate with a control system 70. Other users (not shown) may have similar devices and similarly communicate with the control system. The application may be stored, in whole or in part, on a computer or computers on the mobile device and in a control system that communicates with the device using a communications network 45, such as the Internet. It is contemplated that the application may access all or a portion of the program instructions to implement the method of the present disclosure. The program or application may communicate with a remote computer system via a communications network 45 (e.g., the Internet), access data, and cooperate with programs stored on the remote computer system. Such interactions and mechanisms are described in further detail herein and reference is made to components of a computer system, such as computer-readable storage media, which are illustrated in one embodiment in FIG. 8 and, in this regard, are described in more detail with reference to one or more computer systems 1010.

[0109] Thus, in one example, control system 70 communicates with device 20, which may include application 40. Device 20 communicates with control system 70 using communication network 45.

[0110] In another example, control system 70 may have a front-end computer belonging to one or more users, such as device 20, and a back-end computer embodied as a control system.

[0111] 1, device 20 may include computer 22, computer-readable storage medium 34, an operating system, or program, or a combination thereof, or software application 40, or a combination thereof, where software application 40 may include program instructions executable using processor 15. These features are illustrated herein in FIG. 1 and are also illustrated in one embodiment of a computer system shown in FIG. 8 with reference to one or more computer systems 1010, which may include one or more general computer components.

[0112] The methods according to the present disclosure may include a computer as part of a control system for implementing features of the methods according to the present disclosure. In another example, a computer as part of a control system may work with a mobile device computer to implement features of the methods according to the present disclosure. In another example, a computer for implementing features of the methods may be part of a mobile device, thereby implementing the methods locally.

[0113] 6 and 7 should be understood to be functional representations of features of the present disclosure, which are shown in embodiments of the systems and methods of the present disclosure for illustrative purposes to clarify the functionality of the features of the present disclosure.

[0114] Specifically, with respect to control system 70, devices 20 of one or more users 14 may communicate with control system 70 via communications network 50. In the embodiment of the control system shown in FIG. 1 , control system 70 includes a computer 72 having a database 76 and one or more programs 74 stored on a computer-readable storage medium 73. In the embodiment of the present disclosure shown in FIG. 1 , devices 20 communicate with control system 70 and the one or more programs 74 stored on computer-readable storage medium 73. Control system includes computer 72 having a processor 75, which also has access to database 76.

[0115] The control system 70 may include a storage medium 80 for maintaining a registration 82 of users and their devices for analysis of audio input. Such registration may include a user profile 83, which may include user data provided by the user for account registration and setup. In one embodiment, methods and systems incorporating the present disclosure include a control system (commonly referred to as a backend) in combination with and working with a front end of the method and system, which may be an application 40. In one example, the application 40 is stored on a device, e.g., device 20, and can access data and additional programs in the application's backend, e.g., the control system 70.

[0116] A control system may also be part of a software application implementation, or may represent a software application having a front-end user portion and a back-end portion that provide functionality, or both. In one embodiment, a method and system incorporating the present disclosure includes a control system (which may be generally referred to as a back-end of a software application incorporating a portion of a method and system of an embodiment of the present disclosure) in combination with and working in cooperation with a front-end of a software application incorporating another portion of the method and system of the present disclosure on a device, such as in the example shown in FIG. 1 of device 20 having application 40. Application 40 is stored on device 20 and can access data and additional programs in the back-end of the application, for example, in program 74 stored in control system 70.

[0117] Program 74 may include, in whole or in part, a series of executable steps for implementing the methods of the present disclosure. A program incorporating a method according to the present disclosure may be stored, in whole or in part, on a computer-readable storage medium on the control system or, in whole or in part, on device 20. It is contemplated that control system 70 can not only store user profiles, but also, in one embodiment, interact with a website for viewing on the device's display, or in another example, the Internet, and receive user input related to the methods and systems of the present disclosure. While FIG. 1 shows one or more profiles 83, it is understood that the methods may include multiple profiles, users, registrations, etc. It is contemplated that multiple users or groups of users can register and provide profiles with the control system for use by the methods and systems of the present disclosure.

[0118] With respect to the collection of data pertaining to the present disclosure, such uploading or creation of a profile is voluntary by one or more users and is therefore initiated by and with user authorization. Users may thereby consent to establishing accounts with profiles in accordance with the present disclosure. Similarly, data received by the system or entered or received as input is voluntary by one or more users and therefore initiated by and with user authorization. Users may thereby opt in to entering data in accordance with the present disclosure. Such user authorization also includes the user's option to cancel such profile or account, and / or data entry, thereby opting out of communication and data capture at the user's discretion. It is further understood that any stored or collected data is intended to be securely stored and unavailable without user authorization, and not available to the public and / or unauthorized users. It is understood that such stored data is deleted upon the user's request, and in a secure manner. It is also understood that any use of such stored data, in accordance with this disclosure, is only with the user's authorization and consent.

[0119] In one or more embodiments of the present invention, a user may opt in or register with a control system, voluntarily providing data and / or information with the user's consent and approval, and the data may be stored and used in one or more methods of the present disclosure. The user may also register one or more user electronic devices for use with one or more methods and systems according to the present disclosure. As part of registration, the user may also identify and approve access to one or more activities or other systems (e.g., audio and / or video systems). Such opt-in to registration and approval of data collection and / or storage is voluntary, and the user may request deletion of data (including profile and / or profile data), unsubscription, or opt-out of any registration, or a combination thereof. Such opt-out is understood to include disposal of all data in a secure manner.

[0120] In one example, artificial intelligence (AI) may be used, in whole or in part, to learn models for determining dilation parameters.

[0121] In another example, the control system 70 may be, in whole or in part, an artificial intelligence (AI) system. For example, the control system may be one or more components of an AI system.

[0122] It is also understood that the method 100 according to an embodiment of the present disclosure can be incorporated into an AI (artificial intelligence) device that can communicate with the respective AI system and the respective AI system platform. Thus, such a program or application incorporating the method of the present disclosure can be part of the AI ​​system, as previously discussed. In one embodiment according to the present invention, it is contemplated that the control system can communicate with the AI ​​system, or in another example, can be part of the AI ​​system. The control system can also represent a software application having a front-end user portion and a back-end portion that provide functionality, which in one or more examples can interact with, encompass, or be part of a larger system, such as an AI system. In one example, an AI device can be associated with an AI system that is, in whole or in part, a control system or a content delivery system, or both, and that can be remote from the AI ​​device. Such an AI system can be represented by one or more servers storing a program on a computer-readable medium that can communicate with one or more AI devices. The AI ​​system can communicate with the control system, and in one or more embodiments, the control system can be all or part of the AI ​​system, or vice versa.

[0123] As discussed herein, it is understood that downloads or downloadable data may be initiated using voice commands, or using a mouse, touch screen, or the like. In such examples, a mobile device may be user-initiated, or an AI device may be enabled for use with user consent and permission. Other examples of AI devices include devices that include a microphone, a speaker, and have access to a cellular or mobile network, a communications network, or the Internet, such as a vehicle having a computer and cellular or satellite communications, or, in another example, an IoT (Internet of Things) device such as a home appliance with cellular network or Internet access.

[0124] 1 and 8, a method 500 (FIG. 8) for the system 10 (FIG. 1) according to an embodiment of the present disclosure is provided for speech recognition using dilation of speech content from a separated audio output (also referred to as singular). Referring to FIG. 8, the method includes a series of operational blocks for implementing an embodiment according to the present disclosure. Referring to FIG. 8, the method 100 includes, as in block 504, receiving an audio input at a convolutional neural network (CNN) and receiving predicted modifications for the audio input, where the audio input has speech content from a person, i.e., a human speaker.

[0125] In one example, the audio input may include speech content from a human speaker, e.g., audio input using a computer or device, e.g., using a microphone on the device or a microphone in communication with the device or computer. In one example, the speaker audio input may be represented as block 802 in system 800 shown in Figure 12. In another example, the speaker audio input may be represented as block 52 and used at least in part as audio input 704, as shown in Figures 1 and 11.

[0126] In another example, the audio input may include output from a system or method for recognizing speech. For example, the system and method may use dilation and diarization to recognize speech, such as method 100 shown in FIG. 2. Additionally, in this example, the output of method 100 at block 120 may be used as audio input for a speaker in method 500. Such audio output from the system may also be represented as block 416 in system 800 shown in FIG. 12 and may be used at least in part as audio input in speaker signal 804. In another example, the speaker's audio output from a system using dilation may be represented by block 702 of system 700 shown in FIG. 11 and may be used as at least a portion of audio input 704.

[0127] In each case, the audio input includes speech content from a human speaker, with speech content 708 (FIG. 11) having a plurality of audible tones 706 (FIG. 11).

[0128] 8 and 11 , in one example, the received predicted modifications 712 for the audio input based on external noise 714 may include a set of predicted modifications derived, for example, from a learning model. The learning model may, for example, use one or more external noises and model modifications in the audio input based on the external noise. The external noise may include, for example, but is not limited to, background noise, including ambient sounds, additional speaking noises, for example.

[0129] A CNN (Convolutional Neural Network) 718 can be at least part of deep learning, where a CNN is a class of deep neural networks. A CNN involves a mathematical operation that produces a third function and is commonly defined as two functions called a convolution. A convolution is a specialized type of linear operation. Thus, a convolutional network is a neural network that uses convolution instead of general-purpose matrix multiplication in at least one of its layers.

[0130] 8, method 500 includes, as in block 508, applying diarization 720 to the audio input in a CNN to predict how dilation of speech content from a speaker will modify the audio input to generate CNN output 724. For example, the learned model may include analysis to predict dilation of speech content and how the dilation will modify the speech content.

[0131] Method 500 includes determining a resulting dilation 726 from the CNN output, as in block 512, where the resulting dilation of the CNN output includes isolating sounds 732 of the audio input. For example, the learned model 730 may determine a dilation 734 and a prediction of how the dilation will modify the audio content.

[0132] Method 500 includes determining a word error rate 736 for the dilated CNN output to determine the accuracy of the speech-to-text output, as in block 516. For example, the method may determine an accuracy 740 percentage for the speech-to-text conversion. In another example, the method may determine an accuracy number for the speech-to-text conversion, e.g., different percentages of accuracy for different dilation and prediction models.

[0133] The method 500 includes setting adjustment parameters for varying the range of dilation based on the word error rate 736, as in block 520. For example, one or more adjustment parameters may be used to set the dilation or range of dilation for the audio content. The adjustment parameters 742 may be based on the word error rate 736, e.g., adjusting the dilation 744 in cooperation with the word error rate.

[0134] The method 500 includes adjusting the resulting dilation of the CNN output based on the tuning parameters to reduce the word error rate, as in block 524. In one example, when the word error rate is higher, the dilation may be increased. In another example, a word error rate threshold may be used to trigger a change, e.g., an increase, in the dilation when the word error rate threshold is met (e.g., an unsatisfactory high word error rate). In another example, the word error rate threshold may not be met, indicating an acceptable word error rate.

[0135] 8, the method 500 may have an acceptable word error rate and the method ends at block 526. When the method does not have an acceptable word error rate, the method may return to block 524 and adjust the dilation of the CNN.

[0136] In one embodiment, referring to block 746 of Figure 11 and block 824 of Figure 12, the adjusted resulting dilation 524 may be output for use by a system that uses dilation for speech recognition. The output may be used as input for a system that uses dilation for speech recognition, such as block 748 of Figure 11 and block 828 of Figure 12. In another example, the output at blocks 746 and 824, also referenced in method 500 at block 524 and method 600 at block 648, may be used, at least in part, as input in previously described embodiments, such as block 104 of Figure 2, block 204 of Figure 4, block 60 of Figure 6, and block 404 of Figure 7.

[0137] The method may further comprise identifying speech content from a speaker from the audio input based on the training modifications applying a resulting dilation tailored for the speaker.

[0138] In another example, the method may include generating text from the identified audio content.

[0139] In another example, the audio input and predicted modifications in the method are received without dilation of the audio content.

[0140] In another example, the method further comprises adjusting the resulting dilation of the audio input based on the word error rate using a grid search to reduce the word error rate. For example, a grid search may be used to find optimal hyperparameters of a model, such as a learning algorithm or computer learning model, which may result in more accurate predictions.

[0141] The method may further include receiving at the computer an expected audio input for the speaker, the expected audio input may include speech content for the speaker, an environmental stimulus audio input is generated for the expected audio input, and the method includes predicting changes in the audio input for the speaker based on the environmental stimulus audio input.

[0142] In one example, the method may further include sharing the adjusted resulting dilation of the expected audio input with a social network. A learning modification may be generated from sharing the adjusted resulting dilation of the expected audio input. The learning modification may be applied to the adjusted resulting dilation for the speaker. The method may include identifying speech content from the speaker from the audio input based on the learning modification applied to the adjusted resulting dilation for the speaker. And, the method may include generating text from the identified speech content.

[0143] The method may further comprise sharing the adjusted resulting dilation of the expected audio input with a social network. The method may comprise generating a learned modification from the sharing of the adjusted resulting dilation of the expected audio input and applying the learned modification to the adjusted resulting dilation for the speaker.

[0144] Additionally, the method may comprise identifying speech content from a speaker from the audio input based on a training correction applied to the adjusted resulting dilation for the speaker.

[0145] The method may comprise generating text from the identified audio content.

[0146] The method may comprise receiving, at the CNN, dilation parameters for speech content of one of a plurality of speakers, the dilation parameters being derived from audio input from the plurality of speakers.

[0147] In one example, the audio input from multiple speakers in the method is interlaced audio input.

[0148] 11, a functional system 700 illustrating and supporting embodiments described herein, according to an embodiment of the present disclosure, includes components and operations for speech recognition using dilation of speech content from interlaced audio input. For example, system 700 is representative of the functionality included in embodiments of the present disclosure and includes operations used therein.

[0149] 12, a system 800 illustrating and supporting embodiments described herein, according to an embodiment of the present disclosure, includes components and operations for speech recognition using dilation of speech content from a separated audio input. The system 800 includes a group of human speakers 802 that output audio output. The audio output is received for predicting a speaker's signal, as in block 804. Alternatively, the system 800 may receive an audio output, such as the audio output 416 of the system 400 shown in FIG. 7, to provide a predicted speaker signal using techniques. The system 800 receives an audio input for predicting a speaker's signal, as in block 804.

[0150] In one example, audio input is received from a speaker. In another example, the audio input can be the output of a speech recognition system that uses dilation of speech content from interlaced audio input, such as the dilated speech content in block 332 of Figure 6, or in another example, audio output 416 shown in Figure 7. In another example, the dilated audio input can be the dilated speech content in block 120 of Figure 2, or in another example, the imputed speech content in block 236 of Figure 4, or audio output 416 shown in Figure 7.

[0151] The system 800 includes learning from environmental noise, as in block 806, and predicting the speaker's signal using environmental stimuli, as in block 808.

[0152] As per block 810, both the expected speaker signal with environmental stimuli 808 and the expected speaker signal 804 are received for application to diarization using a CNN.

[0153] A word error rate 812 is determined, as in block 814, and the diarization can be adjusted using a grid search. The diarization with environmental noise is shared with a social network, as in block 816. Active learning adjustments are pushed to the social network, as in block 818. The diarization knowledge is transferred to the speaker, as in block 820.

[0154] In an example, referring to Figure 12, the output 824 for better diarization of speech content may be used as input 828 to a system for speech recognition using dilation of interlaced audio input, e.g., the output may be received in block 414 of Figure 7 as input in block 414 for system 400 to better set the dilation.

[0155] In another example, output 824 may be received in block 320 of FIG. 6 as input for system 300 in block 326 for use by a CNN and used with dilation to produce better resulting voiceprints.

[0156] 9 and 10, a method 600 according to another embodiment of the present disclosure for speech recognition using dilation of speech content from separated (or singular) audio input includes, at block 604, receiving, at a computer, expected audio input for a speaker, the expected audio input including speech content for the speaker.

[0157] The method includes generating an environmental stimulus audio input for the expected audio input, as per block 608 .

[0158] The method includes predicting changes in audio input for a speaker based on the environmental stimulus audio input, as per block 612 .

[0159] The method includes receiving, at a convolutional neural network (CNN), an audio input and a predicted change to the audio input, the audio input having speech content from a speaker.

[0160] The method includes applying diarization to the audio input in a CNN to predict how dilation of speech content from a speaker will modify the audio input and generate a CNN output, as in block 620.

[0161] The method includes determining a resulting dilation from the CNN output, as in block 622, where the resulting dilation of the CNN output includes separating sounds of the audio input.

[0162] The method includes determining a word error rate for the dilated CNN output to determine an accuracy number for the speech-to-text output, as in block 624.

[0163] The method includes setting tuning parameters to change the potential range of dilation based on the word error rate, as in block 628 .

[0164] The method includes adjusting the resulting dilation of the CNN output based on the adjustment parameters to reduce the word error rate, as per block 632 .

[0165] The method includes sharing the adjusted resulting dilation of the expected audio input with a social network, as in block 636 .

[0166] The method includes generating a learned correction from the adjusted resulting dilation share of the expected audio input, as per block 640 .

[0167] The method includes applying the learning corrections to the adjusted resulting dilation for the speaker, as in block 644 .

[0168] The method includes identifying speech content from the speaker from the audio input based on the learning corrections applied to the adjusted resulting dilation for the speaker, as per block 648 .

[0169] The method includes generating text from the identified audio content, as in block 652 .

[0170] 1 and 2, the computer may be a remote computer or part of a remote server, such as remote server 1100 (FIG. 8). In another example, computer 72 may be part of control system 70 and provide performance of the functionality of the present disclosure. In another embodiment, computer 22 may be part of mobile device 20 and provide performance of the functionality of the present disclosure. In yet another embodiment, some performance of the functionality of the present disclosure may be shared between a control system computer and a mobile device computer, e.g., the control system acts as a back end for one or more programs embodying the present disclosure, and the mobile device computer acts as a front end for the one or more programs.

[0171] The computer may be part of the mobile device or a remote computer that communicates with the mobile device. In another example, the mobile device and remote computer may work in cooperation to implement the methods of the present disclosure using stored program code or instructions for performing the method features described herein. In one example, the mobile device 20 may include a computer 22 having a processor 15 and a storage medium 34 that stores an application 40. The application may incorporate program instructions for performing the features of the present disclosure using the processor 15. In another example, the mobile device 20 application 40 may have executable program instructions for a software application front end that incorporates the method features of the present disclosure into the program instructions, while one or more back-end programs 74 of the software application stored on the computer 72 of the control system 70 communicate with the mobile device computer to perform other features of the method. The control system 70 and the mobile device 20 may communicate using a communications network 45, such as the Internet.

[0172] Accordingly, a method 100 according to one embodiment of the present disclosure may be embodied in one or more computer programs or applications 40 stored on an electronic storage medium 34 and executable by a processor 15 as part of a computer on a mobile device 20. For example, a human speaker or user 14 may have a device 20, which may communicate with a control system 70. Another user (not shown) may have a similar device and similarly communicate with the control system. The application may be stored, in whole or in part, on a computer or computers on the mobile device and in a control system that communicates with the device using a communications network 45, such as the Internet. It is contemplated that the application may access all or a portion of the program instructions to implement the method of the present disclosure. The program or application may communicate with a remote computer system via a communications network 45 (e.g., the Internet), access data, and cooperate with programs stored on the remote computer system. Such interactions and mechanisms are described in further detail herein and are referenced with respect to components of a computer system, such as computer-readable storage media, which are shown in one embodiment in FIG. 13 and, in this regard, are described in more detail with reference to one or more computer systems 1010.

[0173] In one example, the control system 70 communicates with the device 20, which may include the application 40. The device 20 communicates with the control system 70 using a communication network 45.

[0174] In another example, control system 70 may have a front-end computer belonging to one or more users, such as device 20, and a back-end computer embodied as a control system.

[0175] 1, device 20 may include computer 22, computer-readable storage medium 34, an operating system, or program, or a combination thereof, or software application 40, or a combination thereof, where software application 40 may include program instructions executable using processor 15. These features are illustrated herein in FIG. 1 and are also illustrated in one embodiment of a computer system shown in FIG. 13 with reference to one or more computer systems 1010, which may include one or more general computer components.

[0176] The methods according to the present disclosure may include a computer as part of a control system for implementing features of the methods according to the present disclosure. In another example, a computer as part of a control system may work with a mobile device computer to implement features of the methods according to the present disclosure. In another example, a computer for implementing features of the methods may be part of a mobile device, thereby implementing the methods locally.

[0177] It should be understood that the features illustrated in Figures 11 and 12 are functional representations of features of the present disclosure, and are shown in embodiments of the systems and methods of the present disclosure for illustrative purposes to clarify the functionality of the features of the present disclosure.

[0178] Specifically, with respect to control system 70, devices 20 of one or more users 14 may communicate with control system 70 via communications network 50. In the embodiment of the control system shown in FIG. 1 , control system 70 includes a computer 72 having a database 76 and one or more programs 74 stored on a computer-readable storage medium 73. In the embodiment of the present disclosure shown in FIG. 1 , devices 20 communicate with control system 70 and the one or more programs 74 stored on computer-readable storage medium 73. Control system includes computer 72 having a processor 75, which also has access to database 76.

[0179] The control system 70 may include a storage medium 80 for maintaining a registration 82 of users and their devices for analysis of audio input. Such registration may include a user profile 83, which may include user data provided by the user for account registration and setup. In one embodiment, methods and systems incorporating the present disclosure include a control system (commonly referred to as a backend) in combination with and working with a front end of the method and system, which may be an application 40. In one example, the application 40 is stored on a device, e.g., device 20, and can access data and additional programs in the application's backend, e.g., the control system 70.

[0180] A control system may also be part of a software application implementation, or may represent a software application having a front-end user portion that provides functionality and a back-end portion, or both. In one embodiment, a method and system incorporating the present disclosure includes a control system (which may be generally referred to as a back-end of a software application incorporating a portion of a method and system of an embodiment of the present disclosure) in combination with and working in cooperation with a front-end of a software application incorporating another portion of the method and system of the present disclosure on a device, such as in the example shown in FIG. 1 of device 20 having application 40. Application 40 is stored on device 20 and can access data and additional programs in the back-end of the application, for example, in program 74 stored in control system 70.

[0181] Program 74 may include, in whole or in part, a series of executable steps for implementing the methods of the present disclosure. A program incorporating a method according to the present disclosure may be stored, in whole or in part, on a computer-readable storage medium on the control system or, in whole or in part, on device 20. It is contemplated that control system 70 can not only store user profiles, but also, in one embodiment, interact with a website for viewing on the device's display, or in another example, the Internet, and receive user input related to the methods and systems of the present disclosure. While FIG. 1 shows one or more profiles 83, it is understood that the methods may include multiple profiles, users, registrations, etc. It is contemplated that multiple users or groups of users can register and provide profiles with the control system for use in accordance with the methods and systems of the present disclosure.

[0182] A set is understood to be a collection of distinct objects or elements. The objects or elements that make up a set can be anything, such as numbers, letters of the alphabet, other sets, etc. It is further understood that a set can be a single element, such as a single thing or number, in other words, a set of single elements.

[0183] Referring to FIG. 13 , one embodiment of a system or computer environment 1000 according to the present disclosure includes a computer system 1010, shown in the form of a generic computing device. The method 100 may be embodied, for example, in a program 1060 including program instructions embodied on a computer-readable storage device or medium, e.g., generally referred to as computer memory 1030 and more specifically as computer-readable storage medium 1050. Such memory and / or computer-readable storage medium may include non-volatile memory or non-volatile storage, also known and referred to as non-transitory computer-readable storage medium or non-transitory computer-readable storage medium. For example, such non-volatile memory may also be a disk storage device, including one or more hard drives. For example, the memory 1030 may include a storage medium 1034, such as random access memory (RAM) or read-only memory (ROM), and a cache memory 1038. The program 1060 is executable by the processor 1020 of the computer system 1010 (to execute program steps, code, or program code). Additional data storage may also be embodied as a database 1110 containing data 1114. The computer system 1010 and the program 1060 are general representations of a computer and program that may be local to a user or provided as a remote service (e.g., as a cloud-based service), and in further examples may be provided using a website accessible using the communications network 1200 (e.g., interacting with a network, the Internet, or a cloud service). It will be understood herein that the computer system 1010 also generally represents a computing device or computer included in a device, such as a laptop or desktop computer, or one or more servers, either alone or as part of a data center.The computer system may include a network adapter / interface 1026 and an input / output (I / O) interface 1022. The I / O interface 1022 allows for input and output of data to and from external devices 1074 that may be connected to the computer system. The network adapter / interface 1026 may provide communication between the computer system and a network, generally depicted as communications network 1200.

[0184] The computer 1010 may be described in the general context of computer system-executable instructions, such as program modules, being executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, and data structures that perform particular tasks or implement particular abstract data types. Method steps and system components and techniques may be embodied in modules of the program 1060 for performing the method steps and respective system tasks. Modules are generally represented in the diagram as program modules 1064. The program 1060 and program modules 1064 may execute particular program steps, routines, subroutines, instructions, or code.

[0185] The methods of the present disclosure may be executed locally on a device, such as a mobile device, or may be executed as a service on a server 1100, which may be remote, for example, and accessible using a communications network 1200. The program or executable instructions may be offered as a service by a provider. The computer 1010 may be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices linked through the communications network 1200. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.

[0186] More specifically, system or computer environment 1000 includes a computer system 1010 shown in the form of a general-purpose computing device with exemplary peripheral devices. Components of computer system 1010 may include, but are not limited to, one or more processors or processing units 1020, a system memory 1030, and a bus 1014 that couples various system components including the system memory 1030 to the processor 1020.

[0187] Bus 1014 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, such architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0188] The computer 1010 may include a variety of computer-readable media. Such media may be any available media accessible by the computer 1010 (e.g., a computer system or server) and may include both volatile and nonvolatile media, as well as removable and non-removable media. The computer memory 1030 may include additional computer-readable media in the form of volatile memory, such as random access memory (RAM) 1034 and / or cache memory 1038. The computer 1010 may further include other removable / non-removable, volatile / non-volatile computer storage media, such as a portable computer-readable storage medium 1072 in one example. In one embodiment, the computer-readable storage medium 1050 may be provided for reading from and writing to non-removable, non-volatile magnetic media. The computer-readable storage medium 1050 may be embodied as, for example, a hard drive. Additional memory and data storage may be provided, for example, as a storage system 1110 (e.g., a database) for storing data 1114 and communicating with the processing unit 1020. The database may be stored on or part of server 1100. Although not shown, a magnetic disk drive for reading from and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from and writing to a removable non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical media, may be provided, in which case each may be connected to bus 1014 by one or more data medium interfaces. As further depicted and described below, memory 1030 may include at least one program product, which may include one or more program modules configured to perform the functions of embodiments of the present invention.

[0189] The methods described in this disclosure may be embodied in one or more computer programs, generally referred to as programs 1060, for example, and stored in memory 1030 within computer-readable storage medium 1050. Program 1060 may include program modules 1064. Program modules 1064 generally may perform the functionality and / or methodology of embodiments of the invention as described herein. The one or more programs 1060 are stored in memory 1030 and executable on processing unit 1020. By way of example, memory 1030 may store an operating system 1052, one or more application programs 1054, other program modules, and program data on computer-readable storage medium 1050. It will be appreciated that program 1060, and the operating system 1052 and application programs 1054 stored on computer-readable storage medium 1050, are similarly executable on processing unit 1020. It is also understood that applications 1054 and programs 1060 are shown generically and may include or be a portion of one or more applications and programs described in this disclosure, or vice versa, i.e., applications 1054 and programs 1060 may be all or a portion of one or more applications or programs described in this disclosure. It is also understood that control system 70, which communicates with a computer system to achieve the control system functionality described in this disclosure, may include all or a portion of computer system 1010 and its components, or the control system may communicate with all or a portion of computer system 1010 and its components as a remote computer system, or both. Control system functionality may include, for example, storing, processing, and executing software instructions to perform the functions of the present disclosure.It is also understood that to accomplish the computer functions described in this disclosure, one or more computers or computer systems shown in FIG. 1 may similarly include all or some of computer system 1010 and its components, or one or more computers may be in communication with all or some of computer system 1010 and its components as remote computer systems, or both.

[0190] In one embodiment according to the present disclosure, one or more programs may be stored on one or more computer-readable storage media such that the programs are embodied and / or encoded on the computer-readable storage media. In one example, the stored programs may include program instructions for execution by a processor or a computer system having a processor to perform a method or to cause the computer system to perform one or more functions. For example, in one embodiment according to the present disclosure, a program embodying a method is embodied or encoded on a computer-readable storage medium, which includes a non-transitory or non-transitory computer-readable storage medium and is defined as a non-transitory or non-transitory computer-readable storage medium. Thus, an embodiment or example according to the present disclosure of a computer-readable storage medium does not include a signal, and the embodiment may include one or more non-transitory or non-transitory computer-readable storage media. Thus, in one example, the program may be recorded on a computer-readable storage medium and be structurally and functionally interrelated with the medium.

[0191] The computer 1010 may communicate with one or more external devices 1074, such as a keyboard, pointing device, display 1080, etc., one or more devices that allow a user to interact with the computer 1010, or any device that allows the computer 1010 to communicate with one or more other computing devices (e.g., a network card, modem, etc.), or a combination thereof. Such communication may occur via an input / output (I / O) interface 1022. Still further, the computer 1010 can communicate with one or more networks 1200, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter / interface 1026. As shown, the network adapter 1026 communicates with other components of the computer 1010 via the bus 1014. It should be understood that other hardware and / or software components, not shown, may be used in conjunction with the computer 1010. Examples include, but are not limited to, microcode, device drivers 1024, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.

[0192] It is understood that a computer, or a program running on computer 1010, may communicate with a server, embodied as server 1100, over one or more communications networks, embodied as communications network 1200. Communications network 1200 may include transmission media and network links, including, for example, wireless, wired, or fiber optic, as well as routers, firewalls, switches, and gateway computers. A communications network may include connections such as wired, wireless communications links, or fiber optic cables. A communications network may represent a worldwide collection of networks and gateways, such as the Internet, that communicate with each other using various protocols, such as Lightweight Directory Access Protocol (LDAP), Transport Control Protocol / Internet Protocol (TCP / IP), Hypertext Transport Protocol (HTTP), Wireless Application Protocol (WAP), etc. A network may include several different types of networks, such as, for example, an intranet, a local area network (LAN), or a wide area network (WAN).

[0193] In one example, a computer may use a network that may access websites on the web (World Wide Web) using the Internet. In one embodiment, a computer 1010 including a mobile device may use a communication system or network 1200 that may include the Internet or a public switched telephone network (PSTN), e.g., a cellular network. The PSTN may include telephone lines, fiber optic cables, microwave transmission links, cellular networks, and communications satellites. The Internet may facilitate numerous searching and texting techniques, such as sending queries to a search engine via text message (SMS), multimedia messaging service (MMS) (related to SMS), email, or a web browser, using, for example, a mobile phone or laptop computer. The search engine may obtain search results, i.e., links to websites, documents, or other downloadable data corresponding to the query, and similarly provide the search results to a user via the device, e.g., as a search result web page.

[0194] 14, an exemplary system 1500 for use with embodiments of the present disclosure is illustrated. The system 1500 includes multiple components and elements connected via a system bus 1504 (also referred to as a bus). At least one processor (CPU) 1510 is connected to the other components via the system bus 1504. A cache 1570, a read-only memory (ROM) 1512, a random access memory (RAM) 1514, an input / output (I / O) adapter 1520, an audio adapter 1530, a network adapter 1540, a user interface adapter 1552, a display adapter 1560, and a display device 1562 are also operatively coupled to the system bus 1504 of the system 1500.

[0195] One or more storage devices 1522 are operatively coupled to the system bus 1504 by an I / O adapter 1520. The storage devices 1522 may be, for example, disk storage devices (e.g., magnetic or optical disk storage devices), solid-state magnetic devices, etc. The storage devices 1522 may be the same type of storage device or different types of storage devices. The storage devices may include, for example, but are not limited to, hard drives or flash memory and may be used to store one or more programs 1524 or applications 1526. The programs and applications are shown as general components executable by the processor 1510. The programs 1524 and / or applications 1526 may include all or part of a program or application discussed in this disclosure, but similarly, the reverse may be true, i.e., the programs 1524 and applications 1526 may be part of other applications or programs discussed in this disclosure. The storage devices may be in communication with the control system 70 having various functions described in this disclosure.

[0196] Speakers 1532 are operatively coupled to the system bus 1504 by an audio adapter 1530. Transceiver 1542 is operatively coupled to the system bus 1504 by a network adapter 1540. Display 1562 is operatively coupled to the system bus 1504 by a display adapter 1560.

[0197] One or more user input devices 1550 are operatively coupled to the system bus 1504 by a user interface adapter 1552. The user input device 1550 may be, for example, any of a keyboard, a mouse, a keypad, an image capture device, a motion sensing device, a microphone, a device incorporating the functionality of at least two of the foregoing devices, etc. Other types of input devices may be used while maintaining the spirit of the invention. The user input devices 1550 may be the same type of user input device or different types of user input devices. The user input devices 1550 are used to input and output information to and from the system 1500.

[0198] The present invention may be a system, method, or computer program product, or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium (or multiple computer-readable storage media) having computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0199] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves on which instructions are recorded, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted through wires.

[0200] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may comprise copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to a computer-readable storage medium in the respective computing / processing device for storage.

[0201] The computer-readable program instructions for carrying out the operations of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk® or C++, and procedural programming languages ​​such as the “C” programming language or similar programming languages. The computer-readable program instructions may run entirely on the user's computer, as a standalone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions to personalize the electronic circuitry by utilizing state information of the computer readable program instructions to perform aspects of the present invention.

[0202] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations or block diagrams or combinations thereof, and combinations of blocks in the flowchart illustrations or block diagrams or combinations thereof, can be implemented by computer-readable program instructions.

[0203] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium capable of instructing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium having stored thereon instructions comprises an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0204] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable apparatus, or other device to perform a series of operational steps to create a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0205] The flowcharts and block diagrams in the figures of this disclosure illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions, that implement a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may actually be realized as a single step, executed concurrently, substantially concurrently, partially, or fully in a time-overlapping manner, or the blocks may possibly be executed in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.

[0206] Although this disclosure includes detailed descriptions of cloud computing, it should be understood that implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0207] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0208] The characteristics are as follows: On-Demand Self-Service: Cloud consumers can unilaterally provision computing capacity, such as server time and network storage, automatically as needed, without the need for human interaction with the service provider. Wide network access: The capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (eg, cell phones, laptops, and PDAs). Resource Pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated according to demand. There is location independence in that consumers generally have no control or knowledge over the exact location of the provided resources, but may be able to specify location at a higher level of abstraction (e.g., country, state, or data center). Rapid Elasticity: This capacity can be rapidly and elastically provisioned, sometimes automatically, to quickly scale out, and rapidly released to quickly scale in. To the consumer, the capacity available for provisioning often appears unlimited, and can be purchased in any quantity at any point in time. Measured Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the services utilized.

[0209] The service model is as follows: Software as a Service (SaaS): The consumer is offered the ability to use a provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through a thin-client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings. Platform as a Service (PaaS): The ability offered to consumers is to deploy applications they create or acquire, written using programming languages ​​and tools supported by the provider, on a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but does control the deployed applications and, in some cases, the application hosting environment configuration. Infrastructure as a Service (IaaS): The ability offered to consumers is to provision processing, storage, network, and other basic computing resources, upon which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but rather controls the operating systems, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0210] The deployment model is as follows: Private Cloud: This cloud infrastructure operates solely for an organization. It may be managed by that organization or a third party and may exist on-premise or off-premise. Community Cloud: This cloud infrastructure is shared by several organizations and supports a specific community with shared interests (e.g., mission, security requirements, policy and compliance considerations). It can be managed by the organization or a third party and can reside on-premises or off-premises. Public cloud: The cloud infrastructure is made available to the general public or large industry organizations and is owned by an organization that sells cloud services. Hybrid cloud: The cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain unique entities but are tied together by standardized or proprietary technologies that allow data and application portability (e.g., cloud bursting for load balancing between clouds).

[0211] Cloud computing environments are service-oriented, emphasizing statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0212] 15 , an exemplary cloud computing environment 2050 is shown. As shown, the cloud computing environment 2050 comprises one or more cloud computing nodes 2010 that may communicate with local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or mobile phone 2054A, a desktop computer 2054B, a laptop computer 2054C, or an automobile computer system 2054N, or a combination thereof. The nodes 2010 may communicate with each other. The nodes 2010 may be physically or virtually grouped in one or more networks (not shown), such as a private cloud, a community cloud, a public cloud, or a hybrid cloud, or a combination thereof, as described herein above. This enables the cloud computing environment 2050 to provide infrastructure, a platform, or software, or a combination thereof, as a service without the cloud consumer having to maintain resources on their local computing device. It should be understood that the types of computing devices 2054A-N shown in FIG. 15 are intended as examples only, and that the computing node 2010 and cloud computing environment 2050 can communicate with any type of computerized device through any type of network or network-addressable connection (e.g., using a web browser) or combination thereof.

[0213] Referring now to Figure 16, a set of functional abstraction layers provided by cloud computing environment 2050 (Figure 15) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 16 are intended to be exemplary only, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0214] Hardware and software layer 2060 includes hardware components and software components. Examples of hardware components include mainframe 2061, RISC (reduced instruction set computer) architecture-based servers 2062, servers 2063, blade servers 2064, storage devices 2065, and networks and networking components 2066. In some embodiments, software components include network application server software 2067 and database software 2068.

[0215] Virtualization layer 2070 provides an abstraction layer from which the following examples of virtual entities may be provided: virtual servers 2071, virtual storage 2072, virtual networks including virtual private networks 2073, virtual applications and operating systems 2074, and virtual clients 2075.

[0216] In one example, management layer 2080 may provide the following functions: Resource provisioning 2081 provides dynamic procurement of computing and other resources used to execute tasks within the cloud computing environment. Metering and pricing 2082 tracks costs as resources are utilized within the cloud computing environment and provides billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection for data and other resources. User portal 2083 provides access to the cloud computing environment to consumers and system administrators. Service level management 2084 provides cloud computing resource allocation and management to ensure required service levels are met. Service level agreement (SLA) planning and fulfillment 2085 provides proactive provisioning and procurement of cloud computing resources to anticipate future requirements according to SLAs.

[0217] Workload tier 2090 provides examples of functionality for which a cloud computing environment may be utilized. Examples of workloads and functionality that may be provided from this tier include mapping and navigation 2091; software development and lifecycle management 2092; virtual classroom instructional delivery 2093; data analytics processing 2094; transaction processing 2095; and speech recognition 2096, e.g., more specifically, using dilation of speech content for speech recognition on an audio input from a separated audio input.

[0218] The descriptions of various embodiments of the present invention have been presented for illustrative purposes and are not intended to be comprehensive or limited to the disclosed embodiments. Similarly, examples of features or functions of embodiments of the present disclosure described herein, whether used to describe a particular embodiment or listed as examples, are not intended to limit the embodiments of the present disclosure described herein or to limit the disclosure to the examples described herein. Such examples are intended to be examples or illustrative and not exhaustive. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein was selected to best explain the principles of the embodiments, practical applications, or technical improvements over technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A computer-implemented method for speech recognition using dilation of speech content from an audio input, comprising: receiving an audio input in a CNN (Convolutional Neural Network) and receiving a predicted change to the audio input based on external noise, the audio input having speech content from a speaker, and the external noise being included in the audio input; applying diarization to the audio input in the CNN to receive dilation parameters for the speaker's speech content and generate a CNN output; determining a resulting dilation from the CNN output, the resulting dilation of the CNN output comprising separating sounds of the audio input; determining a word error rate on the dilated CNN output to determine the accuracy of the speech-to-text output; setting an adjustment parameter to vary the range of the dilation based on the word error rate; adjusting the resulting dilation of the CNN output based on the adjustment parameters to reduce the word error rate; A method for providing the above.

2. 10. The method of claim 1, further comprising identifying speech content from the speaker from the audio input based on a learning correction applied to the resulting dilation adjusted for the speaker.

3. The method of claim 2 further comprising generating text from the identified audio content.

4. The method of claim 1 , wherein the audio input and the predicted modification are received without dilation of the audio content.

5. 5. The method of claim 1, further comprising using a grid search to adjust the resulting dilation of the audio input based on the word error rate to reduce the word error rate.

6. receiving, at a computer, expected audio input for a speaker, the expected audio input including speech content for the speaker; generating an environmental stimulus audio input for the expected audio input; predicting changes in the audio input for the speaker based on environmental stimulus audio input; The method of claim 1 , further comprising:

7. sharing the resulting adjusted dilation of the predicted audio input with a social network; generating a learned correction from said sharing of said adjusted resulting dilation of said expected audio input; applying the learning modifications to the resulting adjusted dilation for the speaker; identifying speech content from the speaker from the audio input based on the learning corrections applied to the resulting dilation adjusted for the speaker; generating text from the identified audio content; The method of claim 6 further comprising:

8. sharing the resulting adjusted dilation of the predicted audio input with a social network; generating a learned correction from said sharing of said adjusted resulting dilation of said expected audio input; applying the learned corrections to the resulting dilation adjusted for the speaker; The method of claim 6 further comprising:

9. 9. The method of claim 8, further comprising identifying speech content from the speaker from the audio input based on the learning modifications applied to the resulting dilation adjusted for the speaker.

10. The method of claim 9 further comprising generating text from the identified audio content.

11. A method according to any one of claims 1 to 10, further comprising deriving the dilation parameters from audio input from multiple speakers.

12. 1. A system for speech recognition using dilation of speech content from an audio input, comprising: a computer system; The present invention includes a computer processor, a computer readable storage medium, and program instructions stored on the computer readable storage medium and executable by the computer processor, the program instructions causing the computer system to perform the following functions: In a CNN (Convolutional Neural Network), a function receives an audio input and receives a predicted change to the audio input based on external noise, the audio input having speech content from a speaker, and the external noise is included in the audio input; a function in the CNN to apply diarization to the audio input and receive dilation parameters for the speaker's speech content to generate a CNN output; a function for determining a resulting dilation from the CNN output, the resulting dilation of the CNN output comprising separating sounds of the audio input; determining a word error rate on the dilated CNN output to determine the accuracy of the speech-to-text output; a function of setting an adjustment parameter to change the range of the dilation based on the word error rate; and adjusting the resulting dilation of the CNN output based on the tuning parameters to reduce the word error rate. A system that executes the following.

13. 13. The system of claim 12, further comprising identifying speech content from the speaker from the audio input based on a learning correction applied to the resulting dilation adjusted for the speaker.

14. The system of claim 13 , further comprising generating text from the identified audio content.

15. 15. The system of claim 12, wherein the audio input and the predicted modifications are received without dilation of the audio content.

16. 16. The system of claim 12, further comprising adjusting the resulting dilation of the audio input based on the word error rate using a grid search to reduce the word error rate.

17. receiving at a computer expected audio input for a speaker, the expected audio input including speech content for the speaker; generating an environmental stimulus audio input for the expected audio input; and Predicting changes in the audio input for the speaker based on environmental stimulus audio input. The system of claim 12 further comprising:

18. Sharing the resulting adjusted dilation of the predicted audio input with a social network; generating a learned correction from said sharing of said adjusted resulting dilation of said expected audio input; applying the learning corrections to the resulting dilation adjusted for the speaker; identifying speech content from the speaker from the audio input based on the learning corrections applied to the adjusted resulting dilation for the speaker; and generating text from the identified audio content; 20. The system of claim 17, further comprising:

19. Sharing the resulting adjusted dilation of the predicted audio input with a social network; generating a learned correction from said sharing of said adjusted resulting dilation of said expected audio input; applying the learning corrections to the resulting dilation adjusted for the speaker.

20. The system of claim 17, further comprising:

20. 20. The system of claim 19, further comprising identifying speech content from the speaker from the audio input based on the learning modifications applied to the adjusted resulting dilation for the speaker.

21. The system of claim 20, further comprising generating text from the identified audio content.

22. A computer program for speech recognition using dilation of speech content from an audio input, the computer program comprising program instructions executable by a computer to cause the computer to perform a function, the function comprising: In a CNN (Convolutional Neural Network), a function receives an audio input and receives a predicted change to the audio input based on external noise, the audio input having speech content from a speaker, and the external noise is included in the audio input; and a function in the CNN to apply diarization to the audio input and receive dilation parameters for the speaker's speech content to generate a CNN output; a function for determining a resulting dilation from the CNN output, the resulting dilation of the CNN output including separating sounds of the audio input; determining a word error rate for the dilated CNN output to determine the accuracy of the speech-to-text output; a function of setting an adjustment parameter to change the range of the dilation based on the word error rate; adjusting the resulting dilation of the CNN output based on the adjustment parameters to reduce the word error rate; a computer program comprising:

23. 23. The computer program of claim 22, further comprising identifying speech content from the speaker from the audio input based on a learning correction applied to the resulting dilation adjusted for the speaker.

24. 24. The computer program of claim 23, further comprising generating text from the identified audio content.

Citation Information

Patent Citations

  • Speech recognition device, program, and speech recognition method

    JP2008076730A

  • System and method for maintaining voice-to-voice translation on-site

    JP2011524991A

  • Full-band scalable audio codec

    JP2012032803A

  • Speech language corpus generation device and its program

    JP2017045027A

  • Speech processor and speech processing method

    JP2020134920A