Data analysis and extended speech recognition using interleaved audio input
By using convolutional neural networks (CNNs) to expand speech signals, the challenges of speech recognition under interleaved audio inputs are addressed, word error rates are reduced, and the accuracy of speech conversion is improved.
Patent Information
- Application Number
- CN202180061515.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-09
- Filing Date
- 2021-08-24
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-08-24
AI Technical Summary
Existing technologies face challenges in speech recognition and conversion when dealing with interleaved audio input, especially in the presence of background noise, multiple overlapping speakers, or changes in sound, resulting in high word error rates.
Speech signal expansion is performed using a convolutional neural network (CNN). Expansion parameters are assigned to each speaker through machine learning, and the expansion parameters are predicted by combining ambient noise and speech type. The parameters are then weighted in a group setting to separate the sound in the audio input.
It reduces the word error rate in the speech recognition process and improves the accuracy of converting interleaved audio input to text.
Smart Images

Figure CN116057625B_ABST
Abstract
Description
Background Technology
[0001] This disclosure relates to techniques for using a computer to perform speech recognition on speech content from audio input. More specifically, the audio input includes interleaved speech content and translation or conversion of the speech content to text.
[0002] Computer-based techniques can be used to convert human speech into text. Human speech can include, for example, words spoken individually or in groups, or singing. The conversion of speech output, or the speech output signal converted into text, during human speech can be challenging. For example, speech recognition and conversion can be challenging when the sound is altered or less typical of the pronunciation of a word (phonetics). For example, the sound may be elongated or mixed with one or more other noises. In one example, there may be background noise while a speaker is speaking. In another example, a group of speakers may be speaking, and there may be speaker overlap. In yet another example, background noise may occur while one or more speakers are speaking. In yet another example, a speaker may unintentionally or intentionally alter the typical pronunciation of one or more words for emphasis, or as part of an unconventional or non-standard speech pattern, or as part of an accent. This altered and / or atypical sound when a speaker is speaking results in challenging speech for speech recognition and conversion from speech to text. Summary of the Invention
[0003] This disclosure recognizes the drawbacks and problems associated with current techniques for speech recognition using dilation of speech content from interleaved audio input.
[0004] This invention analyzes speech content from interleaved audio inputs to perform speech recognition on each of multiple speakers and can provide conversion from speech content to text. For example, the challenges of speech recognition and conversion can be overcome when the speech content includes altered and / or atypical sounds from the speakers for speech-to-text recognition and conversion, or when the sounds are altered or less common than typical word pronunciations.
[0005] For example, when an artist sings a song, a problem may arise where some words might be altered or changed in a way that follows harmonics rather than normal pronunciation. In another example, in noisy environments, the mixing of sound waves and sounds can increase the error rate during word transitions. For instance, during large events, the shouting of a large crowd or the noise of a sporting event can mask the speech signal.
[0006] This invention includes speech recognition that uses an expansion of the speech signal, the speech input, to increase the space between samples or speech samples before attempting to recognize words or analyze speech content to recognize one or more words. In one example according to the invention, a convolutional neural network (CNN) with different expansion parameters can be trained and applied to these problems. Furthermore, predicted ambient noise and speech type can indicate which expansion to use. Alternatively, in another example, expansion parameters can be assigned to each speaker via machine learning. In a grouped setting of a session or song, the expansion parameters can be weighted together by the group based on the amplitude of each speaker.
[0007] In one aspect of the invention, a computer-implemented method is provided for speech recognition using expansions of speech content from interleaved audio input. The method includes initiating a learning model to determine expansion parameters for each of a plurality of audible sounds of speech content received as audio input from a plurality of speakers at a computer. The method includes, as part of the learning model, determining changes in each of the plurality of independent sounds in response to an audio stimulus. The independent sounds are derived from the audio input. The method includes applying the expansion parameters separately based on the changes in each independent sound. A voiceprint for each speaker is constructed separately based on the independent sounds and the expansion parameters. The method includes at least partially attributing the speech content to each of the plurality of speakers based on the voiceprint and the independent sounds.
[0008] One advantage of this invention is a reduced word error rate when using speech recognition with multiple speakers to convert speech content from interleaved audio input into text using the method according to the invention.
[0009] In related aspects, the method also includes generating text from the associated speech content.
[0010] In a related aspect, the method also includes using a computer to display text on a screen or monitor that communicates with the computer and / or device.
[0011] In a related aspect, the method also includes sending text to a computer and / or device via an electronic communication system for display on a screen or monitor communicating with the computer and / or device.
[0012] In a related aspect, the method includes displaying text on a screen or monitor that is communicating with a computer or device.
[0013] In related aspects, audio input may include multiple audible sounds, and audio input is received at the computer, the audible sounds including speech content from multiple speakers.
[0014] In a related aspect, the method also includes using a computer to expand the audio input, which includes separating the sound in the audio input.
[0015] In a related aspect, the method also includes using logs to refine the audio input for each of the multiple speakers.
[0016] In related aspects, the learning models include CNNs (convolutional neural networks) for receiving individual sounds and using logs to determine the changes in each individual sound in response to audio stimuli.
[0017] In a related aspect, the method also includes using the log of the audio input to layer the sounds in the refined audio input into independent sounds.
[0018] In a related aspect, the method further includes receiving, at the computer, the audio input comprising the plurality of audible sounds, wherein the audible sounds include speech content from the plurality of speakers. The method includes using the computer to expand the audio input, and the expansion includes separating the sounds in the audio input. The method includes refining the audio input from each of the plurality of speakers using a log, and using the log of the audio input to layer the sounds in the refined audio input into independent sounds.
[0019] In a related aspect, separating sound in audio input includes distinguishing ambient or background sounds from speech from one of multiple speakers.
[0020] In a related aspect, using the logs to refine the audio input of each of the plurality of speakers includes dividing the audio input into isomorphic segments related to speaker identity.
[0021] According to another aspect of the invention, a system for speech recognition utilizes the expansion of speech content from interleaved audio inputs and includes a computer system. The computer system includes a computer processor, a computer-readable storage medium, and program instructions executable by the processor stored on the computer-readable storage medium, causing the computer system to perform functions including: initiating a learning model to determine expansion parameters for each of a plurality of audible sounds of speech content received as audio input from a plurality of speakers at the computer; as part of the learning model, determining each of a plurality of independent sounds in response to changes in an audio stimulus, said independent sounds being derived from the audio input; applying the expansion parameters respectively based on the changes in each of the independent sounds; constructing a voiceprint for each of the speakers respectively based on the independent sounds and the expansion parameters; and at least partially attributing speech content to each of the plurality of speakers based on the voiceprint and the independent sounds respectively.
[0022] One advantage of the present invention is that when using the system according to the invention, speech recognition using multiple speakers converts speech content from interleaved audio input into text, it reduces the word error rate.
[0023] In related aspects, the system also includes generating text from the associated speech content.
[0024] In related aspects, the system also includes the use of a computer to display text on a screen or monitor that communicates with a computer and / or device.
[0025] In related aspects, the system also includes sending text via an electronic communication system to a computer and / or device for display on a screen or monitor communicating with the computer and / or device.
[0026] In related aspects, audio input may include multiple audible sounds, and audio input is received at the computer, the audible sounds including speech content from multiple speakers.
[0027] In related aspects, the system also includes using a computer to expand the audio input, which includes separating the sound in the audio input.
[0028] In related aspects, the system also includes the use of logs to refine the audio input of each of the multiple speakers.
[0029] In related aspects, the learning models include CNNs (convolutional neural networks) for receiving individual sounds and using logs to determine the changes in each individual sound in response to audio stimuli.
[0030] In related aspects, the system also includes using the log of audio input to layer the sounds in the refined audio input into independent sounds.
[0031] In another aspect of the invention, a computer program product for speech recognition utilizes the expansion of speech content from interleaved audio input and includes a computer-readable storage medium having program instructions embodied therein. The program instructions are executable by a computer to cause the computer to perform functions including: initiating a learning model to determine expansion parameters for each of a plurality of audible sounds of speech content received as audio input from a plurality of speakers at the computer; as part of the learning model, determining, for each of a plurality of independent sounds, in response to changes in an audio stimulus derived from the audio input; applying the expansion parameters respectively based on the changes in each of the independent sounds; constructing a voiceprint for each of the speakers respectively based on the independent sounds and the expansion parameters; and attributing the speech content to each of the plurality of speakers respectively, at least in part based on the voiceprint and the independent sounds.
[0032] One advantage of the present invention is that when using the computer program product according to the invention to convert speech content from interleaved audio input into text using speech recognition with multiple speakers, the word error rate is reduced.
[0033] In related aspects, computer program products also include generating text from the associated speech content.
[0034] In related aspects, computer program products also include the use of a computer to display text on a screen or monitor that communicates with a computer and / or device.
[0035] In related aspects, computer program products also include those that transmit text via electronic communication systems to computers and / or devices for display on screens or monitors in communication with the computers and / or devices. Attached Figure Description
[0036] These and other objects, features, and advantages of the present invention will become apparent from the following detailed description of its illustrative embodiments, which is read in conjunction with the accompanying drawings. The various features in the drawings are not to scale, as the illustrations are provided for clarity and to assist those skilled in the art in understanding the invention in conjunction with the detailed description. The drawings will now be discussed.
[0037] Figure 1 This is a schematic block diagram illustrating an overview of a system, system features or components, and method for speech recognition using expanded speech content from interleaved audio inputs according to embodiments of the present disclosure.
[0038] Figure 2 This illustrates the use according to an embodiment of the invention. Figure 1 The flowchart shown illustrates a method implemented according to an embodiment of the present disclosure for speech recognition using an expansion of speech content from interleaved audio input.
[0039] Figure 3 This is a series of tables illustrating embodiments of the extensions of this disclosure.
[0040] Figure 4 This is a flowchart illustrating another embodiment of the method according to the present invention, which uses... Figure 1 The system shown is used to implement speech recognition using an expansion of speech content from interleaved audio input.
[0041] Figure 5 It is from the embodiment of the present invention Figure 4 The flowchart shown continues the flowchart, which depicts Figure 4 The method shown continues.
[0042] Figure 6This is a functional block diagram illustrating a series of operations and functional methods, used to explain the relationship with... Figure 1 , 2 The embodiments shown in 3, 4 and 5 are intended to guide the functional features of this disclosure for speech recognition using an expansion of speech content from interleaved audio input.
[0043] Figure 7 This is a functional block diagram illustrating a series of operations and functional methods, used to explain the relationship with... Figure 1 , 2 The embodiments shown in 3, 4 and 5 are intended to guide the functional features of this disclosure for speech recognition using an expansion of speech content from interleaved audio input.
[0044] Figure 8 This is a schematic block diagram depicting a computer system according to an embodiment of the present invention, which may be wholly or partially incorporated into... Figure 1 In one or more computers or devices shown, and with Figure 1 , 2 The systems and methods shown in 3, 4, 5, 6 and 7 collaborate.
[0045] Figure 9 This is a schematic block diagram depicting system components interconnected using a bus. Components are intended for use, in whole or in part, with embodiments of this disclosure, according to one or more embodiments of this disclosure.
[0046] Figure 10 This is a block diagram describing a cloud computing environment according to an embodiment of the present invention.
[0047] Figure 11 This is a block diagram illustrating an abstract model layer according to an embodiment of the present invention. Detailed Implementation
[0048] The following description, provided with reference to the accompanying drawings, is intended to aid in a full understanding of exemplary embodiments of the invention as defined by the claims and their equivalents. It includes various specific details to aid understanding, but these details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Furthermore, for clarity and brevity, descriptions of well-known functions and structures may be omitted.
[0049] The terms and words used in the following description and claims are not limited to their literal meaning, but are merely used to enable a clear and consistent understanding of the invention. Therefore, it will be apparent to those skilled in the art that the following description of exemplary embodiments of the invention is for illustrative purposes only and is not intended to limit the invention as defined by the appended claims and their equivalents.
[0050] It should be understood that the singular forms “a,” “an,” and “the” include plural referents unless the context clearly indicates otherwise. Thus, for example, unless the context clearly indicates otherwise, reference to “component surface” includes a reference to one or more such surfaces.
[0051] Embodiments of this disclosure analyze speech content from interleaved audio input to provide speech recognition for each of a plurality of speakers, and thereby provide word recognition as well as recognition and conversion from speech content to text. This disclosure enables speech recognition and conversion from speech to text even when the sounds are altered or less common than typical word pronunciations, such as when the speech content includes altered and / or atypical sounds from the speakers.
[0052] Embodiments of the present invention include speech recognition using an expansion of a speech signal or speech input to increase the space between samples or speech samples before attempting to recognize words or analyze speech content to recognize one or more words. In one example according to this disclosure, a convolutional neural network (CNN) with different expansion parameters can be trained and applied to these problems. In another example, predicted ambient noise and speech type can indicate which expansion to use. In yet another example, expansion parameters can be assigned to each speaker via machine learning. In yet another example, in a group setting for a dialogue or song, the expansion parameters can be weighted together by the group based on the amplitude of each speaker.
[0053] Embodiments of this disclosure may thus provide prediction of environmental noise to set expansion parameters. In another example, embodiments of the invention classify speech types (e.g., singing, spoken language) to contribute to the expansion parameters. In another example, embodiments of the invention adapt speech extensions to other independent models. In yet another example, embodiments of this disclosure may include averaging expansion parameters based on speaker logs and group models. Furthermore, in yet another example, embodiments of this disclosure may include social expansion transfer of knowledge.
[0054] Therefore, embodiments of this disclosure include modeling the expansion based on environmental noise and predicting expansion parameters. The expansion extension can be mapped, and further, social aspects can be combined with expansion measures for each person in the conversation.
[0055] refer to Figure 1 and 2 A reference system 10 according to embodiments of the present disclosure is provided. Figure 1 Method 100 () Figure 2 This is used for speech recognition by expanding upon speech content from interleaved audio input. (See reference) Figure 2The method includes a series of operational blocks for implementing an embodiment of the present disclosure. (See reference...) Figure 2 Method 100 includes initiation learning model 320 (see...) Figure 6 () to determine the expansion parameter 324 of each of the plurality of audible sounds 62 of the speech content 64 from the plurality of human speakers 52, which is received at computer 22 as audio input 60, as in block 104.
[0056] refer to Figure 6 The functional system 300 includes components and operations according to embodiments of the present disclosure, and references are made herein. Figure 1 , 2 The methods and systems shown in 3, 4 and 5 are used.
[0057] In one example, a group of speakers can speak together. The audio output from this group of speakers can be received as audio input using a computer or device, such as by using the device's microphone or communicating with the device or computer.
[0058] In one example, a spectrogram can be generated and used as a visual representation of the spectrum of a signal, such as in an audio signal, because it changes over time. A spectrogram can also be called a spectrograph, acoustic graph, or acoustic map.
[0059] A spectrogram can be created, and the DFT (Discrete Fourier Transform) can be applied to identify potential unique loudspeakers. The DFT transforms a finite sequence of equally spaced samples of a function into a sequence of equally spaced samples of the same length in the Discrete-Time Fourier Transform (DTFT), which is a complex-valued function of frequency. Initial expansion variables can be initialized for each DFT estimate.
[0060] In one example, when it is known who is speaking or singing in a group, expansion parameters can be adjusted or specified based on that information. Such identification information can be collected, for example, from input or observations from social media.
[0061] Multiple audible sounds may include, for example, one or more human speakers 14 or users, as multiple human speakers 52 or users in the vicinity of 50, speaking and producing audible sounds 62. The audible sounds may include, for example, human speech in a conversation, individual speech, singing, a group of speakers singing, etc. And the audible sounds 62 thereby produce and include speech content 64.
[0062] Audible sound can be received as audio input 60 at the computer 22 via a microphone in the computer or device 20 such as a mobile device, and the computer can send audio files, either alone or in combination with control devices of the control system 70 (via a communication network 45, such as the Internet, for processing according to the techniques of this disclosure), to another computer 72 or a server, such as a remote computer or server. In another example, audible sound in the audio file can be processed locally on the computer according to the techniques of this disclosure, and / or in combination with processing on a remote computer or server.
[0063] The learning model 320 may include machine learning using parameters. For example, machine learning can be used to assign expansion parameters to each of multiple speakers or users.
[0064] The expansion of sound for speech recognition can be defined as increasing the space between sounds or sound samples. In this disclosure, the expansion is performed before attempting to recognize words from the sound.
[0065] The expansion parameter can include a specified amount of space between sound samples, or a specified...
[0066] The spatial range between sound samples. Expansion variables can be assigned to each potential speaker and used in the learning model.
[0067] In one example, refer to Figure 3 Representative images 150 in Tables 154, 158, and 162 depict the expansion of image 166, which can be, for example, a sound image layered into blocks or sound samples 168. Image 150 depicts variations in the expansion parameter D. In Table 154, the expansion parameter 172 equals 1, and image 166 has no spacing. In the second table 158, the expansion parameter 174 equals 2, and the image has sound samples 168 with spaces 180 between the samples. In the third table 162, the expansion parameter 176 equals 3, and the image has sound samples 168 with even more space 180 between the samples.
[0068] This method includes, as part of a learning model, determining changes in each of a plurality of independent sounds in response to an audio stimulus derived from an audio input, as shown in block 108. For example, the audio stimulus may include an environmental stimulus. Changes in the sound, or independent sound changes, may be determined in response to the environmental stimulus 322.
[0069] In one embodiment, logging can be used to refine the audio input for each of multiple speakers. For example, logging may include or contain a process of segmenting the input audio stream into segments corresponding to speaker identities, and in one example, the segments may be homogeneous. This logging signal can be used to layer the audio input into independent sounds.
[0070] Logs can be used with deep learning to refine audio inputs attributable to each of multiple speakers. In one example, if errors exist from the DFT of speaker identification and / or the deep learning method, the log parameters can be averaged together.
[0071] In one example, a voiceprint can be constructed. In another example, environmental stimuli can be played, and determinations can be made about how the stratified data changes. The expansion parameters can be modified based on changes in the audio input data (e.g., speech data). For example, if the speech is further prolonged, the expansion parameters can be increased. Additionally, the expansion parameters relative to each speaker can be based on the correlation factor of the independent signals. The correlation coefficient (R-value) is the value given in the summary table in the regression output. The squared R-value is called the coefficient of determination; that is, R multiplied by R' to obtain the squared R-value. The coefficient of determination is the square of the correlation coefficient.
[0072] In one example, the R-squared correlation metric determines how to group the most relevant pairs of speakers together. For instance, the R-squared metric can be shifted between 0 and 0.5, such that the most frequently paired speakers will contribute 50% of the adjusted expansion.
[0073] The method involves applying expansion parameters based on the changes in each individual sound, as shown in block 112.
[0074] The method includes, as in block 116, constructing the soundprint of each speaker separately based on independent sound and expansion parameters.
[0075] The method includes at least in part attributing speech content to each of a plurality of speakers based on voiceprints and individual voices, as in block 120.
[0076] This method may include generating text from the associated speech content, as in block 124.
[0077] If, as determined in block 126, the generated text is to be displayed locally, for example, on the local computer, the method continues to block 130. If, as determined in block 126, the generated text is not to be displayed locally on a device or computer monitor, the method continues to block 128.
[0078] The method includes displaying text on a screen or monitor in communication with a computer or device in response to a display text location determined as in block 126, as in block 130.
[0079] The method may include, in response to the determination in block 126 that the text is not to be displayed locally, sending the text via an electronic communication system to a computer and / or device for display on a screen or monitor in communication with the computer and / or device, as in block 128. The method may also continue to display the text on a screen or monitor in communication with the computer or device, as in block 130.
[0080] The method may include a learning model 320, which includes a CNN 326 (convolutional neural network) for receiving independent sounds and using logs to determine the changes in each independent sound in response to audio stimuli.
[0081] A CNN (Convolutional Neural Network) can be at least a part of deep learning, and CNNs are a class of deep neural networks. A CNN involves mathematical operations, typically defined as producing a third function, called convolution. Convolution is a specialized linear operation. Therefore, a convolutional network is a neural network that uses convolution instead of general matrix multiplication in at least one of its multiple layers.
[0082] The method may include audio input, which may include multiple audible sounds, and the audio input may be received at a computer. Furthermore, the audible sounds may include speech content from multiple speakers.
[0083] This method may include using a computer to expand the audio input. The expansion may include separating the sound in the audio input.
[0084] The method may include using a sound log 308 to refine the expanded audio input 302 for each of the multiple speakers.
[0085] The method may also include using the log of the audio input to layer the sounds in the refined audio input 304 into independent sounds 310.
[0086] The method may include separating sound from audio input, including distinguishing ambient or background sounds from speech from one of a plurality of speakers.
[0087] This method may include using logs to refine the audio input for each of a plurality of speakers, which may include dividing the audio input into isomorphic segments related to speaker identity.
[0088] In another embodiment according to this disclosure, reference is made to Figure 4A computer-implemented method 200 for speech recognition using an expansion of speech content from interleaved audio input includes receiving at a computer an audio input comprising a plurality of audible sounds, the audible sounds including speech content from a plurality of speakers, as in block 204. Figure 4 and 5 The operation block of method 200 shown can be similar to Figure 2 The operation block shown. Figure 4 and 5 The method shown is intended to be another example embodiment that may include the aspects / operations shown and previously discussed in this disclosure.
[0089] Method 200 includes using a computer to expand the audio input, wherein expansion includes separating the sound in the audio input, as in block 208.
[0090] Method 200 includes using logs to refine the audio input of each of the multiple speakers, as in block 212. The method may include using the logs of the audio inputs to layer the sounds in the audio inputs into individual sounds, as in block 216.
[0091] Method 200 includes initiating a learning model to determine the expansion parameters for each audible sound, as in block 220.
[0092] Method 200 may include a learning model, including a CNN (convolutional neural network), for receiving independent sounds and using logs to determine changes in each independent sound in response to audio stimuli, as in block 222.
[0093] Method 200 includes, as part of a learning model, determining each of a plurality of independent sounds in response to changes in an audio stimulus, as in block 224.
[0094] Method 200 includes applying expansion parameters separately based on the changes in each individual sound, as in block 228.
[0095] Method 200 includes, as in block 232, constructing a sound signature 330 for each of the loudspeakers 52 based on the independent sound 310 and the expansion parameter 324, respectively.
[0096] Method 200 includes at least in part attributing speech content to each of a plurality of speakers, based on voiceprints and individual sounds, as described in block 236. The attributed speech content 332 can be used to generate text.
[0097] refer to Figure 5 Method 200 includes generating text 334 from the associated speech content 332, as in block 240.
[0098] Method 200 also includes sending text via an electronic communication system to a computer and / or device for display on a screen or monitor in communication with the computer and / or device, as described in block 244. In another example, communication can be implemented from a group consisting of: SMS, email, instant messaging, and navigation software. These examples are intended to be exemplary and not exhaustive.
[0099] Method 200 may also include displaying text on a screen or monitor in communication with a computer or device, as in block 248.
[0100] refer to Figure 7 According to embodiments of this disclosure and indicating and supporting the functionality of the embodiments discussed herein, system 400 includes components and operations for performing speech recognition using expansions of speech content from interleaved audio inputs. System 400 includes a set of human speakers 402 that output audio outputs. As in block 404, the audio outputs are received to learn each distinct signal using expansions. The system may learn the expansions of the audio input signals based on logs, as in block 406.
[0101] The system includes using logs to hierarchically structure audio input signals, as in block 410. The system includes playing sounds, such as environmental stimuli, to group the logged audio input signals, as in block 412. As in block 414, the system includes setting individual and group expansions based on environmental stimuli. The system includes generating audio output as in block 416. The system includes using audio output 416 to generate text output based on expansions and environmental stimuli, as in block 418.
[0102] In one example, the system may use a speaker signal prediction technique or method / system as described in block 450 to predict the speaker signal, in one example by predicting how the speaker signal will change based on external noise. Such prediction is not the focus of this disclosure.
[0103] exist Figure 1 and 2 In the embodiments shown in this disclosure, the computer may be a remote computer or part of a remote server, such as remote server 1100. Figure 8 In another example, computer 72 may be part of control system 70 and provide the execution of the functions of this disclosure. In another embodiment, computer 22 may be part of mobile device 20 and provide the execution of the functions of this disclosure. In yet another embodiment, the portion of the execution of the functions of this disclosure may be shared between the control system computer and the mobile device computer; for example, the control system may serve as the backend of one or more programs embodying this disclosure, and the mobile device computer may serve as the frontend of one or more programs.
[0104] The computer may be part of a mobile device or a remote computer communicating with the mobile device. In another example, the mobile device and the remote computer may work together to perform features of the methods described herein using stored program code or instructions, thereby implementing the methods of this disclosure. In one example, the mobile device 20 may include a computer 22 having a processor 15 and a storage medium 34 storing an application 40, which may contain program instructions for performing features of this disclosure using the processor 15. In another example, the mobile device 20 application 40 may have program instructions that execute a front-end of a software application for incorporating features of the methods of this disclosure in the program instructions, while one or more back-end programs 74 of the software application stored on a computer 72 of the control system 70 communicate with the mobile device computer and perform other features of the method. The control system 70 and the mobile device 20 may communicate using a communication network 45, such as the Internet.
[0105] Therefore, the method 100 according to an embodiment of the invention can be incorporated into one or more computer programs or applications 40 stored on electronic storage medium 34 and can be executed by processor 15 as part of a computer on mobile device 20. For example, a human speaker or user 14 has device 20, and said device can communicate with control system 70. Other users (not shown) may have similar devices and similarly communicate with control system. The application may be stored wholly or partially on the computer or mobile device, and also on the control system that communicates with the device, for example, using a communication network 45 such as the Internet. It is conceivable that the application can access all or part of the program instructions to implement the methods of this disclosure. The program or application can communicate with and access data to a remote computer system via communication network 45 (e.g., the Internet) and cooperate with programs stored on the remote computer system. Such interactions and mechanisms are described in more detail herein, and reference is made to components of a computer system, such as computer-readable storage media, which... Figure 8 One embodiment is shown, and is described in more detail with reference to one or more computer systems 1010.
[0106] Therefore, in one example, the control system 70 communicates with one or more devices 20, and devices 20 may include an application 40. Devices 20 communicate with the control system 70 using a communication network 45.
[0107] In another example, the control system 70 may have a front-end computer, such as device 20, belonging to one or more users, and a back-end computer embodied as the control system.
[0108] In addition, refer to Figure 1Device 20 may include a computer 22, a computer-readable storage medium 34, an operating system and / or programs and / or software applications 40, which may include program instructions executable using processor 15. These features are... Figure 1 As shown in, and in Figure 8 In an embodiment of a computer system shown with reference to one or more computer systems 1010, it may include one or more general-purpose computer components 1010.
[0109] The method according to this disclosure may include a computer as part of a control system for implementing the features of the method according to this disclosure. In another example, the computer as part of the control system may cooperate with a mobile device computer to implement the features of the method according to this disclosure. In yet another example, the computer for implementing the features of the method may be part of a mobile device and thus implement the method locally.
[0110] It should be understood that Figure 6 and 7 The features shown are functional representations of the features of this disclosure. For illustrative purposes, such features are shown in embodiments of the systems and methods of this disclosure to illustrate the function of the features of this disclosure.
[0111] Specifically, regarding the control system 70, one or more user devices 20 can communicate with the control system 70 via the communication network 50. Figure 1 In the embodiment of the control system shown, the control system 70 includes a computer 72 having a database 76 and one or more programs 74 stored on a computer-readable storage medium 73. Figure 1 In the embodiments of this disclosure shown, device 20 communicates with control system 70 and one or more programs 74 stored on computer-readable storage medium 73. The control system includes computer 72 with processor 75, which also has access to database 76.
[0112] The control system 70 may include a storage medium 80 for maintaining user and device registration 82 for analyzing audio input. Such registration may include a user profile 83, which may include user data provided by the user for referencing registration and account creation. In one embodiment, the methods and systems disclosed herein include a control system (generally referred to as a backend) integrated with and cooperating with a frontend of the methods and systems, which may be an application 40. In one example, application 40 is stored on a device, such as device 20, and has access to data and additional programs at the application backend, such as control system 70.
[0113] The control system can also be part of a software application implementation, and / or represent a software application having a front-end user portion and a back-end portion that provide functionality. In one embodiment, the methods and systems incorporating this disclosure include a control system (which may generally be referred to as the back-end of a software application, which is part of the methods and systems incorporating embodiments of this application) combined and cooperating at the device with a front-end of a software application incorporating another portion of the methods and systems incorporating this application, such as in... Figure 1 The example shown is of a device 20 with application 40. Application 40 is stored on device 20 and has access to data and additional programs at the application backend, such as program 74 stored in control system 70.
[0114] Program 74 (one or more) may include, in whole or in part, a series of executable steps for implementing the methods of this disclosure. Programs incorporating this method may be stored, in whole or in part, on a computer-readable storage medium on the control system or on the device 20. It is contemplated that the control system 70 may not only store user profiles, but in one embodiment, may interact with a website to be viewed on a device display or, in another example, on the Internet, and receive user input related to the methods and systems of this disclosure. It is understood that... Figure 1 One or more profiles 83 are described; however, the method may include multiple profiles, users, registration, etc. It is conceivable that multiple users or a group of users can use the control system used by the method and system according to this disclosure to register and provide profiles.
[0115] Regarding the data collection for this disclosure, the uploading or generation of such profiles is voluntary and therefore initiated and approved by one or more users. Users may thus opt to join and create an account with a profile according to this disclosure. Similarly, data received by the system, entered, or received as input is voluntary and therefore initiated and approved by one or more users. Users may thus opt to join to enter data according to this disclosure. This user approval also includes the option for the user to cancel such profiles or accounts and / or enter data, and thus opt out of capturing communications and data at the user's discretion. Furthermore, any stored or collected data is understood to be intended to be securely stored and unavailable without user authorization, and unavailable to public and / or unauthorized users. Such stored data is understood to be deleted at the user's request and in a secure manner. Furthermore, according to this disclosure, any use of such stored data is understood to be solely with the user's authorization and consent.
[0116] In one or more embodiments of the present invention, a user may choose to join or register with the control system, voluntarily providing data and / or information during the process with the user's consent and authorization, wherein the data is stored and used in one or more methods of this disclosure. Furthermore, a user may register one or more user electronic devices for use with one or more methods and systems according to this disclosure. As part of registration, the user may also identify and authorize access to one or more activities or other systems (e.g., audio and / or video systems). Such optional joining and authorization of data collection and / or storage is voluntary, and the user may request deletion of data (including profiles and / or profile data), unregister, and / or opt out of any registration. It is understood that such opt-out includes handling all data in a secure manner.
[0117] In one example, artificial intelligence (AI) can be used in whole or in part to learn the model and determine the expansion parameters.
[0118] In another example, the control system 70 can be all or part of an artificial intelligence (AI) system. For example, the control system can be one or more components of an AI system.
[0119] It should also be understood that the method 100 according to embodiments of this disclosure can be incorporated into an (artificial intelligence) AI device that can communicate with a corresponding AI system and a corresponding AI system platform. Thus, as described above, such a program or application in conjunction with the methods of this disclosure can be part of an AI system. In one embodiment of the invention, it is conceivable that a control system can communicate with an AI system, or in another example, can be part of an AI system. The control system can also represent a software application having a front-end user portion and a back-end portion providing functionality, which in one or more examples can interact with, contain, or be part of a larger system such as an AI system. In one example, the AI device can be associated with an AI system that can be wholly or partially a control system and / or a content delivery system, and is remote from the AI device. Such an AI system can be represented by one or more servers that store programs on a computer-readable medium that can communicate with one or more AI devices. The AI system can communicate with a control system, and in one or more embodiments, the control system can be wholly or partially of the AI system, or vice versa.
[0120] It should be understood that, as discussed herein, downloads or downloadable data can be initiated using voice commands or by using a mouse, touchscreen, etc. In such examples, the mobile device may be user-initiated, or the AI device may be used with the user's consent and permission. Other examples of AI devices include devices that include microphones, speakers, and access to cellular or mobile networks, communication networks, or the Internet, such as vehicles with computers and cellular or satellite communications, or, in another instance, IoT (Internet of Things) devices with cellular or Internet access, such as appliances.
[0121] As you can understand, the sets used in this article are collections of different objects or elements. The objects or elements that make up a set can be anything, such as numbers, letters of the alphabet, other sets, etc. It should also be understood that a set can be a single element, such as an object or a number; in other words, a collection of a single element.
[0122] refer to Figure 8Embodiments of the system or computer environment 1000 according to this disclosure include a computer system 1010 shown in the form of a general-purpose computing device. Method 100 may be implemented, for example, in program 1060, including program instructions implemented on a computer-readable storage device, or in a computer-readable storage medium, such as generally referred to as computer memory 1030, and more specifically, as computer-readable storage medium 1050. Such memory and / or computer-readable storage medium includes non-volatile memory or non-volatile storage devices, also known and referred to as non-transient computer-readable storage media, or non-transient computer-readable storage medium. For example, such non-volatile memory may also be a disk storage device including one or more hard disk drives. For example, memory 1030 may include storage medium 1034 such as RAM (random access memory) or ROM (read-only memory), and cache memory 1038. Program 1060 may be executed by processor 1020 of computer system 1010 (to execute program steps, code, or program code). Additional data storage devices may also be implemented as a database 1110 including data 1114. Computer system 1010 and program 1060 are general representations of a computer and a program that may be local to a user or provided as a remote service (e.g., as a cloud-based service), and may be provided using a website accessible through communication network 1200 (e.g., interacting with a network, the Internet, or a cloud service), as further exemplified. It should be understood that computer system 1010 herein also generally refers to a computer device or a computer included in a device such as a laptop or desktop computer, or one or more servers, either alone or as part of a data center. The computer system may include network adapter / interface 1026 and one or more input / output (I / O) interfaces 1022. I / O interface 1022 allows input and output of data with external devices 1074 that can be connected to the computer system. Network adapter / interface 1026 can provide communication between the computer system and a network generally represented as communication network 1200.
[0123] Computer 1010 can be described in the general context of executable instructions of a computer system, such as program modules executed by the computer system. Typically, program modules can include routines, programs, objects, components, logic, data structures, etc., that perform specific tasks or implement specific abstract data types. Method steps and system components and techniques can be implemented in modules of program 1060 that perform tasks for each step of the method and system. These modules are generally represented in the diagram as program modules 1064. Program 1060 and program modules 1064 can execute specific steps, routines, subroutines, instructions, or code of a program.
[0124] The methods disclosed herein can run locally on a device such as a mobile device, or as a service on a server 1100, which may be remote and accessible via the communication network 1200. The program or executable instructions may also be provided as a service by a provider. The computer 1010 can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via the communication network 1200. In a distributed cloud computing environment, program modules can reside in local and remote computer system storage media, including memory storage devices.
[0125] More specifically, the system or computer environment 1000 includes a computer system 1010 shown in the form of a general-purpose computing device with illustrative peripherals. Components of the computer system 1010 may include, but are not limited to, one or more processors or processing units 1020, system memory 1030, and a bus 1014 that couples various system components, including system memory 1030, to the processor 1020.
[0126] Bus 1014 represents one or more of several types of bus architectures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, and processor or local buses using any of the various bus architectures. By way of example and not limitation, these architectures include Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MCA) buses, Enhanced ISA (EISA) buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses.
[0127] Computer 1010 may include a variety of computer-readable media. Such media may be any available media accessible by computer 1010 (e.g., a computer system or server) and may include volatile and non-volatile media, as well as removable and non-removable media. Computer memory 1030 may include additional computer-readable media in the form of volatile memory, such as random access memory (RAM) 1034 and / or cache 1038. Computer 1010 may also include other removable / non-removable, volatile / non-volatile computer storage media, such as portable computer-readable storage media 1072 in one example. In one embodiment, computer-readable storage media 1050 may be provided for reading from and writing to non-removable, non-volatile magnetic media. Computer-readable storage media 1050 may be implemented, for example, as a hard disk drive. Additional memory and data storage devices may be provided, for example, as a storage system 1110 (e.g., a database) for storing data 1114 and communicating with processing unit 1020. The database may be stored on or part of server 1100. Although not shown, a disk drive for reading from and writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical media may be provided. In such an example, each may be connected to bus 1014 via one or more data media interfaces. As will be further described below, memory 1030 may include at least one program product that may include one or more program modules configured to perform the functions of embodiments of the present invention.
[0128] For example, the methods described in this disclosure may be embodied in one or more computer programs, generally referred to as program 1060, and may be stored in memory 1030 within computer-readable storage medium 1050. Program 1060 may include program module 1064. Program module 1064 may generally perform the functions and / or methods of embodiments of the invention as described herein. One or more programs 1060 are stored in memory 1030 and may be executed by processing unit 1020. As an example, memory 1030 may store operating system 1052, one or more applications 1054, other program modules, and program data on computer-readable storage medium 1050. It will be understood that program 1060, operating system 1052, and application 1054 stored on computer-readable storage medium 1050 may similarly be executed by processing unit 1020. It should also be understood that application 1054 and program(s) 1060 are generally shown and may include all or part of one or more applications and programs discussed in this disclosure, or vice versa, that is, application 1054 and program 1060 may be all or part of one or more applications or programs discussed in this disclosure. It should also be understood that the control system 70 communicating with the computer system may include all or part of the computer system 1010 and its components, and / or the control system may communicate with the computer system 1010 and its components as a remote computer system to implement the control system functions described in this disclosure. Control system functions may, for example, include storing, processing, and executing software instructions to perform the functions of this disclosure. It should also be understood that... Figure 1 The one or more computers or computer systems shown may similarly include all or part of computer system 1010 and its components, and / or one or more computers may communicate with computer system 1010 and its components as remote computer systems to perform the computer functions described in this disclosure.
[0129] In embodiments according to this disclosure, one or more programs may be stored in one or more computer-readable storage media, such that the program is embodied and / or encoded in the computer-readable storage media. In one example, the stored program may include program instructions for execution by a processor or a computer system having a processor to perform a method or cause the computer system to perform one or more functions. For example, in one embodiment according to this disclosure, a program embodying a method is embodied in or encoded in a computer-readable storage medium, which includes and is defined as a non-transient or non-transient computer-readable storage medium. Therefore, embodiments or examples of computer-readable storage media according to this disclosure do not include signals, and embodiments may include one or more non-transient or non-transient computer-readable storage media. Thus, in one example, a program may be recorded on a computer-readable storage medium and structurally and functionally associated with that medium.
[0130] Computer 1010 can also communicate with: one or more external devices 1074, such as a keyboard, pointing device, display 1080, etc.; one or more devices that enable a user to interact with computer 1010; and / or any device that enables computer 1010 to communicate with one or more other computing devices (e.g., network interface card, modem, etc.). This communication can occur via input / output (I / O) interface 1022. Furthermore, computer 1010 can also communicate with one or more networks 1200 (such as local area networks (LANs), general area networks (WANs), and / or public networks (e.g., the Internet)) via network adapter / interface 1026. As shown, network adapter 1026 communicates with other components of computer 1010 via bus 1014. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with computer 1010. Examples include, but are not limited to: microcode, device drivers 1024, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0131] It should be understood that a computer or a program running on computer 1010 can communicate with a server implemented as server 1100 via one or more communication networks implemented as communication network 1200. Communication network 1200 may include transmission media and network links, including, for example, wireless, wired, or fiber optic connections, as well as routers, firewalls, switches, and gateway computers. Communication network may include connections such as wired, wireless communication links, or fiber optic cables. Communication network may represent a global collection of networks and gateways (e.g., the Internet) that communicate with each other using various protocols such as Lightweight Directory Access Protocol (LDAP), Transmission Control Protocol / Internet Protocol (TCP / IP), Hypertext Transfer Protocol (HTTP), Wireless Application Protocol (WAP), etc. Networks may also include many different types of networks, such as, for example, intranets, local area networks (LANs), or wide area networks (WANs).
[0132] In one example, the computer can use a network that can use the Internet to access websites on the Web (World Wide Web). In one embodiment, the computer 1010, including a mobile device, can use a communication system or network 1200, which may include the Internet or, for example, a Public Switched Telephone Network (PSTN) cellular network. The PSTN may include telephone lines, fiber optic cables, microwave transmission links, cellular networks, and communication satellites. The Internet facilitates many search and text messaging technologies, such as sending queries to search engines using a cellular phone or laptop via text messaging (SMS), multimedia messaging service (MMS) (as opposed to SMS), email, or a web browser. Search engines can retrieve search results, i.e., links to websites, documents, or other downloadable data corresponding to the query, and similarly, provide the search results to the user via the device as, for example, web pages containing the search results.
[0133] refer to Figure 9 This paper depicts an example system 1500 for use with embodiments of the present disclosure. System 1500 includes multiple components and elements connected via a system bus 1504 (also referred to as a bus). At least one processor (CPU) 1510 is connected to other components via the system bus 1504. Cache 1570, read-only memory (ROM) 1512, random access memory (RAM) 1514, input / output (I / O) adapter 1520, sound adapter 1530, network adapter 1540, user interface adapter 1552, display adapter 1560, and display device 1562 are also operatively coupled to the system bus 1504 of system 1500.
[0134] One or more storage devices 1522 are operatively coupled to system bus 1504 via I / O adapter 1520. Storage device 1522 may be, for example, any of the following: disk storage device (e.g., magnetic disk or optical disk storage device), solid-state magnetic device, etc. Storage device 1522 may be the same type of storage device or different types of storage devices. Storage device may include, for example, but not limited to, hard disk drives or flash memory, and may be used to store one or more programs 1524 or applications 1526. Programs and applications are shown as generic components and may be executed using processor 1510. Program 1524 and / or application 1526 may include all or part of the programs or applications discussed in this disclosure, or vice versa, i.e., program 1524 and application 1526 may be part of other applications or programs discussed in this disclosure. Storage devices may communicate with control system 70 having the various functions described in this disclosure.
[0135] Speaker 1532 is operatively coupled to system bus 1504 via sound adapter 1530. Transceiver 1542 is operatively connected to system bus 1504 via network adapter 1540. Display 1562 is operatively coupled to system bus 1504 via display adapter 1560.
[0136] One or more user input devices 1550 are operatively coupled to system bus 1504 via user interface adapter 1552. User input device 1550 may be, for example, a keyboard, mouse, keypad, image capture device, motion sensing device, microphone, or a device combining the functions of at least two of the aforementioned devices. Other types of input devices may also be used while maintaining the spirit of the invention. User input device 1550 may be of the same type or different types. User input device 1550 is used to input information to and output information to system 1500.
[0137] This invention can be a system, method, and / or computer program product at any possible level of technical detail integration. The computer program product may include one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.
[0138] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or recessed structures with instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0139] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the respective computing / processing device.
[0140] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in one or more programming languages, including any combination of object-oriented programming languages (e.g., Smalltalk, C++, etc.) and procedural programming languages (e.g., the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of this invention, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute the computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuits.
[0141] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0142] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0143] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0144] The flowcharts and block diagrams in the accompanying drawings of this disclosure illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the blocks may occur in a different order than indicated in the figures. For example, two blocks shown consecutively may actually be implemented as a single step, executed simultaneously, substantially simultaneously, with partial or complete time overlap, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block illustrated in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0145] It should be understood that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings set forth herein is not limited to a cloud computing environment. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed hereafter.
[0146] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.
[0147] The characteristics are as follows:
[0148] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without requiring manual interaction with the service provider.
[0149] Wide Area Network (WAN) Access: Capabilities are available on the network and accessed through standard mechanisms that facilitate the use of heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0150] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. Location independence has significance because consumers typically do not control or know the exact location of the resources provided, but can specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0151] Rapid Flexibility: In some cases, the ability to rapidly expand outwards and rapidly expand inwards can be provided quickly and flexibly. For consumers, the capacity available to be provided often appears unlimited and can be purchased at any time and in any quantity.
[0152] Measurement services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both the providers and consumers of the services being utilized.
[0153] The service model is as follows:
[0154] Software as a Service (SaaS): The capability offered to consumers is the ability to use the provider's applications running on cloud infrastructure. Applications can be accessed from various client devices through thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.
[0155] Platform as a Service (PaaS): This provides consumers with the ability to deploy consumer-created or acquired applications onto cloud infrastructure using programming languages and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and the configuration of any application hosting environments.
[0156] Infrastructure as a Service (IaaS): This provides consumers with the capability to deliver processing, storage, networking, and other basic computing resources that enable them to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).
[0157] The deployment model is as follows:
[0158] Private cloud: Cloud infrastructure operated solely by an organization. It can be managed by the organization or a third party and can exist inside or outside a building.
[0159] Community cloud: Cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.
[0160] Public cloud: Cloud infrastructure available to the general public or large industrial groups and owned by organizations that sell cloud services.
[0161] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported together (e.g., cloud bursting for load balancing between clouds).
[0162] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure of a network of interconnected nodes.
[0163] Now for reference Figure 10 This describes an illustrative cloud computing environment 2050. As shown, the cloud computing environment 2050 includes one or more cloud computing nodes 2010 to which local computing devices used by cloud computing consumers can communicate, such as, for example, personal digital assistants (PDAs) or cellular phones 2054A, desktop computers 2054B, laptop computers 2054C, and / or automotive computer systems 2054N. The nodes 2010 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 2050 to provide cloud consumers with infrastructure, platforms, and / or software-as-a-service that eliminates the need for them to maintain resources on their local computing devices. It should be understood that... Figure 10 The types of computing devices 2054A-N shown are intended to be illustrative only, and the computing node 2010 and cloud computing environment 2050 can communicate with any type of computing device on any type of network and / or network-addressable connection (e.g., using a web browser).
[0164] Now for reference Figure 11 This demonstrates the 2050 cloud computing environment ( Figure 10 This provides a set of functional abstractions. It should be understood beforehand that... Figure 11 The components, layers, and functions shown are for illustrative purposes only, and embodiments of the invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:
[0165] The hardware and software layer 2060 includes hardware and software components. Examples of hardware components include: a mainframe 2061; a server 2062 based on a RISC (Reduced Instruction Set Computer) architecture; a server 2063; a blade server 2064; a storage device 2065; and network and networking components 2066. In some embodiments, software components include network application server software 2067 and database software 2068.
[0166] The virtualization layer 2070 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 2071; virtual storage 2072; virtual networks 2073, including virtual private networks; virtual applications and operating systems 2074; and virtual clients 2075.
[0167] In one example, the management layer 2080 may provide the functionality described below. Resource Provisioning 2081 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and Pricing 2082 provides cost tracking when utilizing resources in the cloud computing environment, as well as billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, and protection for data and other resources. User Portal 2083 provides access to the cloud computing environment for consumers and system administrators. Service Level Management 2084 provides cloud resource allocation and management to ensure the required service level is met. Service Level Agreement (SLA) Planning and Fulfillment 2085 provides pre-scheduling and procurement of cloud resources, where future needs are anticipated according to the SLA.
[0168] The workload layer 2090 provides examples of functions that can be utilized in a cloud computing environment. Examples of workloads and functions that can be provided from this layer include: mapping and navigation 2091; software development and lifecycle management 2092; virtual classroom education delivery 2093; data analysis and processing 2094; transaction processing 2095; and for speech recognition from audio input, more specifically, using an extension 2096 of speech content from one or more people (or human speakers) from interleaved audio input.
[0169] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Similarly, the examples of features or functions of the embodiments of this disclosure described herein, whether used in the description of specific embodiments or listed as examples, are not intended to limit the embodiments of this disclosure described herein, or to restrict the disclosure to the examples described herein. These examples are intended to be exemplary and not exhaustive. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or improvements to existing technologies in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A computer-implemented method for speech recognition using an expansion of speech content from audio input, comprising: The learning model is initiated to determine the expansion parameters of each of the multiple audible sounds of speech content received as audio input from multiple speakers at the computer. As part of the learning model, the changes in each of a plurality of independent sounds in response to an audio stimulus are determined, the independent sounds being derived from the audio input; The expansion parameters are applied based on the changes in each of the individual sounds; The acoustic signature of each of the loudspeakers is constructed based on the independent sound and the expansion parameters. as well as The speech content is attributed to each of the plurality of speakers, at least in part based on the voiceprint and the individual sound.
2. The method according to claim 1, further comprising: Generate text from the associated speech content.
3. The method according to claim 2, further comprising: Using the computer, the text is displayed on a screen or monitor that communicates with the computer and / or device.
4. The method according to claim 2, further comprising: The text is sent to a computer and / or device via an electronic communication system for display on a screen or monitor in communication with the computer and / or device.
5. The method of claim 1, wherein the audio input comprises the plurality of audible sounds, and the audio input is received at the computer, the audible sounds comprising speech content from the plurality of speakers.
6. The method according to claim 1, further comprising: The computer is used to amplify the audio input, the amplification including separating the sound in the audio input.
7. The method according to claim 1, further comprising: Use logs to refine the audio input for each of the plurality of speakers.
8. The method of claim 7, wherein the learning model comprises a CNN (convolutional neural network) for receiving the independent sounds and using the logs to determine changes in each of the independent sounds in response to the audio stimulus.
9. The method according to claim 7, further comprising: The logs used for the audio input layer the sounds in the refined audio input into independent sounds.
10. The method according to claim 1, further comprising: The computer receives audio input including the plurality of audible sounds, the audible sounds including speech content from the plurality of speakers; Using the computer to amplify the audio input, the amplification including separating the sound in the audio input; Use logs to refine the audio input of each of the multiple speakers; as well as The log of the audio input is used to layer the sounds in the refined audio input into independent sounds.
11. The method of claim 10, wherein separating the sound in the audio input includes distinguishing ambient or background sounds from speech from one of the plurality of speakers.
12. The method of claim 10, wherein using the log to refine the audio input of each of the plurality of speakers includes dividing the audio input into segments related to speaker identity.
13. A system for speech recognition using an expansion of speech content from an audio input, comprising: A computer system includes: a computer processor, a computer-readable storage medium, and program instructions stored on the computer-readable storage medium, the program instructions being executable by the processor to cause the computer system to perform the following functions; The learning model is initiated to determine the expansion parameters of each of the multiple audible sounds of speech content received as audio input from multiple speakers at the computer. As part of the learning model, the changes in each of a plurality of independent sounds in response to an audio stimulus are determined, the independent sounds being derived from the audio input; The expansion parameters are applied based on the changes in each of the individual sounds; The acoustic signature of each of the loudspeakers is constructed based on the independent sound and the expansion parameters. as well as The speech content is attributed to each of the plurality of speakers, at least in part based on the voiceprint and the individual sound.
14. The system of claim 13, further comprising: Generate text from the associated speech content.
15. The system of claim 14, further comprising: Using the computer, the text is displayed on a screen or monitor that communicates with the computer and / or device.
16. The system of claim 14, further comprising: The text is sent to a computer and / or device via an electronic communication system for display on a screen or monitor in communication with the computer and / or device.
17. The system of claim 13, wherein the audio input comprises the plurality of audible sounds, and the audio input is received at the computer, the audible sounds comprising speech content from the plurality of speakers.
18. The system of claim 13, further comprising: The computer is used to amplify the audio input, the amplification including separating the sound in the audio input.
19. The system of claim 13, further comprising: Use logs to refine the audio input for each of the plurality of speakers.
20. The system of claim 19, wherein the learning model comprises a CNN (convolutional neural network) for receiving the independent sounds and using the logs to determine changes in each of the independent sounds in response to the audio stimulus.
21. The system of claim 19, further comprising: The logs used for the audio input layer the sounds in the refined audio input into independent sounds.
22. A computer program product for speech recognition using an extension of speech content from audio input, the computer program product comprising a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a computer to cause the computer to perform functions, wherein... Includes the following features: The learning model is initiated to determine the expansion parameters of each of the multiple audible sounds of speech content received as audio input from multiple speakers at the computer. As part of the learning model, the changes in each of a plurality of independent sounds in response to an audio stimulus are determined, the independent sounds being derived from the audio input; The expansion parameters are applied based on the changes in each of the individual sounds; The acoustic signature of each of the loudspeakers is constructed based on the independent sound and the expansion parameters. as well as The speech content is attributed to each of the plurality of speakers, at least in part based on the voiceprint and the individual sound.
23. The computer program product according to claim 22, further comprising: Generate text from the associated speech content.
24. The computer program product according to claim 23, further comprising: Using the computer, the text is displayed on a screen or monitor that communicates with the computer and / or device.
25. The computer program product according to claim 23, further comprising: The text is sent to a computer and / or device via an electronic communication system for display on a screen or monitor in communication with the computer and / or device.
Citation Information
Patent Citations
Position sensing using loudspeakers as microphones
GB0426448D0
Systems and methods for speech separation and neural decoding of attentional selection in multi-speaker environments
US20190066713A1