Speech or voice recognition in noisy environments
By generating a customized speech recognition model that matches the location, the accuracy problem of speech and voice recognition systems in noisy environments is solved, and efficient recognition is achieved in different noise environments.
Patent Information
- Application Number
- CN202080102085.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-22
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2040-06-22
AI Technical Summary
The accuracy of existing speech and voice recognition systems in noisy environments is affected by changes in environmental noise, making it difficult to adapt to noise variations in different locations.
By using GPS information and environmental noise data, a customized voice recognition model that matches the current location or location category is generated and selected, and the audio input is adjusted to improve recognition accuracy.
It improves the accuracy of speech and voice recognition in different noise environments and reduces the impact of environmental noise on the recognition system.
Smart Images

Figure CN115943689B_ABST
Abstract
Description
[0001] BACKGROUND
[0002] Modern computing devices, such as cellular phones, laptops, tablets, and desktops, use speech and / or voice recognition for various functions. Speech recognition extracts spoken words, while voice recognition (known as speaker identification) identifies the voice that is speaking, not the words spoken. As such, speech recognition determines "what someone said," while voice recognition determines "who said it." Speech recognition facilitates providing spoken commands to a computing device, thereby eliminating the need to touch or directly use a keyboard or touch screen. Voice recognition provides a similar convenience, but can also be used as an identification authentication tool. And, identifying the speaker can improve speech recognition by using a more suitable voice recognition model that is tailored to that speaker. While contemporary software / hardware has improved deciphering the nuances of speech and voice recognition, the accuracy of such systems is often affected by environmental noise. Even systems that attempt to filter out environmental noise have difficulty accounting for variations in environmental noise that occur in different locations or types of locations.
[0003] SUMMARY
[0004] Various aspects include methods and computing devices that implement methods for voice and / or speech recognition in a noisy environment performed by a processor of a computing device. Various aspects can include voice or speech recognition performed by a processor of a computing device, which can include determining a voice recognition model to use for voice and / or speech recognition based on a location at which an audio input was received, and performing voice and / or speech recognition on the audio input using the determined voice recognition model.
[0005] Further aspects can include determining the location at which the audio input was received using global positioning system information, environmental noise, and / or communication network information. In some aspects, determining a voice recognition model to use for voice and / or speech recognition can include selecting the voice recognition model from a plurality of voice recognition models, where each of the plurality of voice recognition models is associated with a different scene category, each scene category having a specified audio profile. In some aspects, performing voice and / or speech recognition on the audio input using the determined voice recognition model can include adjusting the audio input for environmental noise using the determined speech recognition model and performing voice and / or speech recognition on the adjusted audio input.
[0006] Further aspects can include receiving an audio input associated with an environmental noise sample at the location, associating the location or location category with the received audio input, and transmitting the audio input and associated location or location category information to a remote computing device for generating a voice recognition model based on the received audio input for the associated location or location category.
[0007] Further aspects can include compiling an audio profile from audio input associated with ambient noise at the location, associating the location or location category with the compiled audio profile, and transmitting the audio profile associated with the location or location category to a remote computing device for generating the speech recognition model for the location or location category based on the compiled audio profile.
[0008] Various aspects can use a computing device to generate a speech recognition model. The generation of the speech recognition model can include receiving audio input and location information associated with a location at which the audio input was recorded from a user equipment remote from the computing device, generating a speech recognition model associated with the location using the received audio input for use in speech and / or voice recognition, and providing the generated speech recognition model associated with the location to the user equipment.
[0009] In further aspects, receiving the audio input and the location information can also include receiving a plurality of audio inputs, each audio input having location information associated with a different location. Further, generating a speech recognition model associated with the location using the received audio input can also include generating speech recognition models using the received plurality of audio inputs, where each of the generated speech recognition models can be configured for use at a respective one of the different locations.
[0010] Further aspects can also include determining a location category based on the location information received from the user equipment, and associating the generated speech recognition model with the determined location category.
[0011] Further aspects include a computing device including a processor configured with processor-executable instructions to perform operations of any of the methods summarized above. Further aspects include a non-transitory processor-readable storage medium having stored processor-executable software instructions configured to cause a processor to perform operations of any of the methods summarized above. Further aspects include a processing device for use in a computing device and configured to perform operations of any of the methods summarized above. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings, which are incorporated herein and constitute part of this specification, illustrate exemplary embodiments and together with the general description given above and the detailed description given below, serve to explain the features of the various embodiments.
[0014] Figure 1A and 1B are schematic diagrams illustrating example systems configured for speech and / or voice recognition performed by a processor of one or more computing devices.
[0015] Figure 2 is a schematic diagram illustrating components of an example system for speech and / or voice recognition by a processor of a computing device, in accordance with various embodiments.
[0016] Figure 3 shows a component block diagram of an example system configured for speech and / or voice recognition by a processor of a computing device.
[0017] Figure 4A 、 4B , 4C, 4D, 4E, and / or 4F show process flow diagrams of example methods for speech and / or voice recognition by a processor of a computing device, in accordance with various embodiments.
[0018] Figure 5A and / or 5B show process flow diagrams of example methods for speech and / or voice recognition by a processor of a computing device, in accordance with various embodiments.
[0019] Figure 6 is a component block diagram of a network server computing device suitable for use with various embodiments.
[0020] Figure 7 is a component block diagram of a mobile computing device suitable for use with various embodiments.
[0021] DETAILED DESCRIPTION
[0022] Various aspects will be described in detail with reference to the drawings. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts. References made to particular examples and implementations are for illustrative purposes, and are not intended to limit the scope of the various aspects or the claims.
[0023] Various embodiments provide methods for speech and / or voice recognition in varying environments and / or noise environments performed by a processor of a computing device. Various embodiments can determine a speech recognition model to use for speech and / or voice recognition based on a location at which an audio input is received. Speech and / or voice recognition can be performed on the audio input using the determined speech recognition model.
[0024] As used herein, the term "voice recognition model" is used herein to refer to a quantified set of values and / or mathematical descriptions configured to be used in a specified set of circumstances for computer-based predictive analysis of audio signals for automatic voice and / or speech recognition, including translation of spoken language into text and / or identification of a speaker. In voice recognition, the sound of a speaker's voice and specific keywords or phrases can be used to identify and authenticate a speaker, much like a fingerprint sensor or facial recognition process. In speech recognition, the sound of a speaker's voice is transcribed into words (i.e., text) and / or commands that can be processed and stored by a computing device. For example, a user can speak a key phrase to enable voice recognition and authentication of the user's voice to a computing device, after which the user can dictate to the computing device, which uses speech recognition methods to transcribe the user's words. Various embodiments improve both voice recognition and speech recognition by using trained voice recognition models that take into account the ambient sound of the environment in which the speaker is using the computing device. For example, a first voice recognition model can be used to voice and / or speech recognize utterances made by a speaker in a first environment (e.g., in a quiet office), while a second voice recognition model can be used to voice and / or speech recognize utterances made by the same speaker in a second environment that is typically noisier than the first environment or that typically has a different level or type of ambient background noise (e.g., at home with family). Each voice recognition model can take into account the special characteristics of the speaker's voice and / or the typical ambient noise in a particular background, location, or environment, or the characteristics of background noise in a certain category or type of location (e.g., a restaurant, a car, a city street, etc.). In the presence of background noise in a second environment, voice and / or speech recognition can be done more accurately using a voice recognition model that takes into account the background noise and is therefore different from a voice recognition model used in a first environment without the background noise or with different background noise.
[0025] As used herein, the term "computing device" refers to an electronic device that has at least a processor, a communication system, and a memory configured with a contacts database. For example, a computing device can include any or all of the following: a cellular telephone, a smart phone, a portable computing device, a personal or mobile multimedia player, a laptop computer, a tablet computer, a 2-in-l laptop / tablet computer, a smartbook, an ultrabook, a palmtop computer, a wireless electronic mail receiver, an Internet-enabled multimedia cellular phone, a wearable device (including a smart watch), an entertainment device (e.g., a wireless game controller, a music and video player, a satellite radio, etc.), and similar electronic devices that include a memory, a wireless communication component, and a programmable processor. In various embodiments, a computing device can be configured with a memory and / or storage. Additionally, the computing devices referenced in the various example embodiments can be coupled to or can include wired or wireless communication capabilities to implement various embodiments, such as network transceiver(s) and antenna(s) configured to communicate with a wireless communication network.
[0026] The term "system on a chip" (SOC) is used herein to refer to a single integrated circuit (IC) chip that contains multiple resources and / or processors integrated on a single substrate. A single SOC can contain circuitry for digital, analog, mixed-signal, and radio-frequency functions. The single SOC can also include any number of general and / or special-purpose processors (digital signal processors, modem processors, video processors, etc.), memory blocks (e.g., ROM, RAM, Flash, etc.), and resources (e.g., timers, voltage regulators, oscillators, etc.). Each SOC can also include software for controlling the integrated resources and processors, as well as for controlling peripheral devices.
[0027] The term "system in a package" (SIP) can be used herein to refer to a single module or package that contains multiple resources, computing units, cores and / or processors on two or more IC chips, substrates, or SOCs. For example, a SIP can include a single substrate on which multiple IC chips or semiconductor dies are stacked in a vertical configuration. Similarly, a SIP can include one or more multi-chip modules (MCMs) on which multiple ICs or semiconductor dies are packaged into a unified substrate. A SIP can also include multiple independent SOCs coupled together via high-speed communication circuitry and packaged in close proximity, such as on a single motherboard or in a single wireless device. The proximity of the SOCs facilitates high-speed communication and sharing of memory and resources.
[0028] As used herein, the terms "component," "system," "unit," "module," and the like include computer-related entities such as, but not limited to, hardware, firmware, a combination of hardware and software, software, or software in execution, configured to perform particular operations or functions. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a communication device and the communication device can be referred to as a component. One or more components can reside within a process and / or thread of execution and a component can be localized, partially localized, or distributed across two or more processors or cores. Also, these components can execute from various non-transitory computer-readable media having various instructions and / or data structures stored thereon. Components can communicate by way of local and / or remote processes, function or procedure calls, electronic signals, data packets, memory read / writes, and other known computer, processor, and / or process related communication methodologies.
[0029] Voice and speech recognition systems in accordance with various embodiments can employ deep learning techniques and draw from big data (i.e., large collection datasets) to generate voice recognition models that accurately translate speech into text, provide voice activation functionality, and / or determine or confirm the identity of a speaker (i.e., authenticate) in the presence of different types of background noise. By using customized voice recognition models tailored to specific locations or location types, systems employing various embodiments can provide improved voice and / or speech recognition performance by reducing the impact of environmental noise on the accuracy or recognition of voice and / or speech recognition systems.
[0030] In various embodiments, a processor of a computing device can generate voice recognition models that can be used for different environments corresponding to different locations or location types. The voice recognition models can be generated from audio samples or profiles provided from user equipment, crowd-sourced data, and / or other sources. For example, a user can provide samples of environmental noise from a personal environment, such as a home, office, or places frequently visited by the user. Such environmental noise samples can be used to generate a voice recognition model for each respective location or location category. Alternatively or additionally, the processor can generate voice recognition models from crowd-sourced or generalized recordings of environments corresponding to different common locations or location types in which typical user speech is collected. For example, environmental noise on a train, bus, park, restaurant, hospital, etc. can be used to generate a voice recognition model for each respective location or location category.
[0031] Figures 1A-1B Computing devices 110, 150 configured to provide voice and / or speech recognition functionality in accordance with various embodiments are illustrated. In particular, Figures 1A-1BEnvironments 100, 101 are illustrated in which a user operating user equipment 110 at a location A speaks utterances 115 for user equipment 110 to perform speech and / or voice recognition on utterances 115. Environments 100, 101 can include user equipment 110 operated by user 11 and / or one or more remote computing devices 150. User equipment 110 can represent almost any mobile computing device to which a user (e.g., user 11) can speak as a means of user input using speech and / or voice recognition. User equipment 110 can be configured to perform wireless communications, such as communicating with remote computing device(s) 150. In this manner, user equipment 110 can communicate via one or more base stations 55, which in turn can be communicatively coupled to remote computing device(s) 150 through wired and / or wireless connections 57.
[0032] Environments 100, 101 can also include remote computing device(s) 150, which can be part of a cloud-based computing network configured to assist user equipment 110 in improving speech and / or voice recognition by providing different speech recognition models that user equipment 110 can select and use depending on its current location. Remote computing device(s) 150 can compile a plurality of speech recognition models for user equipment 110 to use. Each speech recognition model can be used in a different environment (i.e., location or location category) to convert utterances spoken by user 11 into text and / or to operate user equipment 110.
[0033] Figure 1A Environments 100 are shown in which user 11 has collected (i.e., at a previous point in time) one or more samples of environmental noise through user equipment 110 at a location labeled as location A. User equipment 110 recorded audio input representing environmental sounds at location A from one or more sources 121, 123, 125. Alternatively, the recorded audio input can represent environmental sounds at a location category of which location A is one. The one or more sources 121, 123, 125 can include any element that generates noise at location A, such as machine sounds 121, background conversations 123, music 125, and / or almost anything that makes noise at location A at the time the audio input is recorded by user equipment 110. Thereafter, user equipment 110 transmits the environmental noise sample(s) along with location information to remote computing device 150 for generating a customized speech recognition model associated with location A. Location A can represent a location at which user 11 frequently uses user equipment 110 for speech and / or voice recognition, such as the user's home, office, or similar frequently visited location.
[0034] In various embodiments, the user equipment 110 can transmit the received audio input and associated location information (i.e., identifying location A) to the remote computing device 150 for generating a voice recognition model associated with location A or a location category that includes location A based on the received audio input. Alternatively, the user equipment 110 can transmit an audio profile, which can additionally include characteristics or other information associated with the received audio input. The user equipment 110 can transmit the audio input along with the location or location category information and / or the audio profile to the local base station 55 that is communicatively coupled to the remote computing device(s) 150 using the exchange signal 131.
[0035] In various embodiments, the remote computing device(s) 150 can use the received audio input (i.e., the environmental noise sample) along with the location information and / or the audio profile to generate a voice recognition model associated with location A for future voice and / or speech recognition performed on utterances at location A. This voice recognition model associated with location A can change the way the sounds of the user’s speech are recognized during voice and / or speech recognition at location A. The remote computing device(s) 150 can have transmitted the generated voice recognition model associated with location A back to the user equipment 110 via the exchange signal 133. Further, the user equipment 110 can have stored the generated voice recognition model for use in performing future voice and / or speech recognition on utterances detected at location A.
[0036] In the case that the generated voice recognition model is previously stored, Figure 1A It is further shown that the user 11 and the user equipment 110 are currently at location A. While at location A, the user 11 is speaking (i.e., uttering 115), which can be received by the microphone of the user equipment 110 as audio input along with background environmental noise from one or more sources 121, 123, 125. In response to receiving this audio input at location A, the user equipment 110 can determine a voice recognition model to be used for voice and / or speech recognition based on the location at which the audio input was received. As such, the user equipment 110 can select a voice recognition model associated with location A. Further, the user equipment 110 can perform voice and / or speech recognition on this audio input using the determined voice recognition model for location A.
[0037] Figure 1BThe environment 101 is shown in which the remote computing device 150 has collected (i.e., at a previous point in time) one or more environmental noise samples 170 associated with a location category (i.e., location category B) via crowdsourcing. In this manner, the remote computing device 150 has obtained the environmental noise sample(s) and / or information about the environmental noise sample(s) through a service (paid or unpaid) that typically solicits a large number of people via the Internet. The location category can represent a plurality of different locations whose audio profiles have similar or identical characteristics. For example, the location category can include restaurants, airports, train stations, bus stations, shopping malls, grocery stores, etc. Alternatively, the crowd-sourced environmental noise samples can represent environmental sounds at a particular location (e.g., location A) rather than a location category (e.g., location category B). Additionally or as a further alternative, the location category B-environmental noise samples 170 can have been recorded and / or compiled by a third-party source.
[0038] The location category B-environmental noise samples 170 can represent environmental sounds from one or more sources 141, 143, 145 at one or more locations belonging to the location category B. The one or more sources 141, 143, 145 can include any element that generates noise at locations belonging to the location category B, such as machine sounds 121, background conversations 123, music 125, and / or nearly anything that makes noise at those locations. The location category B can represent a category of locations at which the user 11 uses the user equipment 110 to conduct voice and / or speech recognition.
[0039] In various embodiments, the remote computing device(s) 150 can use the received location category B-environmental noise samples 170 along with location category information to generate a voice recognition model associated with the location category B for use in future voice and / or speech recognition performed on utterances at locations belonging to the location category B. This voice recognition model associated with the location category B can change the way that utterances from the user 11 are recognized during voice and / or speech recognition at the location category B. The remote computing device(s) 150 can have transmitted the generated voice recognition model associated with the location category B to the user equipment 110 via the exchange signals 137. Further, the user equipment 110 can have stored the generated voice recognition model for use in future voice and / or speech recognition performed on utterances detected at locations belonging to the location category B.
[0040] Figure 1BIt is further shown that the user 11 and the user equipment 110 are currently in a location corresponding to location category B. While at location category B, the user 11 is speaking (i.e., uttering speech 116), which can be received by the microphone of the user equipment 110 as audio input along with background environmental noise from one or more sources 141, 143, 145. In response to the current reception of the audio input at location category B, the user equipment 110 can determine a speech recognition model to be used for speech and / or voice recognition based on the location category at which the audio input is received. As such, the user equipment 110 can select a speech recognition model associated with location category B. Further, the user equipment 110 can perform speech and / or voice recognition on the audio input using the determined speech recognition model for location category B.
[0041] Referring to Figures 1A-2 The illustrated example SIP 200 includes two SOCs 202, 204, a clock 205, a voltage regulator 206, a microphone 207, and a wireless transceiver 208. In some embodiments, the first SOC 202 operates as a central processing unit (CPU) of a wireless device that executes arithmetic, logical, control, and input / output (I / O) operations specified by instructions of software applications by executing those instructions. In some embodiments, the second SOC 204 can operate as a specialized processing unit. For example, the second SOC 204 can operate as a specialized 5G processing unit responsible for managing high capacity, high speed (e.g., 5 Gbps or the like), and / or ultra-high frequency short wavelength (e.g., 28 GHz millimeter wave (mmWave) spectrum or the like) communications.
[0042] The first SOC 202 can include a digital signal processor (DSP) 210, a modem processor 212, a graphics processor 214, an application processor 216, one or more coprocessors 218 (e.g., a vector co-processor) connected to one or more of these processors, a memory 220, custom circuitry 222, system components and resources 224, an interconnect / bus module 226, one or more temperature sensors 230, a thermal management unit 232, and a thermal power envelope (TPE) component 234. The second SOC 204 can include a 5G modem processor 252, a power management unit 254, an interconnect / bus module 264, a plurality of millimeter wave transceivers 256, a memory 258, and various additional processors 260 (such as an application processor, a packet processor, or the like).
[0043] Each processor 210, 212, 214, 216, 218, 252, 260 can include one or more cores, and each processor / core can perform operations independently of the other processors / cores. For example, the first SOC 202 can include processors that execute a first type of operating system (e.g., FreeBSD, LINUX, OS X, etc.) and processors that execute a second type of operating system (e.g., MICROSOFT WINDOWS 10). Additionally, any or all of the processors 210, 212, 214, 216, 218, 252, 260 can be included as part of a processor cluster architecture (e.g., a synchronous processor cluster architecture, an asynchronous or heterogeneous processor cluster architecture, etc.).
[0044] The first and second SOCs 202, 204 can include various system components, resources, and custom circuitry for managing sensor data, analog-to-digital conversion, wireless data transmission, and for performing other specialized operations such as decoding data packets and processing encoded audio and video signals for rendering in a web browser. For example, the system components and resources 224 of the first SOC 202 can include power amplifiers, voltage regulators, oscillators, phase-locked loops, peripheral bridges, data controllers, memory controllers, system controllers, access ports, timers, and other similar components used to support the processors and software clients running on the wireless device. The system components and resources 224 and / or custom circuitry 222 can also include circuitry for interfacing with peripheral devices such as cameras, electronic displays, wireless communication devices, external memory chips, etc.
[0045] The first and second SOCs 202, 204 can communicate via the interconnect / bus module 250. The various processors 210, 212, 214, 216, 218 can be interconnected to one or more memory elements 220, system components and resources 224, and custom circuitry 222, and thermal management unit 232 via the interconnect / bus module 226. Similarly, the processor 252 can be interconnected to the power management unit 254, millimeter wave transceiver 256, memory 258, and various additional processors 260 via the interconnect / bus module 264. The interconnect / bus modules 226, 250, 264 can include an array of reconfigurable logic gates and / or implement a bus architecture (e.g., CoreConnect, AMBA, etc.). Communication can be provided by an advanced interconnect such as a high-performance network-on-chip (NoC).
[0046] The first and / or second SOC 202, 204 can further include input / output modules (not illustrated) for communicating with resources external to the SOC, such as a clock 205 and a voltage regulator 206. Resources external to the SOC (e.g., clock 205, voltage regulator 206) can be shared by two or more internal SOC processors / cores.
[0047] In addition to the example SIP 200 discussed above, various embodiments can be implemented in a wide variety of computing systems, which can include a single processor, multiple processors, multi-core processors, or any combination thereof.
[0048] Various embodiments can be implemented using several single- and multi-processor computer systems, including a system on a chip (SOC) or system in a package (SIP). Figure 2 An example computing system or SIP 200 architecture that can be used in a user equipment (e.g., 110), a remote computing device (e.g., 150), or other system for implementing various embodiments is illustrated.
[0049] Figure 3 is a component block diagram illustrating a system 300 configured for speech and / or voice recognition performed by a processor of one or more computing devices, in accordance with various embodiments. In some embodiments, the system 300 can include a user equipment 110 and / or one or more remote computing devices 150. The user equipment 110 can represent almost any mobile computing device to which a user can speak as a means of user input using speech and / or voice recognition. The system 300 can also include remote computing device(s) 150, which can be part of a cloud computing network configured to assist the user equipment 110 in improving speech and / or voice recognition by providing different speech recognition models that the user equipment 110 can select and use depending on its current location. Each speech recognition model can be used in a different environment (i.e., location or location category) to convert utterances (i.e., speech) made by a user into text, identify / authenticate the user, and / or operate other functions of the user equipment 110 in response to spoken commands. The remote computing device(s) 150 can compile multiple speech recognition models for use by the user equipment 110.
[0050] The user equipment can include a microphone 207 for receiving sound (i.e., audio input), which can be digitized into data packets for analysis and / or transmission. The audio input can include ambient sound in the vicinity of the user equipment 110 and / or speech from a user of the user equipment 110. Further, the user equipment 110 can be communicatively coupled to peripheral device(s) (not shown) and configured to communicate with remote computing device(s) 150 and / or external resource 320 using wireless transceiver 208 and communication network 50, such as a cellular communication network. Similarly, remote computing device(s) 150 can be configured to communicate with user equipment 110 and / or external resource 320 using wireless transceiver 208 and communication network 50.
[0051] User equipment 110 can also include electronic storage 325, one or more processors 330, and / or other components. User equipment 110 can include communication circuitry or ports to enable information exchange with a network and / or other computing platforms, such as remote computing device(s) 150. Figure 1A and Figure 1B The illustration of user equipment 110 in FIG. 1 is not intended to be limiting. User equipment 110 can include a number of hardware, software, and / or firmware components operating together to provide the functionality attributed to user equipment 110 herein.
[0052] Remote computing device(s) 150 can include electronic storage 326, one or more processors 331, and / or other components. Remote computing device(s) 150 can include communication circuitry or ports to enable information exchange with a network, other computing platforms, and many user mobile computing devices, such as user equipment 110. Figure 3 The illustration of remote computing device(s) 150 in FIG. 1 is not intended to be limiting. Remote computing device(s) 150 can include a number of hardware, software, and / or firmware components operating together to provide the functionality attributed to remote computing device(s) 150 herein.
[0053] External resource 320 includes a remote server that can receive sound recordings and generate voice recognition models for various locations and location categories, and provide the voice recognition models to computing devices, such as in a download via communication network 50. External resource 320 can receive sound recordings and information from voice and / or speech recognition processing performed in various locations from a number of user equipment and computing devices through a crowdsourcing process.
[0054] The electronic storage 325, 326 can include non-transitory storage media that electronically stores information. The electronic storage media of electronic storage 325, 326 can include one or both of system storage that is provided integrally (i.e., substantially non-removable) with user equipment 110 or remote computing device(s) 150, and / or removable storage that is removably connectable to user equipment 110 or remote computing device(s) 150, e.g., via port (e.g., a universal serial bus (USB) port, a FireWire port, etc.) or drive (e.g., a disk drive, etc.). Electronic storage 325, 326 can include one or more of optically readable storage media (e.g., optical disks, etc.), magnetically readable storage media (e.g., magnetic tape, magnetic hard drive, floppy disk drive, etc.), electrical charge-based storage media (e.g., EEPROM, RAM, etc.), solid-state storage media (e.g., flash drive, etc.), and / or other electronically readable storage media. Electronic storage 325, 326 can include one or more virtual storage resources (e.g., cloud storage, virtual private networks, and / or other virtual storage resources). Electronic storage 325, 326 can store software algorithms, information determined by processor(s) 330, 331, information received from user equipment 110 or remote computing device(s) 150, respectively, that enable user equipment 110 or remote computing device(s) 150, respectively, to function as described herein.
[0055] Processor(s) 330, 331 can be configured to provide information processing capabilities in user equipment 110 or remote computing device(s) 150, respectively. As such, processor(s) 330, 331 can include one or more of a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information. Although processor(s) 330, 331 are each shown as a single entity in FIG. 1, this is for illustrative purposes only. In some implementations, processor(s) 330, 331 can include a plurality of processing units. These processing units can be physically Figure 3
[0056] The user equipment 110 can be configured by machine-readable instructions 335 that can include one or more instruction modules. The instruction modules can include computer program modules. In particular, the instruction modules can include one or more of an audio reception module 340, a location / location category determination module 345, an audio profile compilation module 350, an audio input transmission module 355, a speech recognition model reception module 360, a speech recognition model determination module 365, an audio input adjustment module 370, a speech and / or voice recognition module 375 (i.e., speech / voice recognition module 375), and / or other instruction modules.
[0057] The audio reception module 340 can be configured to receive audio input associated with environmental noise samples and / or user speech at a current location of the user equipment 110. The audio reception module 340 can receive sound (i.e., audio input) from the microphone 207 and digitize them into data packets for analysis and / or transmission. For example, the audio reception module 340 can receive environmental noise from one or more locations as well as speech of a user. The audio input received by the audio reception module 340 can be used in an audio sampling mode and / or a speech / voice recognition mode. In the audio sampling mode, the received environmental noise can be used to train (i.e., generate) a speech recognition model. In the audio sampling mode, the received user speech can include keywords and / or phrases spoken by the user that can be used to train the speech recognition model. Conversely, in the speech / voice recognition mode, the received user utterances (i.e., speech) can be used for speech and / or voice recognition. As a non-limiting example, a device for implementing the machine-readable instructions 335 of the audio reception module 340 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150) that can use the electronic storage 325, 326, external resources 320, one or more sensors (e.g., microphone 207), and an audio profile database for storing received audio input.
[0058] The location / location category determination module 345 can be configured to determine a location and / or a location category of the user equipment 110. The location / location category determination module 345 can determine the location and / or the location category of the user equipment 110 in one or more ways. By accessing a global positioning system, the location / location category determination module 345 can use global positioning system information to determine a location at which an audio input is received. The location / location category determination module 345 can additionally or alternatively use environmental noise to determine a location at which an audio input is received. For example, the location / location category determination module 345 can compare a current sample of environmental noise collected by the audio reception module 340 to one or more audio profiles compiled and maintained by the audio profile compilation module 350. In response to finding an audio profile that matches the current sample of environmental noise, the location / location category determination module 345 can determine a location and / or a location category that will be associated with the audio profile. Also, the location / location category determination module 345 can further or alternatively use communication network information to determine a location at which an audio input is received. For example, some network connections, such as those using WiFi, Bluetooth, and other protocols, can be associated with a fixed location (e.g., a home, an office, etc.). Thus, once the user equipment 110 connects to that kind of network connection, the location / location category determination module 345 can infer a location and / or a location category of the user equipment and the received environmental noise. As a non-limiting example, the means for machine-readable instructions 335 of the machine-readable instructions 335 to implement the location / location category determination module 345 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150) that can use electronic storage 325, 326, external resources 320, one or more sensors (e.g., motion sensors, GPS, microphones, etc.), data input systems (e.g., touch-sensitive input / display and / or connected accessories), and a communication system (e.g., a wireless transceiver) for determining a location or a location category in which the user equipment is located.
[0059] The audio profile compiling module 350 can be configured to compile an audio profile from audio input received by the audio receiving module 340 in the audio sampling mode. The audio profile can include the audio input and / or samples thereof (e.g., environmental noise and / or user keyword utterances) received by the audio receiving module 340. Further, the audio profile can include an indication of the location or location category determined by the location / location category determining module 345. Further, the audio profile compiling module 350 can be configured to analyze, label, and / or convert the received audio input or samples to an appropriate format for transmission to the remote computing device 150. By way of non-limiting example, means for implementing the machine-readable instructions 335 of the audio profile compiling module 350 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150), electronic storage 325, 326, external resources 320, one or more sensors (e.g., microphone 207), and an audio profile database for compiling audio input.
[0060] The audio input transmitting module 355 can be configured to transmit the audio input received by the audio receiving module 340, samples thereof, and / or the audio profile compiled by the audio profile compiling module 350 and received from the audio profile compiling module 350 to the remote computing device 150. Thereby, the audio input transmitting module 355 can transmit the received audio input and associated location or location category to the remote computing device(s) 150 for generating a speech recognition model for the associated location or location category based on the received audio input. Similarly, the audio input transmitting module 355 can transmit the audio profile associated with the location or location category to the remote computing device(s) 150 for generating the speech recognition model for the location or location category based on the compiled audio profile. By way of non-limiting example, means for implementing the machine-readable instructions 335 of the audio input transmitting module 355 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150), electronic storage 325, 326, external resources 320, and a wireless transceiver 208 for transmitting environmental noise samples or profiles.
[0061] The speech recognition model receiving module 360 can be configured to receive, from the remote computing device(s) 150, generated speech recognition models associated with particular locations and / or location categories. As described further below, the remote computing device(s) 150 can generate user-customized speech recognition models based on environmental noise or user keyword utterances received from the user equipment 110. The user- customized speech recognition models can be generated specifically for a user's particular user equipment 110. Further, the remote computing device(s) 150 can generate general-purpose speech recognition models based on crowd-sourced samples and / or information about particular locations or location categories. Each speech recognition model can be associated with a different location or location category. As a non-limiting example, means for machine-readable instructions 335 implementing the speech recognition model receiving module 360 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150), the electronic storage 325, 326, the external resource 320, and a wireless transceiver 208 for receiving speech recognition models.
[0062] The speech recognition model determining module 365 can be used in speech and / or voice recognition mode. The speech recognition model can be configured to determine (e.g., select from a model library) a speech recognition model to use for speech and / or voice recognition based on a location in which an audio input is received. In some embodiments, the speech recognition model can be selected from a plurality of speech recognition models, where each of the plurality of speech recognition models is associated with a different scene category, each scene category having a specified audio profile. As a non-limiting example, means for machine-readable instructions 335 implementing the speech recognition model determining module 365 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150) that can use the electronic storage 325, 326, the external resource 320, and a speech recognition model database for accessing information about various speech recognition models.
[0063] The optional audio input adjustment module 370 can be configured to adjust audio input for environmental noise using the selected voice recognition model. For example, the voice recognition model can be used to filter out environmental noise using samples of environmental noise from the location or location category in which the user equipment 110 is located. With the environmental noise filtered out, the remaining audio input, which can include one or more user utterances, can be processed by the voice and / or speech recognition module 375. By using the optional audio input adjustment module 370, the voice and / or speech recognition module 375 can use a general voice recognition model, as the audio input has already been filtered for typical environmental noise in the determined location or location category. As a non-limiting example, means for implementing the machine-readable instructions 335 of the optional audio input adjustment module 370 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150), electronic storage 325, 326, external resources 320, one or more sensors (e.g., microphone 207), and an audio profile database for storing the adjusted audio input.
[0064] The voice and / or speech recognition module 375 can be configured to perform voice and / or speech recognition on audio input (e.g., user utterances). In particular, the voice and / or speech recognition module 375 can perform voice and / or speech recognition using the voice recognition model determined by the voice recognition model determination module 365. The voice recognition model associated with a particular location or location category can be used for voice and / or speech recognition using different parameters that can be applied to any received audio input for direct voice and / or speech recognition analysis. Alternatively, if the optional audio input adjustment module 370 is included / used, the voice and / or speech recognition module 375 can use a general voice recognition model or at least one voice recognition model tailored for a particular user rather than a particular location or location category for voice and / or speech recognition. As a non-limiting example, means for implementing the machine-readable instructions 335 of the voice and / or speech recognition module 375 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150), electronic storage 325, 326, external resources 320, one or more sensors (e.g., microphone 207), and a voice and / or speech recognition database for storing the results of voice and / or speech recognition.
[0065] The remote computing device(s) 150 can be configured by machine-readable instructions 336 that can include one or more instruction modules. The instruction modules can include computer program modules. In particular, the instruction modules can include one or more of a compiled audio input receiving module 380, a user keyword module 385, a location / location category association module 390, a speech recognition model generation module 395, a speech recognition model transmission module 397, and / or other instruction modules.
[0066] The audio input receiving module 380 can be configured to receive audio inputs, samples thereof, and / or audio profiles (such as audio profiles compiled by the audio profile compiling module 350) from the user equipment 110. Thus, the audio input receiving module 380 can receive environmental noise and / or user keyword utterances and location information associated with the location at which the environmental noise and / or user keyword utterances were recorded. In some embodiments, the audio input receiving module 380 can receive multiple environmental noise samples and / or user keyword utterances, each having location information associated with a different location. As a non-limiting example, means for implementing the machine-readable instructions 336 of the audio input receiving module 380 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150), electronic storage 325, 326, external resources 320, and a transceiver 328 for receiving environmental noise samples or profiles.
[0067] The user keyword module 385 can be configured to maintain information about user keyword utterances that are used for speech and / or voice recognition (i.e., keyword information). The keyword information can be obtained from samples of keywords spoken by a user that are contained in received audio inputs. The keyword information can identify the keywords and include audio characteristics of each keyword that can be used to identify the keyword in an utterance. Each of the speech recognition models generated by the speech recognition model generation module 395 can be associated with a different keyword. Alternatively, a keyword that is unique to a particular user can be associated with all queries and commands. As a non-limiting example, means for implementing the machine-readable instructions 336 of the user keyword module 385 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150), electronic storage 325, 326, external resources 320, one or more sensors (e.g., microphone 207), and a user keyword database for use in generating speech recognition models.
[0068] The location / location category association module 390 can be configured to determine a location or a location category for a received audio input, audio sample, and / or audio profile based on corresponding location information received from a user equipment. As a non-limiting example, means for implementing machine readable instructions 336 of the location / location category association module 390 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150) that can use electronic storage 325, 326, external resources 320, and a location / location category classification database for determining a location or a location category in which an ambient noise was recorded.
[0069] The voice recognition model generation module 395 can be configured to generate voice recognition models associated with a location or a location category for use in voice and / or speech recognition using received audio inputs, audio samples, and / or audio profiles. In some embodiments, the voice recognition model generation module 395 can use samples of keywords spoken by a user contained in a received audio input to generate a voice recognition model. In some embodiments, the voice recognition model generation module 395 can use multiple audio samples for generating different voice recognition models, and each audio sample has location information associated with a different location. In some embodiments, the voice recognition model generation module 395 can use multiple audio samples for generating one voice recognition model associated with a single location or a location category. The location(s) or the location category / locations associated with each voice recognition model can be those determined by the location / location category determination module 345. As a non-limiting example, means for implementing machine readable instructions 336 of the voice recognition model generation module 395 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a processing device (e.g., 110, 150) that can use electronic storage 325, 326, external resources 320, and a voice recognition model database for accessing information about individual voice recognition models and parameters for generating these voice recognition models.
[0070] The voice recognition model transmission module 397 can be configured to provide voice recognition models generated by the voice recognition model generation module 395 and associated with a location or a location category to a user equipment. As a non-limiting example, means for implementing machine readable instructions 336 of the voice recognition model transmission module 397 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331), electronic storage 325, 326, external resources 320, and a transceiver 328 of a processing device (e.g., 110, 150) for receiving voice recognition models.
[0071] User equipment 110 can include one or more processors configured to execute computer program modules similar to those in machine-readable instructions 336 of remote computing device(s) 150 described above. By way of non-limiting example, user equipment can include one or more of: a laptop computer, a handheld computer, a tablet computing platform, a netbook, a smartphone, a gaming console, and / or other mobile computing platform.
[0072] Given remote computing device(s) 150 can include one or more processors configured to execute computer program modules similar to those in machine-readable instructions 335 of user equipment 110 described above. By way of non-limiting example, remote computing device can include one or more of: a server, a desktop computer, a laptop computer, a handheld computer, a tablet computing platform, a netbook, a smartphone, a gaming console, and / or other mobile computing platform.
[0073] Processors 330, 331 can be configured to execute modules 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, and / or 397 and / or other modules. Processors 330, 331 can be configured to execute modules 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, and / or 397 and / or other modules by software; hardware; firmware; some combination of software, hardware, and / or firmware; and / or other mechanisms for configuring processing capability on processors 330, 331. As used herein, the term “module” can refer to
[0074] The description of the functionality provided by the different modules 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, and / or 397 described below is for illustrative purposes as the functionality of one or more of modules 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, and / or 397 can be combined or distributed as desired. For example, one or more of modules 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, and / or 397 can be eliminated, and some or all of its functionality can be provided by other ones of modules 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, and / or 397. As another example, processor(s) 330 can be configured to execute one or more additional modules that can perform some or all of the functionality attributed below to one of modules 340, 345, 350, 355, 360, 365, 370, 375, 380, 385, 390, 395, and / or 397.
[0075] Figure 4A , 4B , 4C, 4D, 4E, and / or 4F illustrate operations of methods 400, 401, 402, 403, 404, and / or 405 for voice and / or speech recognition performed by a processor of a computing device, in accordance with various embodiments. Reference is made to Figures 1A-4A , 4B, 4C, 4D, 4E, and / or 4F, the operations of methods 400, 401, 402, 403, 404, and / or 405 presented below are intended to be illustrative. In some embodiments, method 400, 401, 402, 403, 404, and / or 405 can be accomplished with one or more additional operations not described, or without one or more of the operations discussed. Additionally, the order in which the operations of methods 400, 401, 402, 403, 404, and / or 405 are described is not intended to be limiting. Figure 4A , 4B The order in which the operations of methods 400, 401, 402, 403, 404, and / or 405 illustrated in and described below in 4C, 4D, 4E, and / or 4F is not intended to be limiting.
[0076] In some embodiments, the methods 4A, 4B, 4C, 4D, 4E, and / or 4F can be implemented in one or more processors (e.g., a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information) in response to Figures 1A-4A instructions stored electronically on an electronic storage medium of a computing device. The one or more processors can include one or more devices configured to perform one or more operations of the methods 400, 401, 402, 403, 404, and / or 405 by hardware, firmware, and / or software. For example, with reference to
[0077] Figure 4A A method 400 is illustrated in accordance with one or more implementations.
[0078] At block 410, the processor of the computing device can perform operations including determining a voice recognition model to use for voice and / or speech recognition based on a location at which the audio input is received. At block 410, the processor of the user equipment can use an audio reception module (e.g., 340), a location / location category determination module (e.g., 345), and a voice recognition model determination module 365 to select an appropriate voice recognition model based on a location or location category at which the audio input is received / recoded. For example, the processor can determine that the utterance currently being received was spoken in a user’s home. In such a case, a voice recognition model trained to account for environmental noise in a user’s home can more accurately translate speech and / or identify / authenticate a user from the sound of the user’s voice. As another example, the processor can determine that the utterance currently being received was spoken in a restaurant. In such a case, the restaurant can belong to a location category for which an audio profile crowd-sourced can have been used to generate a voice recognition model that can more accurately translate speech and / or identify / authenticate a user. In some embodiments, means for performing the operations of block 410 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to a microphone (e.g., 207), electronic storage (e.g., 325, 326), an external resource (e.g., 320), and a voice recognition model determination module (e.g., 365).
[0079] At block 415, the processor of the computing device can perform operations including performing speech and / or voice recognition on the audio input using the determined speech recognition model. At block 415, the processor of the user equipment can perform speech and / or voice recognition using an appropriate speech recognition model selected to better account for background environmental noise. For example, where the received audio input is collected from a user's home, some regular background noise (e.g., people talking, children playing, and / or music playing, noisy appliances running, etc.) can have been accounted for in generating the speech recognition model selected to perform speech and / or voice recognition. Alternatively, the processor can use the speech recognition model to adjust the received audio input for predicted environmental noise of the environment (i.e., location and / or location category). In this way, regular background noise can be filtered from the received audio input before a more general model (i.e., even one customized for a particular user) is used for speech and / or voice recognition. In some embodiments, means for performing the operations of block 415 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to the electronic storage (e.g., 325, 326), the external resource (e.g., 320), and the speech and / or voice recognition module (e.g., 375).
[0080] In some embodiments, the processor can repeat any or all of the operations in blocks 410 and 415 to repeatedly or continuously perform speech and / or voice recognition.
[0081] Figure 4B A method 401 is illustrated that can be performed with or as an enhancement to the method 400.
[0082] At block 420, the processor of the computing device can perform operations including determining a location at which the audio input was received using global positioning system (GPS) information. At block 420, the processor of the user equipment can determine the location at which the audio input was received / recorded using an audio reception module (e.g., 340) and a location / location category determination module (e.g., 345). For example, the processor can access a GPS system, providing coordinates, an address, or other location information. Further, the processor can access one or more online databases that can identify a location corresponding to the GPS information. Further, by using contact information stored in the user equipment, the location can be more accurately associated with the user's home, office, or other frequently visited locations. In some embodiments, means for performing the operations of block 420 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to the wireless transceiver (e.g., 208), electronic storage (e.g., 325, 326), external resources (e.g., 320), and the location / location category determination module (e.g., 345). Following the operations in block 420, the processor can determine a voice recognition model to use for voice and / or speech recognition based on the determined location at which the audio input was received in block 410.
[0083] In some embodiments, the processor can repeat any or all of the operations in blocks 410, 415, and 420 to repeatedly or continuously perform voice and / or speech recognition.
[0084] Figure 4C A method 402 is illustrated that can be performed with or as an enhancement to the method 400.
[0085] At block 425, the processor of the computing device can perform operations including determining a location at which the audio input was received using the environmental noise. At block 425, the processor of the user equipment can determine the location at which the audio input was received / recorded using the audio reception module (e.g., 340) and the location / location category determination module (e.g., 345). For example, the processor can compare the environmental noise included in the received audio input to environmental noise samples stored in memory. If the currently received environmental noise matches an environmental noise sample, the processor can assume that the current location is the location associated with the matching environmental noise sample. In some embodiments, means for performing the operations of block 425 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to the wireless transceiver (e.g., 208), electronic storage (e.g., 325, 326), external resources (e.g., 320), and the location / location category determination module (e.g., 345). After the operations in block 425, the processor can determine the voice recognition model to be used for voice and / or speech recognition based on the determined location at which the audio input was received in block 410.
[0086] In some embodiments, the processor can repeat any or all of the operations in blocks 410, 415, and 425 to repeatedly or continuously perform voice and / or speech recognition.
[0087] Figure 4D A method 403 is illustrated that can be performed with or as an enhancement to the method 400.
[0088] At block 430, the processor of the computing device can perform operations including using the communication network information to determine a location at which the audio input was received. At block 430, the processor of the user equipment can use the wireless transceiver 208, an audio reception module (e.g., 340), and a location / location category determination module (e.g., 345) to determine a location at which the audio input was received / recoded. For example, the processor can check for a current local network connection such as to a WiFi, Bluetooth, or other trusted wireless network, where the connection settings are saved in electronic storage (e.g., 325). Such a saved local network connection can be associated with a location such as the user's home, workplace, gym, etc. Thus, by identifying the current local network connection as one that is saved in memory and for which the location is known, the processor can use the communication network information to determine a current location at which the audio input was received. In some embodiments, means for performing the operations of block 430 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to the wireless transceiver (e.g., 208), electronic storage (e.g., 325, 326), external resources (e.g., 320), and the location / location category determination module (e.g., 345). After the operations in block 430, the processor can determine a voice recognition model to use for voice and / or speech recognition based on the determined location at which the audio input was received in block 410.
[0089] In some embodiments, the processor can repeat any or all of the operations in blocks 410, 415, and 430 to repeatedly or continuously perform voice and / or speech recognition.
[0090] Figure 4E A method 404 is illustrated that can be performed with or as an enhancement to the method 400.
[0091] At block 435, the processor of the computing device can receive an audio input. At block 435, the processor of the user equipment can receive the audio input using an audio receiving module (e.g., 340). In some embodiments, means for performing the operations of block 435 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to a microphone (e.g., 207), electronic storage (e.g., 325, 326), external resources (e.g., 320), and an audio receiving module (e.g., 340). After the operations in block 435, the processor can perform the operations in one or more of blocks 420, 425, and 430 to determine a location at which the audio input was received. The selection of the operational block(s) to be performed (i.e., 420, 425, and / or 430) can be based on the availability of those corresponding resources and / or information that can be used to determine a location or location category.
[0092] At determination block 440, the processor of the computing device can determine whether a location or location category of the received audio input has been determined. In other words, has a location or location category been determined by the operations in one or more of blocks 420, 425, and 430? In some embodiments, means for performing the operations of determination block 440 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to electronic storage (e.g., 325, 326), external resources (e.g., 320), and a location / location category determination module (e.g., 345).
[0093] In response to the processor determining that a location or location category of the received audio input has been determined (i.e., determination block 440 = "Yes"), the processor can determine whether the received audio input is part of an audio sampling mode in determination block 450. In some embodiments, the user equipment can operate in at least one of two modes (i.e., an audio sampling mode and a voice / speech recognition mode). The audio sampling mode can be used to train the system (e.g., 300) with one or more environmental noise samples or user utterances of keywords or expressions from a particular location in order to compile a customized voice recognition model. The voice / speech recognition mode can be used to authenticate / identify a speaker of an utterance (i.e., voice recognition) and / or perform speech recognition (i.e., transcribe speech into text and / or recognize and perform spoken commands).
[0094] In response to the processor determining that a location or location category of the received audio input has not been determined (i.e., determination block 440 = "No"), the processor can apply a user input or a default location / location category in block 445.
[0095] In box 445, if the location or location category of the received audio input has been determined to be unknown, the processor of the computing device can access a memory buffer to check whether user input in this regard has been received. For example, before speaking, the user of the user equipment (e.g., 11) may have already entered location information into a field or pop-up screen available for this purpose. Similarly, the user equipment may have a default location stored for all or selected cases in which audio input is received and / or analyzed. In this way, the processor can apply the location information entered by the user or set as the default location as the location or location category of the audio input received in box 435. Optionally, if no location or location category is determined and a default location or location category needs to be used, a failure to determine the location or location category can be reported to the equipment manufacturer, communications provider, or other entity of the user equipment during an error reporting process. Alternatively, the processor can prompt the user for a location and apply the response to the prompt as that location. The prompt may include suggestions, such as a list of recent or favorite locations, a default location, etc. The user's response to the prompt can be regarded as user input providing the location or location category. In some embodiments, means for performing the operation of block 445 may include a user interface coupled to a user interface (e.g., Figure 7 The processors (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of the display 730, electronic storage (e.g., 325, 326), external resources (e.g., 320), and location / location category determination module (e.g., 345).
[0096] In determination box 450, the processor may determine whether the received audio input is part of an audio sampling pattern. The audio sampling pattern may be part of a training routine used by the system (e.g., 300) to collect ambient noise samples and / or keyword utterances from a user's equipment at one or more locations. In the audio sampling pattern, the user may be instructed not to speak while recording sounds naturally heard at the sampling location. The audio sampling pattern may require multiple recordings (i.e., samples) from the same location to detect consistent patterns and / or filter out anomalous sounds that may occur during sampling. In some embodiments, means for performing the operation of determination box 450 may include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to electronic storage (e.g., 325, 326), external resources (e.g., 320), and an audio receiving module (e.g., 340).
[0097] In response to determining that the received audio input is part of an audio sampling mode (i.e., determination block 450 = "Yes"), the processor can associate the determined location or location category with the received audio input in block 455 as part of the audio sampling mode. The audio sampling mode can include taking, saving, and / or transmitting environmental noise samples and / or keyword utterances at a particular location, which can be used to train or compile a customized speech recognition model for that particular location.
[0098] In response to the processor determining that the received audio input is not part of an audio sampling mode (i.e., determination block 450 = "No"), the processor can determine a speech recognition model to use for speech and / or voice recognition based on the determined location at which the audio input was received in block 410 as part of a speech / voice recognition mode. The speech / voice recognition mode can be used for speech and / or voice recognition.
[0099] In block 455, the processor of the computing device can associate the determined location or location category with the received audio input. This determination can be stored in memory and / or become part of the environmental noise sample (e.g., by metadata attachment) whether the location was determined in blocks 420, 425, and / or 430, determined from user input in block 445, or determined from a default location / location category in block 445. In some embodiments, means for performing the operations of block 455 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to electronic storage (e.g., 325, 326), external resources (e.g., 320), and an environmental noise sample / profile compilation module (e.g., 350).
[0100] In block 460, the processor of the computing device can transmit the audio input and associated location or location category information to a remote computing device for generating a speech recognition model for that associated location or location category. For example, the processor can use a wireless transceiver (e.g., 208) to transmit the environmental noise sample associated with the location in block 455 to the remote computing device 150 via a communication network (e.g., 50). In some embodiments, means for performing the operations of block 460 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to a wireless transceiver (e.g., 208), electronic storage (e.g., 325, 326), external resources (e.g., 320), and an audio input transmission module (e.g., 355).
[0101] In some embodiments, the processor can repeat any or all of the operations in blocks 410, 415, 420, 425, 430, 445, 455, and 460, as well as determination blocks 440 and 450, to perform speech and / or voice recognition repeatedly or continuously.
[0102] Figure 4F Method 400 is illustrated that can be performed with or as an enhancement to method 404.
[0103] At block 465, the processor of the computing device can perform operations including compiling an audio profile from the audio input associated with the determined location or location category. For example, the audio profile can include characteristics or other information associated with the received audio input, including location and / or location category information. In some embodiments, means for performing the operations of block 465 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to electronic storage (e.g., 325, 326), external resources (e.g., 320), and an ambient noise sample / profile compilation module (e.g., 350).
[0104] At block 470, the processor of the computing device can perform operations including associating the determined location and / or location category with the compiled audio profile. Whether the location was determined in blocks 420, 425, and / or 430 of methods 401, 402, 403, determined from user input in block 445 of method 404, or determined from a default location / location category in block 445, the determination can be stored in memory and / or become part of the audio profile. In some embodiments, means for performing the operations of block 455 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to electronic storage (e.g., 325, 326), external resources (e.g., 320), and an ambient noise sample / profile compilation module (e.g., 350).
[0105] At block 475, the processor of the computing device can perform operations comprising transmitting the audio profile associated with the location and / or location category to the remote computing device for generating the speech recognition model for the location and / or location category based on the compiled audio profile. For example, the processor can transmit the audio profile compiled in block 465 to the remote computing device 150 using a wireless transceiver (e.g., 208) via a communication network (e.g., 50). In some embodiments, means for performing the operations of block 460 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to the wireless transceiver (e.g., 208), electronic storage (e.g., 325, 326), external resources (e.g., 320), and an audio input transmission module (e.g., 355).
[0106] In some embodiments, the processor can repeat any or all of the operations of blocks 410, 415, 420, 425, 430, 445, 465, 470, and 475, as well as the determination blocks 440 and 450, to repeatedly or continuously perform speech and / or voice recognition.
[0107] Figure 5A and 5B Operations of methods 500 and 501 for speech and / or voice recognition performed by a processor of a computing device in accordance with some embodiments are illustrated. Referring to Figures 1A-5B , the operations of methods 500 and 501 presented below are intended to be illustrative. In some embodiments, method 500 and 501 can be accomplished with one or more additional operations not described, and / or without one or more of the operations discussed. Additionally, Figure 5A and 5B The order in which the operations of methods 500 and / or 501 are illustrated and described below is not intended to be a limitation. For example, some operations can occur in different orders, some operations can be performed concurrently, and some operations can be omitted without departing from the scope of the methods.
[0108] In some embodiments, methods 5A and 5B can be implemented in one or more processors (e.g., a digital processor, an analog processor, a digital circuit designed to process information, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information) in response to Figures 1A-5B , the operations of methods 500 and 501 can be performed by a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) of a computing device (e.g., 110, 150).
[0109] Referring toFigure 5A At block 510, a processor of a computing device (e.g., remote computing device 150) can perform operations including receiving audio input from a user equipment (e.g., 110) that is remote from the computing device (e.g., 150) and location information associated with a location at which the audio input was recorded. The received audio can be part of a plurality of received audio inputs, where each received audio input has location information associated with a different location. At block 510, the processor of the remote computing device can use an audio input receiving module (e.g., 380), a user keyword module (e.g., 385), and a location / location category association module (e.g., 390). For example, after the user equipment collects audio input and transmits the audio input to the remote computing device, the processor of the remote computing device can receive the collected ambient noise samples. In some embodiments, means for performing the operations of block 510 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to a transceiver (e.g., 328), electronic storage (e.g., 326), an external resource (e.g., 320), and an ambient noise sample / profile receiving module (e.g., 380).
[0110] At block 515, the processor of the remote computing device can perform operations including generating, using the received audio input, a voice recognition model associated with the location for use in voice and / or speech recognition. For example, where the received audio input is collected from a user's home or business office, some of the regular background noise (e.g., phone ringing, machines, people talking, etc.) can be taken into account when generating a voice recognition model for that environment. The generated voice recognition model can be configured to be used to adjust received audio input for predicted ambient noise of the environment (i.e., location and / or location category). In this way, regular background noise can be filtered out of the received audio input before a more general model (i.e., even one that is customized for a particular user) is used for voice and / or speech recognition. In some embodiments, a plurality of received ambient noise samples can be used to generate voice recognition models, such that each of the generated voice recognition models can be configured to be used at a respective one of the different locations. In some embodiments, means for performing the operations of block 515 can include a processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to electronic storage (e.g., 325, 326), an external resource (e.g., 320), and a voice recognition model generating module (e.g., 395).
[0111] At block 520, the processor of the remote computing device can perform operations including providing (i.e., transmitting) the generated speech recognition model associated with the location to a remote computing device, such as a user equipment. For example, after the remote computing device generates the speech recognition model in block 515, the speech recognition model can be sent to the user equipment. In some embodiments, means for performing the operations of block 520 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to the transceiver (e.g., 328), electronic storage (e.g., 326), external resource (e.g., 320), and speech recognition model transmission module (e.g., 380).
[0112] In some embodiments, the processor can repeat any or all of the operations in blocks 510, 515, and 520 to repeatedly or continuously perform speech and / or voice recognition.
[0113] Figure 5B Method 501 is illustrated that can be performed as part of or as an improvement to method 500.
[0114] At block 525, after the operations in block 510 of method 500, the processor of the remote computing device can perform operations including determining a location category (of the audio sample) based on location information received from the user equipment. At block 525, the processor of the remote computing device can use a location / location category association module (e.g., 390) to determine a location at which the audio input was received / recorded. For example, the processor can identify a location corresponding to GPS information, environmental noise location identification information, and / or communication network information. In some embodiments, means for performing the operations of block 525 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to the electronic storage (e.g., 325, 326), external resource (e.g., 320), and location / location category determination module (e.g., 345). After the operations in block 525, the processor can use the received audio input to generate a speech recognition model associated with the location for use in speech and / or voice recognition in block 515.
[0115] At block 530, after the operations in block 515, the processor of the remote computing device can perform operations including associating the generated speech recognition model with the determined location category. For example, if the received audio input is associated with a location category (such as "church"), then the speech recognition model generated from that audio input will also be associated with the church location category. In some embodiments, means for performing the operations of block 530 can include the processor (e.g., 210, 212, 214, 216, 218, 252, 260, 330, 331) coupled to the electronic storage (e.g., 325, 326), the external resource (e.g., 320), and the location / location category determination module (e.g., 345).
[0116] In some embodiments, the processor can repeat any or all of the operations in blocks 510, 515, 520, 525, and 530 to perform speech and / or voice recognition repeatedly or continuously.
[0117] Various embodiments, including but not limited to various embodiments discussed above in reference to Figures 1A to 5B may be implemented on a variety of remote computing devices, an example of which is illustrated in the form of a server in Figure 6 . With reference to Figures 1A to 6 , the remote computing device 150 can include a processor 331 coupled to volatile memory 602 and large capacity nonvolatile memory, such as a disk drive 603. The remote computing device 150 can also include a peripheral memory access device coupled to the processor 331, such as a floppy disk drive, a compact disc (CD) or digital video disc (DVD) drive 606. The remote computing device 150 can also include a network access port 604 (or interface) coupled to the processor 331 for establishing data connections with a network, such as the Internet and / or a local area network coupled to other system computers and servers. The remote computing device 150 can include one or more antennas 607 for sending and receiving electromagnetic radiation that can be connected to wireless communication links. The remote computing device 150 can include additional access ports for coupling to peripheral devices, external memory, or other devices, such as USB, Firewire, Thunderbolt, etc.
[0118] Various aspects, including but not limited to various embodiments discussed above in reference to Figures 1A-5B may be implemented on a variety of user equipment, an example of which is illustrated in the form of a mobile computing device in Figure 7 . With reference to Figures 1A-7The user equipment 110 can include a first SoC 202 (e.g., SoC-CPU) coupled to a second SoC 204 (e.g., a 5G capable SOC) and a third SoC 706 (e.g., a C-V2X SoC configured for managing V2V, V2I, and V2P communications over D2D links, such as D2D links established in dedicated intelligent transportation systems (ITS) 5.9 GHz spectrum communications). The first, second, and / or third SoCs 202, 204, and 706 can be coupled to internal memory 716, a display 730, a speaker 714, a microphone 207, and a wireless transceiver 208. Additionally, the mobile computing device 110 can include one or more antennas 704 for sending and receiving electromagnetic radiation that can be connected to a wireless transceiver 208 (e.g., a wireless data link and / or cellular transceiver, etc.) that is coupled to one or more processors in the first, second, and / or third SoCs 202, 204, and 706. The mobile computing device 110 can also include menu selection buttons or switches for receiving user inputs.
[0119] The mobile computing device 110 can additionally include sound encoding / decoding (CODEC) circuitry 710 that digitizes sound received from the microphone 207 into data packets suitable for wireless transmission and decodes received sound data packets to generate analog signals that are provided to a speaker to generate sound and to analyze ambient noise or speech. Also, one or more of the processors in the first, second, and / or third SoCs 202, 204, and 706, the wireless transceiver 208, and the CODEC circuitry 710 can include digital signal processor (DSP) circuitry (not shown separately).
[0120] The processors implementing the various embodiments can be any programmable microprocessor, microcomputer or multiple processor chip or chips that can be configured by software instructions (applications) to perform a variety of functions, including the functions described herein. In some communications devices, multiple processors can be provided, such as one processor dedicated to wireless communication functions and one processor dedicated to running other applications. Typically, software applications can be stored in the internal memory before they are accessed and loaded into the processor. The processor can include internal memory sufficient to store the software instructions.
[0121] As used in this application, the terms "component," "module," "system" and the like are intended to refer to a computer-related entity, either hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to being, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a communication device and the communication device can be referred to as a component. One or more components can reside within a process and / or thread of execution and a component can be localized, partially localized, or distributed across two or more processes or threads of execution. Also, these components can execute from various non-transitory computer-readable media having various instructions stored thereon. Components can communicate via local and / or remote processes, function- or procedure-calls, electronic signals, data packets, memory read / writes, and other known computer, processor, and / or process related communication methodologies.
[0122] A number of different cellular and mobile communication services and standards are available or are contemplated in the future, all of which can implement and benefit from aspects. Such services and standards can include, for example, Third Generation Partnership Project (3GPP), Long Term Evolution (LTE) systems, Third Generation wireless mobile communication technology (3G), Fourth Generation wireless mobile communication technology (4G), Fifth Generation wireless mobile communication technology (5G), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications System (UMTS), 3GSM, General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA) systems (e.g., cdmaOne, CDMA1020™), EDGE, Advanced Mobile Phone System (AMPS), IS-136, Time Division Multiple Access (TDMA), Evolution-Data Optimized (EV-DO), Digital Enhanced Cordless Telecommunications (DECT), Worldwide Interoperability for Microwave Access (WiMAX), Wireless Local Area Network (WLAN), Wi-Fi Protected Access I and II (WPA, WPA2), Integrated Digital Enhanced Network (iden), C-V2X, V2V, V2P, V2I, and V2N, among others. Each of these technologies involves the transmission and reception of, for example, voice, data, signaling, and / or content messages. It should be understood that any reference to terminology and / or technical details related to an individual telecommunication standard or technology is for illustrative purposes only and is not intended to limit the scope of the claims to a particular communication system or technology unless the claim language specifically recites.
[0123] The various aspects illustrated and described are provided only as examples of various features. The features shown and described with respect to any given aspect need not be limited to that aspect and can be used in any other desired aspect or combination. Furthermore, the claims are not intended to be limited to the aspects shown and described. For example, one or more operations of a method can be replaced by one or more other operations or combined with one or more other aspects.
[0124] The methods described above and the process flow diagrams provided are only meant to be illustrative examples and are not intended to require or imply that the operations of aspects must be performed in the order given. As one of skill will appreciate, the order of the operations in the preceding aspects can be performed in any order. Phrases such as "thereafter," "then," "next," etc. are not intended to limit the order of the operations; these phrases are used to guide the reader through the description of the methods. Furthermore, any reference to claim elements in the singular, for example, using the articles "one," "each," or "the," is not to be construed as limiting the element to the singular.
[0125] The various illustrative logical blocks, modules, components, circuits, and algorithm operations described in connection with the aspects disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and operations are described above generally in terms of their functionality, without reference to a specific
[0126] The hardware used to implement various illustrative logics, logical blocks, modules, and circuits described in connection with the aspects disclosed herein can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an ASIC, a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of receiver smart objects, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some or all of the operations or methods can be performed by a state machine that has no
[0127] In one or more aspects, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The operations of a method or algorithm disclosed herein can be embodied in a processor-executable software module or processor-executable instructions, which can reside on a non-transitory computer- or processor- readable storage medium. Non-transitory computer- or processor-readable storage media can be any storage media that can be accessed by a computer or a processor. By way of example but not limitation, such non-transitory computer- or processor-readable storage media can include RAM, ROM, EEPROM, FLASH memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage smart objects, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Disk and disc, as used herein, includes compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of non-transitory computer- and processor-readable media. Additionally, the operations of a method or algorithm can reside in one or any combination of the above memory hardware, which can be located in a computing device or a processor. The processor may
[0128] The preceding description of the disclosed aspects is provided for the purpose of enabling any person skilled in the art to make or use the present claim. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the claim. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the claims and the principles and novel features disclosed herein.
Claims
1. A method of speech or voice recognition performed by a processor of a computing device, comprising: determining, using environmental noise, at which location an audio input was received; determining whether the received audio input is part of an audio sampling mode; in response to determining that the received audio input is not part of the audio sampling mode: determining a speech recognition model to use for speech or voice recognition based on the location at which the audio input was received; and performing speech or voice recognition on the audio input using the determined speech recognition model; in response to determining that the received audio input is part of the audio sampling mode: associating a location or location category with the received audio input or an audio profile compiled from the received audio input associated with the environmental noise at the location; and transmitting the audio input and associated location or location category information or compiled audio profile to a remote computing device for generating a speech recognition model for that associated location or location category.
2. The method of claim 1, further comprising: determining the location at which the audio input was received using global positioning system information.
3. The method of claim 1, wherein, The determination of at which location the audio input was received is further based on communication network information.
4. The method of claim 1, wherein, Determining a speech recognition model to use for speech or voice recognition includes: selecting the speech recognition model from a plurality of speech recognition models, wherein each of the plurality of speech recognition models is associated with a different scene category, each scene category having a specified audio profile.
5. The method of claim 1, wherein, Performing speech or voice recognition on the audio input using the determined speech recognition model includes: adjusting the audio input for environmental noise using the determined speech recognition model; and performing speech and / or voice recognition on the adjusted audio input.
6. A computing device, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to cause the computing device to: determine, using environmental noise, at which location an audio input was received; determine whether the received audio input is part of an audio sampling mode; in response to determining that the received audio input is not part of the audio sampling mode: determine a speech recognition model to use for speech or voice recognition based on the location at which the audio input was received; and perform speech or voice recognition on the audio input using the determined speech recognition model; in response to determining that the received audio input is part of the audio sampling mode: associate a location or location category with the received audio input or an audio profile compiled from the received audio input associated with the environmental noise at the location; and transmit the audio input and associated location or location category information or compiled audio profile to a remote computing device for generating a speech recognition model for that associated location or location category.
7. The computing device of claim 6, wherein the processor is further configured to cause the computing device to determine the location at which the audio input was received using global positioning system information.
8. The computing device of claim 6, wherein the determination of where the audio input was received is further based on communication network information.
9. The computing device of claim 6, wherein, The processor is further configured to cause the computing device to perform the following to determine a speech recognition model to use for speech or voice recognition: selecting the speech recognition model from a plurality of speech recognition models stored in the memory, wherein each of the plurality of speech recognition models is associated with a different scene category, each scene category having a specified audio profile.
10. The computing device of claim 6, wherein, The processor is further configured to cause the computing device to perform the following to perform speech or voice recognition on the audio input using the determined speech recognition model: adjusting the audio input for environmental noise using the determined speech recognition model; and performing speech and / or voice recognition on the adjusted audio input.
11. A non-transitory processor-readable medium having stored thereon processor- executable instructions configured to cause a processor of a computing device to perform operations comprising: determining where an audio input was received using environmental noise; determining whether the received audio input is part of an audio sampling mode; in response to determining that the received audio input is not part of the audio sampling mode: determining a speech recognition model to use for speech or voice recognition based on where the audio input was received; and performing speech or voice recognition on the audio input using the determined speech recognition model; in response to determining that the received audio input is part of the audio sampling mode: associating a location or location category with the received audio input or an audio profile compiled from the received audio input associated with the environmental noise at the location; and transmitting the audio input and associated location or location category information or compiled audio profile to a remote computing device for generating a speech recognition model for that associated location or location category.
12. The non-transitory processor-readable medium of claim 11, wherein the stored processor-executable instructions are configured to cause a processor of a computing device to perform operations further comprising: determining the location where the audio input was received using global positioning system information.
13. The non-transitory processor-readable medium of claim 11, wherein, The determination of where the audio input was received is further based on communication network information.
14. The non-transitory processor-readable medium of claim 11, wherein the stored processor-executable instructions are configured to cause a processor of a computing device to perform operations such that determining a speech recognition model to use for speech or voice recognition comprises: selecting the speech recognition model from a plurality of speech recognition models, wherein each of the plurality of speech recognition models is associated with a different scene category, each scene category having a specified audio profile.
15. The non-transitory processor-readable medium of claim 11, wherein the stored processor-executable instructions are configured to cause a processor of a computing device to perform operations such that performing speech or voice recognition on the audio input using the determined speech recognition model comprises: using the determined speech recognition model to adjust the audio input for environmental noise; and performing speech and / or voice recognition on the adjusted audio input.
Citation Information
Patent Citations
Geotagged environmental audio for enhanced speech recognition accuracy
CN102918591A
Method and apparatus for recognizing speech, and method and apparatus for generating noise-speech recognition model
US20150317998A1
Auditory environment recognition
US9288594B1