Voice Recognition System
The speech recognition system addresses the challenge of inaccurate text conversion by adjusting contextual weights based on speech segments, improving accuracy and reducing the need for repeated input.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2026-03-04
AI Technical Summary
Conventional speech recognition systems face challenges in accurately converting speech input into text due to the lack of dynamic adjustment of contextual weights, leading to potential errors and the need for repeated input.
A speech recognition system that adjusts contextual weights based on segments of speech input, using a context module to enhance transcription accuracy by identifying relevant contexts and updating weights accordingly.
Improves speech recognition accuracy by dynamically adjusting contextual weights, reducing the likelihood of errors and the need for repeated input, thereby enhancing the system's performance.
Smart Images

Figure 0007824250000001 
Figure 0007824250000002 
Figure 0007824250000003
Abstract
Description
[Technical Field]
[0001] This specification relates to speech recognition. [Background technology]
[0002] Conventional speech recognition systems aim to convert speech input from a user into text output. The text output can be used for a variety of purposes, including, for example, search queries, commands, word processing input, etc. In a typical voice search system, a voice interface receives a user's speech input and provides the speech input to a speech recognition engine. The speech recognition engine converts the speech input into a text search query. The voice search system then submits the text search query to a search engine to retrieve one or more search results. Summary of the Invention [Means for solving the problem]
[0003] Generally, one novel aspect of the subject matter described herein can be embodied in a method including receiving data encoding a speech input; determining a transcription for the speech input, the transcription including, for multiple segments of the speech input, obtaining a first candidate transcription for a first segment of the speech input; determining one or more contexts associated with the first candidate transcription; adjusting a respective weight for each of the one or more contexts; and determining a second candidate transcription for a second segment of the speech input based in part on the adjusted weight; and providing a transcription of the multiple segments of the speech input for output. The methods described herein may be embodied as computer-implemented methods. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method. Configuring one or more computer systems to perform particular operations or actions means that software, firmware, hardware, or a combination thereof is installed on the systems that, when operated, causes the systems to perform those operations or actions. One or more computer programs being configured to perform particular operations or actions means that the one or more programs contain instructions that, when executed by a data processing apparatus, cause the apparatus to perform those operations or actions.
[0004] Another novel aspect of the subject matter described herein may be embodied in a computer-readable medium storing software including instructions executable by one or more computers, the instructions, when executed, causing the one or more computers to perform operations including: receiving data encoding a speech input; determining a transcription for the speech input, the transcription including, for multiple segments of the speech input, obtaining a first candidate transcription for a first segment of the speech input; determining one or more contexts associated with the first candidate transcription; adjusting a respective weight for each of the one or more contexts; and determining a second candidate transcription for a second segment of the speech input based in part on the adjusted weights; and providing for output a transcription of the multiple segments of the speech input.
[0005] Each of the above and other embodiments may optionally include one or more of the following features, alone or in any combination. For example, one embodiment includes all of the following features in combination: The method includes obtaining a first candidate transcription for a first segment of speech input, determining that the first segment of speech input satisfies a stability criterion; and obtaining the first candidate transcription for the first segment of speech input in response to determining that the first segment of speech input satisfies the stability criterion. The stability criterion includes one or more semantic characteristics of the first segment of speech input. The stability criterion includes a time delay that occurs after the first segment of speech input. The second segment of speech input occurs after the first segment of speech input. One or more contexts are received from a user device. The one or more contexts include data including a geographic location of a user, a search history of a user, an interest of a user, or an activity of a user. The method includes storing a plurality of scores for a plurality of contexts and updating the adjusted scores for the one or more contexts in response to adjusting a respective weight for each of the one or more contexts. The method further includes providing the output as a search query to, for example, a search engine, where the search engine may provide one or more search results to a user device in response to the search query. The first candidate transcription includes a word, subword, or group of words.
[0006] Particular embodiments of the subject matter described herein can be implemented to achieve one or more of the following advantages: Compared to conventional speech recognition systems, the speech recognition system can provide more accurate text search queries based on segments of speech input. The system can dynamically improve recognition performance because it adjusts contextual weights based on segments of speech input and determines transcriptions of subsequent segments of speech input based in part on the adjusted weights. Thus, the system can improve speech recognition accuracy. This improved accuracy reduces the likelihood that a user will need to repeat the process of providing speech input for processing by the speech recognition system, thereby enabling the speech recognition system to be utilized to process other speech inputs.
[0007] The details of one or more embodiments of the subject matter herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. It is to be appreciated that aspects and implementations may be combined, and that features described in the context of one aspect or implementation may be realized in the context of other aspects or implementations. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 illustrates an overview of an exemplary speech recognition system. [Figure 2] FIG. 1 illustrates an exemplary context. [Figure 3] FIG. 1 illustrates an exemplary process for determining that stability criteria are met. [Figure 4] 1 is a flowchart of an exemplary method for providing a transcription of a speech input. [Figure 5] 1 is a flowchart of an exemplary method for determining a transcription for an audio input. DETAILED DESCRIPTION OF THE INVENTION
[0009] Like reference symbols in the various drawings indicate like elements.
[0010] FIGURE 1 illustrates an overview of an exemplary speech recognition system 100. The speech recognition system 100 includes one or more computers programmed to receive speech input 110 from a user 10 via a user device 120, determine a transcription of the speech input 110, and provide the transcription of the speech input 110 as output. In the example shown in FIGURE 1, the output may be a search query 150 that is provided to a search engine 160 to obtain search results 170 in response to the search query 150. The one or more search results 170 are then provided to the user device 120. The speech recognition system 100 may be implemented on one or more computers, including, for example, a server, or on a user device.
[0011] The speech recognition system 100 includes a speech recognition engine 140 that communicates with user devices 120 over one or more networks 180. The one or more networks 180 may be voice and / or computer networks, including a wireless cellular network, a wireless local area network (WLAN) or Wi-Fi network, a wired Ethernet network, other wired networks, or any other suitable combination thereof. The user devices 120 may be any suitable type of computing device, including, but not limited to, a mobile phone, a smartphone, a tablet computer, a music player, an e-reader, a laptop or desktop computer, a PDA, or other handheld or mobile device that includes one or more processors and computer-readable media.
[0012] User device 120 is configured to receive voice input 110 from user 10. User device 120 may include or be coupled to, for example, an acoustoelectric transducer or sensor (e.g., a microphone). In response to user 10 entering voice input 110, user device 120 may submit the voice input to speech recognition engine 140. (Generally, this may be done by submitting data representing or encoding the voice input to speech recognition engine 140. Speech recognition engine 140 may process the data to extract the voice input from the received data.)
[0013] The speech recognition engine 140 can recognize the speech input sequentially, for example, it can recognize a first portion 111 of the speech input 110 and then a second portion 112 of the speech input 110. One or more portions of the speech input 110 may be recognized as individual segments of the speech input 110 based on certain stability criteria. A portion may include a word, a subword, or a group of words. In some implementations, one or more segments of the speech input 110 can provide intermediate recognition results that can be used to adjust one or more contexts, as described in more detail below.
[0014] Although the example of a search query is used throughout for illustrative purposes, the voice input 110 can represent any type of voice communication, including voice-based commands, search engine query terms, dictation, a dialogue system, or any other input that uses transcribed speech or invokes a software application that uses the transcribed speech to perform an action.
[0015] Speech recognition engine 140 may be a software component of speech recognition system 100 configured to receive and process speech input 110. In the exemplary system shown in FIG. 1 , speech recognition engine 140 converts speech input 110 into a text search query 150 that is provided to search engine 160. Speech recognition engine 140 includes a speech decoder 142, a context module 144, and a context adjustment module 146. Speech decoder 142, context module 144, and context adjustment module 146 may be software components of speech recognition system 100.
[0016] When the speech recognition engine 140 receives the audio input 110, the audio decoder 142 determines a transcription for the audio input 110. The audio decoder 142 then provides the transcription for the audio input 110 as an output, for example, as a search query 150 to be provided to a search engine 160.
[0017] The speech decoder 142 generates candidate transcriptions for the speech input 110 using a language model. The language model includes probability values associated with words or sequences of words. For example, the language model may be an N-gram model. As the speech decoder 142 processes the speech input, intermediate recognition results may be determined. Each intermediate recognition result corresponds to a stable segment of the transcription of the speech input 110. The stability criteria for determining stable segments of the transcription are described in more detail below with respect to FIG. 3.
[0018] The speech decoder 142 provides each stable segment to the context adjustment module 146. The context adjustment module 146 identifies relevant contexts from the context module 144. Each identified context may be associated with a weight. Initially, a base weight for each context may be assigned according to various criteria, for example, based on the popularity of the context, the temporal proximity of the context (i.e., whether a particular context has been actively used in a recent period), or the recent or global usage of the context. The base weight can create an initial bias based on the likelihood that a user input is associated with a particular context. After identifying relevant contexts, the context adjustment module 146 adjusts the weights for the contexts based on one or more stable segments provided by the speech decoder 142. The weights can be adjusted to indicate the extent to which a transcription of the speech input is associated with a particular context.
[0019] The context module 144 stores the contexts 148 and weights associated with the contexts 148. The context module 144 may be a software component of the speech recognition engine 140 configured to cause the computing device to receive one or more contexts 148 from the user device 120. The speech recognition engine 140 may be configured to store the received contexts 148 in the context module 144. In some examples, the context module 144 may be configured to generate one or more contexts 148 customized for the user 10. The speech recognition engine 140 may be configured to store the generated contexts 148 in the context module 144.
[0020] Context 148 may include, for example, (1) data representing user activity, such as the time interval between repeated voice inputs, eye-tracking information reflecting gaze movements from a forward-facing camera near the screen of the user device, (2) data representing the environment in which the voice input is issued, such as the type of mobile application used, the user's location, the type of device used, or the current time, (3) previous voice search queries submitted to a search engine, (4) data representing the type of voice input submitted to a speech recognition engine, such as a command, request, or search query to a search engine, and (5) entities, e.g., members of a particular category, names of places, etc. Context can be formed, for example, from previous search queries, user information, entity databases, etc.
[0021] 2 illustrates exemplary contexts. The speech recognition engine is configured to store a context 210 related to a "tennis player" and a context 220 related to a "basketball player" in a context module, e.g., context module 144. Context 210 includes entities corresponding to specific tennis players, e.g., "Roger Federer," "Rafael Nadal," and "Novak Djokovic." Context 220 includes entities corresponding to specific basketball players, e.g., "Roger Federer," "Rafael Madall," and "Novak Djokovic."
[0022] The context module 144 may be configured to store weights for the contexts 210, 220. These weights may indicate the extent to which one or more transcriptions of the audio input are associated with the contexts 210, 220. Upon identifying a context 210, 220, the context adjustment module 146 also identifies a weight associated with the context 210, 220.
[0023] When the speech decoder 142 obtains the first candidate transcription of "how many wins does tennis player" for the first segment 111 of the speech input 110, the speech decoder 142 provides the first candidate transcription for the first segment 111 to the context adjustment module 146. The context adjustment module 146 identifies contexts 210, 220 as relevant contexts from the context module 144 and weights associated with the contexts 210, 220. The context adjustment module 146 is then configured to adjust the respective weights for the contexts 210, 220 based on the first candidate transcription for the first segment 111 of the speech input 110. In particular, the context adjustment module 146 may adjust the respective weights for the contexts 210, 220 for use in recognizing subsequent segments of the speech input 110.
[0024] The baseline weights for each context may initially bias speech recognition toward the basketball context, which has a heavier initial weight, due to, for example, the historical popularity of speech inputs related to basketball over tennis. However, speech recognition may be adjusted based on intermediate recognition results to bias toward the tennis context. In this example, a first candidate transcription of speech input 110, "how many wins does tennis player," includes the term "tennis player." The context adjustment module 146 may be configured to adjust the weights for one or more of the contexts based on the term "tennis player" in the first candidate transcription. For example, the context adjustment module 146 may boost the weight for context 210, e.g., from "10" to "90," decrement the weight for context 220, e.g., from "90" to "10," or perform a combination of boosting and decrementing weights.
[0025] The speech decoder 142 may be configured to determine a second candidate transcription for the second segment 112 of the speech input 110 based in part on the adjusted weights. The speech recognition engine 140 may be configured to update the adjusted weights for the contexts 210, 220 in the context module 144 in response to adjusting the respective weights for the contexts. In the above example, when determining a second candidate transcription for the second segment 112 of the speech input 110, the speech decoder 142 may give a heavier weight to the context 210 than to the context 220 based on the adjusted weights. The speech decoder 142 may determine "Roger Federer" as the second candidate transcription for the second segment 112 of the speech input 110 based on the weight given to the context 210.
[0026] In contrast, if the context adjustment module 146 does not adjust the weights for the contexts 210, 220 based on the first candidate transcription for the first segment 111, the speech decoder 142 may determine a second candidate transcription for the second segment 112 based on the reference weights for the contexts 210, 220 stored in the context module 144. If the weight for the context 210 is heavier than the weight for the context 220, the speech decoder may determine a basketball player's name, such as "Roger Bederer," as the second candidate transcription for the second segment 112. Thus, the speech decoder 142 may provide an erroneous recognition result.
[0027] After obtaining the entire transcription of the speech input 110, the speech decoder 142 may provide the transcription of the speech input 110 for output. The output may be provided directly to a user device or may be used for further processing. For example, in FIG. 1 , the output recognition is used as a text search query 150. For example, when the speech decoder 142 determines "Roger Federer" as the second candidate transcription for the second segment 112 of the speech input 110, the speech decoder 142 may output the entire transcription "how many wins does tennis player Roger Federer have?" as the search query 150 to the search engine 160.
[0028] The search engine 160 performs a search using the search query 150. The search engine 160 may include a web search engine coupled to the speech recognition system 100. The search engine 160 may determine one or more search results 170 in response to the search query 150. The search engine 160 provides the search results 170 to the user device 120. The user device 120 may include a display interface for presenting the search results 170 to the user 10. In some examples, the user device 120 may include an audio interface for presenting the search results 170 to the user 10.
[0029] 3 illustrates an exemplary process for determining that a stability criterion is met for a given segment. The audio decoder 142 is configured to determine that this portion of the audio input 110 meets the stability criterion.
[0030] When the speech decoder 142 receives a portion 311 of the speech input 310, it may be configured to determine whether the portion 311 of the speech input 310 satisfies a stability criterion, which indicates whether the portion is susceptible to being modified by further speech recognition.
[0031] The stability criterion may include one or more semantic characteristics. If a portion of the speech input is semantically expected to be followed by a word, the speech decoder 142 may determine that the portion does not meet the stability criterion. For example, when the speech decoder 142 receives portion 311 of the speech input 310, it may determine that portion 311 is semantically expected to be followed by a word. The speech decoder 142 then determines that portion 311 does not meet the stability criterion. In some implementations, when the speech decoder 142 receives "mine" as part of the speech input, it may determine that the portion "mine" is semantically not expected to be followed by a word. The speech decoder 142 may then determine that the portion "mine" meets the stability criterion for the segment. The speech decoder 142 may provide the segment to the context adjustment module 146 to adjust the weight for the context.
[0032] The speech decoder 142 may also determine that a portion does not meet the stability criterion if another sub-word is semantically expected to follow the portion. For example, when the speech decoder 142 receives "play" as portion 312 of the speech input 310, it may determine that a word is semantically expected to follow portion 312 because sub-words such as "player-er," "play-ground," and "play-off" can follow portion 312. The speech decoder 142 then determines that portion 311 does not meet the stability criterion. In some implementations, when the speech decoder 142 receives "player" as part of the speech input, it may determine that a word is not semantically expected to follow the portion "player." The speech decoder 142 may then determine that the portion "player" meets the stability criterion for the segment. The speech decoder 142 may provide the segment to the context adjustment module 146 to adjust the weight for the context.
[0033] In some implementations, the stability criterion may include a time delay occurring after a portion of the audio input 310. The audio decoder 142 may determine that this portion of the audio input 310 meets the stability criterion if the time delay after this portion of the audio input 310 has a duration that meets a threshold delay value. When the audio decoder 142 receives this portion of the audio input 310, it may measure the time delay from the moment this portion is received to the moment a subsequent portion of the audio input 310 is received. The audio decoder 142 may determine that this portion meets the stability criterion if the time delay exceeds the threshold delay value.
[0034] 4 is a flowchart of an exemplary method 400 for determining a transcription for a received audio input. For purposes of explanation, the method 400 will be described with respect to a system that performs the method 400.
[0035] The system processes the received speech input in the order in which it was spoken (410) and determines a portion of the speech input as a first segment. The system obtains a first candidate transcription for the first segment of the speech input (420). If the system obtains the first candidate transcription for the first segment, it may determine whether the first segment of the speech input meets a stability criterion. If the first segment of the speech input meets the stability criterion, the system may obtain the first candidate transcription for the first segment. If the first segment of the speech input does not meet the stability criterion, the system may not obtain the first candidate transcription. The system may then receive one or more portions of the speech input, recognize a new first segment of the speech input, and determine whether the new first segment of the speech input meets the stability criterion. The system may determine that the first segment of the speech input meets the stability criterion using process 300, as described above with reference to FIG. 3.
[0036] The system determines one or more contexts associated with the first segment from the set of contexts (430). The specific context associated with the first segment can also be determined based on the context provided by the first segment. For example, specific keywords in the first segment can be identified as keywords associated with a specific context. Referring again to FIG. 2, the system may identify a context associated with a "tennis player" and a context associated with a "basketball player." The tennis player context can be associated with keywords such as "Roger Federer," "Rafael Nadal," and "Novak Djokovic." The basketball player context can be associated with keywords such as "Roger Federer," "Rafael Madall," and "Novak Jocovich." The system may be configured to store a weight for each context. When the system identifies a context, the system may identify a respective weight for the context. Each weight for the context indicates the extent to which one or more transcriptions of the speech input are associated with the context.
[0037] The system adjusts (440) a respective weight for each of the one or more contexts. The system may adjust the respective weight for each context based on a first candidate transcription of the speech input. For example, a first candidate transcription of the speech input, "how many wins does tennis player," includes the term "tennis player." The system may be configured to adjust the weight for the context based on the term "tennis player" in the first candidate transcription. For example, the system can boost the weight for the context, e.g., from "10" to "90," decrement the weight for the context, e.g., from "90" to "10," or perform a combination of boosting and decrementing weights.
[0038] In some implementations, only the weight of the most relevant context is adjusted (e.g., increased) while all other contexts are held constant. In some other implementations, all other contexts are decremented while the most relevant contexts are held constant. Furthermore, any suitable combination of the two can be performed. For example, a relevant context may be increased by a different amount than the amount another context is decremented.
[0039] The system determines a second candidate transcription for a second segment of the speech input based in part on the adjusted weights (450). The system may update the adjusted weights for the contexts in response to adjusting the respective weights for the contexts. For example, the system may assign a heavier weight to a first context identified as being more relevant to the first segment than a second context based on the adjusted weights. The speech decoder may determine a second candidate transcription for the second segment of the speech input based on the adjusted weighted contexts. This process continues until there are no more portions of the speech input to recognize.
[0040] 5 is a flowchart of an exemplary method 500 for conducting a voice search. For purposes of explanation, the method 500 will be described with respect to a system that performs the method 500.
[0041] The system receives speech input (510). The system may be configured to receive speech input from a user. The system may receive each segment of the speech input in real time while the user is speaking.
[0042] Upon receiving the speech input, the system determines a transcription for the speech input (520). The system may determine the transcription as described above with respect to FIG. 4, for example. After the system determines an overall transcription of the speech input, the system provides the transcription of the speech input for output (530). The system may provide the output as a text search query. The system can perform a search using the text search query and obtain search results. The system may provide the search results to the user. In some implementations, the system can provide a display interface for presenting the search results to the user. In other implementations, the system can provide an audio interface for presenting the search results to the user.
[0043] Embodiments of the subject matter and operations described herein can be implemented in digital electronic circuitry, or as computer software, firmware, or hardware including the structures disclosed herein and their structural equivalents, or one or more combinations thereof. Embodiments of the subject matter described herein can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to an appropriate receiver device and execution by the data processing device. The computer storage medium can be, or can be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these. Furthermore, the computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal, rather than a propagated signal. The computer storage medium may be, or may be included in, one or more separate physical components or media, for example, multiple CDs, disks, or other storage devices.
[0044] The operations described herein may be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or on data received from other sources.
[0045] The term "data processing apparatus" encompasses all kinds of apparatus, devices, and machines for processing data, including, by way of example, programmable processing units, computers, systems on chips, personal computer systems, desktop computers, laptops, notebooks, netbook computers, mainframe computer systems, handheld computers, workstations, network computers, application servers, storage devices, or consumer electronic devices such as cameras, camcorders, set-top boxes, mobile devices, video game consoles, handheld video game devices, or peripheral devices such as switches, modems, and routers, or generally any kind of computing or electronic device, or a combination thereof. The apparatus may include special-purpose logic circuitry, e.g., a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program, e.g., code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or one or more combinations thereof. The apparatus and execution environments may implement a variety of different computing model infrastructures, such as web services, distributed computing infrastructures, and grid computing infrastructures.
[0046] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted, declarative or procedural, and can be deployed in any form, such as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program can, but need not, correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple cooperating files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program can be deployed to run on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0047] The processes and logic flows described herein may be performed by one or more programmable processing units that execute one or more computer programs to perform actions by processing input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, for example, an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0048] Processing units suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processing units of any kind of digital computer. Generally, a processing unit receives instructions and data from a read-only memory and / or a random-access memory. The basic elements of a computer are a processing unit for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, e.g., magnetic disks, magneto-optical disks, or optical disks, for storing data, or is operatively coupled to receive data from and / or transfer data to such mass storage devices. However, a computer need not have such devices. Furthermore, a computer can be embedded in another device, e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, a network routing device, or a portable storage device (e.g., a universal serial bus (USB) flash drive). Suitable devices for storing computer program instructions and data include all forms of non-volatile memory, by way of example only, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, as well as media and memory devices, including magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processing unit and memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0049] Embodiments of the subject matter described herein can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to a user to enable interaction with the user, as well as a keyboard and pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to achieve user interaction; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic input, speech input, or tactile input. Furthermore, a computer can interact with a user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0050] Embodiments of the subject matter described herein can be implemented in a computing system that includes back-end components, e.g., data servers, or middleware components, e.g., application servers, or front-end components, e.g., client computers having graphical user interfaces or web browsers that enable users to interact with implementations of the subject matter described herein, or routing devices, e.g., network routers, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks (“LANs”) and wide area networks (“WANs”), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0051] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server sends data (e.g., HTML pages) to client devices (e.g., to display the data to and receive user input from a user interacting with the client device). Data generated at the client device (e.g., a result of a user interaction) can be received at the server from the client device.
[0052] One or more computer systems may be configured to perform certain actions by having software, firmware, hardware, or a combination thereof installed on the system that, when operated, causes the system to perform those actions. One or more computer programs may be configured to perform certain actions by containing instructions that, when executed by a data processing device, cause the device to perform the actions.
[0053] While this specification contains numerous specific implementation details, these should not be construed as limitations on the scope of any invention or claimable therein, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable subcombination. Furthermore, while features are described above as operating in particular combinations, and in some cases may initially be claimed as such, one or more features in a claimed combination may optionally be deleted from the combination, and the claimed combination may be subject to subcombinations or variations of subcombinations.
[0054] Similarly, although operations are shown in a particular order in the figures, this should not be understood as requiring such operations to be performed in the particular order shown or as numbered, nor should it be understood that all of the illustrated operations are required to achieve desired results. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that, in general, the above-described program components and systems may be integrated into a single software product or packaged as multiple software products.
[0055] Specific embodiments of the subject matter have been described above. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results. Furthermore, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or the numbering, to achieve desirable results. In some implementations, multitasking and parallel processing may be advantageous. Accordingly, other embodiments are within the scope of the following claims. [Explanation of symbols]
[0056] 10 users 100 Voice Recognition System 110 Voice Input 111 First Part 112 Second Part 120 user devices 140 Speech Recognition Engine 142 Audio Decoder 144 Context Module 146 Context Adjustment Module 148 Context 150 search queries 160 search engines 170 results 180 Network 210 Context 220 Context 310 Voice Input 311 parts 312 parts
Claims
1. receiving, at an automatic speech recognition (ASR) system implemented on the user device, a speech input spoken by a user of the user device for performing an action; determining, by the ASR system, a specific context associated with a first segment of the speech input, the specific context comprising a list of named entities corresponding to the specific context and stored in a context module of the ASR system; processing a second segment of the speech input that occurs after the first segment based on the particular context using a language model including probability values associated with words or sequences of words, by the ASR system to generate a transcription for the second segment of the speech input, the transcription including one named entity in the list of named entities that corresponds to the particular context; A method comprising:
2. The method of claim 1 , wherein the language model comprises an N-gram language model.
3. The method of claim 1 , wherein the voice input from the user is configured to invoke a software application that uses the transcription of the voice input to perform the action.
4. The method of claim 1 , wherein the particular contexts associated with the speech input include respective weights that indicate the likelihood that the speech input is associated with the particular context.
5. 2. The method of claim 1, wherein determining the particular context associated with the voice input comprises determining the particular context based on a type of software application launched by the voice input that performs the action.
6. 10. The method of claim 1, wherein determining the particular context associated with the speech input comprises determining the particular context based on data describing a type of the speech input received at the ASR system.
7. The method of claim 1 , wherein the particular context is customized for the user.
8. The method of claim 1 , wherein the user device comprises a microphone configured to capture the speech input spoken by the user and provide the speech input to the ASR system.
9. The method of claim 1 , wherein determining the particular context associated with the speech input comprises determining the particular context based on a type of the user device.
10. The method of claim 1 , wherein the user device communicates with the server over a wireless network.
11. 1. A user device, comprising: A microphone and data processing hardware; and memory hardware in communication with the data processing hardware and storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations including: receiving, at an automatic speech recognition (ASR) system implemented on the user device, speech input spoken by a user of the user device for performing an action; determining, by the ASR system, a specific context associated with a first segment of the speech input, the specific context comprising a list of named entities corresponding to the specific context and stored in a context module of the ASR system; and processing, by the ASR system, a second segment of the speech input that occurs after the first segment based on the particular context using a language model including probability values associated with words or sequences of words to generate a transcription for the second segment of the speech input, wherein the generated transcription includes one named entity in the list of named entities that corresponds to the particular context.
12. The user device of claim 11 , wherein the language model comprises an N-gram language model.
13. The user device of claim 11 , wherein the voice input from the user is configured to launch a software application that uses the transcription of the voice input to perform the action.
14. The user device of claim 11 , wherein the particular contexts associated with the speech input include respective weights indicating the likelihood that the speech input is relevant to the particular context.
15. 12. The user device of claim 11, wherein determining the particular context associated with the voice input comprises determining the particular context based on a type of software application launched by the voice input that performs the action.
16. 12. The user device of claim 11, wherein determining the particular context associated with the voice input comprises determining the particular context based on data describing a type of the voice input received at the user device.
17. The user device of claim 11 , wherein the particular context is customized for the user.
18. The user device of claim 11 , wherein the user device comprises a microphone configured to capture the voice input spoken by the user and provide the voice input to the user device.
19. The user device of claim 11 , wherein determining the particular context associated with the voice input comprises determining the particular context based on a type of the user device.
20. The user device of claim 11 , wherein the user device communicates with the server over a wireless network.
Citation Information
Patent Citations
Device and method for translation and recording medium
JP2001101187A
On-vehicle speech recognition apparatus and speech recognition system using the same
JP2004037813A
Distributed speech recognition system and its method
JP2006079079A
Context-based speech recognition grammar selection
JP2011513795A
Dynamic language model
JP2015526797A