System and method for managing voice query using pronunciation information

The system addresses pronunciation loss in voice-to-text conversions by using pronunciation information and metadata to generate accurate text queries, enhancing search accuracy for names with multiple pronunciations.

JP2025102873APending Publication Date: 2025-07-08ADAIR GUYS INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025056544
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-07-31
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing voice query systems often lose pronunciation details during the conversion from voice to text, leading to inaccurate searches when words with multiple pronunciations have different meanings, particularly for names with ambiguous pronunciations.

Method used

A system that utilizes pronunciation information, user context, and metadata to generate text queries, incorporating phonetic representations and alternative spellings to accurately identify entities and retrieve relevant content.

Benefits of technology

Enhances the accuracy of voice query responses by considering pronunciation variations, improving the reachability and relevance of search results, especially for names with multiple pronunciations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102873000001_ABST
    Figure 2025102873000001_ABST
Patent Text Reader

Abstract

To provide a system and a method for managing voice query using preferred pronunciation information.SOLUTION: A system receives a voice query at an audio interface and converts the voice query to a text. The system can determine pronunciation information during conversion and generate metadata indicating pronunciation of one or more words of a query, include phoneme information in the text query, or both. The query includes one or more entities that can be more accurately identified based on the pronunciation. The system searches for information, contents, or both in one or more databases based on the generated text query, the pronunciation information, user profile information, search history or trends, and optionally other information. The system identifies one or more entities or content items that match the text query, retrieves the identified information, and provides it to a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a system for managing voice queries, and more particularly, to a system for managing voice queries based on pronunciation information.

Summary of the Invention

Means for Solving the Problems

[0002] In a conversation system, when a user issues a voice query to the system, the utterance is converted to text using an automatic speech recognition (ASR) module. This text then forms the input to the conversation system, which determines a response to the text. For example, if a user says, "Show me a movie by Tom Cruise," the ASR module converts the user's voice to text and issues it to the conversation system. The conversation system simply acts based on the text it receives from the ASR module. Sometimes, in this process, the conversation system loses details of the pronunciation of words or the sounds contained in the user's query. Pronunciation details can, in particular, provide information that can be useful for searches when the same word has two or more pronunciations and the pronunciations correspond to different meanings.

[0003] This disclosure describes a system and method for performing a search and predicting a search query intended by a user based on a plurality of context inputs when the user speaks a query word. The search can be based on a plurality of context inputs including, for example, the user's search history, the user's likes and dislikes, general trends, pronunciation details of the query word, and any other suitable information. An application receives an audio query and generates a text query representing the audio query. The application uses pronunciation information that may be included in metadata associated with the text query included in the text query or that may be included in metadata of entities in a database to more accurately read out search results. In some embodiments, the application generates metadata based on text-to-speech and speech-to-text conversion to improve the reachability of entities from the search query. The present invention provides, for example, the following. (Item 1) A method for responding to an audio query, the method comprising: receiving an audio query at an audio interface; extracting, using a control circuit, one or more keywords from the audio query; generating, using the control circuit, a text query based on the one or more keywords; identifying an entity, wherein identifying the entity is based on the text query and metadata regarding the entity, the metadata comprising one or more alternative text representations of the entity, the one or more alternative text representations being based on the pronunciation of an identifier associated with the entity; reading out a content item associated with the entity and including. (Item 2) The method according to item 1, wherein the one or more alternative text representations comprise a phoneme representation of the entity. (Item 3) The method according to any one of items 1 and 2, wherein the one or more alternative text representations comprise alternative spellings of the entity based on pronunciation. (Item 4) The method according to any one of items 1-3, wherein the one or more alternative text representations of the entity comprise a text string generated based on previous speech → text conversion. (Item 5) The one or more alternative text representations comprise a plurality of alternative text representations, and each alternative text representation among the plurality of alternative text representations is generated by converting a first text representation into an audio file and converting the audio file into a second text representation, wherein the second text representation is not the same as the first text representation, the method according to any one of items 1-4. (Item 6) (Item 6) The method according to any one of items 1-5, wherein identifying the entity is further based on user profile information. (Item 7) The method according to any one of items 1-6, wherein identifying the entity is further based on popularity information associated with the entity. (Item 8) Identifying the entity comprises identifying the plurality of entities, each metadata of which is stored for each entity among the plurality of entities, and determining a respective score for each respective entity among the plurality of entities based on comparing the respective one or more alternative text representations with the text query, and selecting the entity by determining the maximum score, the method according to any one of items 1-7. (Item 9) Further comprising generating a plurality of text queries, the plurality of text queries comprising the text query, and each text query among the plurality of text queries being generated based on respective settings of the speech-to-text module of the control circuit, the method according to any one of items 1-8. (Item 10) Identifying respective entities based on each text query among the plurality of text queries; Determining respective scores for the respective entities based on a comparison with metadata associated with the respective entities of the respective text queries; Identifying the entity by selecting the maximum score among the respective scores; The method according to item 9, further comprising. (Item 11) A system for responding to a voice query, the system comprising: A memory; Means for implementing the steps of the method according to any one of items 1-10; A system comprising. (Item 12) A non-transitory computer-readable medium having encoded instructions, which when executed by a control circuit enable the control circuit to execute the steps of the method according to any one of items 1-10. (Item 13) A system for responding to a voice query, the system comprising means for implementing the steps of the method according to any one of items 1-10.

Brief Description of the Drawings

[0004] The above and other objects and advantages of the present disclosure will become apparent upon consideration of the following detailed description taken in conjunction with the accompanying drawings in which like reference symbols refer to like parts throughout.

[0005]

Figure 1

[0006]

Figure 2

[0007]

Figure 3

[0008]

Figure 4

[0009]

Figure 5

[0010]

Figure 6

[0011]

Figure 7

[0012]

Figure 8

[0013]

Figure 9

[0014] In some embodiments, the present disclosure is directed to a system configured to receive a voice query from a user, analyze the voice query, and generate a text query (e.g., a transcription) for retrieving content or information. The system responds to the voice query based at least in part on the pronunciation of one or more keywords. For example, in the English language, there are multiple words that have the same spelling but different pronunciations. This can particularly apply to people's names. Some examples include the following. TABLE 1 By way of illustration, a user may speak to the system's audio interface, "Show me an interview with Louis." The system may generate an illustrative text query such as the following. Option 1) "Show me an interview with Louis Freeh from Fraud Magazine" Option 2) "Show me an interview with Lewis Black broadcast on CBS" The resulting text query depends on how the user pronounced the word "Louis". If the user pronounced it as "LOO-ee", the system may select Option 1 or apply a higher weight to Option 1. If the user pronounced it as "LOO-his", the system may select Option 2 or apply a higher weight to Option 2. Without considering pronunciation, it is likely that the system would not be able to accurately respond to the voice query.

[0015] In some situations, a voice query that includes a partial name of a person can cause ambiguity in correctly detecting that person (e.g., referred to as an "ambiguous person search query"). For example, when a user vocalizes "Show me a movie starring Tom" or "Show me an interview with Louis", the system will need to determine whether the user is asking for Tom or Louis / Louie / Lewis. In addition to pronunciation information, the system can analyze one or more context inputs such as, for example, the user's search history (e.g., previous queries and search results), the user's likes / dislikes / preferences (e.g., from user profile information), general trends (e.g., of multiple users), popularity (e.g., among multiple users), any other suitable information, or any combination thereof. After the automatic speech recognition (ASR) process, the system retains the pronunciation information in a suitable form (e.g., in the text query itself or in metadata associated with the text query) so as not to be lost.

[0016] In some embodiments, with respect to the pronunciation information for use by the system, the information fields that the system searches within must include pronunciation information for comparison with the query. For example, the information fields can include information about entities that include pronunciation metadata. The system can perform a phoneme conversion process, which takes the user's voice query as input, converts it to text, and the text, when read back, sounds acoustically correct. The system can be configured to use the output of the phoneme conversion process and the pronunciation metadata to determine search results. In an illustrative example, the pronunciation metadata stored for an entity can include the following. [Table 2]

[0017] In some embodiments, the present disclosure is directed to a system configured to receive an audio query from a user, analyze the audio query, and generate a text query (e.g., a transcription) to search for content or information. The information fields that the system searches include pronunciation metadata, alternative text representations of entities, or both. For example, when a user issues an audio query to the system, the system first uses an ASR module to convert the audio to text. The resulting text then forms an input to a conversation system (e.g., one that performs an action in response to the query). By way of illustration, if a user says, "Show me movies by Tom Cruise," the ASR module converts the user's utterance to text and issues the text query to the conversation system. If an entity corresponding to "Tom Cruise" exists in the data, the system matches it to the text "Tom Cruise" and returns appropriate results (e.g., information about Tom Cruise, content featuring Tom Cruise, or identifiers of that content). An entity may be referred to as "reachable" when it exists in the data (e.g., in an information field) and can be accessed directly using the entity title. Reachability is most important for the system to perform a search operation. For example, if certain data (e.g., a movie, an artist, a TV series, or other entity) exists in the system and associated data is stored, but the user cannot access that information, the entity may be referred to as "unreachable." Unreachable entities in a data system represent a failure of the search system.

[0018] The system can identify one or more entities or content items among a plurality of stored information. In some embodiments, the system generates an audio file based on a first text string representing the entity or content item. Based on the first text string and at least one utterance criterion, the system can generate a second text string based on the audio file using a speech-to-text module. The system compares the text strings and stores the second text string if it is not identical to the first text string. In some embodiments, the system generates metadata including the result from text-speech-text conversion and anticipates possible misidentifications when responding to an audio query during a search operation. The metadata can include alternative representations of the entity to improve reachability.

[0019] Figure 1 shows a block diagram of an illustrative system 100 for generating a text query according to some embodiments of the present disclosure. System 100 includes an ASR module 110, a conversation system 120, pronunciation metadata 150, user profile information 160, and one or more databases 170. For example, the ASR module 110 and the conversation system 120, which may be included together in system 199, can be used to implement a query application.

[0020] A user can voice query 101, which includes the utterance "Show me that interview with Louis from last week", to the audio interface of system 199. ASR module 110 is configured to sample, condition, and digitize the received audio input, analyze the resulting audio file, and generate a text query. In some embodiments, ASR module 110 reads information from user profile information 160 and uses it to help generate the text query. For example, voice recognition information about the user can be stored in user profile information 160, and ASR module 110 can use the voice recognition information to identify the user speaking. In a further example, system 199 can include user profile information 160 stored in a suitable memory. ASR module 110 can determine pronunciation information regarding the uttered word "Louis". Since there are two or more pronunciations for the text word "Louis", system 199 generates a text query based on the pronunciation information. Further, the sound "Loo - his" can be converted to text as "Louis" or "Lewis", and thus, context information can help identify the correct entity for the voice query (e.g., Louis as in Louis Farrakhan as opposed to Lewis as in Lewis Black). In some embodiments, conversation system 120 is configured to generate a text query, respond to the text query, or both, based on recognized words from ASR module 110, context information, user profile information 160, pronunciation metadata 150, one or more databases 170, any other information, or any combination thereof. For example, conversation system 120 can generate a text query and then compare the text query to pronunciation metadata 150 regarding multiple entities to determine a match. In a further example, conversation system 120 can compare one or more recognized words to pronunciation metadata 150 regarding multiple entities to determine a match and then generate a text query based on the identified entity.In some embodiments, the conversation system 120 generates a text query with accompanying pronunciation information. In some embodiments, the conversation system 120 generates a text query with embedded pronunciation information. For example, the text query may include a phonetic representation of a word such as "loo-ee" rather than the correct grammatical expression "Louis". In a further example, the pronunciation metadata 150 may include one or more reference phonetic representations against which the text query can be compared.

[0021] The user profile information 160 includes user identification information (e.g., name, identifier, address, contact information), user search history (e.g., previous voice queries, previous text queries, previous search results, feedback regarding previous search results or queries), user preferences (e.g., search settings, favorite entities, keywords included in two or more queries), things the user likes / dislikes (e.g., entities followed by the user within a social media application, user input information), other users connected to the user (e.g., friends, family, contacts within a social networking application, contacts stored on the user device), user voice data (e.g., audio samples, signatures, speech patterns, or files for identifying the user's voice), any other suitable information about the user, or any combination thereof.

[0022] One or more databases 170 include any suitable information for generating a text query, responding to a text query, or both. In some embodiments, the pronunciation metadata 150, the user profile information 160, or both may be included in one or more databases 170. In some embodiments, one or more databases 170 include statistical information regarding multiple users (e.g., search history, content consumption history, consumption patterns). In some embodiments, one or more databases 170 include information about multiple entities including people, places, objects, events, content items, media content associated with one or more entities, or combinations thereof.

[0023] Figure 2 shows a block diagram of an illustrative system 200 for reading content in response to an audio query, according to some embodiments of the present disclosure. System 200 includes an utterance processing system 210, a search engine 220, an entity database 250, and user profile information 240. The utterance processing system 210 can identify an audio file and can analyze the audio file with respect to phonemes, patterns, words, or other elements for which keywords can be identified. In some embodiments, the utterance processing system 210 can analyze the audio input in the time domain, the spectral domain, or both, and can identify words. For example, the utterance processing system 210 can analyze the audio input in the time domain and determine the periods during which speech occurs (e.g., to exclude periods of pauses or silence). The utterance processing system 210 can then analyze each period in the spectral domain and identify phonemes, patterns, words, or other elements for which keywords can be identified. The utterance processing system 210 can output the generated text query, one or more words, pronunciation information, or a combination thereof. In some embodiments, the utterance processing system 210 can read data from the user profile information 240 for speech recognition, utterance recognition, or both.

[0024] The search engine 220 receives the output from the speech processing system 210 and, in combination with the search settings 221 and the context information 222, generates a response to the text query. The search engine 220 may use the user profile information 240 to generate, modify, or respond to the text query. The search engine 220 uses the text query to search through the data in the database of entities 250. The database of entities 250 may include metadata associated with a plurality of entities, content associated with a plurality of entities, or both. For example, the data may include an identifier for the entity, details describing the entity, a title referring to the entity (which may include a phonetic representation or an alternative representation), phrases associated with the entity (which may include a phonetic representation or an alternative representation), links associated with the entity (such as an IP address, a URL, a hardware address), keywords associated with the entity (which may include a phonetic representation or an alternative representation), any other suitable information associated with the entity, or any combination thereof. When the search engine 220 identifies one or more entities that match the keywords of the text query, identifies one or more content items that match the keywords of the text query, or both, the search engine 220 may then provide the user with information, content, or both as a response 270 to the text query. In some embodiments, the search settings 221 include a database, an entity, a type of entity, a type of content, other search criteria, or any combination thereof that affect the generation of the text query, the retrieval of search results, or both. In some embodiments, the context information 222 includes genre information (e.g., for further narrowing the search field), keywords, database identification (e.g., a database likely to contain target information or content), a type of content (e.g., by date, genre, title, format), any other suitable information, or any combination thereof.Response 270 may include, for example, content (e.g., a displayed video), information, a list of search results, a link to content, any other suitable search results, or any combination thereof.

[0025] Figure 3 shows a block diagram of an illustrative system 300 for generating pronunciation information, according to some embodiments of the present disclosure. System 300 includes a text-to-speech engine 310 and a speech-to-text engine 320. In some embodiments, system 300 determines pronunciation information independently of a text or voice query. For example, system 300 may generate metadata regarding one or more entities (e.g., pronunciation metadata 150 of system 100 or metadata stored in a database of entity 250 of system 200, etc.). The text-to-speech engine 310 may identify a first text string 302 that may include an entity name or other identifier likely to be included in a voice query. For example, since it is more likely that the user speaks a voice query that includes a name rather than a numerical or alphanumeric identifier (e.g., the user speaks "Louis" instead of "WIKI04556"), the text-to-speech engine 310 may identify the "name" field of the entity metadata rather than the "ID" field. Based on the first text string, the text-to-speech engine 310 generates an audio output 312 at a speaker or other audio device. For example, the text-to-speech engine 310 may define voice details (e.g., male / female voice, accent, or other details), playback speed, or any other suitable setting that may affect the generated audio output using one or more settings. The speech-to-text engine 320 receives an audio input 313 from the audio output 312 at a microphone or other suitable device (e.g., in addition to or instead of an audio file that may be stored), and generates a text conversion of the audio input 313 (e.g., in addition to or instead of storing an audio file of the recorded audio). The speech-to-text engine 320 may generate a new text string 322 using processing settings. The new text string 322 is compared to the first text string 302. If the new text string 322 is identical to the text string 302, the metadata need not be generated since the voice query may result in an accurate conversion to a text query.If the new text string 322 is not identical to the text string 302, this indicates that the voice query may have been incorrectly converted to the text query. Thus, if the new text string 322 is not identical to the text string 302, the speech-to-text engine 320 includes the new text string 322 within the metadata associated with the entity to which the text string 302 is associated. The system 300 can identify multiple entities and generate metadata for each entity that includes text strings (e.g., the new text string 322, etc.) resulting from the results from the text-to-speech engine 310 and the speech-to-text engine 320. In some embodiments, for a given entity, the text-to-speech engine 310, the speech-to-text engine 320, or both can use two or more settings to generate two or more new text strings. Thus, since the two or more text strings are different from the text string 302, each new text string can then be stored in the metadata. For example, different pronunciations or interpretations of pronunciations resulting from different settings can generate different new text strings, which can be stored in preparation for voice queries from different users. By generating and storing alternative representations (e.g., the text string 302 and the new text string 322), the system 300 can update the metadata and enable more accurate searches (e.g., improve the reachability and accuracy of entity searches).

[0026] In an illustrative example, with respect to an entity, system 300 identifies a title and associated phrases, passes each phrase through a text-to-speech engine 310, saves each respective audio file, and then passes each respective audio file through a speech-to-text engine 320 to obtain an ASR transcription record (e.g., new text string 322). If the ASR transcription record is different from the original phrase (e.g., text string 302), system 300 adds the ASR transcription record to the entity's associated phrases (e.g., as stored in metadata). In some embodiments, system 300 can be fully automated without requiring any manual operation (e.g., user input is not required). In some embodiments, when a user issues a query and does not obtain a desired result, system 300 is alerted. In response, a person manually identifies what should be the correct entity for the query. Incorrect results are stored and provide information for future queries. System 300 addresses potential inaccuracies at the metadata level rather than the system level. Analysis of text strings 302 for many entities can be comprehensive and automatic such that all incorrect examples are identified and resolved beforehand (e.g., prior to the user's voice query). System 300 does not require the user to provide a voice query to generate incorrect examples (e.g., alternative representations). System 300 can be used to emulate the user's interaction with the query system and anticipate potential sources of error in performing searches.

[0027] The user may access content, an application (e.g., for interpreting voice queries), and other features from one or more of, for example, the device (i.e., the user device or audio device), one or more network-connected devices, one or more electronic devices having a display, or combinations thereof. Any of the illustrative techniques of the present disclosure may be implemented by a user device, a device providing a display to the user, or any other suitable control circuit configured to respond to a voice query and generate display content for the user.

[0028] FIG. 4 shows a generalized embodiment of an exemplary user device. The user equipment system 401 may include a display 412, an audio device 414, and a user input interface 410, or may include a set-top box 416 communicatively coupled thereto. In some embodiments, the display 412 may include a television display or a computer display. In some embodiments, the user input interface 410 is a remote control device. The set-top box 416 may include one or more circuit boards. In some embodiments, the one or more circuit boards include a processing circuit, a control circuit, and a storage device (e.g., RAM, ROM, hard disk, removable disk, etc.). In some embodiments, the circuit board includes input / output paths. Each of the user equipment device 400 and the user equipment system 401 may receive content and data via an input / output (hereinafter “I / O”) path 402. The I / O path 402 may provide content and data to a control circuit 404 including a processing circuit 406 and a storage device 408. The control circuit 404 may be used to transmit and receive commands, requests, and other suitable data using the I / O path 402. The I / O path 402 may connect the control circuit 404 (specifically, the processing circuit 406) to one or more communication paths (described below). The I / O function may be provided by one or more of these communication paths, but is shown as a single path in FIG. 4 so as not to unduly complicate the drawing. Although the set-top box 416 is shown in FIG. 4 for illustration purposes, any suitable computing device having a processing circuit, a control circuit, and a storage device may be used in accordance with the present disclosure. For example, the set-top box 416 may be replaced or supplemented by a personal computer (e.g., notebook, laptop, desktop), a network-based server hosting a user-accessible client device, a non-user-owned device, any other suitable device, or any combination thereof.

[0029] The control circuit 404 can be based on any suitable processing circuit such as the processing circuit 406. As referred to herein, a processing circuit is to be understood to mean a circuit based on one or more microprocessors, microcontrollers, digital signal processors, programmable logic devices, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc., and can include a multi-core processor (e.g., dual-core, quad-core, hexa-core, or any suitable number of cores) or a supercomputer. In some embodiments, the processing circuit is distributed across multiple distinct processors or processing units, e.g., multiple of the same type of processing unit (e.g., two Intel Core i7 processors) or multiple different processors (e.g., an Intel Core i5 processor and an Intel Core i7 processor). In some embodiments, the control circuit 404 executes instructions for an application stored in a memory (e.g., the storage device 408). Specifically, the control circuit 404 can be instructed by an application to perform the functions discussed above and below. For example, the application can provide instructions to the control circuit 404 to generate a media guide display. In some implementations, any action performed by the control circuit 404 can be based on instructions received from an application.

[0030] In some client / server-based embodiments, the control circuit 404 includes a communication circuit suitable for communicating with an application server or other network or server. Instructions for performing the functionality described above may be stored on the application server. The communication circuit may include a cable modem, an integrated services digital network (ISDN) modem, a digital subscriber line (DSL) modem, a telephone modem, an Ethernet® card, or a wireless modem to communicate with other devices or any other suitable communication circuit. Such communication may involve the Internet or any other suitable communication network or path. Additionally, the communication circuit may include a circuit (described in more detail below) that enables peer-to-peer communication of user equipment devices or communication of user equipment devices located remotely from each other.

[0031] The memory may be an electronic storage device such as the storage device 408 that is part of the control circuit 404. As referred to herein, the phrase "electronic storage device" or "storage device" is understood to mean any device for storing electronic data, computer software, or firmware, such as any combination of random access memory, read-only memory, hard drives, optical drives, solid-state devices, quantum memory devices, game consoles, game media, or any other suitable fixed or removable storage device. The storage device 408 may be used to store various types of content described herein and the media guide data described above. Non-volatile memory may also be used (e.g., to boot up routines and other instructions). A cloud-based storage device may be used, for example, to complement or replace the storage device 408.

[0032] A user may send commands to the control circuit 404 using the user input interface 410. The user input interface 410, the display 412, or both may include a touch screen configured to provide a display and receive tactile input. For example, the touch screen may be configured to receive tactile input from a finger, a stylus, or both. In some embodiments, the device 400 may include a front-facing screen and a rear-facing screen, multiple front screens, or multiple angled screens. In some embodiments, the user input interface 410 includes one or more microphones, buttons, keypads, any other component configured to receive user input, or a combination thereof, a remote control device. For example, the user input interface 410 may include an alphanumeric keypad and a handheld remote control device with options. In a further example, the user input interface 410 may include a handheld remote control device having a microphone and a control circuit configured to receive and identify voice commands and transmit information to the set-top box 416.

[0033] The audio device 414 can be provided as integrated with each of the other elements of the user device 400 and the user device system 401, or can be a stand-alone unit. The audio components of the video and other content displayed on the display 412 can be played back through the speakers of the audio device 414. In some embodiments, the audio can be distributed to a receiver (not shown), which processes the audio and outputs it via the speakers of the audio device 414. In some embodiments, for example, the control circuit 404 is configured to use the speakers of the audio device 414 to provide an audio queue to the user or other audio feedback to the user. The audio device 414 can include a microphone configured to receive audio inputs such as voice commands and utterances (including, for example, voice queries). For example, the user can speak characters or words, which are received by the microphone and converted to text by the control circuit 404. In a further example, the user can voice a command, which is received by the microphone and recognized by the control circuit 404.

[0034] (For example, for managing voice queries), an application can be implemented using any suitable architecture. For example, a stand-alone application can be fully implemented on each of the user device 400 and the user equipment system 401. In some such embodiments, instructions for the application are stored locally (e.g., within the storage device 408), and data to be used by the application is downloaded on a periodic basis (e.g., from an out-of-band feed, from an Internet resource, or using another suitable approach). The control circuit 404 can read instructions for the application from the storage device 408, process the instructions, and generate any of the displays discussed herein. Based on the processed instructions, the control circuit 404 can determine the content of the actions to be performed when an input is received from the user input interface 410. For example, the up / down movement of the cursor on the display can be indicated by the processed instructions when the input interface 410 indicates that the up / down button has been selected. An application and / or any instructions for implementing any of the embodiments discussed herein can be encoded on a computer-readable medium. A computer-readable medium includes any medium capable of storing data. A computer-readable medium includes, but is not limited to, propagating electrical or electromagnetic signals, which can be transient, or, but is not limited to, volatile and non-volatile computer memories or storage devices such as hard disks, floppy (registered trademark) disks, USB drives, DVDs, CDs, media cards, register memories, processor caches, random access memory (RAM), etc., which can be non-transient.

[0035] In some embodiments, the application is a client / server-based application. Data for use by a thick or thin client, implemented on each of the user device 400 and the user equipment system 401, is read on demand by issuing a request to a server remote from each of the user equipment device 400 and the user equipment system 401. For example, the remote server may store instructions for the application within a storage device. The remote server may use a circuit (e.g., control circuit 404) to process the stored instructions and generate the displays discussed above and below. The client device may receive the display generated by the remote server and may locally display the content of the display on the user device 400. Thus, while the processing of the instructions is performed remotely by the server, the display resulting as a consequence, which may include text, a keyboard, or other visuals, is provided locally on the user device 400. The user device 400 may receive input from the user via the input interface 410 and may transmit those inputs to the remote server to process and generate corresponding displays. For example, the user device 400 may transmit a communication indicating that the up / down button has been selected via the input interface 410 to the remote server. The remote server may process the instructions according to the input and generate a display of the application corresponding to the input (e.g., a display moving the cursor up / down). The generated display is then transmitted to the user device 400 for presentation to the user.

[0036] In some embodiments, the application is downloaded and interpreted by an interpreter or virtual machine (e.g., launched by control circuit 404), or otherwise launched. In some embodiments, the application is encoded in an ETV binary interchange format (EBIF), received by the control circuit as part of a suitable feed, and can be interpreted by a user agent that launches on control circuit 404. For example, the application can be an EBIF application. In some embodiments, the application can be defined by a series of JAVA (registered trademark)-based files that are received and launched by a local virtual machine or other suitable middleware executed by control circuit 404.

[0037] FIG. 5 shows a block diagram of an illustrative network arrangement 500 for responding to a voice query, according to some embodiments of the present disclosure. The illustrative system 500 can represent a situation where a user provides a voice query at user device 550, views content on a display of user device 550, or both. In system 500, there can be more than two types of user devices, but only one is shown in FIG. 5 to avoid unduly complicating the drawing. Additionally, each user can utilize more than two types of user devices, and more than two of each type of user device can also be utilized. User device 550 can be the same as user device 400 of FIG. 4, user equipment system 401, any other suitable device, or any combination thereof.

[0038] A user device 550 illustrated as a wireless-enabled device can be coupled to a communication network 510 (e.g., connected to the Internet). For example, the user device 550 is coupled to the communication network 510 via a communication path (which may include an access point). In some embodiments, the user device 550 can be a computing device coupled to the communication network 510 via a wired connection. For example, the user device 550 can also include a wired connection to a LAN or any other suitable communication link to the network 510. The communication network 510 can be one or more networks including the Internet, a cellular phone network, a mobile voice or data network (e.g., a 4G or LTE network), a cable network, a public switched telephone network, or other types of communication networks or combinations of communication networks. The communication path can include one or more communication paths such as a satellite path, an optical fiber-based path, a cable path, a path supporting Internet communication, a free space connection (e.g., for broadcast or other wireless signals), or any other suitable wired or wireless communication path or combination of such paths. Although the communication path is not depicted between the user device 550 and the network device 520, these devices can communicate directly with each other via communication paths such as those described above, and other short-range two-point communication paths such as a USB cable, an IEEE 1394 cable, a wireless path (e.g., Bluetooth®, infrared, IEEE 802-11x, etc.), or other short-range communication via a wired or wireless path. BLUETOOTH® is a certification mark owned by Bluetooth® SIG, Inc. The devices can also communicate directly with each other through an indirect path via the communication network 510.

[0039] A system 500 as illustrated includes a network device 520 (e.g., a server or other suitable computing device) coupled to a communication network 510 via a suitable communication path. Communication between the network device 520 and the user device 550 can be exchanged via one or more communication paths, but for the sake of avoiding unduly complicating the drawing, in FIG. 5 it is shown as a single path. The network device 520 can include a database and one or more applications (e.g., as an application server, a host server). Multiple network entities can exist and communicate with the network 510, but only one is shown in FIG. 5 for the sake of avoiding unduly complicating the drawing. In some embodiments, the network device 520 can include one source device. In some embodiments, the network device 520 implements an application that communicates with instances of applications on many user devices (e.g., user device 550). For example, an instance of a social media application can be implemented on the user device 550, and the application information is communicated to and from the network device 520 that can store profile information about the user (e.g., so that the current social media feed is available on a device other than the user device 550). In a further example, an instance of a search application can be implemented on the user device 550, and the application information is communicated to and from the network device 520 that can store profile information about the user, search history from multiple users, entity information (e.g., content and metadata), any other suitable information, or any combination thereof.

[0040] In some embodiments, the network device 520 includes one or more types of stored information, such as, for example, entity information, metadata, content, historical communications and search records, user preferences, user profile information, any other suitable information, or any combination thereof. The network device 520 may include an application host database or server, a plugin, a software development kit (SDK), an application programming interface (API), or other software tools configured to provide software (such as, for example, downloaded to a user device), remotely launch software (such as, for example, hosting an application accessed by a user device), or otherwise provide application support to an application on the user device 550. In some embodiments, information from the network device 520 is provided to the user device 550 using a client / server approach. For example, the user device 550 may pull information from the server, or the server may push information to the user device 550. In some embodiments, an application client resident on the user device 550 initiates a session with the network device 520 and may obtain information as needed (such as, for example, when data becomes stale or when the user device receives a request from the user to receive data). In some embodiments, the information may include user information (such as, for example, user profile information, user-generated content). For example, the user information may include current and / or historical user activity information such as content transactions the user is involved in, searches the user has performed, content the user has consumed, whether the user interacts with a social network, any other suitable information, or any combination thereof. In some embodiments, the user information may identify a pattern of a given user over a period of time. As illustrated, the network device 520 includes entity information regarding a plurality of entities.Entity information 521, 522, and 523 includes metadata regarding each entity. Entities for which metadata is stored in network device 520 can be linked to each other, can reference each other, can be described by one or more tags within the metadata, or can be a combination thereof.

[0041] In some embodiments, the application can be implemented on user device 550, network device 520, or both. For example, the application can be implemented as a set of software or executable instructions that are stored in the storage device of user device 550, network device 520, or both and can be executed by the control circuit of each device. In some embodiments, the application can include an audio recording application, a speech-to-text application, a text-to-speech application, a voice-recognition application, or a combination thereof that is implemented as a client / server-based application, where only the client application resides on user device 550 and the server application resides on a remote server (e.g., network device 520). For example, the application can be implemented partially as a client application on user device 550 (e.g., by the control circuit of user device 550) and partially as a server application that is launched on a remote server by the control circuit of the remote server (e.g., the control circuit of network device 520). When executed by the control circuit of the remote server, the application can generate a display and instruct the control circuit to transmit the generated display to user device 550. The server application can instruct the control circuit of the remote device to transmit data for storage on user device 550. The client application can instruct the control circuit of the receiving user device to generate an application display.

[0042] In some embodiments, the arrangement of system 500 is a cloud-based arrangement. The cloud provides access to services such as, among other examples, information storage, search, messaging, or social networking services, and access to any content described above with respect to user devices. The services can be provided within the cloud through a cloud-computing service provider or through other providers of online services. For example, cloud-based services can include storage services, sharing sites, social networking sites, search engines, or other services that are distributed for viewing by others on a device to which user source content is connected. These cloud-based services can enable a user device to store information in the cloud and receive information from the cloud, rather than storing information locally and accessing locally stored information. Cloud resources can be accessed by a user device using, for example, a web browser, a messaging application, a social media application, a desktop application, or a mobile application, and can include an audio recording application, a speech-to-text application, a text-to-speech application, a speech-recognition application, and / or any combination of those access applications. The user device 550 can be a cloud client that relies on cloud computing for application delivery, or the user device 550 can have some functionality without access to cloud resources. For example, some applications launched on the user device 550 can be cloud applications (e.g., applications delivered as a service via the Internet), while other applications can be stored and launched on the user device 550. In some embodiments, the user device 550 can receive information from multiple cloud resources simultaneously.

[0043] In an illustrative example, a user may speak a voice query into user device 550. The voice query is recorded by the audio interface of user device 550, sampled and digitized by application 560, and converted to a text query by application 560. Application 560 may also include pronunciation along with the text query. For example, one or more words of the text query may be represented by phonetic symbols rather than proper spelling. In a further example, pronunciation metadata may be stored along with the text query that includes the phonetic representation of one or more words of the text query. In some embodiments, application 560 transmits the text query and any suitable pronunciation information to network device 520 to search within a database of entities, content, metadata, or combinations thereof. Network device 520 may identify an entity associated with the text query, content associated with the text query, or both, and provide that information to user device 550.

[0044] For example, a user may speak into the microphone of user device 550 “Show me movies by Tom Cruise”. Application 560 may generate the text query “Movies by Tom Cruise” and transmit the text query to network device 520. Network device 520 may identify the entity “Tom Cruise” and then may identify movies linked to the entity. Network device 520 may then transmit content (e.g., video files, trailers, or clips), content identifiers (e.g., movie titles and images), content addresses (e.g., URLs, websites, or IP addresses), any other suitable information, or any combination thereof to user device 550. Since the pronunciations of “Tom” and “Cruise” are generally not ambiguous, application 560 need not generate pronunciation information in this situation.

[0045] In a further example, the user may speak into the microphone of user device 550, “Show me the interview with Louis,” and the user pronounces the name Louis as “loo-ee” rather than “loo-ihs.” In some embodiments, application 560 may generate a text query “Interview with Louis” and transmit the text query to network device 520 along with metadata including the phonetic representation as “loo-ee.” In some embodiments, application 560 may generate a text query “Interview with Loo-ee” and transmit the text query to network device 520, and the text query itself may include pronunciation information (e.g., in this example, the phonetic representation). Since the name Louis is common, there may be many entities that include this identifier. In some embodiments, network device 520 may identify an entity having metadata including a pronunciation tag having “loo-ee” as the phonetic representation. In some embodiments, network device 520 may read trend searches, the user's search history, or other context information and identify the entity that the user is likely referring to. For example, the user may have previously searched for “FBI,” and the entity Louis Freeh (e.g., the former director of the FBI) may include metadata including tags related to “FBI.” Once the entity is identified, network device 520 may then transmit to user device 550 content (e.g., a video file or clip of the interview), content identifiers (e.g., a file title and still image from the interview), a content address (e.g., a URL, website, or IP address for streaming one or more video files of the interview), any other suitable information related to Louis Freeh, or any combination thereof. Since the pronunciation of “Louis” can be ambiguous, application 560 may generate pronunciation information in such situations.

[0046] In an illustrative example, a user may speak to the microphone of user device 550 as “William Djoko”. Application 560 may generate a text query that may not correspond to the correct spelling of the entity. For example, the voice query “William Djoko” may be converted to text as “William gjoka”. This incorrect text conversion may pose difficulties in identifying the correct entity. In some embodiments, the metadata associated with the entity William Djoko includes alternative representations based on pronunciation. The metadata regarding the entity “William Djoko” may include a pronunciation tag (e.g., “related phrases”), as shown in Table 1.

Table 3

[0047] In the above illustrative example, the reachability of the entity William Djoko is improved by memorizing alternative representations, especially since the ASR process can result in a text transformation that is not grammatically correct for the entity name.

[0048] In an illustrative example, the metadata may be generated based on pronunciation for later reference (e.g., by a text query or other search and retrieval processes), rather than in response to a user's voice query. In some embodiments, the network device 520, the user device 550, or both may generate metadata based on pronunciation information. For example, the user device 550 may receive user input of alternative representations of an entity (e.g., based on previous search results or speech-to-text conversion). In some embodiments, the network device 520, the user device 550, or both may automatically generate metadata regarding an entity using a text-to-speech module and a speech-to-text module. For example, the application 560 may identify a text representation of an entity (e.g., a text string of the name of the entity), input the text representation into the text-to-speech module, and generate an audio file. In some embodiments, the text-to-speech module includes one or more settings or criteria used to generate the audio file. For example, the settings or criteria may include language (e.g., English, Spanish, Mandarin), accent (e.g., regional or language-based), voice (e.g., a particular person's voice, male voice, female voice), speed (e.g., the playback time of a relevant portion of the audio file), pronunciation (e.g., regarding multiple phonetic variants), any other suitable settings or criteria, or any combination thereof. The application 560 then inputs the audio file into the speech-to-text module and generates the resulting text representation. If the resulting text representation is not identical to the original text representation, the application 560 may store the resulting text representation in the metadata associated with the entity. In some embodiments, the application 560 repeats this process for various settings or criteria and thus may generate various text representations that may be stored in the metadata. The resulting metadata includes the original text representation, along with variants generated using text-speech-text conversion to anticipate likely variants.Thus, when application 560 receives a voice query from a user and the conversion to text does not exactly match the entity identifier, application 560 can still identify the correct entity. Further, since the metadata includes variants, application 560 does not need to analyze the text query with respect to pronunciation information (e.g., the analysis is performed beforehand rather than in real time).

[0049] Application 560 can include any suitable functionality such as, for example, audio recording, speech recognition, speech-to-text conversion, text-to-speech conversion, query generation, search engine functionality, content reading, display generation, content presentation, metadata generation, database functionality, or combinations thereof. In some embodiments, aspects of application 560 are implemented across two or more devices. In some embodiments, application 560 is implemented on a single device. For example, entity information 521, 522, and 523 can be stored in the memory storage device of user device 550 and accessed by application 560.

[0050] FIG. 6 shows a flowchart of an illustrative process 600 for responding to a voice query based on pronunciation information, according to some embodiments of the present disclosure. For example, a query application can implement process 600 implemented on any suitable hardware such as user device 400 of FIG. 4, user equipment system 401 of FIG. 4, user device 550 of FIG. 5, network device 520 of FIG. 5, any other suitable device, or any combination thereof. In a further example, the query application can be an instance of application 560 of FIG. 5.

[0051] In step 602, the query application receives an audio query. In some embodiments, an audio interface (e.g., audio device 414, user input interface 410, or a combination thereof) may include a microphone or other sensor that receives an audio input and generates an electronic signal. In some embodiments, the audio input is received by an analog sensor, which provides an analog signal that is conditioned, sampled, and digitized to generate an audio file. The audio file can then be analyzed by the query application in steps 604 and 606. In some embodiments, the audio file is stored in a memory (e.g., storage device 408). In some embodiments, the query application includes a user interface (e.g., user input interface 410), which enables the user to record, play back, modify, crop, visualize, or otherwise manage the audio recording. For example, in some embodiments, the audio interface is configured to always receive an audio input. In a further example, in some embodiments, the audio interface is configured to receive an audio input when the user provides an instruction to the user input interface (e.g., by selecting a soft button on a touch screen to start an audio recording). In a further example, in some embodiments, the audio interface is configured to start recording when an audio input is received and a speech or other suitable audio signal is detected. The query application may include any suitable conditioning software or hardware to convert the audio input into a stored audio file. For example, the query application may apply one or more filters (e.g., low-pass, high-pass, notch, or band-pass filters), amplifiers, digitizers, or other conditioning to generate the audio file.In a further example, the query application may apply any suitable processing such as compression, transformation (e.g., spectral transformation, wavelet transformation), normalization, equalization, truncation (e.g., in the time or spectral domain), any other suitable processing, or any combination thereof to the conditioned signal to generate an audio file. In some embodiments, at step 602, the control circuit receives an audio file from a separate application, from a separate module of the query application, based on user input, or in any combination thereof. For example, at step 602, the control circuit may receive an audio query as an audio file stored in a storage device (e.g., storage device 408) for further processing (e.g., steps 604 - 612 of process 600).

[0052] In step 604, the query application extracts one or more keywords from the voice query of step 602. In some embodiments, the one or more keywords may represent the complete voice query. In some embodiments, the one or more keywords may include only important words or parts of utterances. For example, in some embodiments, the query application may identify words within the utterance and select some of those words as keywords. For example, the query application may identify words and select words that are not prepositions from among them. In a further example, the query application may identify only words of at least three characters in length as keywords. In a further example, the query application may identify a keyword as a phrase including two or more words (e.g., to be more descriptive and provide more context), which may be useful for narrowing down potential search fields for relevant content. In some embodiments, the query application uses any suitable criteria for identifying keywords from the audio input, such as keywords like words, phrases, names, locations, channels, media asset titles, or other keywords. The query application may process words using any suitable word detection technique, utterance detection technique, pattern recognition technique, signal processing technique, or any combination thereof. For example, the query application may compare a series of signal templates with a portion of the audio signal to find whether there is a match (e.g., whether a particular word is included in the audio signal). In a further example, the query application may apply a learning technique to better recognize words within the voice query. For example, the query application may collect feedback from the user regarding a plurality of requested content items in relation to a plurality of queries, and thus, may use past data as a training set to make recommendations and read out content. In some embodiments, the query application may store snippets (i.e., short-duration clips) of the recorded audio within the detected utterance and process the snippets.In some embodiments, the query application stores relatively large segments of speech (e.g., greater than 10 seconds) as audio files and processes the files. In some embodiments, the query application can detect words by processing the speech and using continuous calculations. For example, a wavelet transform can be performed on the speech in real time and provide continuous calculations of the speech pattern (e.g., comparable to a reference for identifying words) with some time lag. In some embodiments, the query application can detect words and the user who uttered the words (e.g., speech recognition) in accordance with the present disclosure.

[0053] In some embodiments, at step 604, the query application adds the detected words to a list of words detected within the query. In some embodiments, the query application can store these detected words in memory. For example, the query application can store the words in memory as a set of ASCII characters (i.e., 8-bit codes), patterns (e.g., indicating speech signal criteria used to match words), identifiers (e.g., codes for words), strings, any other data type, or any combination thereof. In some embodiments, the media guide application can add words to memory as they are detected. For example, the media guide application can append a newly detected word to a string of previously detected words, add the newly detected word to an array of cells of previously detected words (e.g., incrementing the cell array size by 1), create a new variant corresponding to the newly detected word, create a new file corresponding to the newly created word, or store one or more words detected at step 604.

[0054] In step 606, the query application determines pronunciation information regarding one or more of the keywords of step 604. In some embodiments, the pronunciation information includes a phonetic representation of one or more of the keywords (e.g., using International Phonetic Alphabet). In some embodiments, the pronunciation information includes one or more alternative spellings of one or more of the keywords for incorporating pronunciation. In some embodiments, in step 606, the control circuit generates metadata associated with a text query that includes the phonetic representation.

[0055] In step 608, the query application generates a text query based on one or more of the keywords of step 604 and the pronunciation information of step 606. The query application may generate the text query by arranging one or more of the keywords in a suitable order (e.g., the order in which they were spoken). In some embodiments, the query application may omit one or more words of the voice query (e.g., short words, prepositions, or any other words determined to be relatively less important). The text query may be generated as a file (e.g., a text file) and stored in a suitable storage device (e.g., storage device 408).

[0056] In step 610, the query application identifies an entity among a plurality of entities in the database based on the text query and the stored metadata regarding the entity. In some embodiments, the metadata includes a pronunciation tag. In some embodiments, the query application may identify the entity by identifying the metadata tag of the content item corresponding to the entity. For example, the content item may include a movie having tags regarding the actors in the movie. If the text query includes an actor, the query application may determine a match and, based on the match, identify the entity as being associated with the content item. By way of illustration, the query application may first identify the entity (e.g., search among the entities) and then read the content associated with the entity, or the query application may first identify the content (e.g., search among the content) and determine whether the entity associated with the content matches the text query. A database arranged by entity, by content, or both may be searched by the query application.

[0057] In some embodiments, the query application identifies the entity based on user profile information. For example, the query application may identify the entity based on previously identified entities from previous voice queries. In a further example, the query application may identify the entity based on popularity information associated with the entity (e.g., based on searches regarding a plurality of users). In some embodiments, the query application identifies the entity based on the user's preferences. For example, if one or more keywords match a preferred entity name or identifier in the user profile information, the query application may identify that entity or weight it more heavily.

[0058] In some embodiments, the query application identifies entities by identifying a plurality of entities (e.g., using metadata stored for each entity), determining a respective score for each of the plurality of entities based on comparing each respective pronunciation tag to the text query, determining a maximum score, and selecting an entity by determining the entity with the maximum score. The score may be based on the number of matches identified between the keywords of the text query and the metadata associated with the entity or content item.

[0059] In some embodiments, the query application identifies two or more entities (e.g., associated metadata) among a plurality of entities based on the text query. The query application may identify content items associated with some or all of the entities of the query. In some embodiments, the query application identifies an entity by comparing at least a portion of the text query to tags of the metadata stored for each entity and identifying a match.

[0060] In step 612, the query application reads content items associated with the entity. In some embodiments, the query application may identify the content items, download the content items, stream the content items, generate the content items for display, or combinations thereof. For example, a voice query may include "Show me the latest Tom Cruise movies", and the query application may provide a link to the movie "Mission Impossible: Fallout" that the user may select to view the video content. In some embodiments, the query application may read multiple contents associated with the entity that match the text query. For example, the query application may read multiple links, video files, audio files, or other contents, or a list of identified content items, according to the present disclosure.

[0061] FIG. 7 shows a flowchart of an illustrative process 700 for responding to a voice query based on an alternative representation, according to some embodiments of the present disclosure. For example, the query application may implement the process 700 implemented on any suitable hardware such as the user device 400 of FIG. 4, the user equipment system 401 of FIG. 4, the user device 550 of FIG. 5, the network device 520 of FIG. 5, any other suitable device, or any combination thereof. In a further example, the query application may be an instance of the application 560 of FIG. 5.

[0062] In step 702, the query application receives an audio query. In some embodiments, the audio interface (e.g., audio device 414, user input interface 410, or a combination thereof) may include a microphone or other sensor that receives an audio input and generates an electrical signal. In some embodiments, the audio input is received at an analog sensor, which provides an analog signal that is conditioned, sampled, and digitized to generate an audio file. The audio file can then be analyzed by the query application in step 704. In some embodiments, the audio file is stored in a memory (e.g., storage device 408). In some embodiments, the query application includes a user interface (e.g., user input interface 410), which enables the user to record, play back, modify, crop, visualize, or otherwise manage the audio recording. For example, in some embodiments, the audio interface is configured to always receive an audio input. In a further example, in some embodiments, the audio interface is configured to receive an audio input when the user provides an instruction to the user interface (e.g., by selecting a soft button on a touch screen to start an audio recording). In a further example, in some embodiments, the audio interface is configured to receive an audio input and start recording when speech or other suitable audio signal is detected. The query application may include any suitable conditioning software or hardware for converting the audio input into a stored audio file. For example, the query application may apply one or more filters (e.g., low-pass, high-pass, notch filter, or band-pass filter), an amplifier, a demimeter, or other conditioning to generate the audio file.In a further example, the query application may apply any suitable processing such as compression, conversion (e.g., spectral conversion, wavelet conversion), normalization, equalization, truncation (e.g., in the time or spectral domain), any other suitable processing, or any combination thereof to the conditioned signal to generate an audio file. In some embodiments, at step 702, the control circuit receives an audio file from a separate application, from a separate module of the query application, based on user input, or in any combination thereof. For example, step 702 may include receiving an audio query as an audio file stored in a storage device (e.g., storage device 408) for further processing (e.g., steps 704-710 of process 700).

[0063] In step 704, the query application extracts one or more keywords from the voice query of step 702. In some embodiments, the one or more keywords may represent the complete voice query. In some embodiments, the one or more keywords include only important words or parts of utterances. For example, in some embodiments, the query application may identify words within the utterance and select some of those words as keywords. For example, the query application may identify words and select words that are not prepositions from among them. In a further example, the query application may identify only words that are at least three characters long as keywords. In a further example, the query application may identify a keyword as a phrase that includes two or more words (e.g., to be more descriptive and provide more context), which may be useful for narrowing down potential search fields for relevant content. In some embodiments, the query application uses any suitable criteria for identifying keywords from the audio input, such as keywords like words, phrases, names, locations, channels, media asset titles, or other keywords. The query application may process words using any suitable word detection technique, utterance detection technique, pattern recognition technique, signal processing technique, or any combination thereof. For example, the query application may compare a series of signal templates with a portion of the audio signal to find out if there is a match (e.g., whether a particular word is included in the audio signal). In a further example, the query application may apply learning techniques to better recognize words within the voice query. For example, the query application may collect feedback from the user regarding a plurality of requested content items in relation to a plurality of queries, and thus, may use past data as a training set to make recommendations and read out content. In some embodiments, the query application may store snippets (i.e., short-duration clips) of the recorded audio within the detected utterance and process the snippets.In some embodiments, the query application stores relatively large segments of the utterance (e.g., greater than 10 seconds) as audio files and processes the files. In some embodiments, the query application may detect words by processing the utterance and using continuous computations. For example, a wavelet transform may be performed on the fly on the utterance, providing continuous computations of the utterance pattern (e.g., comparable to a reference for identifying words) with some latency. In some embodiments, the query application may detect words and the user who uttered the words (e.g., voice recognition) according to the present disclosure.

[0064] In some embodiments, at step 704, the query application adds the detected words to a list of words detected within the query. In some embodiments, the query application may store these detected words in memory. For example, the query application may store the words in memory as a set of ASCII characters (i.e., 8-bit codes), patterns (e.g., indicating the utterance signal criteria used to match the words), identifiers (e.g., codes for the words), strings, any other data type, or any combination thereof. In some embodiments, the media guide application may add words to memory as they are detected. For example, the media guide application may append the newly detected word to a string of previously detected words, add the newly detected word to an array of cells of previously detected words (e.g., incrementing the array cell size by 1), create a new variant corresponding to the newly detected word, create a new file corresponding to the newly created word, or store one or more words detected at step 704.

[0065] In step 706, the query application generates a text query based on one or more keywords from step 704. The query application may generate the text query by arranging the one or more keywords in a suitable order (e.g., the order in which they were spoken). In some embodiments, the query application may omit one or more words of the voice query (e.g., short words, prepositions, or any other words determined to be relatively less important). The text query is generated as a file (e.g., a text file) and may be stored in a suitable storage device (e.g., storage device 408).

[0066] In step 708, the query application identifies an entity based on the text query and metadata regarding the entity. The metadata includes alternative text representations of the entity based on pronunciation. In some embodiments, the query application may identify the entity by identifying metadata tags of content items corresponding to alternative representations of the entity. For example, the content item may include a movie having tags regarding an actor in the movie, and the tags include alternative spellings (e.g., derived from a system such as system 300 or otherwise included in the metadata). When the text query includes an actor, the query application may determine a match and, based on the match, identify the entity as being associated with the content item. By way of illustration, the query application may first identify the entity (e.g., search within the entity) and then read the content associated with the entity, or the query application may first identify the content (e.g., search within the content) and determine whether the entity associated with the content matches the text query. A database arranged by entity, by content, or both may be searched by the query application. The query application may determine a match when one or more words of the text query match an alternative representation of the entity (e.g., as stored in the metadata associated with the entity).

[0067] In some embodiments, the query application identifies an entity based on user profile information. For example, the query application may identify an entity based on an already identified entity from a previous voice query. In a further example, the query application may identify an entity based on popularity information associated with the entity (e.g., based on searches for multiple users). In some embodiments, the query application identifies an entity based on user preferences. For example, if one or more keywords match an alternative representation of a preferred entity name or identifier in the user profile information, the query application may identify that entity or weight it more heavily.

[0068] In some embodiments, the query application identifies a plurality of entities (e.g., with metadata stored for each entity), determines a respective score for each of the plurality of entities based on comparing each respective metadata with a text query, determines a maximum score, and selects an entity by identifying the entity having the maximum score. The score may be based on the number of matches identified between the keywords of the text query and the metadata associated with the entity or content item.

[0069] In some embodiments, the query application identifies two or more entities (e.g., associated metadata) out of a plurality of entities based on a text query. The query application may identify content items associated with some or all of the entities of the query. In some embodiments, the query application identifies an entity by comparing at least a portion of the text query with tags of the metadata stored for each entity and identifying a match.

[0070] In step 710, the query application reads content items associated with the entity. In some embodiments, the query application may identify content items, download content items, stream content items, generate content items for display, or combinations thereof. For example, an audio query may include "show me the latest Tom Cruise movies", and the query application may provide a link to the movie "Mission Impossible: Fallout" that the user may select to view video content. In some embodiments, the query application may read multiple contents associated with an entity that matches the text query. For example, the query application may read multiple links, video files, audio files, or other contents, or a list of identified content items, in accordance with the present disclosure.

[0071] FIG. 8 shows a flowchart of an illustrative process 800 for generating metadata regarding an entity based on pronunciation, according to some embodiments of the present disclosure. For example, the application may implement process 800 implemented on any suitable hardware such as the user device 400 of FIG. 4, the user equipment system 401 of FIG. 4, the user device 550 of FIG. 5, the network device 520 of FIG. 5, any other suitable device, or any combination thereof. In a further example, the application may be an instance of the application 580 of FIG. 5. In a further example, the system 300 of FIG. 3 may implement the illustrative process 800.

[0072] In step 802, the application identifies the entity in which information of a plurality of entities is stored. In some embodiments, the application selects an entity based on a predetermined order. For example, the application may select entities in alphabetical order and perform a part of process 800. In some embodiments, the application identifies an entity when metadata regarding the entity is created. For example, the application may identify an entity when the entity is added to a database (e.g., a database of entities). In some embodiments, the application identifies an entity when a search operation misidentifies the entity and thus an alternative representation is desired to prevent further misidentification. In some embodiments, the application identifies an entity based on user input. For example, the user may indicate an entity to the application based on incorrect search results, unreachable entities, or errors observed within the search results (e.g., in a suitable user interface). In some embodiments, the application need not identify an entity in response to an error in the search results or a predetermined order. For example, the application may randomly select an entity from the entity database and proceed to step 804. In some embodiments, the application may identify an entity based on the popularity of the entity within a search query. For example, greater search effectiveness may be achieved by determining alternative representations for more common entities such that more search queries are correctly answered. In a further example, the application may identify less common or more ambiguous entities and prevent their unreachability since very few search queries may define those entities. The application may apply any suitable criteria to determine the entity to be identified. In some embodiments, the application may identify two or more entities in step 802 and thus steps 804 - 810 may be performed for each identified entity.In some embodiments, the application may identify content items, rather than or in addition to, entities. For example, the application may identify an entity such as a movie and then identify all other important entities associated with that entity and may receive steps 804 - 810.

[0073] In step 804, the application generates an audio file based on a first text string and at least one speaking criterion. The first text string describes the entity identified in step 802. For example, as illustrated in FIG. 3, the application may include a text-to-speech engine 310, which may be configured to generate an audio file. The application may generate audio output from a speaker or other suitable sound generating device that can be detected by a microphone or other suitable detection device. The application may apply one or more settings or speaking criteria in generating and outputting the audio file. For example, aspects of the generated "voice" may be adjusted or otherwise selected based on any suitable criteria. In some embodiments, the at least one speaking criterion includes a pronunciation setting (e.g., how one or more syllables, groups of characters, or words are pronounced, or the phonemes to be used). In some embodiments, the at least one speaking criterion includes a language setting (e.g., defining a language, accent, regional accent, or other language information).

[0074] In an illustrative example that includes multiple speaking criteria, the application generates respective audio files based on the first text string and respective speaking criteria, generates respective second text strings based on the respective audio files, compares the respective second text strings with the first text string, and may store (e.g., within metadata associated with the entity) the respective second text strings if they are not identical to the first text string.

[0075] In an illustrative example, the application converts a first text string into a first audio signal, generates speech at a speaker based on the audio signal, uses a microphone to detect the speech, generates a second audio signal, processes the audio signal, and may generate an audio file. In some embodiments, the application generates speech at the speaker based on at least one speech setting of the text-to-speech module.

[0076] In step 806, the application generates a second text string based on the audio file. The second text string should match the first text string, apart from differences that may arise from text-to-speech conversion or speech-to-text conversion, and should describe the entity identified in step 802. For example, as illustrated in FIG. 3, the application may include a speech-to-text engine 320, which may be configured to receive an audio input or its generated file and convert the audio into a written record (e.g., a text string). The application may receive the audio input at a microphone or other suitable sound detection device. The application may receive, adjust, and apply one or more settings in converting the audio file to text. For example, aspects of adjusting and converting the detected "speech" may be adjusted or otherwise selected based on any suitable criteria.

[0077] In an illustrative example, the application generates playback of the audio file at a speaker, uses a microphone to detect the playback, generates an audio signal, and converts the audio signal into a second text string by identifying one or more words. In some embodiments, the application converts the audio signal into a second text string based on at least one text setting of the speech-to-text module.

[0078] In step 808, the application compares the second text string with the first text string. In some embodiments, the application compares each character of the first and second text strings and determines a match. In some embodiments, the application determines the degree to which the first text string and the second text string match (e.g., the ratio of matching text strings, the number of differences present, the number of keywords that match or do not match). The application may use any suitable technique to determine whether the first and second text strings are identical, similar, or different and the degree to which they are similar or different.

[0079] In step 810, if the application determines that the second text string is not identical to the first text string, the application stores the second text string. In some embodiments, the application stores the second text string in metadata associated with the entity. In some embodiments, step 810 includes the application updating existing metadata based on one or more text queries. For example, when a query is responded to and the search results are evaluated, the application may update the metadata to reflect new learning. If the second text string is determined to be identical to the first text string, no new information is obtained by storing the second text string. However, the comparison instruction of step 808 may be stored in the metadata and increase the confidence in the reachability of the entity via voice queries. For example, if the second text string is identical to the first text string, it may serve to verify existing metadata regarding voice-based queries.

[0080] FIG. 9 shows a flowchart of an illustrative process 900 for reading content associated with an entity of an audio query according to some embodiments of the present disclosure. For example, a query application may implement process 900 implemented on any suitable hardware such as the user device 400 of FIG. 4, the user equipment system 401 of FIG. 4, the user device 550 of FIG. 5, the network device 520 of FIG. 5, any other suitable device, or any combination thereof. In a further example, the query application may be an instance of application 560 of FIG. 5.

[0081] In step 902, the query application receives an audio signal at the audio interface. The system may include a microphone or other audio detection device and may record an audio file based on the audio input to the device.

[0082] In step 904, the query application analyzes the audio signal of step 902 to identify speech. The query application may apply any suitable decimation, conditioning (e.g., amplification, filtering), processing (e.g., in the time or spectral domain), pattern recognition, algorithms, transformation, any other suitable action, or any combination thereof. In some embodiments, the query application uses any suitable technique to identify words, sounds, phrases, or combinations thereof.

[0083] In step 906, the query application determines whether a voice query has been received. In some embodiments, the query application determines that a voice query has been received based on the parameters of the audio signal. For example, a period without speech before and after the query can delimit the scope of the voice query within the recording. In some embodiments, the query application identifies keywords in the order in which they are spoken, applies a sentence or query template to the keywords, and extracts a text query. For example, the placement of nouns, proper nouns, verbs, adjectives, adverbs, and other parts of speech can provide indications of the start and end of the voice query. The query application can apply any suitable criteria and extract text when analyzing the audio signal. In step 908, the query application generates a text query based on the results of steps 904 and 906. In some embodiments, in step 908, the query application can store the text query in a suitable storage device (e.g., storage device 408). In step 906, if the query application determines that no voice query has been received or, alternatively, that a text query cannot be generated based on the audio analyzed in step 904, the query application can return to step 902 and proceed to the step of detecting audio until a voice query is received.

[0084] In step 910, the query application accesses a database regarding entity information. The query application uses the text query of step 908 to search through the information in the database. The query application can apply any suitable search algorithm to identify information, entities, or content in the database.

[0085] In step 912, the query application determines whether the entities in the database of step 910 match the text query of step 908. The query application can identify and evaluate multiple entities and find a match. In some embodiments, the text query includes two or more entities, and the query application searches through the content and determines content items having entities associated within the metadata (e.g., by comparing the text query with the metadata tags of the content items). In some situations, the query application may be unable to identify a match and, in response, may continue the search, search in another database, modify the text query (e.g., return to step 908 (not shown)), return to step 904 and modify the settings used in step 904 (not shown), return an indication that no search results were found, perform any other suitable response, or any combination thereof. In some embodiments, the query application can identify multiple entities, content, or both that match the text query. Step 914 includes the query application identifying the content associated with the text query of step 908. In some embodiments, steps 914 and 910 can be reversed, and the query application can search through the content based on the text query. In some embodiments, the entity can include a content identifier, and thus, steps 910 and 914 can be combined.

[0086] In step 916, the query application reads out the content associated with the text query of step 908. In step 916, for example, the query application can identify the content item, download the content item, stream the content item, generate a content item or a list of content items (e.g., or a list of links to the content items) for display, or any combination thereof.

[0087] The embodiments described above of the present disclosure are presented for illustrative purposes and not for limitation, and the present disclosure is limited only by the following claims. Further, note that the features and limitations described in any one embodiment may be applicable to any other embodiment herein, that a flowchart or example relating to one embodiment may be combined with any other embodiment in a suitable manner, may be performed in a different order, or may be performed in parallel. Additionally, the systems and methods described herein may be implemented in real time. Note also that the systems and / or methods described above may be applied to or used in accordance with other systems and / or methods. This specification discloses embodiments including, but not limited to, the following: (Item 1) A method for responding to an audio query, the method comprising: Receiving an audio query at an audio interface; Extracting one or more keywords from the audio query using a control circuit; Determining pronunciation information regarding the one or more keywords using a control circuit; Generating a text query based on the one or more keywords and the pronunciation information using a control circuit; Identifying an entity among a plurality of entities in a database based on the text query and stored metadata regarding the entity, the metadata comprising a pronunciation tag; Reading out a content item associated with the entity; A method comprising the above. (Item 2) The method according to item 1, wherein the pronunciation information comprises a phoneme of one of the one or more keywords. (Item 3) The method according to item 1, wherein identifying the entity is further based on user profile information. (Item 4) The method according to item 3, wherein identifying the entity is based on a previously identified entity from a previous audio query. (Item 5) Identifying an entity is the method described in Item 1, further based on popularity information associated with the entity. (Item 6) Identifying an entity is identifying a plurality of entities, wherein metadata for each is stored for each entity among the plurality of entities, and determining a respective score for each entity among the plurality of entities based on comparing each respective pronunciation tag to a text query, and selecting an entity by determining a maximum score, the method described in Item 1, including. (Item 7) The entity is a first entity and further includes identifying a second entity among a plurality of entities based on a text query and second metadata regarding a second entity, and the content item is associated with the first entity and the second entity, the method described in Item 1. (Item 8) Identifying an entity among a plurality of entities in a database includes comparing at least a portion of the text query to tags of stored metadata and identifying a match, the method described in Item 1. (Item 9) A first keyword among one or more keywords is associated with two or more pronunciations of the first keyword, the method described in Item 1. (Item 10) The pronunciation information comprises a phonetic representation of a first keyword among one or more keywords, the method described in Item 1. (Item 11) A system for responding to an audio query, the system comprising an audio interface for receiving an audio query, and a control circuit coupled to the audio interface and comprising the control circuit extracts one or more keywords from the audio query, and determines and extracts pronunciation information regarding the one or more keywords generating and extracting a text query based on one or more keywords and pronunciation information; identifying and extracting an entity from a plurality of entities in a database based on the text query and stored metadata regarding the entity, the metadata comprising a pronunciation tag; reading out a content item associated with the entity; A system configured to perform the above. (Item 12) The system according to Item 11, wherein the pronunciation information includes a phoneme of one of the one or more keywords. (Item 13) The system according to Item 11, wherein the control circuit is further configured to identify an entity based on user profile information. (Item 14) The system according to Item 13, wherein the control circuit is further configured to identify an entity based on a previously identified entity from a previous voice query. (Item 15) The system according to Item 11, wherein the control circuit is further configured to identify an entity based on popularity information associated with the entity. (Item 16) The control circuit is identifying a plurality of entities, each metadata being stored for each of the plurality of entities; determining a respective score for each of the plurality of entities based on comparing each respective pronunciation tag with the text query; selecting an entity by determining a maximum score; The system according to Item 11, wherein the control circuit is further configured to identify an entity by the above method. (Item 17) The entity is a first entity, and the control circuit is further configured to identify a second entity among a plurality of entities based on a text query and second metadata regarding the second entity, and the content item is associated with the first entity and the second entity, the system according to item 11. (Item 18) The control circuit is further configured to identify an entity among a plurality of entities in a database by comparing at least a part of the text query with tags of the stored metadata and identifying a match, according to item 11. (Item 19) The first keyword among one or more keywords is associated with two or more pronunciations of the first keyword, the system according to item 11. (Item 20) The pronunciation information includes a phonetic representation of the first keyword among one or more keywords, the system according to item 11. (Item 21) A non-transitory computer-readable medium having encoded instructions that, when executed by a control circuit, receive an audio query at an audio interface, extract one or more keywords from the audio query, determine pronunciation information regarding the one or more keywords, generate a text query based on the one or more keywords and the pronunciation information, identify an entity among a plurality of entities in a database based on the text query and the stored metadata regarding the entity, the metadata including pronunciation tags, read out a content item associated with the entity and cause the control circuit to do so, a non-transitory computer-readable medium. (Item 22) The pronunciation information includes one phoneme of one of the one or more keywords, the non-transitory computer-readable medium according to item 21. (Item 23) The non-transitory computer-readable medium according to item 21, further comprising encoded instructions that, when executed by a control circuit, cause the control circuit to identify an entity based on user profile information. (Item 24) The non-transitory computer-readable medium according to item 23, further comprising encoded instructions that, when executed by a control circuit, cause the control circuit to identify an entity based on a previously identified entity from a previous voice query. (Item 25) The non-transitory computer-readable medium according to item 21, further comprising encoded instructions that, when executed by a control circuit, cause the control circuit to identify an entity based on popularity information associated with the entity. (Item 26) The non-transitory computer-readable medium according to item 21, further comprising encoded instructions that, when executed by a control circuit, identifying a plurality of entities, each metadata of which is stored for each entity of the plurality of entities, determining a respective score for each entity of the plurality of entities based on comparing each respective pronunciation tag with a text query, selecting an entity by determining a maximum score and causing the control circuit to identify the entity. (Item 27) The entity is a first entity, and the non-transitory computer-readable medium according to item 21, further comprising encoded instructions that, when executed by a control circuit, cause the control circuit to identify a second entity among a plurality of entities based on a text query and second metadata regarding the second entity, and the content item is associated with the first entity and the second entity. (Item 28) Further comprising encoded instructions which, when executed by a control circuit, cause the control circuit to compare at least a part of a text query with tags of stored metadata and identify an entity among a plurality of entities in a database by identifying a match, the non-transitory computer-readable medium according to Item 21. (Item 29) The non-transitory computer-readable medium according to Item 21, wherein a first keyword among one or more keywords is associated with two or more pronunciations of the first keyword. (Item 30) The non-transitory computer-readable medium according to Item 21, wherein the pronunciation information comprises a phonetic representation of a first keyword among one or more keywords. (Item 31) A system for responding to an audio query, the system comprising: means for receiving an audio query; means for extracting one or more keywords from the audio query; means for determining pronunciation information regarding one or more keywords; means for generating a text query based on one or more keywords and the pronunciation information; means for identifying an entity among a plurality of entities in a database based on the text query and stored metadata regarding the entity, the metadata comprising pronunciation tags; means for reading out content items associated with the entity; and comprising. (Item 32) The system according to Item 31, wherein the pronunciation information comprises one phoneme of one of the one or more keywords. (Item 33) The system according to Item 31, wherein the means for identifying an entity comprises means for identifying an entity based on user profile information. (Item 34) The system according to Item 33, wherein the means for identifying an entity comprises means for identifying an entity based on a previously identified entity from a previous audio query. (Item 35) The means for identifying an entity comprises the means for identifying an entity based on popularity information associated with the entity, of the system according to Item 31. (Item 36) The means for identifying an entity is a means for identifying a plurality of entities, wherein respective metadata is stored for each of the plurality of entities, a means for determining a respective score for each of the plurality of entities based on comparing respective pronunciation tags with a text query, and a means for selecting an entity by determining a maximum score of the system according to Item 31. (Item 37) The entity is a first entity, and further comprises a means for identifying a second entity among the plurality of entities based on a text query and second metadata regarding a second entity, and the content item is associated with the first entity and the second entity, of the system according to Item 31. (Item 38) The means for identifying an entity among a plurality of entities in a database comprises a means for comparing at least a part of the text query with tags of stored metadata and identifying a match, of the system according to Item 31. (Item 39) A first keyword among one or more keywords is associated with two or more pronunciations of the first keyword, of the system according to Item 31. (Item 40) The pronunciation information comprises a phonetic representation of a first keyword among one or more keywords, of the system according to Item 31. (Item 41) A method for responding to an audio query, the method comprising receiving the audio query at an audio interface, extracting one or more keywords from the audio query using a control circuit, determining pronunciation information regarding the one or more keywords using the control circuit, Using a control circuit, generating a text query based on one or more keywords and pronunciation information; Identifying an entity among a plurality of entities in a database based on the text query and stored metadata regarding the entity, the metadata comprising a pronunciation tag; Reading out a content item associated with the entity; A method comprising the above. (Item 42) The method according to item 41, wherein the pronunciation information comprises a phoneme of one of the one or more keywords. (Item 43) The method according to any one of items 41 - 42, wherein identifying the entity is further based on user profile information. (Item 44) The method according to any one of items 41 - 43, wherein identifying the entity is based on a previously identified entity from a previous voice query. (Item 45) The method according to any one of items 41 - 44, wherein identifying the entity is further based on popularity information associated with the entity. (Item 46) Identifying the entity comprises: Identifying a plurality of entities, wherein respective metadata is stored for each of the plurality of entities; Determining a respective score for each of the plurality of entities based on comparing each respective pronunciation tag with the text query; Selecting an entity by determining a maximum score; The method according to any one of items 41 - 45, comprising the above. (Item 47) The entity is a first entity, and further comprising identifying a second entity among the plurality of entities based on the text query and second metadata regarding a second entity, the content item being associated with the first entity and the second entity; the method according to any one of items 41 - 46. (Item 48) Identifying an entity among a plurality of entities in a database includes comparing at least a portion of a text query with tags of stored metadata and identifying a match, the method according to any of Items 41-47. (Item 49) The first keyword among one or more keywords is associated with two or more pronunciations of the first keyword, the method according to any of Items 41-48. (Item 50) The pronunciation information comprises a phonetic representation of the first keyword among one or more keywords, the method according to any of Items 41-49. (Item 51) A method for responding to an audio query, the method comprising: Receiving an audio query at an audio interface; Using a control circuit to extract one or more keywords from the audio query; Using a control circuit to generate a text query based on the one or more keywords; Identifying an entity based on the text query and metadata regarding the entity, the metadata comprising one or more alternative text representations of the entity based on the pronunciation of an identifier associated with the entity; Reading out a content item associated with the entity; and including. (Item 52) The one or more alternative text representations comprise a phonetic representation of the entity, the method according to Item 51. (Item 53) The one or more alternative text representations comprise an alternative spelling of the entity based on pronunciation, the method according to Item 51. (Item 54) The one or more alternative text representations of the entity comprise a text string generated based on previous speech-to-text conversion, the method according to Item 51. (Item 55) The one or more alternative text representations comprise a plurality of alternative text representations, and each alternative text representation among the plurality of alternative text representations: Converting a first text representation into an audio file; Converting an audio file into a second text representation, where the second text representation is not the same as the first text representation, and The method according to item 51, generated by (Item 56) Identifying an entity, the method according to item 51, further based on user profile information. (Item 57) Identifying an entity, the method according to item 51, further based on popularity information associated with the entity. (Item 58) Identifying an entity Is to identify a plurality of entities, each metadata of which is stored for each entity among the plurality of entities, and Based on comparing each one or more alternative text representations with a text query, determining a respective score for each entity among the plurality of entities, and Selecting an entity by determining a maximum score The method according to item 51, including (Item 59) Further including generating a plurality of text queries, the plurality of text queries comprising text queries, and each text query among the plurality of text queries is generated based on respective settings of the speech-to-text module of the control circuit. The method according to item 51. (Item 60) Identifying each entity based on each text query among the plurality of text queries, and Determining a respective score for each entity based on a comparison with the metadata associated with each entity of each text query, and The method according to item 59, further including identifying an entity by selecting the maximum score of each score. (Item 61) A system for responding to an audio query, the system comprising An audio interface for receiving an audio query, and A control circuit and is provided, wherein the control circuit extracts one or more keywords from an audio query, generates a text query based on the one or more keywords, identifies an entity based on the text query and metadata regarding the entity, the metadata comprising one or more alternative text representations of the entity based on the pronunciation of an identifier associated with the entity, reads out a content item associated with the entity, and is configured to perform the above. (Item 62) The system according to item 61, wherein the one or more alternative text representations comprise a phonetic representation of the entity. (Item 63) The system according to item 61, wherein the one or more alternative text representations comprise an alternative spelling of the entity based on pronunciation. (Item 64) The system according to item 61, wherein the one or more alternative text representations of the entity comprise a text string generated based on previous speech → text conversion. (Item 65) The one or more alternative text representations comprise a plurality of alternative text representations, and the control circuit converts a first text representation into an audio file, converts the audio file into a second text representation, the second text representation being different from the first text representation, and is configured to generate each alternative text representation among the plurality of alternative text representations thereby. (Item 66) The system according to item 61, wherein the control circuit is further configured to identify an entity based on user profile information. (Item 67) The system according to item 61, wherein the control circuit is further configured to identify an entity based on popularity information associated with the entity. (Item 68) The control circuit identifying a plurality of entities, each metadata of which is stored for each entity among the plurality of entities determining a respective score for each respective entity among the plurality of entities based on comparing each one or more alternative text representations with a text query selecting an entity by determining a maximum score The system according to item 61, further configured to identify an entity thereby (Item 69) The control circuit is further configured to generate a plurality of text queries, the plurality of text queries comprising the text query, the control circuit comprising a speech-to-text module, and each text query among the plurality of text queries is generated based on respective settings of the speech-to-text module. The system according to item 61 (Item 70) The control circuit identifying respective entities based on each text query among the plurality of text queries determining a respective score for each entity based on a comparison with metadata associated with each entity of each text query The system according to item 69, further configured to identify an entity by selecting a maximum score of each score (Item 71) A non-transitory computer-readable medium having encoded instructions that, when executed by a control circuit receive a voice query at an audio interface extract one or more keywords from the voice query generate a text query based on the one or more keywords Identifying an entity based on a text query and metadata regarding the entity, the metadata comprising one or more alternative text representations of the entity based on the pronunciation of an identifier associated with the entity, and reading out a content item associated with the entity and causing a control circuit to perform, a non-transitory computer-readable medium. (Item 72) The non-transitory computer-readable medium according to item 71, wherein the one or more alternative text representations comprise a phonetic representation of the entity. (Item 73) The non-transitory computer-readable medium according to item 71, wherein the one or more alternative text representations comprise an alternative spelling of the entity based on pronunciation. (Item 74) The non-transitory computer-readable medium according to item 71, wherein the one or more alternative text representations of the entity comprise a text string generated based on previous speech-to-text conversion. (Item 75) The one or more alternative text representations comprise a plurality of alternative text representations and further comprise encoded instructions that, when executed by a control circuit, cause the control circuit to convert a first text representation into an audio file and convert the audio file into a second text representation, the second text representation being different from the first text representation, and thereby generate each alternative text representation among the plurality of alternative text representations, the non-transitory computer-readable medium according to item 71. (Item 76) The non-transitory computer-readable medium according to item 71, further comprising encoded instructions that, when executed by a control circuit, cause the control circuit to identify the entity based on user profile information. (Item 77) The non-transitory computer-readable medium according to item 71, further comprising encoded instructions which, when executed by a control circuit, cause the control circuit to identify an entity based on popularity information associated with the entity. (Item 78) The non-transitory computer-readable medium according to item 71, further comprising encoded instructions which, when executed by a control circuit, identify a plurality of entities, each metadata of which is stored for each entity of the plurality of entities, determine a respective score for each entity of the plurality of entities based on comparing each one or more alternative text representations with a text query, select an entity by determining a maximum score, and cause the control circuit to identify the entity. (Item 79) The non-transitory computer-readable medium according to item 71, further comprising encoded instructions which, when executed by a control circuit, cause the control circuit to generate a plurality of text queries, each of the plurality of text queries comprising a text query and each text query of the plurality of text queries being generated based on respective settings of the control circuit's speech-to-text module. (Item 80) The non-transitory computer-readable medium according to item 71, further comprising encoded instructions which, when executed by a control circuit, identify respective entities based on respective text queries of the plurality of text queries, determine a respective score for each entity based on a comparison with metadata associated with the respective entity of each text query, and cause the control circuit to identify an entity by selecting a maximum score of the respective scores. (Item 81) A system for responding to a voice query, the system comprising: means for receiving a voice query at an audio interface; means for extracting one or more keywords from the voice query; means for generating a text query based on the one or more keywords; means for identifying an entity based on the text query and metadata regarding the entity, the metadata comprising one or more alternative text representations of the entity based on the pronunciation of an identifier associated with the entity; means for reading out a content item associated with the entity; and a system comprising the same. (Item 82) The system according to item 81, wherein the one or more alternative text representations comprise a phoneme representation of the entity. (Item 83) The system according to item 81, wherein the one or more alternative text representations comprise an alternative spelling of the entity based on pronunciation. (Item 84) The system according to item 81, wherein the one or more alternative text representations of the entity comprise a text string generated based on previous speech → text conversion. (Item 85) The one or more alternative text representations comprise a plurality of alternative text representations, and each alternative text representation among the plurality of alternative text representations means for converting a first text representation into an audio file; means for converting the audio file into a second text representation, the second text representation being different from the first text representation; and a system according to item 81 generated by the same. (Item 86) The system according to item 81, wherein the means for identifying an entity further comprises means for identifying the entity based on user profile information. (Item 87) The system according to item 81, wherein the means for identifying an entity further comprises means for identifying the entity based on popularity information associated with the entity. (Item 88) The means for identifying an entity is means for identifying a plurality of entities, each metadata of which is stored for each entity among the plurality of entities, means for determining a respective score for each entity among the plurality of entities based on comparing each of one or more alternative text representations with a text query, means for selecting an entity by determining a maximum score The system according to item 81, comprising (Item 89) The system according to item 81, further comprising means for generating a plurality of text queries, the plurality of text queries comprising text queries, each text query among the plurality of text queries being generated based on respective settings of the speech-to-text module of the control circuit. (Item 90) means for identifying respective entities based on each text query among the plurality of text queries, means for determining a respective score for each entity based on comparison with metadata associated with the respective entity of each text query, The system according to item 89, further comprising means for identifying an entity by selecting a maximum score of each score. (Item 91) A method for responding to a voice query, the method comprising receiving a voice query at an audio interface, extracting one or more keywords from the voice query using a control circuit, generating a text query based on the one or more keywords using a control circuit, identifying an entity based on the text query and metadata regarding the entity, the metadata comprising one or more alternative text representations of the entity based on the pronunciation of an identifier associated with the entity, Reading content items associated with an entity and A method including the above. (Item 92) The method according to item 91, wherein one or more alternative text representations comprise a phonetic representation of the entity. (Item 93) The method according to any one of items 91-92, wherein one or more alternative text representations comprise an alternative spelling of the entity based on pronunciation. (Item 94) The method according to any one of items 91-93, wherein one or more alternative text representations of the entity comprise a text string generated based on previous speech→text conversion. (Item 95) One or more alternative text representations comprise a plurality of alternative text representations, and each alternative text representation among the plurality of alternative text representations Converting a first text representation into an audio file, and Converting the audio file into a second text representation, wherein the second text representation is not the same as the first text representation, and The method according to any one of items 91-94, generated by the above. (Item 96) Identifying the entity is further based on user profile information, the method according to any one of items 91-95. (Item 97) Identifying the entity is further based on popularity information associated with the entity, the method according to any one of items 91-96. (Item 98) Identifying the entity is Identifying a plurality of entities, wherein metadata for each of the plurality of entities is stored for each entity among the plurality of entities, and Determining a respective score for each entity among the plurality of entities based on comparing each of the respective one or more alternative text representations with a text query, and Selecting the entity by determining the maximum score The method according to any one of items 91-97, including the above. (Item 99) The method according to any one of Items 91-98, further comprising generating a plurality of text queries, the plurality of text queries including text queries, and each text query among the plurality of text queries being generated based on respective settings of the speech-to-text module of the control circuit. (Item 100) Identifying each entity based on each text query among the plurality of text queries; Determining a respective score for each entity based on a comparison with metadata associated with each entity of each text query; The method according to Item 99, further comprising identifying an entity by selecting a maximum score among the respective scores. (Item 101) A method for generating entity metadata for a voice query, the method comprising: Identifying an entity in which information among a plurality of entities is stored; Generating an audio file using a text-to-speech module based on a first text string and at least one speech criterion, the first text string describing the entity; Generating a second text string based on the audio file using a speech-to-text module; Comparing the second text string with the first text string; If not identical to the first text string, storing the second text string in the metadata associated with the entity; and including. (Item 102) The method according to Item 101, wherein at least one speech criterion comprises a pronunciation setting. (Item 103) The method according to Item 101, wherein at least one speech criterion comprises a language setting. (Item 104) At least one speech criterion comprises a plurality of speech criteria, and the method comprises: Using the text-to-speech module, generating respective audio files based on a first text string and respective speech criteria, Using the speech-to-text module, generating respective second text strings based on respective audio files, Comparing each of the second text strings with the first text string, If not identical to the first text string, storing each of the second text strings in metadata associated with the entity The method according to item 101, further comprising. (Item 105) The method according to item 101, further comprising updating metadata based on one or more text queries. (Item 106) The method according to item 101, further comprising storing a phonetic representation of the first text string in metadata associated with the entity. (Item 107) Generating an audio file based on the first text string Converting the first text string into a first audio signal, Generating speech at a speaker based on the audio signal, Detecting the speech using a microphone and generating a second audio signal, Processing the audio signal and generating an audio file The method according to item 101, comprising. (Item 108) Generating speech at a speaker is further based on at least one speech setting of the text-to-speech module, the method according to item 107. (Item 109) Generating a second text string based on an audio file Generating playback of the audio file at a speaker, Detecting the playback using a microphone and generating an audio signal, Converting the audio signal into a second text string by identifying one or more words The method according to item 101, comprising (Item 110) Converting an audio signal into a second text string, which is the method according to item 109 based on at least one text setting of the speech-to-text module. (Item 111) A system for generating entity metadata regarding an audio query, the system comprising a control circuit, The control circuit identifying an entity in which information of a plurality of entities is stored, generating an audio file based on a first text string and at least one speech criterion using an audio interface coupled to the control circuit, wherein the first text string describes the entity, generating a second text string based on the audio file using the audio interface, comparing the second text string with the first text string, if not identical to the first text string, storing the second text string in the metadata associated with the entity and configured to perform. A system. (Item 112) The system according to item 111, wherein at least one speech criterion comprises a pronunciation setting. (Item 113) The system according to item 111, wherein at least one speech criterion comprises a language setting. (Item 114) At least one speech criterion comprises a plurality of speech criteria, and the control circuit generates respective audio files based on the first text string and respective speech criteria using an audio device, generates respective second text strings based on the respective audio files using the audio device, compares each second text string with the first text string, If not identical to the first text string, store each second text string in the metadata associated with the entity, and The system according to item 111, further configured to perform. (Item 115) The control circuit is further configured to update the metadata based on one or more text queries, the system according to item 111. (Item 116) The control circuit is further configured to store the phonetic representation of the first text string in the metadata associated with the entity, the system according to item 111. (Item 117) The audio device includes a speaker and a microphone, and the control circuit Convert the first text string into a first audio signal, and Generate speech at the speaker based on the audio signal, and Detect speech using the microphone and generate a second audio signal, and Process the audio signal and generate an audio file, and By, further configured to generate an audio file based on the first text string, the system according to item 111. (Item 118) The control circuit is further configured to generate speech at the speaker based on at least one speech setting, the system according to item 117. (Item 119) The audio device includes a speaker and a microphone, and the control circuit Generate playback of the audio file at the speaker, and Detect playback at the microphone and generate an audio signal, and Convert the audio signal into a second text string by identifying one or more words, and By, further configured to generate a second text string based on the audio file, the system according to item 111. (Item 120) The system according to item 119, wherein the control circuit is further configured to convert an audio signal into a second text string based on at least one text setting of the speech → text module. (Item 121) A non-transitory computer-readable medium having encoded instructions that, when executed by a control circuit, identifies the entity in which information of a plurality of entities is stored, generates an audio file based on a first text string and at least one speech criterion, wherein the first text string describes the entity, generates a second text string based on the audio file, compares the second text string with the first text string, and stores the second text string in the metadata associated with the entity if it is not identical to the first text string and causes the control circuit to perform. A non-transitory computer-readable medium. (Item 122) The non-transitory computer-readable medium according to item 121, wherein the at least one speech criterion comprises a pronunciation setting. (Item 123) The non-transitory computer-readable medium according to item 121, wherein the at least one speech criterion comprises a language setting. (Item 124) The at least one speech criterion comprises a plurality of speech criteria and further comprises encoded instructions that, when executed by a control circuit, generate respective audio files based on the first text string and respective speech criteria, generate respective second text strings based on the respective audio files, compare the respective second text strings with the first text string, and store the respective second text strings in the metadata associated with the entity if they are not identical to the first text string A non-transitory computer-readable medium according to item 121, causing a control circuit to perform (Item 125) Further comprising encoded instructions that, when executed by a control circuit, cause the control circuit to update metadata based on one or more text queries, the non-transitory computer-readable medium according to item 121. (Item 126) Further comprising encoded instructions that, when executed by a control circuit, cause the control circuit to store a phonetic representation of a first text string in metadata associated with an entity, the non-transitory computer-readable medium according to item 121. (Item 127) Further comprising encoded instructions that, when executed by a control circuit, convert a first text string into a first audio signal, generate speech at a speaker based on the audio signal, detect the speech using a microphone and generate a second audio signal, process the audio signal and generate an audio file, and cause a control circuit to perform, the non-transitory computer-readable medium according to item 121. (Item 128) Further comprising encoded instructions that, when executed by a control circuit, cause the control circuit to generate speech at a speaker based on at least one speech setting of a text-to-speech module, the non-transitory computer-readable medium according to item 127. (Item 129) Further comprising encoded instructions that, when executed by a control circuit, generate playback of an audio file at a speaker, detect the playback using a microphone and generate an audio signal, convert the audio signal into a second text string by identifying one or more words, and cause a control circuit to perform, the non-transitory computer-readable medium according to item 121. (Item 130) Further comprising an encoded instruction, which, when executed by a control circuit, causes the control circuit to convert an audio signal into a second text string based on at least one text setting of the speech-to-text module, the non-transitory computer-readable medium according to item 129. (Item 131) A system for generating entity metadata regarding a voice query, the system comprising means for identifying an entity in which information of a plurality of entities is stored; means for generating an audio file based on a first text string and at least one speech criterion, the first text string describing the entity, means for generating a second text string based on the audio file, means for comparing the second text string with the first text string; means for storing the second text string in metadata associated with the entity if it is not identical to the first text string A system comprising. (Item 132) The system according to item 131, wherein at least one speech criterion comprises a pronunciation setting. (Item 133) The system according to item 131, wherein at least one speech criterion comprises a language setting. (Item 134) At least one speech criterion comprises a plurality of speech criteria, and the system comprises means for generating respective audio files based on the first text string and each speech criterion, means for generating respective second text strings based on each audio file, means for comparing each second text string with the first text string, means for storing each second text string in metadata associated with the entity if it is not identical to the first text string The system according to item 131, further comprising. (Item 135) The system according to Item 131, further comprising means for updating metadata based on one or more text queries. (Item 136) The system according to Item 131, further comprising means for storing a phonetic representation of a first text string in metadata associated with an entity. (Item 137) The means for generating an audio file based on a first text string, means for converting the first text string into a first audio signal, means for generating speech at a speaker based on the audio signal, means for detecting the speech using a microphone and generating a second audio signal, and means for processing the audio signal and generating an audio file The system according to Item 131, comprising. (Item 138) The means for generating speech at a speaker further comprises means for generating speech at a speaker based on at least one speech setting of a text-to-speech module. The system according to Item 137. (Item 139) The means for generating a second text string based on an audio file, means for generating playback of the audio file at a speaker, means for detecting the playback using a microphone and generating an audio signal, and means for converting the audio signal into a second text string by identifying one or more words. The system according to Item 131, comprising. (Item 140) The means for converting an audio signal into a second text string comprises means for converting an audio signal into a second text string based on at least one text setting of a speech-to-text module. The system according to Item 139. (Item 141) A method for generating entity metadata for a voice query, the method comprising: identifying an entity in which information of a plurality of entities is stored, Using a text-to-speech module to generate an audio file based on a first text string and at least one speech criterion, wherein the first text string describes an entity, and Using a speech-to-text module to generate a second text string based on the audio file, and Comparing the second text string with the first text string, and If not identical to the first text string, storing the second text string in the metadata associated with the entity A method comprising the above steps. (Item 142) The method according to item 141, wherein at least one speech criterion comprises a pronunciation setting. (Item 143) The method according to any one of items 141-142, wherein at least one speech criterion comprises a language setting. (Item 144) At least one speech criterion comprises a plurality of speech criteria, and the method further comprises: Using a text-to-speech module to generate respective audio files based on the first text string and each speech criterion, and Using a speech-to-text module to generate respective second text strings based on each audio file, and Comparing each second text string with the first text string, and If not identical to the first text string, storing each second text string in the metadata associated with the entity The method according to any one of items 141-143, further comprising the above steps. (Item 145) The method according to any one of items 141-144, further comprising updating the metadata based on one or more text queries. (Item 146) The method according to any one of items 141-145, further comprising storing a phonetic representation of the first text string in the metadata associated with the entity. (Item 147) Generating an audio file based on the first text string is converting the first text string into a first audio signal, and generating speech at a speaker based on the audio signal, and detecting the speech using a microphone and generating a second audio signal, and processing the audio signal to generate an audio file and includes the method according to any one of Items 141-146. (Item 148) Generating speech at a speaker is the method according to Item 147, further based on at least one speech setting of the text-to-speech module. (Item 149) Generating a second text string based on an audio file is generating playback of the audio file at a speaker, and detecting the playback using a microphone and generating an audio signal, and converting the audio signal into a second text string by identifying one or more words and includes the method according to any one of Items 141-148. (Item 150) Converting an audio signal into a second text string is the method according to Item 149, based on at least one text setting of the speech-to-text module.

Claims

【Claim 1】 The invention described in this specification.

Citation Information

Patent Citations

  • Voice processing method and device

    JP2008145456A

  • Dynamic language model

    JP2015526797A

  • Search engine and method for implementing the same

    JP2017010514A

  • Contextual search on multimedia content

    JP2019032876A

  • Content Analysis to Enhance Voice search

    US20170147576A1