Adaptive text-to-speech output

By estimating language proficiency and analyzing context, the complexity of text-to-speech output is adaptively adjusted, solving the problem that TTS systems cannot adapt to the language proficiency of different users, and improving user comprehension and system adaptability.

CN116504221BActive Publication Date: 2026-07-21GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GOOGLE LLC
Filing Date
2016-12-29
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing TTS systems cannot adapt to the different language proficiency levels of different users, making it difficult for users with limited language proficiency to understand complex text-to-speech output. Furthermore, users' comprehension ability varies with specific contexts, especially in environments with background noise.

Method used

The language proficiency estimator determines the user's language proficiency, adjusts the complexity of the text-to-speech output, and generates synthesized speech that adapts to the user's language proficiency and context. This includes selecting or modifying the complexity and structure of text fragments and making adaptive adjustments using the user's historical data and context information.

Benefits of technology

It increases the likelihood that users will understand the text-to-speech output, enhances the user experience in different language proficiency and contexts, and improves the adaptability and effectiveness of the TTS system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116504221B_ABST
    Figure CN116504221B_ABST
Patent Text Reader

Abstract

The present invention relates to adaptive text-to-speech output. In some implementations, one or more computers determine a language proficiency of a user of a client device. The one or more computers then determine a text segment for output by a text-to-speech module based on the determined language proficiency of the user. After determining the text segment for output, the one or more computers generate audio data of a synthesized utterance that includes the text segment. The audio data of the synthesized utterance that includes the text segment is then provided to the client device for output. An improved user interface is provided through better text-to-speech conversion.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Case Analysis

[0002] This application is a divisional application of Chinese Invention Patent Application No. 201680080197.1, filed on December 29, 2016.

[0003] Cross-reference to related applications

[0004] This application claims priority to U.S. Application Serial No. 15 / 009,432, filed January 28, 2016, entitled “ADAPTIVE TEXT-TO-SPEECH-OUTPUTS,” the entire contents of which are incorporated herein by reference. Technical Field

[0005] This specification generally relates to electronic communications. Background Technology

[0006] Speech synthesis refers to the artificial creation of human speech. Speech synthesizers can be implemented in software or hardware components to generate speech output corresponding to text. For example, text-to-speech (TTS) systems typically convert normal spoken text into speech by concatenating recorded speech segments stored in a database. Summary of the Invention

[0007] As a larger portion of electronic computing has shifted from desktop to mobile environments, voice synthesis has become increasingly important for user experience. For example, the growing use of smaller mobile devices without displays has led to a surge in the use of text-to-speech (TTS) systems to access and use content displayed on mobile devices.

[0008] This specification discloses an improved user interface, particularly enhancing computer-to-user communication through improved TTS (Text-to-Speech).

[0009] A particular problem with existing TTS systems is their inability to adapt to varying language proficiency levels among users. This lack of flexibility often hinders users with limited language proficiency from understanding complex text-to-speech output. For example, non-native speakers using TTS systems may struggle to understand text-to-speech output due to their limited language proficiency. Another issue with existing TTS systems is that a user's immediate ability to understand text-to-speech output can vary depending on the specific user context. For instance, some user contexts include background noise, which can make it even more difficult to understand longer or more complex text-to-speech output.

[0010] In some implementations, the system adjusts the text used for text-to-speech output based on the user's language proficiency to increase the likelihood that the user can understand the output. For example, the system can infer the user's language proficiency from prior user activity and use this information to adjust the text-to-speech output to an appropriate complexity commensurate with that proficiency. In some examples, the system obtains multiple candidate text fragments corresponding to different levels of language proficiency. The system then selects the candidate text fragment that best matches and most closely corresponds to the user's language proficiency and provides a synthesized utterance of the selected text fragment for output to the user. In other examples, the system modifies the text in a text fragment to better correspond to the user's language proficiency before generating the text-to-speech output. Various aspects of the text fragment can be adjusted, including vocabulary, sentence structure, length, etc. The system then provides a synthesized utterance of the modified text fragment for output to the user.

[0011] In situations where systems discussed in this paper collect or may use personal information about users, users can be given the opportunity to control whether programs or features collect personal information (e.g., information about a user's social networks, social actions or activities, occupation, user preferences, or current location), or to control whether and / or how content that may be more relevant to the user is received from content servers. Furthermore, certain data can be anonymized in one or more ways before storage or use, thereby removing personally identifiable information. For example, a user's identity can be anonymized to make it impossible to determine the user's personally identifiable information, or, where location information is available, the user's geographic location can be generalized (e.g., to the city, zip code, or state level) to make it impossible to determine the user's specific location. Therefore, users can control how information about themselves is collected and how it is made available to content servers.

[0012] In one aspect, a computer-implemented method may include: determining the language proficiency of a user of a client device by one or more computers; determining, based on the determined language proficiency of the user, a text segment for output by a text-to-speech module by the one or more computers; generating audio data of synthesized speech including the text segment by the one or more computers; and providing the audio data of the synthesized speech including the text segment to the client device by the one or more computers.

[0013] Other versions include corresponding system and computer programs configured to perform actions of methods coded on computer storage devices.

[0014] One or more implementations may include the following optional features. For example, in some implementations, the client device displays a mobile application using a text-to-voice interface.

[0015] In some implementations, determining the user's language proficiency includes inferring the user's language proficiency based at least on previous queries submitted by the user.

[0016] In some implementations, determining the text segment to be output by the text-to-speech module includes: identifying a plurality of text segments as candidates for the user's text-to-speech output, the plurality of text segments having different levels of language complexity; and selecting from the plurality of text segments based at least on the user's determined language proficiency on the client device.

[0017] In some implementations, selecting from the plurality of text segments includes: determining a language complexity score for each of the plurality of text segments; and selecting the text segment whose language complexity score best matches a reference score describing the language proficiency of the user of the client device.

[0018] In some implementations, determining the text segment to be output by the text-to-speech module includes: identifying the text segment to be output to the user; calculating a complexity score for the text segment to be output to the user; and modifying the text segment to be output to the user based at least on the user's determined language proficiency and the complexity score of the text segment to be output to the user.

[0019] In some implementations, modifying the text segment for text-to-speech output to the user includes: determining an overall complexity score for the user based at least on the user's determined language proficiency; determining the complexity scores of individual portions within the text segment of the text-to-speech output to the user; identifying one or more individual portions within the text segment that have a complexity score greater than the user's overall complexity score; and modifying the one or more individual portions within the text segment to reduce their complexity scores to below the overall complexity score.

[0020] In some implementations, modifying the text segment used for the text-to-speech output to the user includes: receiving data indicating a context associated with the user; determining an overall complexity score of the context associated with the user; determining that the complexity score of the text segment exceeds the overall complexity score of the context associated with the user; and modifying the text segment so that the complexity score is reduced to below the overall complexity score of the context associated with the user.

[0021] On the other hand, a computer program includes machine-readable instructions that, when executed by a computing device, cause it to perform any of the methods described above.

[0022] In another general aspect, a computer-implemented method includes: receiving data indicating a context associated with a user; determining an overall complexity score of the context associated with the user; identifying a text segment for text-to-speech output to the user; determining that the complexity score of the text segment exceeds the overall complexity score of the context associated with the user; and modifying the text segment such that the complexity score is reduced to below the overall complexity score of the context associated with the user.

[0023] In some implementations, determining the overall complexity score of the context associated with the user includes: identifying terms included in a query previously submitted by the user when the user is determined to be in the context; and determining the overall complexity score of the context associated with the user based at least on the identified terms.

[0024] In some implementations, the data indicating the context associated with the user includes queries previously submitted by the user.

[0025] In some implementations, the data indicating the environment associated with the user includes GPS signals indicating the current location associated with the user.

[0026] Details of one or more embodiments are set forth in the accompanying drawings and the following description. Other possible features and advantages will become apparent from the specification, drawings, and claims.

[0027] Other implementations of these aspects include corresponding systems, apparatuses, and computer programs configured to perform actions of methods encoded on computer storage devices. Attached Figure Description

[0028] Figure 1 This is a diagram illustrating an example of a process for generating text-to-speech output based on language proficiency.

[0029] Figure 2 This is a diagram illustrating an example of a system for generating adaptive text-to-speech output based on user context.

[0030] Figure 3 This is a diagram illustrating an example of a system used to modify the sentence structure within text-to-speech output.

[0031] Figure 4This is a block diagram illustrating an example of a system for generating adaptive text-to-speech output based on clustering techniques.

[0032] Figure 5 This is a flowchart illustrating an example of a process for generating adaptive text-to-speech output.

[0033] Figure 6 It is a block diagram of a computing device or a part thereof capable of implementing the process described herein.

[0034] In the accompanying drawings, similar reference numerals are used to indicate corresponding parts in each drawing. Detailed Implementation

[0035] Figure 1 This diagram illustrates examples of processes 100A and 100B for generating text-to-speech output based on language proficiency. Processes 100A and 100B are used to generate different text-to-speech outputs for a user 102a with high language proficiency and a user 102b with low language proficiency, respectively, in response to text query 104. As shown, after receiving query 104 on user devices 106a and 106b, process 100A generates a high-complexity text-to-speech output 108a for user 102a, while process 100B generates a low-complexity output 108b for user 102b. Furthermore, the TTS system executing processes 100A and 100B can include a language proficiency estimator 110 and a text-to-speech engine 120. The text-to-speech engine 120 can further include a text analyzer 122, a linguistic analyzer 124, and a waveform generator 126.

[0036] Generally, the content of the text used to generate the text-to-speech output can be determined based on the user's language proficiency. Alternatively, the text used to generate the text-to-speech output can be determined based on the user's context, such as the user's location or activity, the presence of background noise, and the user's current task. Furthermore, by using other information—such as indications that the user has failed to complete a task or is repeating an action—the text to be converted into an audible form can be adjusted or determined.

[0037] In this example, two users—user 102a and user 102b—provide the same query 104 on user devices 106a and 106b, respectively, as input to an application, web page, or other search function. For example, query 104 could be a voice query sent to user devices 106a and 106b to determine the day's weather forecast. Query 104 is then passed to text-to-speech engine 120 to generate text-to-speech output in response to query 104.

[0038] The language proficiency estimator 110 can be a software module within a TTS system, which determines a language proficiency score associated with a specific user (e.g., user 102a or user 102b) based on highly complex text-to-speech output 108a. The language proficiency score can be an estimate of a user's ability to understand communication in a particular language—specifically, their ability to understand spoken language. One measure of language proficiency is a user's ability to successfully complete voice-controlled tasks. Many types of tasks, such as setting appointments or finding directions, follow a series of interactions involving spoken communication between the user and the device. The percentage of users who successfully complete these task workflows via the voice interface is a significant indicator of a user's language proficiency. For example, a user who completes nine out of ten user-initiated voice tasks may have high language proficiency. On the other hand, a user who fails to complete most user-initiated voice tasks may be inferred to have low language proficiency because the user may not fully understand communication from the device or may not be able to provide appropriate spoken responses. As discussed further below, when users fail to complete workflows that include standard TTS outputs, resulting in low language proficiency scores, TTS can use adapted, simplified outputs, which can improve users' ability to understand and complete various tasks.

[0039] As shown in the figure, the high-complexity text-to-speech output 108a can include the words used in the user-submitted prior text query, an indication of whether English or any other language used by the TTS system is the user's native language, and a set of activities and / or behaviors reflecting the user's language comprehension skills. For example, Figure 1 As shown, a user's typing speed can be used to determine the user's language fluency. Furthermore, based on associating a predetermined complexity with the words the user used in previous text queries, a language vocabulary complexity score or language proficiency score can be assigned to the user. In another example, the number of misidentified words in previous queries can also be used to determine the language proficiency score. For example, a large number of misidentified words can be used to indicate low language proficiency. In some implementations, the language proficiency score is determined by looking up a stored score associated with the user, which was determined for the user before submitting query 104.

[0040] Although Figure 1 The language proficiency estimator 110 is depicted as a component separate from the TTS engine 120, but in some implementations, such as Figure 2 As shown, the language proficiency estimator 110 can be an integrated software module within the TTS engine 120. In this case, operations involving language proficiency estimation can be directly controlled by the TTS engine 120.

[0041] In some implementations, the language proficiency score assigned to a user can be based on a specific user context estimated for that user. For example, as more specifically referred to Figure 2 The user context determination can be used to identify context-specific language proficiency that allows a user to temporarily have limited language comprehension. For example, if the user context indicates significant background noise or if the user is engaged in a task such as driving, the language proficiency score can be used to indicate a temporary decrease in the user's current language comprehension relative to other user contexts.

[0042] In some implementations, instead of inferring language proficiency based on previous user activity, language proficiency scores can be directly provided to the TTS engine 120 without using the language proficiency estimator 110. For example, during the registration process for a specified user's language proficiency level, a user's language proficiency score can be assigned based on user input. For instance, during registration, the user can provide a selection of a specified skill level, which can then be used to calculate the user's appropriate language proficiency. In other examples, the user can provide other types of information, such as group characteristics, education level, place of residence, etc., which can be used to specify the user's language proficiency level.

[0043] In the examples above, language proficiency scores can be either a discrete set of values ​​periodically adjusted based on recently generated user activity data, or continuous scores initially assigned during the registration process. In the first case, the language proficiency score can be biased based on one or more factors indicating a potential weakening of the user's current language comprehension and proficiency (e.g., a user environment with significant background noise). In the second case, the language proficiency score can be preset after initial calculation and adjusted only after specific milestone events indicating an improvement in the user's language proficiency (e.g., an increase in typing speed or a decrease in correction rate for a given language). In other cases, a combination of these two techniques can be used to variably adjust text-to-speech output based on specific text input. In such cases, multiple language proficiency scores, each representing a specific aspect of the user's language skills, can be used to determine how best to adjust the text-to-speech output for the user. For example, one language proficiency score could represent the complexity of the user's vocabulary, while another could be used to represent the user's grammatical skills.

[0044] The TTS engine 120 can use a language proficiency score to generate text-to-speech output adapted to the user's language proficiency level indicated by that score. In some cases, the TTS engine 120 adapts the text-to-speech output by selecting a specific TTS string from a set of candidate TTS strings for the text query 104. In such cases, the TTS engine 120 selects the specific TTS string based on predicting the likelihood that the user will accurately understand each candidate TTS string using the user's language proficiency score. (See also...) Figure 2 A more specific description of these techniques is provided. Alternatively, in other cases, the TTS engine 120 can select a baseline TTS string and adjust the structure of the TTS string based on the user's language proficiency score. In such cases, the TTS engine 120 can adjust the syntax of the baseline TTS string, provide word substitutions, and / or reduce sentence complexity to generate an adapted TTS string that the user is more likely to understand. (See also...) Figure 3 A more specific description of these technologies is provided.

[0045] Still refer to Figure 1 The TTS engine 120 can generate different text-to-speech outputs for users 102a and 102b because the users have different language proficiency scores. For example, in process 100A, a language proficiency score 106a indicates high English language proficiency, which is inferred from a highly complex text-to-speech output 108a indicating the following: user 102a has a complex vocabulary, speaks English as their first language, and has a relatively high number of words per minute in previous user queries. Based on the value of the language proficiency score 106a, the TTS engine 120 generates a highly complex text-to-speech output 108a that includes complex grammatical structures. As shown, the highly complex text-to-speech output 108a includes a separate clause describing today's weather forecast as sunny, as well as subordinate clauses containing additional information about the day's high and low temperatures.

[0046] In the example of process 100B, the language proficiency score 106b indicates low English language proficiency, which is inferred from user activity data 108b indicating that user 102b has a simple vocabulary, speaks English as a second language, and has previously provided ten incorrect queries. In this example, the TTS engine 120 generates low-complexity text-to-speech output 108b, which includes a simpler grammatical structure compared to high-complexity text-to-speech output 108a. For example, instead of including multiple clauses within a single sentence, text-to-speech output 108b includes a single independent clause that conveys the same main information as high-complexity text-to-speech output 108a (e.g., today's weather forecast is sunny), but does not include additional information about the day's high and low temperatures.

[0047] Text adaptation for TTS output can be performed using various devices and software modules. For example, a server system's TTS engine may include functionality that adjusts text based on language proficiency scores and then outputs audio of synthesized speech containing the adjusted text. Alternatively, a server system's preprocessing module may adjust the text and pass it to the TTS engine for speech synthesis. Another example is a user device that may include a TTS engine or a TTS engine and a text preprocessor to generate appropriate TTS output.

[0048] In some implementations, the TTS system may include software modules configured to exchange communication with third-party mobile applications or web pages on client devices. For example, the system's TTS functionality may be made available to third-party mobile applications via an Application Package Interface (API). The API may include a defined set of protocols that applications or websites can use to request TTS audio from a server system running the TTS engine 120. In some implementations, the API may enable TTS functionality to run locally on the user device. For example, the API may be available to applications or web pages via inter-process communication (IPC), remote procedure calls (RPC), or other system calls or functions. The TTS engine, along with associated language proficiency analysis or text preprocessing, may run locally on the user device to determine appropriate text based on the user's language proficiency and also generate synthesized speech audio.

[0049] For example, third-party applications or web pages can use APIs to generate a set of voice commands for users based on the speech interface of the third-party application or web page. The API can specify that the application or web page should provide text to be converted into speech. In some cases, other information, such as user identifiers or language proficiency scores, can be provided.

[0050] In an implementation where the TTS engine 120 communicates with a third-party application via an API, the TTS engine 120 can be used to determine whether text segments from the third-party application should be adjusted before generating text-to-speech output of the text. For example, the API can include a computer-implemented protocol specifying the conditions within the third-party application for initiating the generation of adaptive text-to-speech output.

[0051] As an example, an API can allow an application to submit multiple different text snippets as candidates for TTS output, where the different text snippets correspond to different levels of language proficiency. For example, candidates could be text snippets with equivalent meaning but different levels of complexity (e.g., high-complexity response, medium-complexity response, and low-complexity response). The TTS engine 120 can then determine the language proficiency required to understand each candidate, determine the appropriate language proficiency score for the user, and select the candidate text that best corresponds to that language proficiency score. The TTS engine 120 then provides synthesized audio of the selected text back to the application, for example, via a network using the API. In some cases, the API can be locally available on user devices 106a and 106b. In such cases, the API can be accessed via various types of inter-process communication (IPC) or via system calls. For example, the output of the API on user devices 106a and 106b could be the text-to-speech output of the TTS engine 120 because the API operates locally on user devices 106a and 106b.

[0052] In another example, the API allows third-party applications to provide a single text fragment along with an indication of permission for the TTS engine 120 to modify the text fragment to generate text fragments with varying complexities. If the app or web page indicates permission to change, the TTS engine 120 can make various changes to the text, such as reducing the text complexity when a language proficiency score indicates that the original text is more complex than the user can understand in a spoken response. In other examples, the API allows third-party applications to also provide user data (e.g., prior user queries submitted on the third-party application) along with the text fragment, enabling the TTS engine 120 to determine the user context associated with the user and adjust the generation of specific text-to-speech output based on the determined user context. Similarly, the API allows applications to provide contextual data from the user's device (e.g., GPS signal, accelerometer data, ambient noise levels, etc.) or indications of the user context, allowing the TTS engine 120 to adjust the text-to-speech output that will ultimately be provided to the user through the third-party application. In some cases, third-party applications can also provide the API with data that can be used to determine the user's language proficiency.

[0053] In some implementations, the TTS engine 120 can adjust the text-to-speech output for a user query without using the user's language proficiency or without determining the context associated with the user. In such implementations, the TTS engine 120 can determine that the initial text-to-speech output is too complex for the user based on a signal that the user has misunderstood the output (e.g., repeatedly retries the same query or task). In response, the TTS engine 120 can reduce the complexity of subsequent text-to-speech responses to retried queries or related queries. Therefore, when the user fails to complete an action successfully, the TTS engine 120 can progressively reduce the amount of detail or language proficiency required to understand the TTS output until it reaches a level that the user can understand.

[0054] Figure 2 This is a diagram illustrating an example of a system 200 that adaptively generates text-to-speech output based on user context. In short, system 200 can include a TTS engine 210, which includes a query analyzer 211, a language proficiency estimator 212, an interpolator 213, a linguistic analyzer 214, a re-ranking unit 215, and a waveform generator 216. System 200 also includes a context repository 220 storing a set of context profiles 232 and a user history manager 230 storing query logs 234. In some cases, the TTS engine 210 corresponds to, as shown in the reference... Figure 1 The TTS engine 120 mentioned above.

[0055] In this example, user 202 initially submits query 204 on user device 208, which includes a request for information related to the user's first appointment of the day. User device 208 is then able to transmit query 204 and contextual data 206 associated with user 202 to query analyzer 211 and language proficiency estimator 212, respectively. The same technique can be used to adapt other types of TTS output that are not responses to queries, such as calendar reminders, notifications, task workflows, etc.

[0056] Context data 206 can include information about the specific context associated with user 202, such as the time interval between repeated text queries, GPS data indicating the location, speed, or movement pattern associated with user 202, prior text queries submitted to TTS engine 210 within a specific time period, or other types of background information that can indicate user activity associated with TTS engine 210. In some cases, context data 206 can indicate the type of query 204 submitted to TTS engine 210, such as whether query 204 is a text fragment associated with a user action or an instruction sent to TTS engine 210 to generate text-to-speech output.

[0057] Upon receiving query 204, query analyzer 211 parses query 204 to identify information responding to query 204. For example, in some cases where query 204 is a speech query, query analyzer 211 initially generates a transcript of the speech query and then processes the individual words or segments within query 204 to determine information responding to query 204, for example, by providing the query to a search engine and receiving the search results. The transcript of query 204 and the identified information can then be transmitted to linguistic analyzer 214.

[0058] The language proficiency estimator 212 is described below. After receiving the context data 206, it uses a reference... Figure 1 In the described technique, the language proficiency estimator 212 calculates the language proficiency of the user 202 based on the received context data 206. Specifically, the language proficiency estimator 212 parses various context profiles 232 stored in a repository 220. The context profiles 232 may be an archive containing relevant types of information associated with a specific user context and capable of being included in text-to-speech output. The context profiles 232 additionally specify values ​​associated with each type of information, representing the degree to which the user 202 might understand each type of information when the user 202 is currently in the context associated with the context profile 232.

[0059] exist Figure 2 In the example shown, the context profile 232 specifies that user 202 is currently in a context indicating that user 202 is commuting to and from get off work daily. Furthermore, the context profile 232 also specifies the values ​​of various words and phrases that user 202 is likely to understand. For example, data or time information is associated with a value of "0.9" for "SINCE," indicating that user 202 is more likely to understand broader information associated with an appointment (e.g., the time of the next upcoming appointment) 204, rather than more detailed information associated with the appointment (e.g., the participants or location of the appointment). In this example, the difference in values ​​indicates a difference in the user's ability to understand specific types of information, as the user's ability to understand complex or detailed information decreases.

[0060] The values ​​associated with each word and phrase can be determined based on user activity data from previous user sessions in which user 202 was previously in the context indicated by context data 206. For example, historical user data can be transferred from user history manager 230, which retrieves data stored in query log 234. In this example, the values ​​of date and time information can be increased based on the determination that the user typically accesses date and time information associated with a meeting more frequently than the meeting location.

[0061] After the language proficiency estimator 212 selects a specific context profile 232 corresponding to the received context data 206, it transmits the selected context profile 232 to the interpolator 213. The interpolator 213 parses the selected context profile 232 and extracts the included words and phrases and their associated values. In some cases, the interpolator 213 directly transmits different types of information and associated values ​​to the linguistic analyzer 214 to generate a list 240a of text-to-speech output candidates. In such cases, the interpolator 213 extracts specific types of information and associated values ​​from the selected context profile 232 and transmits them to the linguistic analyzer 214. In other cases, the interpolator 213 can also transmit the selected context profile 232 to the re-ranking unit 215.

[0062] In some cases, a collection of structured data (e.g., fields from calendar events) can be provided to the TTS engine 210. In such cases, the interpolator 213 can transform the structured data into text that matches the user's proficiency level indicated by the context profile 232. For example, the TTS engine 210 can access data indicating one or more grammar points—which indicate different levels of detail or complexity in expressing the information in the structured data—and select appropriate grammar points based on the user's language proficiency score. Similarly, the TTS engine 210 can use a dictionary to select appropriate words based on the language proficiency score.

[0063] Linguistic analyzer 214 performs processing operations such as standardization on the information included in query 204. For example, query analyzer 211 can assign phonetic transcription to each word or snippet included in query 204 and use text-to-speech conversion to segment query 204 into prosodic units such as phrases, clauses, and sentences. Linguistic analyzer 214 also generates list 240a, which includes multiple text-to-speech output candidates identified as responses to query 204. In this example, list 240a includes multiple text-to-speech output candidates with different levels of complexity. For example, the response “At 12:00PM with Mr. John near Dupont Circle” is the most complex response because it identifies the time of the meeting, the location of the meeting, and the individual to be met. In contrast, the response “In three hours” is the least complex because it only identifies the time of the meeting.

[0064] Listing 240a also includes a baseline ranking of text-to-speech candidates based on the likelihood that each text-to-speech output candidate might respond to query 204. In this example, Listing 240a indicates that the most complex text-to-speech output candidate is most likely to respond to query 204 because it includes the largest amount of information associated with the content of query 204.

[0065] After the linguistic analyzer generates a list 240a of text-to-speech output candidates, the re-ranking unit 215 generates a list 240b of adjusted rankings of the text-to-speech output candidates based on the received context data 206. For example, the re-ranking unit 215 can adjust the rankings based on scores associated with specific types of information included in the selected context profile 232.

[0066] In this example, the re-ranker 215 ranks the simplest text-to-speech output as the highest based on the context profile 232, which indicates that, given the user's current commuting situation, the user 202 is likely to understand the date and time information within the text-to-speech response, but is unlikely to understand the participant names or location information within the text-to-speech response. In this regard, the received context data 206 can be used to adjust the selection of specific text-to-speech output candidates to increase the likelihood that the user 202 will understand the content of the text-to-speech output 204c from the TTS engine 210.

[0067] Figure 3 This is a diagram illustrating an example of a system 300 used to modify the sentence structure within text-to-speech output. In short, the TTS engine 310 receives a query 302 from a user (e.g., user 202) and a language proficiency profile 304. The TTS engine 310 then performs operations 312, 314, and 316 to generate an adjusted text-to-speech output 302c in response to the query 302. In some cases, the TTS engine 310 corresponds to a reference... Figure 1 The TTS engine 120 or reference Figure 2 The TTS engine 210 mentioned above.

[0068] Generally, the TTS engine 310 can modify the sentence structure of the baseline text-to-speech output 306a for query 302 by using different types of adjustment techniques. For example, the TTS engine 310 can replace words or phrases in the baseline text-to-speech output 306a based on the determination that the complexity score associated with each word or phrase is greater than the threshold score indicated by the user's language proficiency profile 304. As another example, the TTS engine 310 can rearrange the sentence segments to reduce the overall complexity of the baseline text-to-speech output 306a to a satisfactory level based on the language proficiency profile 304. The TTS engine 310 can also reorder words, split or combine sentences, and make other changes to adjust the complexity of the text.

[0069] More specifically, during operation 312, the TTS engine 310 initially generates a baseline text-to-speech output 306a in response to query 302. The TTS engine 310 then parses the baseline text-to-speech output 306a into segments 312a to 312c. The TTS engine 310 also detects punctuation marks (e.g., commas, periods, semicolons, etc.) that indicate breakpoints between the segments. The TTS engine 310 also calculates a complexity score for each of segments 312a to 312c. In some cases, the complexity score can be calculated based on the frequency of a particular word within a particular language. Alternative techniques may include calculating the complexity score based on the frequency of user usage or its occurrence in the user's access history (e.g., news reports, web pages, etc.). In each of these examples, the complexity score can be used to indicate words that are likely to be understood by the user and other words that are unlikely to be understood by the user.

[0070] In this example, segments 312a and 312b are determined to be relatively complex based on highly complex terms such as “FORECAST” and “CONSISTENT”, respectively. However, segment 312c is determined to be relatively simple because the terms it includes are relatively simple. This determination is represented by segments 312a and 312b having higher complexity scores (e.g., 0.83, 0.75) compared to the complexity score of segment 312c (e.g., 0.41).

[0071] As mentioned above, the language proficiency profile 304 can be used to calculate a threshold complexity score, which indicates the maximum complexity a user can comprehend. In this example, the threshold complexity score can be calculated as "0.7", making it unlikely that TTS 310 will be able to comprehend segments 312a and 312b.

[0072] After identifying each segment of the associated complexity score that has a complexity score greater than the threshold complexity score indicated by the language proficiency profile 304, during operation 314, the TTS engine 310 replaces the identified words with alternative items predicted to be more likely to be understood by the user. As Figure 3 shown, "FORECAST" can be replaced with "WEATHER () weather", and "CONSISTENT" can be replaced with "CHANGE (change)". In these examples, segments 314a and 314b represent simpler alternatives with a lower complexity score than the threshold complexity score indicated by the language proficiency profile 304.

[0073] In some embodiments, the TTS engine 310 can use a trained skip-gram model to handle word replacement of highly complex words, which uses unsupervised techniques to determine appropriately complex words to replace highly complex words. In some cases, the TTS engine 310 can also use a thesaurus or synonym data to handle word replacement of highly complex words.

[0074] Now introduce operation 316. Based on calculating the complexity associated with a specific statement structure and based on the language proficiency indicated by the language proficiency profile 304 to determine whether the user will be able to understand the statement structure, the statement clauses of the query can be adjusted.

[0075] In this example, based on determining that the baseline text-to-speech output 306a includes three statement clauses (e.g., "today’s forecast is sunny (今日天气预报为晴)", "but not consistent (而不持续)", and "and warm (并且温暖)"), the TTS engine 310 determines that the baseline text-to-speech output 306a has a high statement complexity. In response, the TTS engine 310 can generate adjusted statement parts 316a and 316b, which combine the subordinate clause and the independent clause into a single clause without the punctuation for division. As a result, the adjusted text-to-speech output 306b includes simpler vocabulary (e.g., "WEATHER", "CHANGE") and a simpler statement structure (e.g., no clause division), which increases the likelihood that the user will understand the adjusted text-to-speech output 306b. Then, the adjusted text-to-speech output 306b is generated for the TTS engine 310 to output as output 306c.

[0076] In some implementations, the TTS engine 310 can perform statement structure adjustments based on a user-specific restructuring algorithm that includes using weighting factors to adjust the baseline query 302a to avoid being identified as a specific statement structure that is problematic to the user. For example, the user-specific restructuring algorithm can specify options to reduce the weight of clauses containing subordinate clauses or increase the weight of clauses with a simple subject-verb-object order.

[0077] Figure 4 This is a block diagram illustrating an example of a system 400 that adaptively generates text-to-speech output using clustering techniques. System 400 includes a language proficiency estimator 410, a user similarity determiner 420, a complexity optimizer, and a machine learning system 440.

[0078] In short, the language proficiency estimator 410 receives data from multiple users 402. The language proficiency estimator 410 then estimates a set of language complexity profiles 412 for each of the multiple users 402 and sends them to the user similarity determiner 420. The user similarity determiner 420 identifies user clusters 424 of similar users. Then, the complexity optimizer 430 and the machine learning system 440 analyze the language complexity profiles 412 of each user within the user clusters 424, as well as the contextual data received from the multiple users 402, to generate a complexity map 442.

[0079] Generally, System 400 can be used to analyze the relationship between active and passive language complexity of a user group. Active language complexity refers to the detected language input provided by the user (e.g., text query, voice input, etc.). Passive language complexity refers to the user's ability to understand or comprehend the speech signals provided to the user. In this regard, System 400 can use the determined relationship between the active and passive language complexity of multiple users to determine the appropriate passive language complexity for each individual user, where a particular user has the highest probability of understanding text-to-speech output.

[0080] Multiple users 402 can be multiple users using an application associated with a TTS engine (e.g., TTS engine 120). For example, multiple users 402 can be a group of users using a mobile application that utilizes the TTS engine to provide text-to-speech features to users through the mobile application's user interface. In such a case, data from multiple users 402 (e.g., prior user queries, user selections, etc.) can be tracked by the mobile application and aggregated for analysis by the language proficiency estimator 410.

[0081] The language proficiency estimator 410 can use the same reference as above. Figure 1The substantially similar technique is used to initially measure the passive language complexity of multiple users 402. Then, the language proficiency estimator 410 is able to generate a language complexity profile 412, which includes an individual language complexity profile for each of the multiple users 402. Each individual language complexity profile includes data indicating the passive and active language complexity of each of the multiple users 402.

[0082] User similarity determiner 420 uses the linguistic complexity data included in the set of linguistic complexity profiles 412 to identify similar users among multiple users 402. In some cases, user similarity determiner 420 can group users with similar active linguistic complexity (e.g., providing similar language input, voice queries, etc.). In other cases, user similarity determiner 420 can determine similar users by comparing words included in queries submitted by previous users, specific user behaviors on the mobile application, or user locations. User similarity determiner 420 then clusters similar users to generate user clusters 424.

[0083] In some implementations, the user similarity determiner 420 generates a user cluster 424 based on stored cluster data 422, which includes aggregated data of users in a specified cluster. For example, the cluster data 422 can be grouped by specific parameters indicating the passive language complexity associated with multiple users 402 (e.g., the number of incorrect query responses, etc.).

[0084] After generating user clusters 424, the complexity optimizer 430 modifies the complexity of the language output by the TTS system and measures the passive language complexity of the users using a set of parameters indicative of user performance (e.g., comprehension rate, speech action flow completion rate, or response success rate). These parameters indicate the user's ability to understand the language output by the TTS system. For example, these parameters can be used to characterize the degree to which users within each cluster 424 understand a given text-to-speech output. In this case, the complexity optimizer 430 can initially provide users with low-complexity speech signals and recursively provide additional speech signals within a certain complexity range.

[0085] In some implementations, the complexity optimizer 430 is also able to determine the optimal passive language complexity for each user context associated with each user cluster 424. For example, after measuring the user's language proficiency using this set of parameters, the complexity optimizer 430 can then classify the measured data using context data received from multiple users 402 to enable the determination of the optimal passive language complexity for each user context.

[0086] After aggregating performance data across the range of passive language complexity, the machine learning system 440 then determines a specific passive language complexity where the performance parameters indicate the user's strongest language comprehension. For example, the machine learning system 440 aggregates performance data from all users within a specific user cluster 424 to determine the relationship between active language complexity, passive language complexity, and the user context.

[0087] Then, the aggregated data of user cluster 424 can be compared with the individual data of each user within user cluster 424 to determine the actual language complexity score of each user within user cluster 424. For example, Figure 4 As shown, the complexity mapping 442 can represent the relationship between active language complexity and passive language complexity to infer the actual language complexity, which corresponds to the active language complexity mapped to the optimal passive language complexity.

[0088] Complexity mapping 442 represents the relationship between active language complexity, TTS complexity, and passive language complexity for all user clusters within multiple users 402, which can then be used to predict the appropriate TTS complexity for subsequent queries by individual users. For example, as described above, user input (e.g., queries, text messages, emails, etc.) can be used to group similar users into user clusters 424. For each cluster, the system provides TTS outputs requiring different levels of language proficiency to understand. The system then evaluates the responses received from users and the task completion rate of the different TTS outputs to determine the appropriate level of language complexity for users in each cluster. The system stores a mapping 442 between cluster identifiers and TTS complexity scores corresponding to the identified clusters. The system then uses complexity mapping 442 to determine the appropriate level of complexity for the TTS outputs for users. For example, the system identifies clusters representing users' active language proficiency, looks up the corresponding TTS complexity score for that cluster in mapping 442 (e.g., indicating the passive language comprehension level), and generates TTS outputs with the complexity level indicated by the retrieved TTS complexity score.

[0089] Then, by using a reference Figures 1 to 3 The described technique enables the TTS system to be tuned using the actual language complexity determined for the user. In this regard, aggregated language complexity data from a group of similar users (e.g., user cluster 424) can be used to intelligently adjust the performance of the TTS system for individual users.

[0090] Figure 5This is a flowchart illustrating an example of a process 500 for adaptively generating text-to-speech output. In short, process 500 can include determining the language proficiency of a user of a client device (510), determining text segments for output by the text-to-speech module (520), generating audio data of synthesized speech including the text segments (530), and providing the audio data to the client device (540).

[0091] More specifically, process 500 may include determining the language proficiency of the user on the client device (510). For example, as referenced Figure 1 The language proficiency estimator 110 is capable of using various techniques to determine a user's language proficiency. In some cases, language proficiency can represent an assignment score indicating the level of language proficiency. In other cases, language proficiency can represent an assignment category from multiple language proficiency categories. In still other cases, language proficiency can be determined based on user input and / or behavior that indicate the user's proficiency level.

[0092] In some implementations, language proficiency can be inferred from different user signals. For example, as shown in [reference] Figure 1 The system can infer language proficiency based on factors such as the complexity of the user's input vocabulary, the user's data entry rate, the number of misidentified words from the voice input, the number of voice actions completed at different levels of TTS complexity, or the complexity level of the text viewed by the user (e.g., books, articles, text on web pages, etc.).

[0093] Process 500 can include determining a text segment for output by the text-to-speech module (520). For example, the TTS engine can adjust a baseline text segment based on determining the user's language proficiency. In some cases, such as referring to... Figure 2 The aforementioned method can adjust the text fragments used for output based on the user context associated with the user. In other cases, such as referring to... Figure 3 As described above, by reducing the complexity of text fragments through word substitution or sentence restructuring, the text fragments used for output can also be adjusted. For example, adjustment can be based on the rarity of individual words included in the text fragment, the type of verbs used (e.g., compound verbs or verb tenses), and the linguistic structure of the text fragment (e.g., the number of subordinate clauses, the spacing between related words, the degree of phrase nesting, etc.). In other examples, adjustment can also be based on the aforementioned linguistic measures and reference measurements of linguistic characteristics (e.g., the average spacing between subjects and verbs, the spacing between adjectives and nouns, etc.). In such examples, reference measurements can represent averages or may include ranges or examples for different levels of complexity.

[0094] In some implementations, determining the text fragments for output can include selecting text fragments with scores that best match a reference score describing the user's language proficiency level. In other implementations, individual words or phrases can be scored based on complexity, and the most complex words can then be replaced, deleted, or reconstructed to make the overall complexity appropriate for the user's level.

[0095] Process 500 can include generating audio data of synthesized speech including text fragments (530).

[0096] Process 500 can include providing audio data to a client device (540).

[0097] Figure 6 This is a block diagram of computing devices 600, 650 capable of acting as a client or as one or more servers to implement the systems and methods described herein. Computing device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 650 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. Additionally, computing devices 600 or 650 may include a Universal Serial Bus (USB) flash drive. The USB flash drive can store an operating system and other applications. The USB flash drive may include input / output components, such as a wireless transmitter or USB connector capable of being plugged into the USB port of another computing device. The components shown herein, their connections and relationships, and their functions are intended to be exemplary only and not to limit the embodiments of the invention described herein and / or claimed.

[0098] Computing device 600 includes a processor 602, a memory 604, a storage device 606, a high-speed interface 608 connected to the memory 604 and a high-speed expansion port 610, and a low-speed interface 612 connected to a low-speed expansion port 614 and the storage device 606. Each of components 602, 604, 606, 608, 610, and 612 is interconnected using various buses and can be mounted on a common motherboard or otherwise, as appropriate. Processor 602 is capable of processing instructions for execution within computing device 600, including instructions stored in memory 604 or on storage device 606, to display graphical information for a GUI on an external input / output device such as a display 616 coupled to high-speed interface 608. In other embodiments, multiple processors and / or multiple buses can be used in conjunction with multiple memories and memory types, as appropriate. Furthermore, multiple computing devices 600 can be connected, each providing multiple parts of the required operation, for example, as a server group, a blade server group, or a multiprocessor system.

[0099] Memory 604 stores information within computing device 600. In one embodiment, memory 604 is one or more volatile memory cells. In another embodiment, memory 604 is one or more non-volatile memory cells. Memory 604 can also be another form of computer-readable medium, such as a magnetic disk or optical disk.

[0100] Storage device 606 provides mass storage for computing device 600. In one embodiment, storage device 606 may be or include: computer-readable media, such as floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, flash memory or other similar solid-state storage devices, or device arrays including devices in a storage area network or other configuration. A computer program product may be tangibly embodied in an information carrier. A computer program product may also contain instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 604, storage device 606, or memory on processor 602.

[0101] High-speed interface 608 manages bandwidth-intensive operations for computing device 600, while low-speed interface 612 manages less bandwidth-intensive operations. This functional allocation is merely exemplary. In one embodiment, high-speed interface 608 is coupled to memory 604, display 616 (e.g., via a graphics processor or accelerator), and high-speed expansion port 610 capable of accepting various expansion cards (not shown). In this embodiment, low-speed interface 612 is coupled to storage device 606 and low-speed expansion port 614. The low-speed expansion port, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as keyboards, pointing devices, microphone / speaker pairs, scanners, or network devices such as switches or routers via, for example, network adapters. As shown, computing device 600 can be implemented in several different forms. For example, it can be implemented as a standard server 620 or multiple implementations in such a server group. It can also be implemented as part of rack-mount server system 624. Furthermore, it can be implemented in a personal computer such as laptop computer 622. Alternatively, components from computing device 600 can be combined with other components from mobile device (not shown), such as device 650. Each of such devices can contain one or more of computing devices 600, 650, and the entire system can consist of multiple computing devices 600, 650 communicating with each other.

[0102] As shown, computing device 600 can be implemented in several different forms. For example, it can be implemented as a standard server 620 or multiple implementations within such a server group. It can also be implemented as part of a rack server system 624. Furthermore, it can be implemented in a personal computer such as a laptop computer 622. Alternatively, components from computing device 600 can be combined with other components in a mobile device (not shown), such as device 650. Each of such devices can contain one or more of computing devices 600, 650, and the entire system can consist of multiple computing devices 600, 650 communicating with each other.

[0103] Computing device 650 includes a processor 652, memory 664, input / output devices such as a display 654, a communication interface 666, and a transceiver 668, as well as other components. Device 650 may also have storage devices, such as microdrives or other devices, for providing additional storage. Each of components 650, 652, 664, 654, 666, and 668 is interconnected using respective buses, and some of these components may be mounted on a common motherboard or otherwise, as appropriate.

[0104] Processor 652 is capable of executing instructions within computing device 650, including instructions stored in memory 664. The processor can be implemented as a chipset comprising multiple separate analog and digital processors. Additionally, the processor can be implemented using any of several architectures. For example, processor 652 can be a CISC (Complex Instruction Set Computer) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimum Instruction Set Computer) processor. For example, the processor can provide cooperation with other components of device 650, such as controls for the user interface, applications running on device 650, and wireless communications of device 650.

[0105] Processor 652 can communicate with the user via control interface 658 and display interface 656 coupled to display 654. For example, display 654 can be a TFT (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display or other suitable display technology. Display interface 656 can include appropriate circuitry for driving display 654 to present graphical and other information to the user. Control interface 658 can receive commands from the user and translate them for submission to processor 652. Furthermore, an external interface 662 can be provided to communicate with processor 652 to enable near-field communication between device 650 and other devices. For example, external interface 662 can provide wired communication in some embodiments or wireless communication in others, and multiple interfaces can also be used.

[0106] Memory 664 stores information within computing device 650. Memory 664 can be implemented as one or more computer-readable media, one or more volatile memory cells, or one or more non-volatile memory cells. An expansion memory 674 can also be provided and connected to device 650 via an expansion interface 672, for example, which can include a SIMM (Single In-line Memory Module) card interface. Such an expansion memory 674 can provide additional storage space for device 650, or it can store applications or other information for device 650. Specifically, expansion memory 674 can include instructions for performing or supplementing the above processes, and it can also include security information. Therefore, for example, expansion memory 674 can be provided as a security module for device 650 and can be programmed with instructions that authorize secure use of device 650. Furthermore, secure applications along with additional information, such as placing identification information on the SIMM card in a non-hackable manner, can be provided via a SIMM card.

[0107] For example, the memory can include flash memory and / or NVRAM memory, as discussed below. In one embodiment, the computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 664, extended memory 674, or memory on processor 652, that can be received, for example, via transceiver 668 or external interface 662.

[0108] Device 650 can perform wireless communication via communication interface 666, which may include digital signal processing circuitry if necessary. Communication interface 666 can provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS or MMS message sending and receiving, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, etc. For example, such communication can be performed via radio frequency transceiver 668. Furthermore, short-range communication is possible, such as using Bluetooth, Wi-Fi, or other transceivers (not shown). Additionally, GPS (Global Positioning System) receiver module 670 can provide additional navigation and location-related wireless data to device 650, which can be used by applications running on device 650 as appropriate.

[0109] Device 650 can also use audio codec 660 for audible communication, which receives spoken information from a user and converts it into usable digital information. Audio codec 660 can also generate audible sounds for the user, such as through a speaker, for example, in the handheld device of device 650. Such sounds can include sounds from voice telephone calls, recorded sounds, such as voice messages, music files, etc., and sounds generated by applications operating on device 650.

[0110] As shown in the figure, the computing device 650 can be implemented in several different forms. For example, it can be implemented as a cellular phone 480. It can also be implemented as a smartphone 682, a personal digital assistant, or part of other similar mobile devices.

[0111] Various embodiments of the systems and methods described herein can be implemented in digital electronic circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations of these embodiments. These various embodiments can include implementations in one or more executable and / or interpretable computer programs on a programmable system, said programmable system comprising at least one programmable processor, which may be dedicated or general-purpose, coupled to receive and send data and instructions to a storage system, a storage system, at least one input device, and at least one output device.

[0112] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level programming languages ​​and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device for providing machine instructions and / or data to a programmable processor, such as a disk, optical disk, memory, programmable logic device (PLD), including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0113] To provide interaction with the user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any type of sensory feedback, such as visual, auditory, or tactile feedback; and the input from the user can be received in any form, including sound, voice, or tactile input.

[0114] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), middleware components (e.g., application servers), frontend components (e.g., client computers with a graphical user interface or web browser that a user can interact with by means of an implementation of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium, such as a communication network. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), and the Internet.

[0115] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is implemented using computer programs that run on the respective computers and have a client-server relationship with each other.

[0116] Several embodiments have been described herein. However, it should be understood that various modifications can be made without departing from the spirit and scope of the invention. Furthermore, the logical flow depicted in the accompanying drawings does not require a specific order or sequence to obtain the desired result. Moreover, additional steps can be provided from the flow or some steps can be omitted, and other components can be added to the system or some components can be removed from the system. Therefore, other embodiments fall within the scope of the appended claims.

Claims

1. A computer-implemented method, which, when executed on data processing hardware, causes the data processing hardware to perform operations, the operations including: Obtain the previous text query submitted by the user on the client device; The user's language proficiency is determined based on the previous text query; Receive the query from the user to the client device; as well as In response to the query and based on the language proficiency determined for the user, a specific text fragment is generated, the specific text fragment including one of the following: A first text fragment, when the language proficiency determined for the user includes a first level of language proficiency, includes key information in response to the query; or When the language proficiency determined for the user includes a second level of language proficiency, the second text fragment includes additional information not included in the first text fragment in response to the query.

2. The computer-implemented method according to claim 1, wherein, The operation further includes: Generate audio data, the audio data including synthesized utterances of the specific text segment in response to the query; and The audio data is provided for audible output by the client device.

3. The computer-implemented method according to claim 1, wherein: The first text fragment includes corresponding independent clauses conveying the main information in response to the query; and The second text segment includes a corresponding independent clause and one or more subordinate clauses, the one or more subordinate clauses of the second text segment conveying additional information not included in the first text segment in response to the query.

4. The computer-implemented method according to claim 3, wherein, The corresponding independent clauses of the second text fragment convey the same main information in response to the query as the first text fragment.

5. The computer-implemented method according to claim 3, wherein, The corresponding independent clause of the second text segment includes at least one term that is different from the corresponding independent clause of the first text segment.

6. The computer-implemented method according to claim 1, wherein, The operation further includes, before generating the specific text fragment: Identify multiple candidate text fragments in response to the query, each candidate text fragment being associated with a different level of language proficiency; and Based on the language proficiency determined for the user, the specific text segment responding to the query is selected from the plurality of candidate text segments.

7. The computer-implemented method according to claim 6, wherein, Select from the plurality of candidate text fragments including: For each of the plurality of candidate text segments, a language complexity score is determined; and The text fragment associated with the language complexity score that best matches the reference score describing the user's determined language proficiency is selected as the specific text fragment.

8. The computer-implemented method according to claim 1, wherein, The operation further includes, before generating the specific text fragment: Obtain the baseline text fragment in response to the query; and The specific text fragment is generated by increasing the complexity level of the baseline text fragment based on the language proficiency assigned to the user.

9. The computer-implemented method according to claim 1, wherein, The operation further includes, before generating the specific text fragment: Obtain the baseline text fragment in response to the query; and The specific text segment is generated by reducing the complexity level of the baseline text segment based on the language proficiency assigned to the user.

10. The computer-implemented method according to claim 1, wherein: The second level of language proficiency includes a level of language proficiency higher than the first level; and The second text fragment is associated with a more complex grammatical structure than the one associated with the first text fragment.

11. A system for providing audio data, the system comprising: Data processing hardware; as well as Memory hardware that communicates with the data processing hardware and stores instructions, which, when executed by the data processing hardware, cause the data processing hardware to perform operations, including: Obtain the previous text query submitted by the user on the client device; The user's language proficiency is determined based on the previous text query; Receive queries from the user to the client device; and In response to the query and based on the language proficiency determined for the user, a specific text fragment is generated, the specific text fragment including one of the following: A first text fragment, when the language proficiency determined for the user includes a first level of language proficiency, includes key information in response to the query; or When the language proficiency determined for the user includes a second level of language proficiency, the second text fragment includes additional information not included in the first text fragment in response to the query.

12. The system according to claim 11, wherein, The operation further includes: Generate audio data, the audio data including synthesized utterances of the specific text segment in response to the query; and The audio data is provided for audible output by the client device.

13. The system according to claim 11, wherein: The first text fragment includes corresponding independent clauses conveying the main information in response to the query; and The second text segment includes a corresponding independent clause and one or more subordinate clauses, the one or more subordinate clauses of the second text segment conveying additional information not included in the first text segment in response to the query.

14. The system according to claim 13, wherein, The corresponding independent clauses of the second text fragment convey the same main information in response to the query as the first text fragment.

15. The system according to claim 13, wherein, The corresponding independent clause of the second text segment includes at least one term that is different from the corresponding independent clause of the first text segment.

16. The system according to claim 11, wherein, The operation further includes, before generating the specific text fragment: Identify multiple candidate text fragments in response to the query, each candidate text fragment being associated with a different level of language proficiency; and Based on the language proficiency determined for the user, the specific text segment responding to the query is selected from the plurality of candidate text segments.

17. The system according to claim 16, wherein, Select from the plurality of candidate text fragments including: For each of the plurality of candidate text segments, a language complexity score is determined; and The text fragment associated with the language complexity score that best matches the reference score describing the user's determined language proficiency is selected as the specific text fragment.

18. The system according to claim 11, wherein, The operation further includes, before generating the specific text fragment: Obtain the baseline text fragment in response to the query; and The specific text fragment is generated by increasing the complexity level of the baseline text fragment based on the language proficiency assigned to the user.

19. The system according to claim 11, wherein, The operation further includes, before generating the specific text fragment: Obtain the baseline text fragment in response to the query; and The specific text segment is generated by reducing the complexity level of the baseline text segment based on the language proficiency assigned to the user.

20. The system according to claim 11, wherein: The second level of language proficiency includes a level of language proficiency higher than the first level; and The second text fragment is associated with a more complex grammatical structure than the one associated with the first text fragment.

21. A computer-implemented method, which, when executed on data processing hardware, causes the data processing hardware to perform operations, the operations including: During the registration process on the client device: Receive the group characteristic information of the users of the client device; as well as Based on the received group characteristic information, a language proficiency level is assigned to the user. The language proficiency level assigned to the user includes either a first level of language proficiency or a second level of language proficiency that is different from the first level of language proficiency. Receive voice queries from the user to the client device; In response to the voice query and based on the language proficiency assigned to the user, audio data is generated, the audio data comprising synthesized utterances of specific text segments, the specific text segments including one of the following: When the language proficiency level assigned to the user includes a first level of language proficiency, the first text segment includes corresponding independent clauses conveying key information in response to the voice query; or When the language proficiency assigned to the user includes a second level of language proficiency, the second text fragment includes a corresponding independent clause and one or more subordinate clauses, the one or more subordinate clauses of the second text fragment conveying additional information not included in the first text fragment in response to the voice query; as well as The audio data is provided for audible output by the client device.

22. The computer-implemented method according to claim 21, wherein, The corresponding independent clauses of the second text segment convey the same main information in response to the voice query as the first text segment.

23. The computer-implemented method according to claim 21, wherein, The corresponding independent clause of the second text segment includes at least one term that is different from the corresponding independent clause of the first text segment.

24. The computer-implemented method according to claim 21, wherein, Before generating the audio data of the synthesized utterance including the specific text fragment: Obtain one or more search results from the search engine in response to the voice query; and The specific text fragment is determined based on the one or more search results and the language proficiency assigned to the user.

25. The computer-implemented method according to claim 21, wherein, The operation further includes, prior to generating the audio data of the synthesized utterance including the specific text fragment: Identify multiple candidate text segments in response to the voice query, each candidate text segment being associated with a different level of language proficiency; and Based on the user's language proficiency, the specific text segment responding to the voice query is selected from the plurality of candidate text segments.

26. The computer-implemented method according to claim 25, wherein, Select from the plurality of candidate text fragments including: For each of the plurality of candidate text segments, a language complexity score is determined; and The text fragment associated with the language complexity score that best matches the reference score describing the language proficiency assigned to the user is selected as the specific text fragment.

27. The computer-implemented method according to claim 21, wherein, The operation further includes, prior to generating the audio data of the synthesized utterance including the specific text fragment: Obtain the baseline text fragment in response to the voice query; and The specific text fragment is generated by increasing the complexity level of the baseline text fragment based on the language proficiency assigned to the user.

28. The computer-implemented method according to claim 21, wherein, The operation further includes, prior to generating the audio data of the synthesized utterance including the specific text fragment: Obtain the baseline text fragment in response to the voice query; and The specific text segment is generated by reducing the complexity level of the baseline text segment based on the language proficiency assigned to the user.

29. The computer-implemented method according to claim 21, wherein: The second level of language proficiency includes a level of language proficiency higher than the first level; and The second text fragment is associated with a more complex grammatical structure than the one associated with the first text fragment.

30. A system for providing audio data, the system comprising: Data processing hardware; as well as Memory hardware that communicates with the data processing hardware and stores instructions, which, when executed by the data processing hardware, cause the data processing hardware to perform operations, including: During the registration process on the client device: Receive the group characteristic information of the users of the client device; and Based on the received group characteristic information, a language proficiency level is assigned to the user. The language proficiency level assigned to the user includes either a first level of language proficiency or a second level of language proficiency that is different from the first level of language proficiency. Receive voice queries from the user to the client device; In response to the voice query and based on the language proficiency assigned to the user, audio data is generated, the audio data comprising synthesized utterances of specific text segments, the specific text segments including one of the following: When the language proficiency level assigned to the user includes a first level of language proficiency, the first text segment includes corresponding independent clauses conveying key information in response to the voice query; or A second text segment is provided when the user's language proficiency level includes a second level of language proficiency. This second text segment includes a corresponding independent clause and one or more subordinate clauses, wherein the one or more subordinate clauses of the second text segment convey additional information not included in the first text segment in response to the voice query; and The audio data is provided for audible output by the client device.

31. The system according to claim 30, wherein, The corresponding independent clauses of the second text segment convey the same main information in response to the voice query as the first text segment.

32. The system according to claim 30, wherein, The corresponding independent clause of the second text segment includes at least one term that is different from the corresponding independent clause of the first text segment.

33. The system according to claim 30, wherein, Before generating the audio data of the synthesized utterance including the specific text fragment: Obtain one or more search results from the search engine in response to the voice query; and The specific text fragment is determined based on the one or more search results and the language proficiency assigned to the user.

34. The system according to claim 30, wherein, The operation further includes, prior to generating the audio data of the synthesized utterance including the specific text fragment: Identify multiple candidate text segments in response to the voice query, each candidate text segment being associated with a different level of language proficiency; and Based on the user's language proficiency, the specific text segment responding to the voice query is selected from the plurality of candidate text segments.

35. The system according to claim 34, wherein, Select from the plurality of candidate text fragments including: For each of the plurality of candidate text segments, a language complexity score is determined; and The text fragment associated with the language complexity score that best matches the reference score describing the language proficiency assigned to the user is selected as the specific text fragment.

36. The system according to claim 30, wherein, The operation further includes, prior to generating the audio data of the synthesized utterance including the specific text fragment: Obtain the baseline text fragment in response to the voice query; and The specific text fragment is generated by increasing the complexity level of the baseline text fragment based on the language proficiency assigned to the user.

37. The system according to claim 30, wherein, The operation further includes, prior to generating the audio data of the synthesized utterance including the specific text fragment: Obtain the baseline text fragment in response to the voice query; and The specific text segment is generated by reducing the complexity level of the baseline text segment based on the language proficiency assigned to the user.

38. The system according to claim 30, wherein: The second level of language proficiency includes a level of language proficiency higher than the first level; and The second text fragment is associated with a more complex grammatical structure than the one associated with the first text fragment.

39. A method for providing audio data, the method comprising: Data is received at the data processing hardware from a client device associated with the user, the data indicating: The voice query is input by the user into the client device; as well as An indication of the user's language proficiency level, wherein the language proficiency level assigned to the user includes either a first level of language proficiency or a second level of language proficiency that is different from the first level of language proficiency; In response to the voice query and based on the language proficiency assigned to the user, audio data is generated via the data processing hardware. The audio data includes synthesized utterances of specific text segments, which include one of the following: When the language proficiency level assigned to the user includes a first level of language proficiency, the first text fragment includes first information in response to the voice query; or A second text fragment is provided when the language proficiency level assigned to the user includes a second level of language proficiency; the second text fragment includes second information in response to the voice query; and The audio data is provided to the client device associated with the user via the data processing hardware, wherein: The first text segment includes corresponding independent clauses that convey key information in response to the voice query; and The second text segment includes a corresponding independent clause and one or more subordinate clauses, the one or more subordinate clauses of the second text segment conveying additional information not included in the first text segment in response to the voice query.

40. A system for providing audio data, the system comprising: Data processing hardware; as well as Memory hardware that communicates with the data processing hardware and stores instructions, which, when executed by the data processing hardware, cause the data processing hardware to perform operations, including: Data is received from a client device associated with the user, the data indicating: The voice query is input by the user into the client device; and An indication of the user's language proficiency level, wherein the language proficiency level assigned to the user includes either a first level of language proficiency or a second level of language proficiency that is different from the first level of language proficiency; In response to the voice query and based on the language proficiency assigned to the user, audio data is generated, the audio data comprising synthesized utterances of specific text segments, the specific text segments including one of the following: When the language proficiency level assigned to the user includes a first level of language proficiency, the first text fragment includes first information in response to the voice query; or A second text fragment is provided when the language proficiency level assigned to the user includes a second level of language proficiency; the second text fragment includes second information in response to the voice query; and The audio data is provided to the client device associated with the user, wherein: The first text segment includes corresponding independent clauses that convey key information in response to the voice query; and The second text segment includes a corresponding independent clause and one or more subordinate clauses, the one or more subordinate clauses of the second text segment conveying additional information not included in the first text segment in response to the voice query.