Language determination method and device and electronic equipment

By obtaining text and response text during text to pronunciation and determining the language of the target statement, the problem of inaccurate recognition of text languages ​​such as numbers and serial numbers in the text in the prior art is solved, and the accuracy of speech synthesis is improved.

CN120020941APending Publication Date: 2025-05-20BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311540579.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In the process of text-to-speech, it is difficult to accurately determine the language of text such as numbers and serial numbers contained in the text, resulting in inaccurate language recognition of speech synthesis.

Method used

By obtaining the first text and its response text, the target statement is determined in the response text, and the target language when the target statement is converted into pronunciation is determined in combination with the language of the first text and the language above of the target statement.

Benefits of technology

It improves the accuracy of language recognition, ensures the language consistency of pronunciation during text-to-speech, and thus improves the accuracy of pronunciation synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020941A_ABST
    Figure CN120020941A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a language determination method and apparatus, and an electronic device. The method comprises the steps of obtaining a first text and a response text of the first text; determining a target statement in the response text; determining a first language of the first text and a second language of a previous text of the target statement; and based on the target statement, the response text, the first language and the second language, determining a target language when the target statement is converted into voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of text processing, and in particular, to a method, apparatus, and electronic device for determining a language type. Background Art

[0002] An electronic device can convert the text output by a language model into speech based on a speech synthesis algorithm. During the process of text-to-speech conversion, the electronic device needs to accurately identify the language type of the text to be converted into speech.

[0003] Currently, an electronic device can identify the language type of the text output by a language model. Then, when performing text-to-speech conversion, the electronic device can determine that the language type of the speech is also the same as that of the text. For example, if the text output by the language model is in Chinese, the electronic device can convert the text into Chinese speech. However, when the text includes numbers, serial numbers, etc., it is difficult for the electronic device to accurately determine the language type of the text. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, and electronic device for determining a language type, which are used to solve one or more technical problems in the prior art.

[0005] In a first aspect, the present disclosure provides a method for determining a language type, the method including:

[0006] Obtain a first text and a response text of the first text;

[0007] Determine a target sentence in the response text;

[0008] Determine a first language type of the first text and a second language type of the context above the target sentence;

[0009] Based on the target sentence, the response text, the first language type, and the second language type, determine a target language type when converting the target sentence into speech.

[0010] In a second aspect, the present disclosure provides a device for determining a language type, the device for determining a language type including an obtaining module, a first determining module, a second determining module, and a third determining module, where:

[0011] The obtaining module is configured to obtain a first text and a response text of the first text;

[0012] The first determining module is configured to determine a target sentence in the response text;

[0013] The second determining module is configured to determine a first language type of the first text and a second language type of the context above the target sentence;

[0014] The third determination module is configured to determine a target language for converting the target statement into speech based on the target statement, the response text, the first language, and the second language.

[0015] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;

[0016] The memory stores computer-executable instructions;

[0017] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the language determination method as described in the first aspect and various possible aspects related to the first aspect above.

[0018] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the language determination method as described in the first aspect and various possible aspects related to the first aspect above is implemented.

[0019] The present disclosure provides a language determination method, apparatus, and electronic device. The electronic device can obtain a first text and a response text of the first text, determine a target statement in the response text, determine a first language of the first text and a second statement above the target statement, and determine a target language for converting the target statement into speech based on the target statement, the response text, the first language, and the second language. In the above method, since the electronic device can combine the language of the first text and the language of the text above the target statement in the response text of the first text to accurately determine the target language for converting each text in the target statement into speech, the electronic device can improve the accuracy of language recognition and further improve the accuracy of the synthesized speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 FIG. is a schematic diagram of an application scenario provided by an embodiment of the present disclosure;

[0022] Figure 2 FIG. is a flowchart of a language determination method provided by an embodiment of the present disclosure;

[0023] Figure 3 FIG. is a schematic diagram of a process of obtaining a response text of a first text provided by an embodiment of the present disclosure;

[0024] Figure 4 A schematic diagram of a first language and a second language provided by an embodiment of the present disclosure;

[0025] Figure 5 A schematic diagram of a method for determining a target language provided by an embodiment of the present disclosure;

[0026] Figure 6 A schematic diagram of determining a target language provided by an embodiment of the present disclosure;

[0027] Figure 7 Another schematic diagram of determining a target language provided by an embodiment of the present disclosure;

[0028] Figure 8 Another schematic diagram of determining a target language provided by an embodiment of the present disclosure;

[0029] Figure 9 Another schematic diagram of determining a target language provided by an embodiment of the present disclosure;

[0030] Figure 10 A schematic diagram of the process of a method for determining a language provided by an embodiment of the present disclosure;

[0031] Figure 11 A schematic diagram of the structure of a language determination device provided by an embodiment of the present disclosure; and,

[0032] Figure 12 A schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0034] For ease of understanding, the concepts related to the embodiments of the present disclosure will be described below.

[0035] Electronic device: A device with wireless transceiver functions. The electronic device can be deployed on land, including indoor or outdoor, handheld, wearable or vehicle-mounted. The electronic device can be a mobile phone, a tablet computer (Pad), a computer with wireless transceiver functions, a virtual reality (VR) electronic device, an augmented reality (AR) electronic device, a wireless terminal in industrial control, a vehicle-mounted electronic device, a wireless terminal in self-driving, a wireless electronic device in remote medical, a wireless electronic device in smart grid, a wireless electronic device in transportation safety, a wireless electronic device in smart city, a wireless electronic device in smart home, a wearable electronic device, etc. The electronic device involved in the embodiments of the present disclosure can also be referred to as a terminal, a user equipment (UE), an access electronic device, a vehicle-mounted terminal, an industrial control terminal, a UE unit, a UE station, a mobile station, a mobile unit, a remote station, a remote electronic device, a mobile device, a UE electronic device, a wireless communication device, a UE agent or a UE device, etc. The electronic device can also be fixed or mobile.

[0036] In the related art, the electronic device can convert the text output by the language model into speech based on the speech synthesis algorithm. And during the process of text-to-speech conversion, the electronic device needs to accurately identify the language of the text converted into speech. Currently, the electronic device can directly identify the language of the text output by the language model. Then, when performing text-to-speech conversion, the electronic device can determine the language of the text as the language of the speech. For example, if the language of the text is Chinese, the speech generated by the electronic device is Chinese speech; if the language of the text is English, the speech generated by the electronic device is English speech. However, when the text includes numbers, symbols, formulas, etc., the electronic device cannot accurately determine the language of the text (for example, numbers can be played based on English or Chinese). As a result, the electronic device cannot accurately convert the text into speech.

[0037] To solve the problems in the related art, an embodiment of the present disclosure provides a language determination method. An electronic device can obtain a first text and a response text of the first text, and determine a target sentence in the response text. The electronic device can determine a first language of the first text and a second language of the context of the target sentence, and identify the text in the target sentence to determine the sentence type of the target sentence. The electronic device can obtain the number of characters in the response text, and based on the number of characters, the sentence type, the first language, and the second language, determine the target language when converting the target sentence into speech. In this way, the electronic device can flexibly determine the language when converting the target sentence into speech based on the number of characters and the sentence type of the target sentence. Moreover, the electronic device can combine the language of the first text and the language of the context of the target sentence to accurately determine the target language when converting each text in the target sentence into speech, thereby improving the accuracy of language recognition and the accuracy of text-to-speech conversion.

[0038] Next, in conjunction with Figure 1 , the application scenarios of the embodiments of the present disclosure will be described.

[0039] Figure 1 FIG. is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. Please refer to Figure 1 , including: voice, language model, and speech synthesis module. Among them, the speech synthesis module can be set in the electronic device, and the language model can also be set in the electronic device. The embodiments of the present disclosure do not limit this. The electronic device ( Figure 1 not shown) can input voice to the language model. The language model can determine the first text corresponding to the voice and generate a response text corresponding to the first text. The electronic device can identify the language of the response text and input the response text and the language of the response text to the speech synthesis module. The speech synthesis module can generate the voice corresponding to the response text. In the above method, the electronic device can accurately determine the language of each text in the response text based on the language of the voice and the language of the response text, thereby improving the accuracy of language recognition and the accuracy of the generated voice.

[0040] It should be noted 1, Figure 1 is only an example of the application scenario of the embodiments of the present disclosure, and is not a limitation on the application scenario of the embodiments of the present disclosure.

[0041] Next, specific embodiments will be used to describe in detail the technical solutions of the present disclosure and how the technical solutions of the present disclosure solve the above technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Next, the embodiments of the present disclosure will be described in conjunction with the drawings.

[0042] Figure 2Schematic flowchart of a language determination method provided by an embodiment of the present disclosure. Please refer to Figure 2 The method may include:

[0043] S201. Obtain a first text and a response text of the first text.

[0044] The execution subject of the embodiment of the present disclosure may be an electronic device or a language determination device provided in the electronic device. Among them, the language determination device may be implemented based on software, or may be implemented based on the combination of software and hardware. The embodiment of the present disclosure does not limit this.

[0045] Optionally, the first text may be the text of the questioner in a conversation, and the response text may be the text answering the first text. For example, in a conversation scenario, the conversation may include a questioner and an answerer. The first text may be the text associated with the questioner, and the response text may be the text associated with the answerer. For example, if text 1 is: What is the temperature today, and text 2 may be: 10 degrees Celsius. In this conversation segment, text 1 may be the first text, and text 2 may be the response text of the first text.

[0046] Optionally, the electronic device may determine the first text based on the acquired voice. For example, the electronic device may collect the voice emitted by the user, and then perform text recognition processing on the voice to obtain the first text. For example, the electronic device may receive the voice sent by other devices, and then perform text recognition processing on the voice to obtain the first text.

[0047] It should be noted that the electronic device may also obtain the first text based on any other feasible implementation method (for example, the electronic device may receive the first text sent by other devices). The embodiment of the present disclosure does not limit this.

[0048] Optionally, the electronic device may obtain the response text of the first text based on the following feasible implementation method: in response to a touch operation on the voice collection control, obtain the voice, perform voice recognition processing on the voice to obtain the first text corresponding to the voice, and input the first text into the language model to obtain the response text of the first text.

[0049] Optionally, the electronic device may display a Q&A page, which may include a voice collection control. When the user clicks the voice collection control, the electronic device may obtain the voice. The electronic device may determine the text associated with the voice (i.e., the first text) based on voice recognition technology, and input the first text into the language model. The language model may determine the answer associated with the first text based on the question associated with the first text, and then may obtain the response text of the first text.

[0050] Next, in combination with Figure 3, the process of obtaining the response text of the first text is described.

[0051] Figure 3 This is a schematic diagram of the process of obtaining the response text of the first text provided by the embodiments of the present disclosure. Please refer to Figure 3 , including: an electronic device, a speech-to-text module, and a language model. The display page of the electronic device may include a voice collection control. When the user clicks the voice collection control, the electronic device can record voice. At the end of voice collection, the electronic device can input the collected voice to the speech-to-text module, and the speech-to-text module can process the voice to obtain the first text and input the first text to the language model. The language model can output the response text corresponding to the first text. In this way, the user can ask questions to the electronic device in the form of voice, and the electronic device can generate the response text corresponding to the voice, improving the diversity of interaction and the user experience.

[0052] It should be noted that the speech-to-text module and the language model can be set in the electronic device (for example, the electronic device can be a device with end-side computing capabilities such as a computer or a server), and the embodiments of the present disclosure do not limit this.

[0053] S202. Determine the target sentence in the response text.

[0054] Optionally, the target sentence can be the sentence to be synthesized into voice. For example, if the electronic device synthesizes voice for sentence 1, the electronic device can determine sentence 1 as the target sentence; if the electronic device synthesizes voice for sentence 2, the electronic device can determine sentence 2 as the target sentence.

[0055] After the electronic device obtains the response text, it can perform clause segmentation on the response text to obtain multiple sentences, and the electronic device can determine the target sentence among the multiple sentences. For example, the response text can be the text output by the large language model, and the text output by the large language model can be streaming text (text output character by character). Therefore, the electronic device can perform clause segmentation on the streaming text, and then can obtain multiple complete sentences.

[0056] It should be noted that the electronic device can perform clause segmentation on the streaming text based on any feasible implementation method, and the embodiments of the present disclosure do not limit this.

[0057] S203. Determine the first language of the first text and the second language of the context of the target sentence.

[0058] Among them, the first language can be the language of the first text. For example, the first language can be the language of the voice input by the user. For example, when the user emits Chinese voice, the electronic device can convert the Chinese voice into Chinese text, that is, the first language of the first text can be Chinese.

[0059] Optionally, the second language may be the language of the text above the target sentence. For example, the second language may be the language of the first sentence in the response text, or the second language may be the language of the sentence immediately preceding the target sentence. The embodiments of the present disclosure do not limit this. For example, if the target sentence is the third sentence in the response text and the language of the first sentence in the response text is Chinese, the second language may be Chinese. For example, if the language of the sentence immediately preceding the target sentence is English, the second language may be English.

[0060] Next, the first language and the second language will be described in conjunction with Figure 4 .

[0061] Figure 4 FIG. Figure 4 is a schematic diagram of a first language and a second language provided by an embodiment of the present disclosure. Please refer to

[0062] Please refer to Figure 4 . An electronic device ( Figure 4 not shown) can determine that the language of the text "What hardware is image display related to?" is Chinese. Therefore, the electronic device can determine that the first language is Chinese. The electronic device can determine that the target sentences in the response text may be "1. GPU" and "2. Screen". Since the language of the text "Image display is mainly related to the following hardware" above the target sentences is Chinese, the electronic device can determine that the second language is Chinese.

[0063] It should be noted that when the electronic device determines the first language and the second language, if a sentence includes multiple languages, the electronic device may determine the language of the sentence as the language with the largest number of text characters. For example, if the first text is a mixed text of Chinese and English, and the number of Chinese phrases is greater than the number of English words, the electronic device may determine that the language of the first text is Chinese.

[0064] It should be noted that the electronic device may also determine the first language of the first text and the second language of the text above the target sentence based on any feasible implementation manner. The embodiments of the present disclosure do not limit this.

[0065] S204. Determine the target language for converting the target sentence into speech based on the target sentence, the response text, the first language, and the second language.

[0066] Among them, the target language can be the language when the text in the target statement is converted into speech. For example, the target statement may include multiple texts, and for each text when it is converted into speech, there is a corresponding language. For example, if the target statement is: Lunch starts at 12:00, then the target language of the text "Lunch" in this statement is Chinese, and the target language of the text "starts" is also Chinese. However, the electronic device needs to accurately identify the target language of the text "12:00".

[0067] Among them, the electronic device can determine the target language of each text in the target statement when it is converted into speech based on the following feasible implementation methods: identify the text in the target statement, determine the statement type of the target statement, and based on the statement type, response text, first language, and second language, determine the target language of each text in the target statement when it is converted into speech. In this way, the electronic device can flexibly determine the languages corresponding to non-standard words, formulas, and serial numbers in the target statement in combination with the statement type, thereby improving the accuracy of language determination.

[0068] Among them, the statement type can include at least one of the following: non-standard word type, formula type, number type, and standard word type. In this way, the electronic device can flexibly determine the target language for text-to-speech in the target statement in combination with the statement type.

[0069] Optionally, the non-standard word type can indicate that the target statement includes non-standard words. Among them, non-standard words can be words composed of other symbols except the characters and punctuation marks of this language. For example, non-standard words can be words composed of symbols such as Arabic numerals, currency symbols, mathematical symbols, and physical symbols, and non-standard words cannot be pronounced based on normal pronunciation rules. For example, if the statement is: The lunch break is at 12:00, then this statement includes the non-standard word "12:00". For example, if the target statement includes non-standard words such as ">" and "12:00", the electronic device can determine that the statement type of this target statement is the non-standard word type.

[0070] Optionally, the formula type can indicate that the target statement includes formulas. Among them, the formulas can be mathematical formulas, physical formulas, etc., and this disclosure embodiment does not limit this. The formulas can include mathematical symbols, physical symbols, etc. For example, if the target statement includes mathematical formulas, physical formulas, etc., the electronic device can determine that the statement type of this target statement is the formula type.

[0071] Optionally, the number type can indicate that the target statement includes numbers. Optionally, the numbers can be serial numbers in the target statement. For example, each statement in the response text generated by the large language model may include a serial number, and this serial number can be the number in the target statement. For example, if the target statement includes serial numbers such as "1、" and "2、", the electronic device can determine that the statement type of the target statement is the number type.

[0072] It should be noted that the target statement may also include any other words that cannot be pronounced according to normal pronunciation rules. The embodiments of the present disclosure do not limit this.

[0073] Optionally, the electronic device may perform text recognition processing on the target statement, and then obtain the statement type of the target statement. The electronic device may also process the target statement based on any other feasible implementation manner to obtain the statement type of the target statement. The embodiments of the present disclosure do not limit this.

[0074] Optionally, if the statement type is a non-standard word type and the number of words in the response text is small, the electronic device may determine that the language of the non-standard word in the target statement is the first language. If the statement type is a formula type and / or a number type, the electronic device may determine that the language of the formula in the target statement is the second language and determine that the language of the serial number in the target statement is the second language.

[0075] Optionally, after the electronic device determines the target language for each text in the target statement to be converted into speech, the above language determination method further includes: performing text-to-speech processing on the target statement based on the target language when converting the target statement into speech to obtain the target speech.

[0076] Among them, the electronic device may process the target statement based on a text-to-speech (TTS) module to obtain the target speech corresponding to the target statement. For example, if the target statement is: one<two, the electronic device may determine that the target languages of one and two are English, and determine that the target language of the “<” symbol is English, and then generate English speech. In this way, the electronic device can combine the TTS module to accurately synthesize the speech corresponding to the target statement and improve the accuracy of speech synthesis.

[0077] An embodiment of the present disclosure provides a language determination method. An electronic device can obtain a first text and a response text of the first text, and determine a target sentence in the response text. The electronic device can determine a first language of the first text and a second language of the context of the target sentence, and identify the text in the target sentence to determine the sentence type of the target sentence. The sentence type may include at least one of the following: non-standard word type, formula type, number type, and standard word type. The electronic device can determine a target language for converting the target sentence into speech based on the sentence type, the response text, the first language, and the second language, and perform text-to-speech processing on the target sentence based on the target language to obtain a target speech. In this way, the electronic device can flexibly and accurately determine the languages of non-standard words, formulas, and serial numbers in the target sentence in combination with the sentence type of the target sentence, improving the flexibility and accuracy of language determination. Moreover, since the languages of non-standard words, formulas, and serial numbers in the target sentence are highly accurate, the electronic device can improve the accuracy of the speech associated with the generated target sentence.

[0078] Based on the embodiment shown in Figure 2 , the following further describes the method for determining the target language for converting each text in the target sentence into speech in the above language determination method based on the sentence type, the response text, the first language, and the second language. Figure 5

[0079] Figure 5 It is a schematic diagram of a method for determining a target language provided by an embodiment of the present disclosure. Please refer to Figure 5 , and the method flow includes:

[0080] S501. Obtain the number of characters in the response text.

[0081] The number of characters may be the total number of characters in the response text. For example, if the response text is: The weather is nice today, the electronic device can determine that the number of characters in the response text is 7. For example, if the response text is: 12:00, the electronic device can determine that the number of characters in the response text is 5 (where ":" is also a character).

[0082] It should be noted that the electronic device can obtain the number of characters in the response text based on any feasible implementation manner, and the embodiments of the present disclosure do not limit this.

[0083] S502. Determine the target language for converting the target sentence into speech based on the number of characters, the sentence type, the first language, and the second language.

[0084] Among them, the electronic device determines the target language for each text in the target sentence to be converted into speech based on the number of characters, the type of sentence, the first language, and the second language. There are the following two cases:

[0085] Case 1: The number of characters is less than or equal to a preset threshold.

[0086] Among them, if the number of characters is less than or equal to the preset threshold, then it is determined whether the sentence type is a non-standard word type to obtain a judgment result, and based on the judgment result, the target language is determined. For example, if the number of characters is less than or equal to the preset threshold, it means that the response text of the first text has fewer characters, and the electronic device cannot determine the language of the target sentence based on the context information. The electronic device can accurately determine the target language corresponding to each text in the sentence in combination with the judgment result of whether there are non-standard words in the sentence.

[0087] It should be noted that in the embodiments of the present disclosure, the preset threshold can be 5. In this way, for the response text regarding the answering time, the electronic device can accurately identify the language of non-standard words. The preset threshold can also be any number, and the embodiments of the present disclosure do not limit this.

[0088] Optionally, the electronic device determines the target language based on the judgment result. Specifically, if the judgment result is that the sentence type is a non-standard word type, then the target language of the non-standard word in the target sentence is determined to be the first language. If the judgment result is that the sentence type is a standard word type, then the language associated with the text in the target sentence is determined to be the target language when the text is converted into speech.

[0089] For example, if the sentence type is a non-standard word type and the number of characters in the response text is less than or equal to 5, it means that the electronic device cannot determine the language of the non-standard word based on the context information. The electronic device can determine the first language as the language of the non-standard word. For example, the first text can be: What time to have dinner, and the corresponding response text of the first text can be: 12:00. In this case, the electronic device cannot determine the language of the text "12:00", but since the language of the first text is Chinese, the electronic device can determine the language of the text "12:00" as Chinese.

[0090] For example, if the sentence type is a standard word type and the number of characters in the response text is small, the electronic device can identify the language associated with the text in the target sentence. When the text is converted into speech, the language of the text in the speech can be the language associated with the text. For example, if the target sentence is: Hello, the electronic device can identify the language of "Hello" as Chinese, and then can determine the target language of the text "Hello" as Chinese.

[0091] Next, in combination with Figures 6 - 7, the process of determining the target language of the target statement in this case is described.

[0092] Figure 6 A schematic diagram for determining the target language provided by an embodiment of the present disclosure. Please refer to Figure 6 , including: a conversation page. Among them, the conversation page may include a questioner and an answerer. Among them, the first text input by the questioner may be: What time is it now? The answerer may output a response text: 12:00. Since the number of characters in the response text is small (5), and the statement in the response text includes a non-standard word "12:00", the electronic device cannot directly determine the target language of the statement based on "12:00". The electronic device may determine the language of the first text as the language of the text "12:00", that is, the target language is Chinese.

[0093] Figure 7 Another schematic diagram for determining the target language provided by an embodiment of the present disclosure. Please refer to Figure 7 , including: a conversation page. Among them, the conversation page may include a questioner and an answerer. Among them, the first text input by the questioner may be: What's the weather today? The answerer may output a response text: Sunny. Since the number of characters in the response text is small (5), and the words in the statement in the response text are all standard words, the electronic device can directly identify the language associated with the text in the statement, and then determine the language as the target language for converting the statement into speech, that is, the target language is Chinese.

[0094] In this way, when the short answer includes non-standard words, the electronic device can determine the target language for converting the non-standard words in the short answer into speech based on the language of the text of the question. When the short answer includes standard words, the electronic device can directly identify the language of the standard words, and then determine the language as the target language for converting the standard words into speech. In this way, not only can the flexibility of language determination be improved, but also the language accuracy of non-standard words can be improved.

[0095] Case 2: The number of characters is greater than the preset threshold.

[0096] If the number of characters is greater than the preset threshold, the target language is determined based on the statement type and the second language. For example, when the number of characters is greater than the preset threshold, it means that the number of characters in the response text is large. The electronic device can accurately determine the target language of the target statement in combination with the language of the above text of the target statement, improving the accuracy of the target language.

[0097] Among them, the electronic device determines the target language based on the statement type and the second language. Specifically, it can be: if the statement type is a formula type, the target language of the target statement is determined as the second language; if the statement type is a number type, the numbers in the target statement are determined as the second language, and the target language of the other texts in the target statement is the language associated with the other texts; if the statement type is a standard word type, the language associated with the text in the target statement is determined as the target language when the text is converted to speech.

[0098] For example, when the number of characters in the response text is large, if the statement type is a formula type, it indicates that there is a formula in the statement. Therefore, the electronic device can determine the second language above the formula as the language for converting each text in the formula to speech.

[0099] It should be noted that in the actual application process, when the electronic device clauses the response text, it can divide the formula into one statement, or divide consecutive non-standard words (such as 12:00) into one statement. The embodiments of the present disclosure do not limit this.

[0100] For example, when the number of characters in the response text is large, if the statement type is a number type, it indicates that there is a serial number in the statement and there is text after the serial number. Therefore, the electronic device can determine the second language above the statement as the language of the serial number, and determine the language associated with the text after the serial number as the target language of the text after the serial number. For example, if the target statement is: The weather today may be 1, clear day, 2, rainy day, then the electronic device can determine the target language of clear day and rainy day as English, and determine the target language of "1" and "2" as Chinese ("The weather today may be" is Chinese).

[0101] For example, when the number of characters in the response text is large, if the statement type is a standard word type, the electronic device can directly identify the language associated with the standard word in the statement and determine this language as the target language when the standard word is converted to speech. For example, if the target statement is: The weather today is really good, then the electronic device can determine the target language of this target statement as Chinese.

[0102] Next, in combination with Figures 8 - 9 , the process of determining the target language in this case will be described.

[0103] Figure 8 This is another schematic diagram for determining the target language provided by the embodiments of the present disclosure. Please refer to Figure 8, including: a conversation page. Among them, the conversation page may include a questioner and an answerer. Among them, the first text input by the questioner may be: Which hardware is image display related to? The answerer may output a response text: Image display is mainly related to the following hardware: 1. GPU, 2. Screen.

[0104] Please refer Figure 8 , an electronic device ( Figure 8 not shown) can determine that statement 1 in the response text is: Image display is mainly related to the following hardware, statement 2 is: 1. GPU, and statement 3 is: 2. Screen. Since there are serial numbers in statement 2 and statement 3, the statement types of statement 2 and statement 3 are numeric types.

[0105] Please see Figure 8 , since the language of statement 1 is Chinese, the electronic device can determine that the language of the serial numbers of statement 2 and statement 3 is Chinese. That is, the electronic device can determine that the target languages of "1" and "2" are Chinese (in the prior art, since GPU follows 1, the prior art would determine the target language of 1 to be English, resulting in a poor voice synthesis effect), the target language of "GPU" is English, and the target language of "screen" is Chinese.

[0106] Figure 9 This is another schematic diagram for determining the target language provided by an embodiment of the present disclosure. Please see Figure 9 , including: a conversation page. Among them, the conversation page may include a questioner and an answerer. Among them, the first text input by the questioner may be: Where is suitable to play today? The answerer may output a response text: It is a clear day today, and we can go camping. The electronic device ( Figure 8 not shown) can determine that statement 1 in the response text is: It is a clear day today, and statement 2 is: We can go camping. Since the statement types of statement 1 and statement 2 are both standard word types, the electronic device can determine the target languages of the texts in each statement based on the languages associated with the texts. That is, the target language of "clear day" is English, and the target languages of other texts are Chinese.

[0107] In this way, when the number of characters in the response text is relatively large, if there are serial numbers in the target statement, the electronic device can determine the language of the serial number based on the language above, thereby improving the accuracy of the language and the accuracy of voice synthesis.

[0108] An embodiment of the present disclosure provides a method for determining a target language. The method includes: obtaining the number of characters in a response text; if the number of characters is less than or equal to a preset threshold, determining whether the statement type is a non-standard word type to obtain a determination result, and determining the target language based on the determination result; when the number of characters is greater than the preset threshold, if the statement type is a formula type, determining that the target language of the target statement is a second language, if the statement type is a number type, determining that the numbers in the target statement are in the second language, and the target language of the other text in the target statement is the language associated with the other text, and if the statement type is a standard word type, determining the language associated with the text in the target statement as the target language when the text is converted into speech. In this way, the electronic device can flexibly determine the target language of each text in the target statement when the text is converted into speech, and moreover, the electronic device can accurately determine the languages of non-standard words, formulas, and serial numbers in the target statement, improving the accuracy of the determined language, and further improving the accuracy of speech synthesis.

[0109] Based on any of the above embodiments, below, in combination with Figure 10 , the process of the above language determination method will be described.

[0110] Figure 10 FIG. is a schematic diagram of the process of a language determination method provided by an embodiment of the present disclosure. Please refer to Figure 10 , which includes an electronic device. The display page of the electronic device may include a voice collection control. When the user clicks the voice collection control, the electronic device can record voice. When the voice collection ends, the electronic device can obtain the first text of the voice and display the first text on the conversation page.

[0111] Please refer to Figure 10 , the conversation page may include a questioner and an answerer. The first text input by the questioner may be: What hardware is image display related to? The answerer may output a response text: Image display is mainly related to the following hardware: 1. GPU, 2. Screen. The electronic device can determine that statement 1 in the response text is: Image display is mainly related to the following hardware, statement 2 is: 1. GPU, and statement 3 is: 2. Screen.

[0112] Please refer to Figure 10 , since there are serial numbers in statement 2 and statement 3, the statement types of statement 2 and statement 3 are number types. Since the language of statement 1 is Chinese, the electronic device can determine that the languages of the serial numbers in statement 2 and statement 3 are Chinese. That is, the electronic device can determine that the target languages of "1" and "2" are Chinese, the target language of "GPU" is English, and the target language of "screen" is Chinese.

[0113] Please refer to Figure 10, the electronic device can generate three segments of speech based on the target language of each text in the response text. The first segment of speech can include the Chinese speech "Image display is mainly related to the following hardware", the second segment of speech can include the Chinese speech "1" and the English speech "GPU", and the third segment of speech can include the Chinese speech "2" and the Chinese speech "screen". In this way, the electronic device can not only flexibly determine the target language of the target sentence in the response text, but also accurately determine the target language of non-standard words, formulas, and serial numbers in the target sentence, thereby improving the accuracy of speech synthesis.

[0114] Figure 11 FIG. is a schematic structural diagram of a language determination device provided by an embodiment of the present disclosure. Please refer to Figure 11 , the language determination device 110 includes an acquisition module 111, a first determination module 112, a second determination module 113, and a third determination module 114, where:

[0115] The acquisition module 111 is configured to acquire a first text and a response text of the first text;

[0116] The first determination module 112 is configured to determine a target sentence in the response text;

[0117] The second determination module 113 is configured to determine a first language of the first text and a second language of the context of the target sentence;

[0118] The third determination module 114 is configured to determine a target language when converting the target sentence into speech based on the target sentence, the response text, the first language, and the second language.

[0119] According to one or more embodiments of the present disclosure, the third determination module 114 is specifically configured to:

[0120] Identify the text in the target sentence to determine the sentence type of the target sentence, where the sentence type includes at least one of the following: non-standard word type, formula type, number type, and standard word type;

[0121] Based on the sentence type, the response text, the first language, and the second language, determine the target language of each text in the target sentence when converted into speech.

[0122] According to one or more embodiments of the present disclosure, the third determination module 114 is specifically configured to:

[0123] Acquire the number of characters in the response text;

[0124] Based on the number of characters, the statement type, the first language, and the second language, determine the target language for each piece of text in the target statement when converted to speech.

[0125] According to one or more embodiments of the present disclosure, the third determination module 114 is specifically configured to:

[0126] If the number of characters is less than or equal to a preset threshold, then determine whether the statement type is a non-standard word type to obtain a determination result, and based on the determination result, determine the target language;

[0127] If the number of characters is greater than the preset threshold, then determine the target language based on the statement type and the second language.

[0128] According to one or more embodiments of the present disclosure, the third determination module 114 is specifically configured to:

[0129] If the determination result is that the statement type is a non-standard word type, then determine that the target language of the non-standard word in the target statement is the first language;

[0130] If the determination result is that the statement type is a standard word type, then determine the language associated with the text in the target statement as the target language when the text is converted to speech.

[0131] According to one or more embodiments of the present disclosure, the third determination module 114 is specifically configured to:

[0132] If the statement type is the formula type, then determine that the target language of the target statement is the second language;

[0133] If the statement type is the number type, then determine that the numbers in the target statement are in the second language, and the target language of the other text in the target statement is the language associated with the other text;

[0134] If the statement type is the standard word type, then determine the language associated with the text in the target statement as the target language when the text is converted to speech.

[0135] According to one or more embodiments of the present disclosure, the acquisition module 111 is specifically configured to:

[0136] In response to a touch operation on the voice acquisition control, acquire voice;

[0137] Perform speech recognition processing on the voice to obtain a first text corresponding to the voice;

[0138] Input the first text into a language model to obtain a response text of the first text.

[0139] According to one or more embodiments of the present disclosure, the third determination module 114 is further configured to:

[0140] Perform text-to-speech processing on the target statement based on the target language when converting the target statement into speech, to obtain target speech.

[0141] The language determination device provided in this embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0142] Figure 12 It is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. Please refer to Figure 12 , which shows a schematic structural diagram of an electronic device 1200 suitable for implementing the embodiments of the present disclosure. Among them, the electronic device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (Personal Digital Assistant, abbreviated as PDA), tablet computers (Portable Android Device, abbreviated as PAD), portable multimedia players (Portable Media Player, abbreviated as PMP), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 12 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0143] As Figure 12 shown, the electronic device 1200 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (Read Only Memory, abbreviated as ROM) 1202 or a program loaded from a storage device 1208 into a random access memory (Random Access Memory, abbreviated as RAM) 1203. In the RAM 1203, various programs and data required for the operation of the electronic device 1200 are also stored. The processing device 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0144] Typically, the following devices can be connected to the I / O interface 1205: an input device 1206 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1207 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1208 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1209. The communication device 1209 can allow the electronic device 1200 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 12 the electronic device 1200 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.

[0145] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1209, or installed from the storage device 1208, or installed from the ROM 1202. When the computer program is executed by the processing device 1201, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0146] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0147] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.

[0148] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above-mentioned embodiments.

[0149] The embodiments of the present disclosure provide a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the methods that may be involved in various embodiments above are implemented.

[0150] The embodiments of the present disclosure provide a computer program product, including a computer program, and when the computer program is executed by the processor, the methods that may be involved in various embodiments above are implemented.

[0151] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, execute as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0153] The units described in the embodiments of this disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0154] The functions described above in this document may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0155] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0156] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0157] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0158] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0159] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message. As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0160] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0161] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions. The data may include information, parameters, messages, etc., such as the cross-flow indication information.

[0162] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0163] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0164] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for determining a language, characterized in that: include: Acquire a first text and a response text to the first text; determining a target sentence in the response text; Determining a first language of the first text and a second language of the context of the target sentence; A target language when converting the target sentence into speech is determined based on the target sentence, the response text, the first language, and the second language.

2. The method according to claim 1, characterized in that The step of determining a target language when converting the target sentence into speech based on the target sentence, the response text, the first language, and the second language includes: Recognize the text in the target sentence and determine the sentence type of the target sentence, wherein the sentence type includes at least one of the following: a non-standard word type, a formula type, a number type, and a standard word type; A target language when converting the target sentence into speech is determined based on the sentence type, the response text, the first language, and the second language.

3. The method according to claim 2, characterized in that The step of determining a target language when converting the target sentence into speech based on the sentence type, the response text, the first language, and the second language includes: Get the number of characters in the response text; A target language when converting the target sentence into speech is determined based on the number of characters, the sentence type, the first language, and the second language.

4. The method according to claim 3, characterized in that The step of determining a target language when converting the target sentence into speech based on the number of characters, the sentence type, the first language, and the second language includes: If the number of characters is less than or equal to a preset threshold, determining whether the sentence type is a non-standard word type, obtaining a determination result, and determining the target language based on the determination result; If the number of characters is greater than a preset threshold, the target language is determined based on the sentence type and the second language.

5. The method according to claim 4, characterized in that The determining the target language based on the judgment result includes: If the judgment result is that the sentence type is a non-standard word type, determining that the target language of the non-standard words in the target sentence is the first language; If the judgment result is that the sentence type is a standard word type, the language associated with the text in the target sentence is determined as the target language when the text is converted into speech.

6. The method according to claim 4 or 5, characterized in that: The determining the target language based on the sentence type and the second language includes: If the sentence type is the formula type, determining that the target language of the target sentence is the second language; If the sentence type is the number type, determining that the number in the target sentence is the second language, and the target language of other text in the target sentence is the language associated with the other text; If the sentence type is the standard word type, the language associated with the text in the target sentence is determined as the target language when the text is converted into speech.

7. The method according to any one of claims 1 to 5, characterized in that: Get the response text of the first text, including: Acquiring voice in response to a touch operation on the voice acquisition control; Performing speech recognition processing on the speech to obtain a first text corresponding to the speech; The first text is input into a language model to obtain a response text of the first text.

8. The method according to any one of claims 1 to 5, characterized in that: After determining the target language when converting the target sentence into speech, the method further includes: Based on the target language when converting the target sentence into speech, the target sentence is subjected to text-to-speech processing to obtain a target speech.

9. A language determination device, characterized in that: It includes an acquisition module, a first determination module, a second determination module and a third determination module, wherein: The acquisition module is used to acquire a first text and a response text of the first text; The first determination module is used to determine the target sentence in the response text; The second determination module is used to determine a first language of the first text and a second language of the context of the target sentence; The third determination module is used to determine the target language when converting the target sentence into speech based on the target sentence, the response text, the first language and the second language.

10. An electronic device, characterized in that: include: Processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the language determination method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the language determination method according to any one of claims 1 to 8 is implemented.