Method and device for converting text into voice, medium, electronic equipment and program product
By combining the first and second texts for language identification and using voting to determine the target language, the problem of inaccurate playback sequence language in large-scale model dialogue scenarios is solved, thus improving the accuracy and reliability of voice playback.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2026-03-10
AI Technical Summary
In large-scale model-based dialogue scenarios, it is difficult to accurately determine which language to use when playing numbered text.
Language identification is performed by combining the first and second texts, and the target language is determined by voting. The serial number is then played in audio format.
The accuracy and reliability of the language when the sequence number is played in voice format have been improved.
Smart Images

Figure CN121640985A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of electronic information technology, and in particular, to a text-to-speech method, device, medium, electronic equipment and program product. BACKGROUND
[0002] In a large model-based dialogue scenario, it is a common scenario to output text organized by serial numbers.
[0003] Currently, in a large model-based dialogue scenario, text is supported to be played in the form of speech, and when playing text organized by serial numbers in the form of speech, it is necessary to determine in which language the serial numbers are played. SUMMARY
[0004] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter's scope.
[0005] In a first aspect, the present disclosure provides a text-to-speech method, comprising: obtaining conversation content, the conversation content comprising first text and second text output by a dialogue model for the first text, the second text comprising text organized by serial numbers; identifying a target language for playing the serial numbers in the form of speech according to the first text and the second text; playing the serial numbers in the form of speech using the target language.
[0006] In a second aspect, the present disclosure provides a text-to-speech device, comprising: an obtaining module configured to obtain conversation content, the conversation content comprising first text and second text output by a dialogue model for the first text, the second text comprising text organized by serial numbers; an identifying module configured to identify a target language for playing the serial numbers in the form of speech according to the first text and the second text; a playing module configured to play the serial numbers in the form of speech using the target language.
[0007] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, the computer program being executed by a processing device to implement the steps of the method of the first aspect.
[0008] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0009] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0010] The above technical solution identifies the language based on the first text and the second text, thereby obtaining the target language for playing the serial number in voice form. This simultaneously considers the influence of the first text and the second text in the conversation content on the language used when playing the serial number in voice form, thus improving the accuracy and reliability of the language used when playing the serial number in voice form.
[0011] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart illustrating a text-to-speech method according to an exemplary embodiment of the present disclosure.
[0013] Figure 2 This is a schematic diagram illustrating a voting process according to an exemplary embodiment of the present disclosure.
[0014] Figure 3 This is a block diagram illustrating a text-to-speech apparatus according to an exemplary embodiment of the present disclosure.
[0015] Figure 4 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0017] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0018] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0020] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0021] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0022] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0023] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0024] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0026] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0027] In dialogue scenarios based on large models, at least the text output by the large model should be played in the form of speech. When playing text organized by sequence numbers, it is necessary to determine which language to use for playing the sequence numbers.
[0028] In related technologies, language identification is performed solely based on the text output by a large model to determine the language sampled for the corresponding playback sequence number. However, this method often struggles to accurately determine the language when the input information and the output text of the large model involve multiple languages. For example, in a conversation: First text: "Introduces three fruits, with the descriptions in English."
[0029] Second text: "The following are three kinds of fruits for you:" 1. Apple Apples are a globally popular fruit known for their round shape, smooth skin, and sweet to tart taste; 2. Banana Bananas are elongated, curved fruits that are typically eaten raw; 3. Oranges Oranges are a type of citrus fruit known for their vibrant orangecolor and sweet, tangy flavor.” In the above conversation, both the second and first texts contain language information in Chinese and English. Only the second text is used to accurately determine whether the number "1" should be read in Chinese or English.
[0030] Therefore, how to correctly play the sequence number in voice format is a technical problem that urgently needs to be solved.
[0031] Figure 1 This is a flowchart illustrating a text-to-speech method according to an exemplary embodiment of the present disclosure. This text-to-speech method can be applied to an electronic device and can be executed by a text-to-speech apparatus, wherein the text-to-speech apparatus can be implemented in software and / or hardware and can be configured in an electronic device. (Refer to...) Figure 1 The text-to-speech method may include the following steps: Step 110: Obtain the conversation content, which includes a first text and a second text output by the dialogue model in response to the first text. The second text includes text organized by sequence numbers.
[0032] The first text can refer to the text input by the user into the dialogue model, or it can be the text obtained by converting the speech input by the user into the dialogue model.
[0033] The dialogue model can be a dialogue model based on a large language model (LLM). This dialogue model can receive a first text, perform text understanding on the first text, and output a second text given in response to the first text.
[0034] The serial number can refer to a number that represents a sequence, such as a numerical serial number or a Roman numeral serial number.
[0035] Step 120: Based on the first text and the second text, language identification is performed to identify the target language used to play the serial number in speech form.
[0036] It should be noted that the target language of the serial numbers in the following text is understood to be the language used when the serial numbers are played in audio form.
[0037] The dialogue model can incorporate language recognition functionality to identify the target language used when playing the sequence number; in other embodiments, other trained language recognition models can also be used to achieve language recognition.
[0038] In this embodiment, language identification can be performed on the first text and the second text respectively to obtain the corresponding language identification results. Then, all language identification results are combined to determine the target language for the serial number. For example, step 120 above can be implemented as follows: language identification is performed on the first text to obtain the first language identification result for the serial number; language identification is performed on the second text to obtain the second language identification result for the serial number; based on the first and second language identification results, a voting process is performed on the language identification results of the serial number to obtain the target language for playing the serial number in voice form. The implementation methods for language identification and for determining the target language based on the language identification results can be referred to the following related embodiments, which will not be elaborated upon here.
[0039] Step 130: Play the serial number in the target language in voice format.
[0040] It should be understood that this embodiment can also simultaneously play other content in the second text, excluding the serial number, in the form of voice.
[0041] The above technical solution identifies the language based on the first text and the second text, thereby obtaining the target language for playing the serial number in voice form. This simultaneously considers the influence of the first text and the second text in the conversation content on the language used when playing the serial number in voice form, thus improving the accuracy and reliability of the language used when playing the serial number in voice form.
[0042] In some embodiments, the step of performing language identification based on the first text to obtain the first language identification result of the serial number can be implemented in the following manner: performing language identification based on the target information of the first text to obtain the first language identification result of the serial number, wherein the target information includes language information and / or semantic information.
[0043] In this context, the language information of the first text refers to the language used in the first text. For example, if the first text is "Introducing three types of fruit, with the introduction part in English," its corresponding language information is Chinese; if the first text is "Introduce three types of fruit, with the introduction part in English.", its corresponding language information is English. When using language information for language identification, the language information of the first text can be determined as the first language identification result of the sequence number. Continuing with the above example, if the first text is "Introducing three types of fruit, with the introduction part in English," its corresponding language information is Chinese, meaning the first language identification result of the sequence number is Chinese.
[0044] The semantic information of the first text is obtained through text understanding of its content. The semantic information of the first text can be represented by the language involved in its semantics. When using semantic information for language identification, the semantic information of the first text is determined as the first language identification result for the sequence number. Continuing the example above, if the first text is "Introducing three kinds of fruit, with the introduction in English," its corresponding semantic information includes introducing the relevant parts in English. Therefore, the first language identification result for the sequence number can be English.
[0045] In this method, both language information and semantic information of the first text can be used for language identification simultaneously. The first language identification result is determined based on the priority of different target information, which can be set according to the actual situation. For example, the priority of language information can be set to be greater than the priority of semantic information, that is, the language identification result determined based on language information is used as the first language identification result. For example, continuing the above example, if the first text is "Introducing three kinds of fruits, the introduction part is in English", the language identification result determined based on language information is Chinese, while the language identification result determined based on semantic information is English. Since the priority of language information is set to be greater than the priority of semantic information, in this embodiment, the language identification result determined based on language information can be used as the first language identification result.
[0046] It's worth noting that, besides being a specific language type, the first language identification result can also be "cannot be determined." For example, if the first text is a formula, and formulas cannot be interpreted as a language, then the first language identification result will be "cannot be determined."
[0047] The above method is used to determine the first language of the serial number by utilizing the language information and / or semantic information of the first text.
[0048] In some embodiments, the second language identification result includes the first sub-language identification result and the second sub-language identification result. The step of performing language identification based on the second text to obtain the second language identification result of the serial number can be implemented in the following manner: performing language identification based on the context preceding the first serial number in the second text to obtain the first sub-language identification result of the serial number; performing language identification based on the context following the first serial number in the second text to obtain the second sub-language identification result of the serial number.
[0049] It is worth noting that in this embodiment, the first sequence number refers to the first sequence number in a text organized by sequence number.
[0050] In this process, when identifying the language based on the context of the first serial number, the language used in the context is determined as the corresponding sub-language identification result. Continuing with the example above, the second text is “The following introduces three kinds of fruit: 1. Apple. Apples are a globally popular fruit known for their round shape, smoothskin, and sweet to tart taste…”. The text preceding the first serial number is “The following introduces three kinds of fruit,” which uses Chinese. Therefore, the first sub-language identification result for the serial number is Chinese. The text following the first serial number is “Apples are a globally popular fruit known for their round shape,” which uses English. Therefore, the second sub-language identification result for the serial number is English.
[0051] Similar to the first language identification result, the first sub-language identification result and the second sub-language identification result can also be indeterminate.
[0052] By using the above method, the context of the second text is used to determine the corresponding sub-language identification results, thereby realizing the judgment of the language corresponding to the serial number.
[0053] Furthermore, it should be understood that the first number in a text organized by sequence number can be regarded as a transitional element, and in the scenario of a large model, it is necessary to maintain the consistency of the sequence number reading. Therefore, the language can be determined by using only the context of the first number, thereby reducing the amount of computation involved in language recognition.
[0054] As described above, the language identification result for the sequence number can include candidate languages or "cannot be determined." A candidate language is a specific type of language, such as Chinese or English, while "cannot be determined" indicates that the language could not be identified. It should be understood that the language identification result in this embodiment includes a first language identification result, a first sub-language identification result, and a second sub-language identification result. The following example illustrates how the target language of the sequence number is obtained by combining the first language identification result, the first sub-language identification result, and the second sub-language identification result through a voting process.
[0055] First, in some embodiments, the step of voting on the language identification results of the serial number based on the first language identification result and the second language identification result to obtain the target language for playing the serial number in voice form can be implemented in the following way: voting on the language identification results of the serial number based on the first language identification result, the first sub-language identification result, and the second sub-language identification result to obtain the voting result; if the number of votes for the same candidate language in the voting result is greater than or equal to the preset number of votes, the candidate language is determined as the target language for playing the serial number in voice form.
[0056] The preset number of votes can be set based on the total number of votes in the voting results. For example, the preset number of votes can be an integer value that is higher than half of the total number of votes. For instance, in this embodiment, there are a total of three votes for the first language identification result, the first sub-language identification result, and the second sub-language identification result. Therefore, the preset number of votes can be set to 2.
[0057] Figure 2 This is a schematic diagram illustrating a voting process according to an exemplary embodiment of the present disclosure. In this diagram, the voting process is illustrated using Chinese or English as examples of candidate languages. (Refer to...) Figure 2 If there are more than 2 votes for Chinese or English, that is, the number of votes for the same candidate language is greater than or equal to the preset number of votes, then the candidate language can be determined as the target language for playing the serial number in voice form. In other words, Chinese or English with more than 2 votes is determined as the target language for the serial number.
[0058] In some embodiments, the step of voting on the language identification results of the serial number based on the first language identification result and the second language identification result to obtain the target language for playing the serial number in voice form may further include the step of: when there is only one vote as a candidate language in the voting results and the other votes are all undeterminable, determining the candidate language as the target language for playing the serial number in voice form.
[0059] Continue to refer to Figure 2 If Chinese or English equals 1 vote, and the remaining 2 votes are undetermined, then only 1 vote is a candidate language and the other votes are undetermined. Therefore, the candidate language can be determined as the target language for playing the serial number in voice form. It should be understood that the candidate language is English or Chinese determined based on language identification.
[0060] In some embodiments, the step of voting on the language identification results of the serial number based on the first language identification results and the second language identification results to obtain the target language for playing the serial number in voice form may further include the step of: if the voting results meet preset conditions, determining the third language identification result determined according to a preset fallback strategy as the target language for playing the serial number in voice form, wherein the preset conditions include all votes in the voting results being undeterminable or the votes in the voting results being different from each other.
[0061] The preset fallback strategy can determine the third language recognition result based on the second text. As an example, determining the third language recognition result based on the second text and the preset fallback strategy can be implemented as follows: determine the position of the first sequence number in the second text, determine the context based on the position of the first sequence number, and determine the third language recognition result based on the language used in the context. The implementation method for determining the language based on the context can refer to the relevant embodiments described above, and will not be repeated here.
[0062] Continue to refer to Figure 2 One vote is given for Chinese and one for English, and the remaining vote is undetermined, which can be considered as the votes being distinct. If all votes in the voting results are either undetermined or distinct, then the third language recognition result determined by the preset fallback strategy can be selected as the target language for playing the sequence number in speech.
[0063] It should be understood that in some cases, the third language recognition result may also be undeterminable. Therefore, in order to further improve the practicality of this disclosure, when the third language recognition result is undeterminable, the default language can be set as the target language for playing the serial number in voice form, as a further fallback strategy to avoid the situation where the serial number cannot be played in voice form.
[0064] By using the above method, a decision-making strategy for the target language is set according to the different number of votes in the voting results, thereby improving the accuracy and reliability of the target language.
[0065] Based on the same concept, this disclosure provides a text-to-speech apparatus. Figure 3 This is a block diagram illustrating a text-to-speech apparatus according to an exemplary embodiment of the present disclosure. (Refer to...) Figure 3 The text-to-speech device may include Acquisition module 301 is used to acquire conversation content, the conversation content including first text and second text output by the dialogue model for the first text, the second text including text organized by sequence number; The recognition module 302 is used to identify the target language for playing the serial number in speech form based on the first text and the second text; The playback module 303 is used to play the serial number in the target language in the form of speech.
[0066] In some embodiments, the identification module 302 includes: The first identification submodule is used to perform language identification based on the first text to obtain the first language identification result of the serial number; The second identification submodule is used to perform language identification based on the second text to obtain the second language identification result of the serial number; The voting submodule is used to perform voting processing on the language recognition result of the serial number based on the first language recognition result and the second language recognition result, so as to obtain the target language for playing the serial number in voice form.
[0067] In some embodiments, the first identification submodule is further configured to: Language identification is performed based on the target information of the first text to obtain the first language identification result of the serial number. The target information includes language information and / or semantic information.
[0068] In some embodiments, the second language identification result includes a first sub-language identification result and a second sub-language identification result, and the first identification submodule is further configured to: Based on the preceding text of the first serial number in the second text, the language is identified to obtain the first sub-language identification result of the serial number; Based on the language identification of the text following the first serial number in the second text, the second sub-language identification result of the serial number is obtained.
[0069] In some embodiments, the language identification result of the sequence number includes candidate languages or cannot be determined, and the voting submodule is further used for: Based on the first language identification result, the first sub-language identification result, and the second sub-language identification result, a voting process is performed on the language identification result of the sequence number to obtain the voting result; If the number of votes for the same candidate language in the voting results is greater than or equal to the preset number of votes, the candidate language will be determined as the target language for playing the serial number in voice form.
[0070] In some embodiments, the voting submodule is further configured to: If, in the voting results, there is only one vote for the candidate language and all other votes are for the language that cannot be determined, then the candidate language is determined as the target language for playing the serial number in voice form.
[0071] In some embodiments, the voting submodule is further configured to: If the voting results meet the preset conditions, the third language identification result determined according to the preset fallback strategy will be determined as the target language for playing the serial number in voice form. The preset conditions include that all votes in the voting results are either "cannot be determined" or that the votes in the voting results are different from each other.
[0072] The implementation methods of each module in the above-mentioned text-to-speech device 300 can refer to the above method embodiments, and will not be repeated here.
[0073] Based on the same inventive concept, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described method.
[0074] Based on the same inventive concept, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0075] Based on the same inventive concept, this disclosure also provides an electronic device, including: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the above method.
[0076] The following is for reference. Figure 4 This diagram illustrates a structural schematic of an electronic device 400 suitable for implementing embodiments of the present disclosure. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0077] like Figure 4As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. Processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0078] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0079] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined in the methods of embodiments of this disclosure.
[0080] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0081] In some implementations, electronic devices can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0082] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0083] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire session content, the session content including first text and second text output by a dialogue model for the first text, the second text including text organized by serial numbers; identify a target language for playing the serial numbers in speech based on the first text and the second text; and play the serial numbers in speech using the target language.
[0084] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0086] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0087] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0088] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0089] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0090] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0091] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A method of text-to-speech, characterized by, The method comprises: obtaining session content, the session content comprising first text and second text output by a dialogue model for the first text, the second text comprising text organized by serial numbers; identifying a target language for playing the serial numbers in a voice form according to the first text and the second text; playing the serial numbers in the voice form using the target language.
2. The method of claim 1, wherein, The identifying the target language for playing the serial numbers in the voice form according to the first text and the second text comprises: performing language identification on the first text to obtain a first language identification result of the serial numbers; performing language identification on the second text to obtain a second language identification result of the serial numbers; performing voting processing on the language identification result of the serial numbers according to the first language identification result and the second language identification result to obtain the target language for playing the serial numbers in the voice form.
3. The method of claim 2, wherein, The performing language identification on the first text to obtain the first language identification result of the serial numbers comprises: performing language identification on target information of the first text to obtain the first language identification result of the serial numbers, the target information comprising language information and / or semantic information.
4. The method of claim 2, wherein, The second language identification result comprises a first sub-language identification result and a second sub-language identification result, and the performing language identification on the second text to obtain the second language identification result of the serial numbers comprises: performing language identification on a preceding text of a first serial number in the second text to obtain the first sub-language identification result of the serial numbers; performing language identification on a following text of the first serial number in the second text to obtain the second sub-language identification result of the serial numbers.
5. The method of claim 4, wherein, The language identification result of the serial numbers comprises a candidate language or an undetermined result, and the performing voting processing on the language identification result of the serial numbers according to the first language identification result and the second language identification result to obtain the target language for playing the serial numbers in the voice form comprises: performing voting processing on the language identification result of the serial numbers according to the first language identification result, the first sub-language identification result and the second sub-language identification result to obtain a voting result; in a case where a number of votes for a same candidate language in the voting result is greater than or equal to a preset number of votes, determining the candidate language as the target language for playing the serial numbers in the voice form.
6. The method of claim 5, wherein, The performing voting processing on the language identification result of the serial numbers according to the first language identification result and the second language identification result to obtain the target language for playing the serial numbers in the voice form further comprises: in a case where there is only one vote for the candidate language in the voting result and other votes are all for the undetermined result, determining the candidate language as the target language for playing the serial numbers in the voice form.
7. The method of claim 5, wherein, The performing voting processing on the language identification result of the serial numbers according to the first language identification result and the second language identification result to obtain the target language for playing the serial numbers in the voice form further comprises: In a case where the voting result meets a preset condition, a third language recognition result determined according to a preset bottom strategy is determined as a target language used for playing the serial number in a voice form, and the preset condition includes that all the votes in the voting result are the unable-to-judge or the votes in the voting result are different from each other.
8. An apparatus for text-to-speech, the apparatus comprising: Comprising: An obtaining module, configured to obtain session content, the session content comprising a first text and a second text output by a dialogue model for the first text, the second text comprising texts organized by serial numbers; An identifying module, configured to identify, according to the first text and the second text, a target language used for playing the serial numbers in a voice form; A playing module, configured to play the serial numbers in a voice form by using the target language.
9. A computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by a processing device to implement the steps of the method in any one of claims 1-7.
10. An electronic device, comprising: Comprising: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method in any one of claims 1-7.
11. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processor to implement the steps of the method in any one of claims 1-7.