Text-to-Speech Method, Device, Storage Medium and Computer Equipment
By performing multi-dimensional emotional recognition of text information and calling preset voice conversion models, the problem of unclear voice communication and difficult text communication to convey emotions in noisy environments is solved, and high-quality communication and convenient information acquisition are achieved.
Patent Information
- Application Number
- CN202111620527.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-27
AI Technical Summary
When performing voice communication in noisy environments, the voice input is unclear, which affects the quality of communication. It is difficult to accurately convey emotions and tone of text communication, especially for people with low cultural level or visual impairment.
By obtaining the text information to be converted, multi-dimensional emotional recognition is performed, including intention, emotion and tone recognition, and matching preset speech conversion model is called to convert text information into speech information.
It has achieved the improvement of communication quality in noisy environments, conveying emotions and tone, and facilitates recipients with low cultural level or visual impairment to obtain communication content.
Smart Images

Figure CN114299919B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a text-to-speech method, device, storage medium and computer equipment. Background Art
[0002] In today's society, communication and exchanges between people are becoming more frequent and closer. The birth of smartphones has also made communication tools based on wireless communication possible, and voice communication has become the main mode of communication.
[0003] At present, when the message initiator is in a noisy environment, because the voice input will be unclear and affect the communication quality, the message initiator usually switches to text communication. However, this text communication method is difficult to accurately convey the message initiator's emotions or tone at the time. In addition, if the message receiver has a low level of education, cannot read text messages, or has visual impairments, this method will cause inconvenience to the message receiver. Summary of the invention
[0004] The present invention provides a text-to-speech method, device, storage medium and computer equipment, which are mainly capable of converting text information in the communication process into voice information, thereby avoiding inconvenience to the information receiver.
[0005] According to a first aspect of the present invention, there is provided a text-to-speech method, comprising:
[0006] Get the text information to be converted;
[0007] Performing multi-dimensional emotion recognition on the text information to obtain intention recognition results, emotion recognition results, and tone recognition results corresponding to the text information;
[0008] A preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result is called to convert the text information into speech information.
[0009] Optionally, calling a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result to convert the text information into speech information includes:
[0010] Obtaining an initial speech conversion model and its corresponding multiple sets of model parameters;
[0011] Determine a target model parameter corresponding to the intention recognition result, the emotion recognition result and the tone recognition result from the multiple groups of model parameters, wherein each group of emotion recognition results corresponds to a group of model parameters, and the group of emotion recognition results includes the intention recognition result, the emotion recognition result and the tone recognition result;
[0012] Add the target model parameters to the initial voice conversion model to obtain a preset voice conversion model that matches the intent recognition result, the emotion recognition result, and the tone recognition result;
[0013] Call the matching preset voice conversion model to convert the text information into voice information.
[0014] Optionally, the obtaining of the text information to be converted includes:
[0015] Receive the text information input by the sender;
[0016] The calling of the preset voice conversion model that matches the intent recognition result, the emotion recognition result, and the tone recognition result to convert the text information into voice information includes:
[0017] If the text information contains special characters, verify the identity of the sender, and call the preset voice conversion model that matches the intent recognition result, the emotion recognition result, the tone recognition result, and the identity information of the sender to convert the text information input by the sender into voice information.
[0018] Optionally, the receiving of the text information input by the sender includes:
[0019] Detect the sound decibel of the environment where the sender is currently located;
[0020] If the sound decibel is greater than the preset sound decibel, provide a text input interface and receive the text information input by the sender based on the text input interface; or
[0021] Obtain the historical conversation record between the sender and the receiver;
[0022] Based on the historical conversation record, determine the input mode selected by the sender during the last conversation with the receiver;
[0023] If the input mode selected by the sender during the last conversation with the receiver is the text mode, output a text input interface and receive the text information input by the sender based on the text input interface.
[0024] Optionally, before receiving the text information input by the sender, the method further includes:
[0025] Collect the voice information of the sender reading multiple groups of scenario sentences, where the intents, emotions, or tones corresponding to different groups of scenario sentences are different;
[0026] Determine the text information corresponding to each group of voice information read by the sender. Based on the multiple groups of voice information and their corresponding text information, train the initial voice conversion model bound to the identity information of the sender and its corresponding multiple groups of model parameters.
[0027] Optionally, after training the initial voice conversion model bound to the identity information of the sender and its corresponding multiple groups of model parameters based on the multiple groups of voice information and their corresponding text information, the method further includes:
[0028] Collect the real-time voice information of the sender through a predetermined device;
[0029] Optimize the multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the collected real-time voice information; and / or
[0030] Obtain the specific text input by the sender and the voice information input by the sender for the specific text;
[0031] Optimize the multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the specific text and its corresponding voice information.
[0032] Optionally, after collecting the real-time voice information of the sender through a predetermined device, the method further includes:
[0033] Perform text conversion on the real-time voice information to obtain the real-time text information corresponding to the real-time voice information;
[0034] Perform intention, emotion, and tone recognition on the real-time text information respectively to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the real-time text information;
[0035] If it is determined that the emotion of the sender is abnormal according to the intention recognition result corresponding to the real-time text information, or it is determined that the emotion recognition result corresponding to the real-time text information does not match the tone recognition result, then stop collecting the real-time voice information of the sender.
[0036] Optionally, after training the initial voice conversion model bound to the identity information of the sender and its corresponding multiple groups of model parameters based on the multiple groups of voice information and their corresponding text information, the method further includes:
[0037] Obtain the trial reading statement input by the sender;
[0038] Using multiple sets of model parameters of the trained initial voice conversion model bound to the identity information of the sender, convert the trial reading statement into voice information and play it to the sender, and output a selection and correction interface corresponding to the trial reading statement;
[0039] Obtain the revised voice corresponding to the target character in the trial reading statement selected by the sender based on the selection and correction interface;
[0040] Optimize multiple sets of model parameters of the initial voice conversion model bound to the identity information of the sender based on the revised voice.
[0041] Optionally, after receiving the text information input by the sender, the method further includes:
[0042] When the text information is in dialect, obtain the output voice type selected by the sender;
[0043] If the output voice type is dialect voice, call the preset dialect voice conversion model bound to the identity information of the sender to convert the dialect input by the sender into dialect voice information;
[0044] If the output voice type is standard voice, use a preset dialect word library to convert the dialect into standard text information;
[0045] Perform multi-dimensional emotion recognition on the standard text information to obtain the emotion recognition result of the standard text information in multiple dimensions;
[0046] Call the preset voice conversion model that matches the emotion recognition result in multiple dimensions and the identity information of the sender to convert the standard text information into standard voice information.
[0047] Optionally, after calling the preset voice conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result to convert the text information into voice information, the method further includes:
[0048] In response to the received sound wave conversion instruction, convert the audio sound wave in the voice information into a bone conduction sound wave.
[0049] Optionally, after calling the preset voice conversion model that matches the intention recognition result, the emotion recognition result, the tone recognition result, and the identity information of the sender to convert the text information input by the sender into voice information, the method further includes:
[0050] In response to the received background sound addition instruction, output and display a list of background sounds, obtain a selection instruction for selecting a target background sound from the list of background sounds, and add the target background sound to the voice information by superimposing the sound waves in the target background sound on the sound waves in the voice information; or
[0051] Obtain the current conversation record between the sender and the receiver, and determine the current scene where the sender is located according to the conversation record; determine a target background sound that matches the current scene where the sender is located, and add the target background sound to the voice information by superimposing the sound waves in the target background sound on the sound waves in the voice information.
[0052] Optionally, after invoking a preset voice conversion model that matches the intent recognition result, the emotion recognition result, the tone recognition result, and the identity information of the sender, and converting the text information input by the sender into voice information, the method further includes:[[]]
[0053] In response to the encrypted voice addition instruction triggered by the sender, obtain the encrypted voice of the sender;
[0054] Adjust the frequency of the sound waves in the encrypted voice to a specific frequency that cannot be recognized by the human ear;
[0055] Add the adjusted encrypted voice to the voice information by superimposing the sound waves in the adjusted encrypted voice on the sound waves in the voice information.
[0056] Optionally, the method further includes:[[]]
[0057] Collect the operation data of the sender on the communication device during this communication process;
[0058] Match the operation data in this communication process with the historical operation data;
[0059] If the operation data in this communication process does not match the historical operation data, verify the identity information of the sender by turning on the camera device.
[0060] According to the second aspect of the present invention, there is provided a text-to-voice device, including:[[]]
[0061] An acquisition unit for acquiring text information to be converted;
[0062] An identification unit for performing multi-dimensional emotion recognition on the text information to obtain an intent recognition result, an emotion recognition result, and a tone recognition result corresponding to the text information;
[0063] A conversion unit, configured to call a preset voice conversion model that matches the intent recognition result, the emotion recognition result, and the tone recognition result, and convert the text information into voice information.
[0064] Optionally, the conversion unit includes: a first acquisition module, a first determination module, an addition module, and a conversion module.
[0065] The first acquisition module is configured to acquire an initial voice conversion model and its corresponding multiple groups of model parameters.
[0066] The first determination module is configured to determine, from the multiple groups of model parameters, target model parameters commonly corresponding to the intent recognition result, the emotion recognition result, and the tone recognition result, where each group of emotion recognition results corresponds to a group of model parameters, and the group of emotion recognition results includes an intent recognition result, an emotion recognition result, and a tone recognition result.
[0067] The addition module is configured to add the target model parameters to the initial voice conversion model to obtain a preset voice conversion model that matches the intent recognition result, the emotion recognition result, and the tone recognition result.
[0068] The conversion module is configured to call the matching preset voice conversion model to convert the text information into voice information.
[0069] Optionally, the acquisition unit is specifically configured to receive text information input by a sender.
[0070] The conversion unit is specifically configured to, if the text information contains special characters, verify the identity of the sender, and call a preset voice conversion model that matches the intent recognition result, the emotion recognition result, the tone recognition result, and the identity information of the sender, and convert the text information input by the sender into voice information.
[0071] Optionally, the acquisition unit includes: a detection module, a receiving module, a second acquisition module, and a second determination module.
[0072] The detection module is configured to detect the sound decibel of the environment where the sender is currently located.
[0073] The receiving module is configured to, if the sound decibel is greater than a preset sound decibel, provide a text input interface and receive the text information input by the sender based on the text input interface.
[0074] The second acquisition module is configured to acquire the historical conversation record between the sender and the receiver.
[0075] The second determination module is configured to determine the input mode selected by the sender during the last conversation with the receiver based on the historical conversation record;
[0076] The receiving module is further configured to, if the input mode selected by the sender during the last conversation with the receiver is the text mode, output a text input interface and receive the text information input by the sender based on the text input interface.
[0077] Optionally, the apparatus further includes: a collection unit and a training unit,
[0078] The collection unit is configured to collect the voice information of the sender reading multiple groups of scenario sentences, where the intents, emotions, or tones corresponding to different groups of scenario sentences are different;
[0079] The training unit is configured to determine the text information corresponding to each of the multiple groups of voice information read by the sender, and based on the multiple groups of voice information and their corresponding text information, train an initial voice conversion model bound to the identity information of the sender and its corresponding multiple groups of model parameters.
[0080] Optionally, the apparatus further includes: an optimization unit,
[0081] The collection unit is further configured to collect the real-time voice information of the sender through a predetermined device;
[0082] The optimization unit is configured to optimize multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the collected real-time voice information;
[0083] The obtaining unit is configured to obtain the specific text input by the sender and the voice information input by the sender for the specific text;
[0084] The optimization unit is further configured to optimize multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the specific text and its corresponding voice information.
[0085] Optionally, the apparatus further includes: an interruption unit,
[0086] The conversion unit is further configured to perform text conversion on the real-time voice information to obtain the real-time text information corresponding to the real-time voice information;
[0087] The recognition unit is further configured to perform intent, emotion, and tone recognition on the real-time text information respectively to obtain an intent recognition result, an emotion recognition result, and a tone recognition result corresponding to the real-time text information;
[0088] The interruption unit is configured to interrupt the acquisition of the real-time voice information of the sender if it is determined according to the intention recognition result corresponding to the real-time text information that the emotion of the sender is abnormal, or if it is determined that the emotion recognition result corresponding to the real-time text information does not match the tone recognition result.
[0089] Optionally, the acquisition unit is further configured to acquire a trial reading statement input by the sender;
[0090] The conversion unit is further configured to use multiple groups of model parameters of the initial voice conversion model trained and bound to the identity information of the sender to convert the trial reading statement into voice information and play it to the sender, and output a selection correction interface corresponding to the trial reading statement;
[0091] The acquisition unit is further configured to acquire the revised voice corresponding to the target character in the trial reading statement selected by the sender based on the selection correction interface;
[0092] The optimization unit is further configured to optimize multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the revised voice.
[0093] Optionally, the acquisition unit is further configured to acquire the output voice type selected by the sender when the text information is in dialect;
[0094] The conversion unit is further configured to, if the output voice type is dialect voice, call a preset dialect voice conversion model bound to the identity information of the sender to convert the dialect input by the sender into dialect voice information;
[0095] The conversion unit is further configured to, if the output voice type is standard voice, use a preset dialect word library to convert the dialect into standard text information;
[0096] The recognition unit is further configured to perform multi-dimensional emotion recognition on the standard text information to obtain the emotion recognition result of the standard text information in multiple dimensions;
[0097] The conversion unit is further configured to call a preset voice conversion model that matches the emotion recognition result in multiple dimensions and the identity information of the sender to convert the standard text information into standard voice information.
[0098] Optionally, the conversion unit is further configured to convert the audio sound wave in the voice information into a bone conduction sound wave in response to a received sound wave conversion instruction.
[0099] Optionally, the device further includes: a superimposing unit,
[0100] The superimposing unit is configured to, in response to a received background sound addition instruction, output and display a list of background sounds, obtain a selection instruction for selecting a target background sound from the list of background sounds, and add the target background sound to the voice information by superimposing the sound waves in the target background sound on the sound waves in the voice information; or obtain the current conversation record between the sender and the receiver, and determine the scene where the sender is currently located according to the conversation record; determine a target background sound that matches the scene where the sender is currently located, and add the target background sound to the voice information by superimposing the sound waves in the target background sound on the sound waves in the voice information.
[0101] Optionally, the obtaining unit is further configured to, in response to an encrypted voice addition instruction triggered by the sender, obtain the encrypted voice of the sender.
[0102] The superimposing unit is further configured to adjust the frequency of the sound waves in the encrypted voice to a specific frequency that cannot be recognized by the human ear; add the adjusted encrypted voice to the voice information by superimposing the sound waves in the adjusted encrypted voice on the sound waves in the voice information.
[0103] Optionally, the apparatus further includes: a matching unit.
[0104] The collecting unit is further configured to collect the operation data of the sender on the communication device during the current communication process.
[0105] The matching unit is configured to match the operation data during the current communication process with the historical operation data; if the operation data during the current communication process does not match the historical operation data, verify the identity information of the sender by turning on the camera device.
[0106] According to a third aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned upper and lower body movement matching method is implemented.
[0107] According to a fourth aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the above-mentioned upper and lower body movement matching method is implemented.
[0108] A method, device, storage medium, and computer device for converting text to speech provided by the present invention can, compared with the current way of communicating using text, obtain the text information to be converted; perform multi-dimensional emotion recognition on the text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information; at the same time, call a preset speech conversion model that matches the intention recognition result, emotion recognition result, and tone recognition result to convert the text information into speech information. Through the converted speech information, the recipient can feel the mood, tone, and intention of the sender at that time. In addition, for recipients with low educational levels or visual impairments, it is also more convenient to obtain the communication content. BRIEF DESCRIPTION OF THE DRAWINGS
[0109] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0110] Figure 1 shows a schematic flowchart of a method for converting text to speech provided by an embodiment of the present invention;
[0111] Figure 2 shows a schematic flowchart of another method for converting text to speech provided by an embodiment of the present invention;
[0112] Figure 3 shows a schematic structural diagram of a device for converting text to speech provided by an embodiment of the present invention;
[0113] Figure 4 shows a schematic structural diagram of another device for converting text to speech provided by an embodiment of the present invention;
[0114] Figure 5 shows a schematic physical structure diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0115] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0116] Currently, this way of text communication cannot convey the mood or tone of the information sender at that time. In addition, if the recipient of the information has a low educational level and cannot read the text information, or has a visual impairment, this way will cause inconvenience to the recipient of the information.
[0117] To solve the above problems, an embodiment of the present invention provides a method for converting text to speech, as Figure 1 shown, the method includes:
[0118] 101. Obtain the text information to be converted.
[0119] Among them, the text information to be converted is the text information sent by the sender to the receiver through the communication software. The embodiments of the present invention are mainly applicable to the scenario of converting text information in the communication process into voice information. The execution subject of the embodiments of the present invention is a device or equipment capable of performing voice conversion on text information, specifically, it can be a client or a server.
[0120] In a specific application scenario, there are usually two input methods for the sender's client. One is the text input method, and the other is the voice input method. When the sender enters text information on the client and selects voice conversion, the sender's client will obtain the text information entered by the sender and use it as the text information to be converted. Then, the sender's client directly converts the text information into voice information and sends it to the receiver's client. In addition, the sender can also directly send the entered text information to the receiver. After the receiver's client receives the text information, the receiver can select voice conversion in the client. At this time, the receiver's client will obtain the text information received by the receiver and convert the text information into voice information to play for the receiver.
[0121] Furthermore, the sender's client can also send the text information entered by the sender to the server. After the server receives the text information sent by the sender, it will send an information prompt to the receiver. At the same time, it directly converts the text information sent by the sender into voice information on the server side and stores the voice information and the text information in correspondence. When the receiver sees the prompt information and knows that there is communication information from the sender, the receiver will send an information acquisition request to the server. The server will send the text information or voice information to the receiver's client based on the information reception mode selected by the receiver.
[0122] It can be seen from this that the acquisition process and voice conversion process of the text information to be converted can be executed either in the client or in the server. The embodiments of the present invention do not make specific limitations on this.
[0123] 102. Perform multi-dimensional emotion recognition on the text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information.
[0124] Among them, multi-dimensional emotion recognition includes intention recognition, emotion recognition, and tone recognition. The intention recognition result includes requests, daily chats, others asking for help from me, etc. The tone recognition result includes statements, questions, imperatives, exaggerations, prayers, assumptions, emphases, rhetorical questions, euphemisms, etc. The emotion recognition result includes gratitude, pleasure, love, complaints, anger, disgust, fear, etc.
[0125] For the embodiments of the present invention, in order to enable the receiving party to perceive the tone, emotion, and intention of the sending party at that time, during the process of voice conversion, it is necessary to first perform intention recognition, emotion recognition, and tone recognition on the text information. For the specific process of intention recognition, as an alternative embodiment, the method includes: determining the semantic information vectors of each word segment corresponding to the text information; inputting the semantic information vectors corresponding to each word segment into a preset intention recognition model for intention recognition to obtain the intention recognition result corresponding to the text information. Further, the determining the semantic information vectors of each word segment corresponding to the text information includes: determining the query vector, key vector, and value vector corresponding to any one word segment among the word segments; multiplying the query vector corresponding to the any one word segment by the key vectors corresponding to each word segment to obtain the attention scores of each word segment for the any one word segment; multiplying and summing the attention scores and value vectors corresponding to each word segment to obtain the semantic information vector corresponding to any one word segment. Among them, the preset intention recognition model can specifically be a multi-layer perceptron.
[0126] Specifically, the text information can be first segmented to obtain each word segment corresponding to the text information, and then the embedding vectors corresponding to each word segment are determined by using the word2vec method, and the embedding vectors corresponding to each word segment are input into the attention layer of the encoder for feature extraction. During the processing of the attention layer, first, different linear transformations are performed on the embedding vectors to obtain the query vector, key vector, and value vector corresponding to each word segment, and then, according to the query vector, key vector, and value vector corresponding to each word segment, the semantic information vector corresponding to each word segment is determined. Further, after determining the semantic information vectors corresponding to each word segment, the semantic information vectors corresponding to each word segment are input into a multi-layer perceptron for intention recognition. The process of using the multi-layer perceptron for intention recognition is actually a classification process. The multi-layer perceptron will finally output the probability values of the text information belonging to different intentions, and the intention corresponding to the maximum probability value is determined as the target intention corresponding to the text information.
[0127] Further, when performing tone recognition on the text information, the tone corresponding to the text information can be determined according to the tone words, punctuation, and information input speed included in the text information. For example, if the text information contains an exclamation mark, the tone recognition result is exclamation; if the text information contains the tone word "pay attention", the tone recognition result is emphasis; if the input speed of the text information is relatively slow, the tone recognition result is statement.
[0128] Furthermore, in the process of emotion recognition for text information, the multi-head attention layer and the feed-forward neural network layer of the encoder are mainly used to extract feature vectors of the text information. Then, the extracted feature vectors are input into the softmax layer for emotion recognition to obtain the emotion recognition result corresponding to the text information. Based on this, the method includes: performing word segmentation on the text information to obtain each word corresponding to the text information; inputting the embedding vector corresponding to any one of the words into different attention subspaces in the attention layer of the encoder for feature extraction to obtain the first feature vector of the any one word under the different attention subspaces; multiplying the first feature vector of the any one word under the different attention subspaces by the weights corresponding to the different attention subspaces and summing them to obtain the output vector of the attention layer corresponding to the any one word; adding the output vector of the attention layer and the first feature vector to obtain the second feature vector corresponding to the any one word; inputting the second feature vector into the feed-forward neural network layer of the encoder for feature extraction to obtain the third feature vector corresponding to any one word; inputting the third feature vectors corresponding to each word into the softmax layer for emotion recognition to obtain the emotion recognition result corresponding to the text information.
[0129] Furthermore, the step of inputting the embedding vector corresponding to any one of the words into different attention subspaces in the attention layer of the encoder for feature extraction to obtain the first feature vector of the any one word under the different attention subspaces includes: determining the query vector, key vector, and value vector of the any one word under the different attention subspaces according to the embedding vector corresponding to the any one word; multiplying the query vector of the any one word under the different attention subspaces by the key vectors of each word under the different attention subspaces to obtain the attention scores of each word under the different attention subspaces for the any one word; multiplying the attention scores of each word under the different attention subspaces by the key vectors and summing them to obtain the first feature vector corresponding to the any one word.
[0130] Thus, the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information are obtained in the above manner. It should be noted that the multi-dimensional emotion recognition in the embodiments of the present invention is not limited to intention recognition, emotion recognition, and tone recognition, and may also include other dimensions of emotion recognition.
[0131] 103. Invoke a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result, and convert the text information into speech information.
[0132] In order to enable the recipient to perceive the current emotions, intentions, and tones of the sender, the embodiments of the present invention need to call a preset speech conversion model that matches the intention recognition result, emotion recognition result, and tone recognition result of the text information for speech conversion, so as to ensure that the converted speech information contains the current emotions, intentions, and tones of the sender. Based on this, step 103 specifically includes: obtaining an initial speech conversion model and its corresponding multiple groups of model parameters; determining, from the multiple groups of model parameters, the target model parameters commonly corresponding to the intention recognition result, the emotion recognition result, and the tone recognition result; adding the target model parameters to the initial speech conversion model to obtain a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result; and calling the matching preset speech conversion model to convert the text information into speech information. Among them, each group of emotion recognition results corresponds to a group of model parameters, and the group of emotion recognition results includes an intention recognition result, an emotion recognition result, and a tone recognition result.
[0133] For example, the initial speech conversion model does not include model parameters. The intention recognition result corresponding to the text information is "seeking something", the tone recognition result is "praying", and the emotion recognition result is "gratitude". Determine the target model parameters corresponding to the group of emotion recognition results of "seeking something", "praying", and "gratitude" from the multiple groups of model parameters, and then add the target model parameters to the initial speech conversion model to obtain a preset speech conversion model. Use this preset speech conversion model to convert the text information into speech information. Based on the converted speech information, the recipient can perceive that the intention of the sender is "seeking something", the tone is "praying", and the emotion is "gratitude".
[0134] Furthermore, in the embodiments of the present invention, the emotion recognition results, intention recognition results, and tone recognition results listed in step 102 can also be refined according to degrees. For example, the emotion recognition results include "gratitude" and "not gratitude", and "gratitude" can be refined into "especially grateful", "comparatively grateful", and "generally grateful" according to degrees. Specifically, the degrees can be refined according to the interval where the output value is located. For example, when the output value is greater than 0.5, it is gratitude; when the output value is between 0.5 and 0.6, the emotion recognition result is "generally grateful"; when the output value is between 0.6 and 0.8, the emotion recognition result is "comparatively grateful"; when the output value is above 0.8, the emotion recognition result is "especially grateful". Different degrees of emotion recognition results will also lead to different model parameters. For example, "seeking something", "praying", and "especially grateful" correspond to group A model parameters, and "seeking something", "praying", and "generally grateful" correspond to group B model parameters. Thus, the emotion recognition results can be divided into finer granularities in the above manner, improving the determination accuracy of the target model parameters.
[0135] After determining the target parameter model and obtaining the preset speech conversion model, the text information is input into the preset speech conversion model for speech conversion to obtain speech information. Specifically, the preset speech conversion model can be a Tacotron model, which mainly includes an encoder, an attention mechanism-based decoder, and a post-processing network. Specifically, first, the encoder is used to extract the feature vector corresponding to the text information, then the decoder is used to convert the feature vector into spectrogram data, and finally, the post-processing network is used to convert the spectrogram data into a waveform, so as to output speech information.
[0136] In a specific application scenario, in order to facilitate the hearing-impaired recipient to obtain the communication content, the audio sound wave in the speech information can also be converted into a bone conduction sound wave. The hearing-impaired recipient can capture the bone conduction sound wave with the help of a specific hardware device, and then can obtain the corresponding communication content. Among them, the audio sound wave is a sound wave that can be normally received by the human ear.
[0137] A text-to-speech method provided by an embodiment of the present invention, compared with the current method of communicating using text, the present invention can obtain the text information to be converted; perform multi-dimensional emotion recognition on the text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information; at the same time, call a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result, and convert the text information into speech information. Through the converted speech information, the recipient can feel the sender's mood, tone, and intention at that time. In addition, for recipients with low educational levels or visual impairments, it is also more convenient to obtain the communication content.
[0138] Further, to better illustrate the above text-to-speech process, as a refinement and extension of the above embodiment, the embodiment of the present invention provides another text-to-speech method, as Figure 2 shown, the method includes:
[0139] 201. Receive the text information input by the sender.
[0140] For the embodiments of the present invention, when the sender sends information to the receiver, the client of the sender will detect the sound decibel of the current environment. When the sound decibel exceeds a certain value, the client will automatically switch to the text input interface and prompt the sender to input text information. Based on this, the method includes: detecting the sound decibel of the environment where the sender is currently located; if the sound decibel is greater than the preset sound decibel, providing a text input interface and receiving the text information input by the sender based on the text input interface; if the sound decibel is less than or equal to the preset sound decibel, outputting a voice recording interface and receiving the voice information input by the sender based on the voice recording interface. Among them, the preset sound decibel can be set according to actual business requirements.
[0141] For example, the preset sound decibel is 70 decibels. The client of the sender can detect the sound decibel of the current environment with the help of a sensor. If it is detected that the sound decibel of the environment where the sender is currently located is 90, since it is greater than 70 decibels, the client will output a text input interface and obtain the text information input by the sender through this text input interface; if it is detected that the sound decibel of the environment where the sender is currently located is 50 decibels, since it is less than 70 decibels and will not interfere with the sender's voice recording in the current environment, the client will output a voice recording interface and receive the voice information input by the sender based on this voice recording interface.
[0142] In a specific application scenario, the client can also recommend corresponding input modes for the sender according to the historical conversation records between the sender and a specific receiver. Based on this, the method includes: obtaining the historical conversation records between the sender and the receiver; based on the historical conversation records, determining the input mode selected by the sender during the last conversation with the receiver; if the input mode selected by the sender during the last conversation with the receiver is the text mode, outputting a text input interface and receiving the text information input by the sender based on the text input interface; if the input model selected by the sender during the last conversation with the receiver is the voice mode, outputting a voice recording interface and receiving the voice information input by the sender based on the voice recording interface.
[0143] Furthermore, to ensure the security of information transmission, when the sender inputs text information on the client, the operation data of the sender on the communication device during this communication process will be collected, and the operation data during this communication process will be matched with the historical operation data. If the operation data during this communication process does not match the historical operation data, the identity information of the sender will be verified by turning on the camera device. Among them, the historical operation data includes the sender's text input habits, such as frequently misspelled words or frequently used phrases, and can also include the position where the sender habitually holds the mobile phone.
[0144] Specifically, when the sender inputs text information on the client side, the client will collect the operation data of the sender during this communication process for the communication device, including the mobile phone holding position, the misspelled text, the phrases used, etc. If the mobile phone holding position collected by the client this time is different from the sender's habitual mobile phone holding position, or the misspelled text collected this time is different from the sender's habitual misspelled text, then the identity of the sender needs to be verified. For example, the client collects a picture of the current mobile phone holder through the camera device and compares this picture with the picture of the client login user. If the pictures are the same, the identity verification of the sender passes, and the text information is converted into voice information and sent to the recipient; if the pictures are different, the identity verification of the sender fails, and the text information is intercepted, that is, the text information is not converted into voice.
[0145] Furthermore, when the identity verification of the sender fails, not only can the text information be intercepted, but the text information can also be normally converted into voice information. Before sending the voice information to the recipient, a voice prompt can be sent to the recipient first. The specific content can be "The information sender is not the mobile phone holder", so that the recipient can receive this prompt information before receiving the voice information, thus avoiding the situation of fraudulently receiving the recipient. It should be noted that the identity verification of the sender can be carried out not only before the voice information is sent, but also after the voice information is sent. For example, after the voice information is sent to the recipient, the identity of the sender is verified. If the sender fails the identity verification, a voice prompt message is sent to the recipient.
[0146] 202. Perform multi-dimensional emotion recognition on the text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information.
[0147] For the embodiments of the present invention, the specific processes of intention recognition, emotion recognition, and tone recognition of text information are generally similar to those in step 102 and will not be elaborated here.
[0148] 203. If the text information contains special characters, verify the identity of the sender and call a preset voice conversion model that matches the intention recognition result, the emotion recognition result, the tone recognition result, and the identity information of the sender to convert the text information input by the sender into voice information.
[0149] Among them, special characters include characters related to financial transactions such as money, password verification, account passwords, etc. For the embodiments of the present invention, if the text information input by the sender contains the above special characters related to financial transactions, in order to avoid fraud against the recipient and thus damage the recipient's interests, it is necessary to authenticate the sender's identity and call a preset voice model that matches both the emotion recognition result and the sender's identity information to convert the text information into voice information. Among them, the preset voice model that matches the sender's identity information is the sender's original voice model, that is, after receiving the voice information, the recipient can identify the sender's identity through the voice, thus ensuring the security of communication information and avoiding damage to the recipient's interests.
[0150] Specifically, first determine the initial voice conversion model that matches the sender's identity information and its corresponding multiple sets of model parameters, and then determine the target model parameters that match the intent recognition result, emotion recognition result, and tone recognition result from the multiple sets of model parameters, and add the target model parameters to the matching initial voice conversion model to obtain a preset voice conversion model that matches both the sender's emotion recognition result and identity information. Finally, use this preset voice conversion model to convert the text information into voice information, and the voice information played by the recipient client is the sender's original voice, so that the recipient can identify the sender's identity.
[0151] For the embodiments of the present invention, in order to more realistically reflect the scene where the sender is located when the text is input, corresponding background sounds can also be added to the converted voice information. As an optional implementation manner of adding background sounds, the method includes: responding to the received background sound addition instruction, outputting and displaying a background sound list, obtaining a selection instruction for selecting a target background sound from the background sound list, and adding the target background sound to the voice information by superimposing the sound waves in the target background sound and the sound waves in the voice information. For example, the sender selects the background sound of "Merry Christmas" from the background sound list, and the client superimposes the sound waves in the background sound of "Merry Christmas" and the sound waves in the voice information, and sends the superimposed voice information to the recipient, and the recipient can hear the background sound of "Merry Christmas" during the process of listening to the voice information.
[0152] Further, as another alternative implementation of adding background sound, the method further includes: obtaining the current conversation record between the sender and the receiver, and determining the current scene where the sender is located according to the conversation record; determining a target background sound that matches the current scene where the sender is located, and adding the target background sound to the voice message by superimposing the sound waves in the target background sound and the sound waves in the voice message. For example, the client determines that the sender is currently at the seaside by obtaining the conversation record between the sender and the receiver, so the sound of the waves is added as the background sound to the voice message, and the receiver can hear the sound of the waves during the process of listening to the voice message, so as to know that the sender is currently at the seaside.
[0153] In a specific application scenario, the sender can also embed encrypted voice in the voice message to prevent the voice message from being maliciously used. For example, when the sender sends the identity document information to the receiver, the encrypted voice "This identity information is only used for credit card business handling" can be embedded in the voice message. Thus, the voice can become evidence, which is convenient for the sender to present evidence and prevent its interests from being damaged. Based on this, the method includes: in response to the encrypted voice addition instruction triggered by the sender, obtaining the encrypted voice of the sender; adjusting the frequency of the sound waves in the encrypted voice to a specific frequency that cannot be recognized by the human ear; adding the adjusted encrypted voice to the voice message by superimposing the sound waves in the adjusted encrypted voice and the sound waves in the voice message.
[0154] Specifically, since the encrypted voice may affect the receiver's reception of the normal voice message, the frequency of the sound waves in the encrypted voice can be adjusted to a specific frequency that cannot be recognized by the human ear first, and then the adjusted encrypted voice is added to the voice message. The receiver cannot hear the adjusted encrypted voice during the process of listening to the voice message and can only hear the normal voice message, so it will not affect the receiver.
[0155] For the embodiments of the present invention, before performing voice conversion using the initial voice conversion model bound to the identity information of the sender and its corresponding multiple sets of model parameters, it is necessary to train the initial voice conversion model bound to the identity information of the sender and its corresponding multiple sets of model parameters. Based on this, the method includes: collecting the voice information of the sender reading multiple sets of scenario sentences, where the intentions, emotions, or tones corresponding to different groups of scenario sentences are different; determining the text information corresponding to the multiple sets of voice information read by the sender, and based on the multiple sets of voice information and their corresponding text information, training the initial voice conversion model bound to the identity information of the sender and its corresponding multiple sets of model parameters. Among them, one set of intention, emotion, and tone corresponds to one set of model parameters. Thus, through the voice information of multiple sets of scenario sentences read by the sender, multiple sets of model parameters of the initial voice conversion model bound to the identity information of the sender can be trained.
[0156] Further, during the process of collecting voice information, it is necessary to verify the reliability of the collected voice information. If it is found that there is a problem with the collected voice information, the collection is immediately interrupted. Based on this, the method includes: performing text conversion on the real-time voice information to obtain the real-time text information corresponding to the real-time voice information; respectively performing intention, emotion, and tone recognition on the real-time text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the real-time text information; if it is determined according to the intention recognition result corresponding to the real-time text information that the emotion of the sender is abnormal, or it is determined that the emotion recognition result corresponding to the real-time text information does not match the tone recognition result, then interrupt the collection of the real-time voice information of the sender.
[0157] For example, if the intention recognition result is blackmail, it indicates that the emotion of the sender is abnormal, and the collection of voice information is interrupted. Another example is that if the emotion recognition result is anger and the tone recognition result is statement, it indicates that the emotion recognition result does not match the tone recognition result, and the collection of voice information is interrupted.
[0158] Further, after training the initial voice conversion model bound to the identity information of the sender and its corresponding multiple sets of model parameters, the voice intonation of the voice conversion model can also be corrected, that is, the model parameters are optimized. As an optional implementation manner of model parameter optimization, the method includes: collecting the real-time voice information of the sender through a predetermined device; based on the collected real-time voice information, optimizing the multiple sets of model parameters of the initial voice conversion model bound to the identity information of the sender.
[0159] Specifically, the real-time voice information of the sender can be collected, and the collected real-time voice information can be converted into text information. The real-time voice information of the sender and its corresponding text information are used as a sample training set to optimize multiple groups of model parameters of the initial voice conversion model, so that the intonation of the converted voice information is closer to the actual intonation of the sender.
[0160] Further, as an alternative implementation of model parameter optimization, the method includes: obtaining specific text input by the sender and voice information recorded by the sender for the specific text; optimizing multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the specific text and its corresponding voice information. Specifically, the sender can input some uncommon text or text with specific pronunciations on the client side and read them aloud. The client collects the voice information of the specific text read by the sender and uses it to optimize multiple groups of model parameters of the initial voice conversion model, thereby improving the voice conversion effect of the model.
[0161] In a specific application scenario, the trained initial voice conversion model and multiple groups of model parameters can also be used to convert the trial reading sentence input by the sender into corresponding voice information for playback. The sender can revise the voice information, and at the same time, optimize multiple groups of model parameters of the initial voice conversion model based on the revised voice. Based on this, the method includes: obtaining the trial reading sentence input by the sender; using multiple groups of model parameters of the trained initial voice conversion model bound to the identity information of the sender to convert the trial reading sentence into voice information and play it to the sender, and outputting a selection correction interface corresponding to the trial reading sentence; obtaining the revised voice corresponding to the target character in the trial reading sentence selected by the sender based on the selection correction interface; optimizing multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the revised voice. Among them, the target character can be any one, two, or more characters in the trial reading sentence.
[0162] For example, if the trial reading sentence is "Today we go to a seaside", after the sender hears the voice information corresponding to the trial reading sentence and is dissatisfied with the pronunciation of "seaside", the sender can select the target character "seaside" in the selection correction interface corresponding to the trial reading sentence and at the same time correct the intonation of "seaside" to obtain the revised voice corresponding to the trial reading sentence. Optimize multiple groups of model parameters corresponding to the initial voice conversion model based on the revised voice, so that the voice conversion effect of the optimized model parameters is closer to the real voice of the sender.
[0163] In a specific application scenario, a preset dialect speech conversion model bound to the identity information of the sender can also be pre-constructed. When the text information input by the sender is in dialect, the preset dialect speech conversion model can be called to convert the text information into dialect speech information. Based on this, the method includes: when the text information is in dialect, obtaining the output speech type selected by the sender; if the output speech type is dialect speech, calling the preset dialect speech conversion model bound to the identity information of the sender to convert the dialect input by the sender into dialect speech information; if the output speech type is standard speech, using a preset dialect word library to convert the dialect into standard text information, obtaining the standard text information corresponding to the dialect; performing multi-dimensional sentiment recognition on the standard text information to obtain the sentiment recognition result of the standard text information in multiple dimensions; calling a preset speech conversion model that matches the sentiment recognition result in multiple dimensions and the identity information of the sender to convert the standard text information into standard speech information. Among them, the preset dialect word library includes various dialect phrases and their corresponding standard phrases.
[0164] Specifically, when the text information input by the sender is in dialect, the sender can select the output speech type. When the output speech type is dialect speech, the preset dialect speech conversion model bound to the identity information of the sender is called to convert the dialect into dialect speech information. Among them, the preset dialect speech conversion model is trained by collecting the dialect speech information read by the sender and its corresponding dialect text information. Further, when the output speech type is standard speech, the input dialect can be converted into standard text information using the preset dialect word library, and then the matching preset speech conversion model is called for speech conversion to obtain standard speech information. The specific conversion process of the standard speech is exactly the same as the above process in step 203 and will not be elaborated here.
[0165] Another text-to-speech method provided by an embodiment of the present invention, compared with the current way of communicating using text, the present invention can obtain the text information to be converted; and perform multi-dimensional sentiment recognition on the text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information; at the same time, call a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result to convert the text information into speech information. Through the converted speech information, the receiver can feel the sender's mood, tone, and intention at that time. In addition, for receivers with low educational levels or visual impairments, it is also more convenient to obtain the communication content.
[0166] Further, as Figure 1 a specific implementation of Figure 3As shown, the device includes: an acquisition unit 31, an identification unit 32, and a conversion unit 33.
[0167] The acquisition unit 31 can be used to acquire the text information to be converted.
[0168] The identification unit 32 can be used to perform multi-dimensional sentiment recognition on the text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information.
[0169] The conversion unit 33 can be used to call a preset speech conversion model that matches the intention recognition result, emotion recognition result, and tone recognition result to convert the text information into speech information.
[0170] In a specific application scenario, the conversion unit 33, as Figure 4 shown, includes: a first acquisition module 331, a first determination module 332, an addition module 333, and a conversion module 334.
[0171] The first acquisition module 331 can be used to acquire the initial speech conversion model and its corresponding multiple groups of model parameters.
[0172] The first determination module 332 can be used to determine the target model parameters jointly corresponding to the intention recognition result, emotion recognition result, and tone recognition result from the multiple groups of model parameters, where each group of sentiment recognition results corresponds to a group of model parameters, and the group of sentiment recognition results includes the intention recognition result, emotion recognition result, and tone recognition result.
[0173] The addition module 333 can be used to add the target model parameters to the initial speech conversion model to obtain a preset speech conversion model that matches the intention recognition result, emotion recognition result, and tone recognition result.
[0174] The conversion module 334 can be used to call the matching preset speech conversion model to convert the text information into speech information.
[0175] In a specific application scenario, the acquisition unit 31 can specifically be used to receive the text information input by the sender.
[0176] The conversion unit 33 can specifically be used to, if the text information contains special characters, verify the identity of the sender, and call a preset speech conversion model that matches the intention recognition result, emotion recognition result, tone recognition result, and identity information of the sender to convert the text information input by the sender into speech information.
[0177] In a specific application scenario, the acquisition unit 31, asFigure 4 As shown in the figure, it includes: a detection module 311, a receiving module 312, a second acquisition module 313, and a second determination module 314.
[0178] The detection module 311 can be used to detect the sound decibel of the current environment where the sender is located.
[0179] The receiving module 312 can be used to provide a text input interface if the sound decibel is greater than a preset sound decibel, and receive the text information input by the sender based on the text input interface.
[0180] The second acquisition module 313 can be used to acquire the historical conversation record between the sender and the receiver.
[0181] The second determination module 314 can be used to determine the input mode selected by the sender during the last conversation with the receiver based on the historical conversation record.
[0182] The receiving module 312 can also be used to output a text input interface if the input mode selected by the sender during the last conversation with the receiver is the text mode, and receive the text information input by the sender based on the text input interface.
[0183] In a specific application scenario, the device further includes: a collection unit 34 and a training unit 35.
[0184] The collection unit 34 can be used to collect the voice information of the sender reading multiple groups of scenario sentences, where the intentions, emotions, or tones corresponding to different groups of scenario sentences are different.
[0185] The training unit 35 can be used to determine the text information corresponding to the multiple groups of voice information read by the sender, and train the initial voice conversion model bound to the identity information of the sender and its corresponding multiple groups of model parameters based on the multiple groups of voice information and their corresponding text information.
[0186] In a specific application scenario, the device further includes: an optimization unit 36.
[0187] The collection unit 34 can also be used to collect the real-time voice information of the sender through a predetermined device.
[0188] The optimization unit 36 can be used to optimize the multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the collected real-time voice information.
[0189] The acquisition unit 31 can also be used to acquire the specific text input by the sender, and the voice information input by the sender for the specific text.
[0190] The optimization unit 36 can also be used to optimize multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the specific text and its corresponding voice information.
[0191] In a specific application scenario, the device further includes: an interruption unit 37.
[0192] The conversion unit 33 can also be used to perform text conversion on the real-time voice information to obtain the real-time text information corresponding to the real-time voice information.
[0193] The recognition unit 32 can also be used to respectively perform intention, emotion, and tone recognition on the real-time text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the real-time text information.
[0194] The interruption unit 37 can be used to interrupt the acquisition of the real-time voice information of the sender if it is determined according to the intention recognition result corresponding to the real-time text information that the emotion of the sender is abnormal, or if it is determined that the emotion recognition result corresponding to the real-time text information does not match the tone recognition result.
[0195] In a specific application scenario, the acquisition unit 31 can also be used to acquire the trial reading statement input by the sender.
[0196] The conversion unit 33 can also be used to use multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender that have been trained to convert the trial reading statement into voice information and play it to the sender, and output a selection correction interface corresponding to the trial reading statement.
[0197] The acquisition unit 31 can also be used to acquire the revised voice corresponding to the target character in the trial reading statement selected by the sender based on the selection correction interface.
[0198] The optimization unit 36 can also be used to optimize multiple groups of model parameters of the initial voice conversion model bound to the identity information of the sender based on the revised voice.
[0199] In a specific application scenario, the acquisition unit 31 can also be used to acquire the output voice type selected by the sender when the text information is in a dialect.
[0200] The conversion unit 33 can also be used to call a preset dialect voice conversion model bound to the identity information of the sender to convert the dialect input by the sender into dialect voice information if the output voice type is dialect voice.
[0201] The conversion unit 33 may also be configured to, if the output voice type is standard voice, use a preset dialect word library to convert the dialect into Mandarin to obtain the standard text information corresponding to the dialect.
[0202] The recognition unit 32 may also be configured to perform multi-dimensional emotion recognition on the standard text information to obtain the emotion recognition result of the standard text information in multiple dimensions.
[0203] The conversion unit 33 may also be configured to call a preset voice conversion model that matches the emotion recognition result in multiple dimensions and the identity information of the sender, and convert the standard text information into standard voice information.
[0204] In a specific application scenario, the conversion unit 33 may also be configured to respond to a received sound wave conversion instruction, and convert the audio sound wave in the voice information into a bone conduction sound wave.
[0205] In a specific application scenario, the device further includes: a superimposing unit 38.
[0206] The superimposing unit 38 may be configured to respond to a received background sound addition instruction, output and display a background sound list, obtain a selection instruction for selecting a target background sound from the background sound list, and add the target background sound to the voice information by superimposing the sound wave in the target background sound and the sound wave in the voice information; or obtain the current conversation record between the sender and the receiver, and determine the scene where the sender is currently located according to the conversation record; determine a target background sound that matches the scene where the sender is currently located, and add the target background sound to the voice information by superimposing the sound wave in the target background sound and the sound wave in the voice information.
[0207] In a specific application scenario, the acquisition unit 31 may also be configured to respond to an encrypted voice addition instruction triggered by the sender, and acquire the encrypted voice of the sender.
[0208] The superimposing unit 38 may also be configured to adjust the frequency of the sound wave in the encrypted voice to a specific frequency that cannot be recognized by the human ear; add the adjusted encrypted voice to the voice information by superimposing the sound wave in the adjusted encrypted voice and the sound wave in the voice information.
[0209] In a specific application scenario, the device further includes: a matching unit 39.
[0210] The acquisition unit 34 may also be configured to acquire the operation data of the sender on the communication device during the current communication process.
[0211] The matching unit 39 can be used to match the operation data in the current communication process with the historical operation data; if the operation data in the current communication process does not match the historical operation data, the identity information of the sender is verified by turning on the camera device.
[0212] It should be noted that for other corresponding descriptions of each functional module involved in the text-to-speech device provided in the embodiments of the present invention, reference can be made to Figure 1 the corresponding description of the method shown, which will not be elaborated here.
[0213] Based on the above as Figure 1 shown in the method, correspondingly, the embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the following steps are implemented: obtaining the text information to be converted; performing multi-dimensional sentiment recognition on the text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information; calling a preset voice conversion model that matches the intention recognition result, emotion recognition result, and tone recognition result, and converting the text information into voice information.
[0214] Based on the above as Figure 1 shown in the method and as Figure 3 shown in the embodiment of the device, the embodiments of the present invention also provide a physical structure diagram of a computer device, as Figure 5 shown, the computer device includes: a processor 41, a memory 42, and a computer program stored on the memory 42 and executable on the processor, wherein both the memory 42 and the processor 41 are arranged on a bus 43. When the processor 41 executes the program, the following steps are implemented: obtaining the text information to be converted; performing multi-dimensional sentiment recognition on the text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information; calling a preset voice conversion model that matches the intention recognition result, emotion recognition result, and tone recognition result, and converting the text information into voice information.
[0215] Through the technical solution of the present invention, it is possible to obtain the text information to be converted; perform multi-dimensional sentiment recognition on the text information to obtain the intention recognition result, emotion recognition result, and tone recognition result corresponding to the text information; at the same time, call a preset voice conversion model that matches the intention recognition result, emotion recognition result, and tone recognition result, and convert the text information into voice information. Through the converted voice information, the receiving party can feel the emotions, tones, and intentions of the sending party at that time. In addition, for receiving parties with low educational levels or visual impairments, it is also more convenient to obtain the communication content.
[0216] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to be implemented. In this way, the present invention is not limited to any specific combination of hardware and software.
[0217] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for text-to-speech conversion, characterized in that, it includes: Obtaining the text information to be converted; Performing multi-dimensional emotion recognition on the text information to obtain an intention recognition result, an emotion recognition result, and a tone recognition result corresponding to the text information; Invoking a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result to convert the text information into speech information; Among them, the step of invoking a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result to convert the text information into speech information includes: Obtaining an initial speech conversion model and its corresponding multiple groups of model parameters; Determining the target model parameters jointly corresponding to the intention recognition result, the emotion recognition result, and the tone recognition result from the multiple groups of model parameters, where each group of emotion recognition results corresponds to a group of model parameters, and the group of emotion recognition results includes an intention recognition result, an emotion recognition result, and a tone recognition result; Adding the target model parameters to the initial speech conversion model to obtain a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result; Invoking the matching preset speech conversion model to convert the text information into speech information; During the process of collecting speech information for training the preset speech conversion model, performing intention, emotion, and tone recognition on the real-time text information corresponding to the real-time speech information respectively to obtain an intention recognition result, an emotion recognition result, and a tone recognition result corresponding to the real-time text information; if it is determined that the emotion of the sender is abnormal according to the intention recognition result corresponding to the real-time text information, or it is determined that the emotion recognition result corresponding to the real-time text information does not match the tone recognition result, then the collection of the real-time speech information is interrupted.
2. The method according to claim 1, characterized in that, the step of obtaining the text information to be converted includes: Receiving the text information input by the sender; the step of invoking a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result to convert the text information into speech information includes: If the text information contains special characters, verifying the identity of the sender and invoking a preset speech conversion model that matches the intention recognition result, the emotion recognition result, the tone recognition result, and the identity information of the sender to convert the text information input by the sender into speech information.
3. The method according to claim 1, characterized in that, after the step of invoking a preset speech conversion model that matches the intention recognition result, the emotion recognition result, and the tone recognition result to convert the text information into speech information, the method further includes: Responding to the received sound wave conversion instruction, converting the audio sound wave in the speech information into a bone conduction sound wave.
4. The method according to claim 2, characterized in that, After invoking a preset speech conversion model that matches the intent recognition result, the emotion recognition result, the tone recognition result, and the identity information of the sender, and converting the text information input by the sender into speech information, the method further includes: In response to a received background sound addition instruction, output and display a list of background sounds, obtain a selection instruction for selecting a target background sound from the list of background sounds, and add the target background sound to the speech information by superimposing the sound waves in the target background sound with the sound waves in the speech information; or Obtain the current conversation record between the sender and the receiver, and determine the current scene where the sender is located according to the conversation record; determine a target background sound that matches the current scene where the sender is located, and add the target background sound to the speech information by superimposing the sound waves in the target background sound with the sound waves in the speech information.
5. The method according to claim 2, wherein, After invoking a preset speech conversion model that matches the intent recognition result, the emotion recognition result, the tone recognition result, and the identity information of the sender, and converting the text information input by the sender into speech information, the method further includes: In response to an encrypted speech addition instruction triggered by the sender, obtain the encrypted speech of the sender; Adjust the frequency of the sound waves in the encrypted speech to a specific frequency that cannot be recognized by the human ear; Add the adjusted encrypted speech to the speech information by superimposing the sound waves in the adjusted encrypted speech with the sound waves in the speech information.
6. The method according to claim 2, wherein, The method further includes: Collect the operation data of the sender on the communication device during this communication process; Match the operation data during this communication process with the historical operation data; If the operation data during this communication process does not match the historical operation data, verify the identity information of the sender by turning on the camera device.
7. A text-to-speech device, wherein, comprising: An acquisition unit for acquiring text information to be converted; An identification unit for performing multi-dimensional emotion recognition on the text information to obtain an intent recognition result, an emotion recognition result, and a tone recognition result corresponding to the text information; A conversion unit for invoking a preset speech conversion model that matches the intent recognition result, the emotion recognition result, and the tone recognition result, and converting the text information into speech information; The conversion unit is specifically configured to obtain an initial speech conversion model and its corresponding multiple groups of model parameters; determine target model parameters corresponding to the intent recognition result, the emotion recognition result, and the tone recognition result from the multiple groups of model parameters, where each group of emotion recognition results corresponds to a group of model parameters, and the group of emotion recognition results includes an intent recognition result, an emotion recognition result, and a tone recognition result; add the target model parameters to the initial speech conversion model to obtain a preset speech conversion model that matches the intent recognition result, the emotion recognition result, and the tone recognition result; and call the matching preset speech conversion model to convert the text information into speech information. The acquisition unit is configured to, during the process of acquiring speech information for training the preset speech conversion model, perform intent, emotion, and tone recognition on the real-time text information corresponding to the real-time speech information respectively, to obtain the intent recognition result, the emotion recognition result, and the tone recognition result corresponding to the real-time text information; and if it is determined according to the intent recognition result corresponding to the real-time text information that the emotion of the sender is abnormal, or it is determined that the emotion recognition result corresponding to the real-time text information does not match the tone recognition result, then interrupt the acquisition of the real-time speech information.
8. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Speech synthesis method and device with mood, computing equipment and storage medium
CN111161703A
Voice synthesis method and voice synthesis device
CN111192568A
Audio human interactive proof based on text-to-speech and semantics
US20130218566A1
Voice synthetic server and terminal
WO2020141643A1