A voice processing method and apparatus
By calling the trained speech synthesis model on the server side, the synthetic speech data corresponding to the target text is directly generated, which solves the problem of complex interaction process in the recording stage in the prior art, and simplifies user operations and good speech synthesis effects are achieved.
Patent Information
- Application Number
- CN202110813868.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-07-19
AI Technical Summary
The existing personalized voice synthesis system has a complex interactive process during the recording stage, which affects the user experience and leads to user churn.
By calling the trained speech synthesis model on the server side, the synthetic speech data corresponding to the target text is directly generated, simplifying the user's operation process on the front-end.
While simplifying user operations in voice synthesis services, it ensures good voice synthesis effects and improves user experience.
Smart Images

Figure CN113823258B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a voice processing method and apparatus. Background Art
[0002] With the continuous development and progress of science and technology, personalized speech synthesis technology has become one of the research hotspots in the field of speech synthesis. It refers to a technology that can construct a speech synthesis system with the timbre characteristics of a user by collecting a small amount of the user's speech data.
[0003] In the currently common personalized speech synthesis systems in the industry, in order to ensure the quality of speech synthesis, users are required to provide high-quality recordings. For example, for environmental noise detection, recordings must be made in a quiet environment; the recorded data must be exactly the same as the reference text that prompts the user's recording content, and if there are errors, repeated recordings are required; in addition, gender information, etc. also need to be provided. The above series of requirements will result in a complex interaction process during the recording stage at the front end, affecting the user experience. The complex recording operations at the front end often lead to serious user loss. Therefore, how to simplify the operations of users in speech synthesis services and ensure good speech synthesis effects has become an urgent problem to be solved. Summary of the Invention
[0004] Embodiments of this application provide a voice processing method and apparatus, which can simplify the operations of users during speech synthesis and provide good speech synthesis effects.
[0005] Embodiments of this application provide a voice processing method, and the method includes:
[0006] Receiving a target text sent by a user terminal.
[0007] Invoking a speech synthesis model to process the target text to generate synthetic speech data corresponding to the target text, where the speech synthesis model is trained based on the user's speech data, text feature information of the speech data, and identity feature information of the user.
[0008] Sending the synthetic speech data corresponding to the target text to the user terminal.
[0009] Embodiments of this application provide a voice processing method, and the method includes:
[0010] In response to a user's application start instruction, displaying a voice recording interface of a speech synthesis application, where the voice recording interface includes one or more of a recording progress indication area, a reference text display area, and a recording control operation area.
[0011] Obtain the voice data input by the user through the voice input interface, where the voice data includes at least one voice segment input by the user based on at least one reference text.
[0012] Send the voice data to the server so that the server trains a voice synthesis model based on the voice data, the text feature information of the voice data, and the identity feature information of the user.
[0013] An embodiment of the present application provides a voice processing device, and the device includes:
[0014] A receiving module, configured to receive a target text sent by a user terminal.
[0015] A processing module, configured to call a voice synthesis model to process the target text and generate synthesized voice data corresponding to the target text, where the voice synthesis model is trained based on the user's voice data, the text feature information of the voice data, and the identity feature information of the user.
[0016] A sending module, configured to send the synthesized voice data corresponding to the target text to the user terminal.
[0017] An embodiment of the present application provides a voice processing device, and the device includes:
[0018] A display module, configured to display a voice input interface of a voice synthesis application in response to a user's application start instruction, where the voice input interface includes one or more of a recording progress indication area, a reference text display area, and a recording control operation area.
[0019] An obtaining module, configured to obtain the voice data input by the user through the voice input interface, where the voice data includes at least one voice segment input by the user based on at least one reference text.
[0020] A sending module, configured to send the voice data to the server so that the server trains a voice synthesis model based on the voice data, the text feature information of the voice data, and the identity feature information of the user.
[0021] An embodiment of the present application provides a server, and the server includes a processor, a network interface, and a storage device. The processor, the network interface, and the storage device are interconnected. Among them, the network interface is controlled by the processor to send and receive data, the storage device is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the voice processing method described in the first aspect.
[0022] An embodiment of the present application provides a user terminal, which includes a processor, a storage device, a display device, and a communication device. The processor, the storage device, the display device, and the communication device are interconnected. The storage device is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the voice processing method described in the second aspect.
[0023] An embodiment of the present application further provides a computer-readable storage medium. The computer storage medium stores a computer program, and the computer program includes program instructions. The program instructions are executed by a processor to execute the voice processing method described in the first aspect or the second aspect.
[0024] The present application discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the voice processing method described in the first aspect or the second aspect above.
[0025] In an embodiment of the present application, the server can receive the target text sent by the user terminal, call the speech synthesis model to process the target text, and generate the synthesized speech data corresponding to the target text. Among them, the speech synthesis model is trained according to the user's speech data, the text feature information of the speech data, and the user's identity feature information. Then, the server sends the synthesized speech data corresponding to the target text to the user terminal. It can be seen that the user can directly submit the text that needs to be synthesized into speech to the background server, and the background server can quickly generate the corresponding synthesized speech data by using the corresponding speech synthesis model, which simplifies the user operation in the speech synthesis service and also ensures a good speech synthesis effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0028] Figure 1 It is a schematic diagram of the architecture of a voice processing system provided by an embodiment of the present application;
[0029] Figure 2 It is a schematic flowchart of a voice processing method provided by an embodiment of the present application;
[0030] Figure 3 It is a schematic flowchart of another voice processing method provided by an embodiment of the present application;
[0031] Figure 4a It is a schematic flowchart of a noise reduction process provided by an embodiment of the present application;
[0032] Figure 4b It is a schematic flowchart of a voice recognition process provided by an embodiment of the present application;
[0033] Figure 4c It is a schematic flowchart of an identity feature information recognition provided by an embodiment of the present application;
[0034] Figure 4d It is a schematic overall implementation flowchart of a voice processing provided by an embodiment of the present application;
[0035] Figure 5 It is a schematic flowchart of yet another voice processing method provided by an embodiment of the present application;
[0036] Figure 6a It is a schematic diagram of a voice input interface provided by an embodiment of the present application;
[0037] Figure 6b It is a schematic diagram of another voice input interface provided by an embodiment of the present application;
[0038] Figure 7 It is a schematic diagram of the structure of a voice processing device provided by an embodiment of the present application;
[0039] Figure 8 It is a schematic diagram of the structure of another voice processing device provided by an embodiment of the present application;
[0040] Figure 9 It is a schematic diagram of the structure of a server provided by an embodiment of the present application;
[0041] Figure 10 It is a schematic diagram of the structure of a user terminal provided by an embodiment of the present application. Detailed implementation manners
[0042] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0043] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0044] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0045] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and speech has become one of the most promising human-computer interaction methods in the future.
[0046] The solutions provided in the embodiments of the present application involve technologies such as speech recognition and text-to-speech in artificial intelligence, and will be specifically described through the following embodiments:
[0047] Please refer to Figure 1 , which is a schematic diagram of the architecture of a speech processing system provided in the embodiments of the present application. The data processing system includes a server 10 and one or more user terminals 20. Among them:
[0048] Server 10 can provide a speech synthesis service for converting the text content submitted by user terminal 20 into synthesized speech data, and the synthesized speech data matches the gender, voice color, etc. of the user corresponding to user terminal 20. For example, a typical application scenario of the speech synthesis service can be parent-child reading aloud, that is, telling stories to children in the voices of parents. Parents submit the text content of the story through user terminal 20, and server 10 can use the voice characteristics of the parents to generate synthesized speech data corresponding to the text content of the story. User terminal 20 can play the synthesized speech data to realize telling stories to children in the voices of parents.
[0049] In some feasible implementation manners, server 10 can provide a personalized speech synthesis service to the corresponding user through training a speech synthesis model. A speech synthesis application can be installed on user terminal 20. Through the speech synthesis application, the user can submit recording data and the text content for which speech needs to be synthesized. The user can directly submit the recording data to server 10 through the application interface of the speech synthesis application. Server 10 can extract the corresponding text feature information and the user's identity feature information according to the user's recording data. The identity feature information can specifically be gender. Using the text feature information and the user's identity feature information, a speech synthesis model that conforms to the features such as the gender and voice color of the user can be trained, and a personalized speech synthesis service can be provided to the user through this speech synthesis model. In the speech synthesis service provided by this application, the user does not need to set gender information at the front end (i.e., user terminal 20), nor does the front end need to perform verification on the consistency of the recording data and the text content, enabling the user to quickly complete the input of speech data with simple operations. The background (i.e., server 10) can determine the gender of the user by automatically analyzing the recording data, and train a speech synthesis model that matches the user in combination with the extracted text feature information, greatly simplifying the interaction process at the front end and reducing the operations of the user in the recording stage, thereby improving the usage experience of the personalized speech synthesis system. At the same time, in the background training stage, corresponding technical means such as gender analysis are introduced to repair the recording quality and obtain gender auxiliary information, ensuring a good speech synthesis effect while simplifying the front-end user interaction process.
[0050] Among them, server 10 can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. User terminal 20 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, an in-vehicle intelligent terminal, etc., but is not limited thereto. User terminal 20 and server 10 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0051] The implementation details of the technical solution of the embodiments of the present application are elaborated in detail as follows:
[0052] Please refer to Figure 2 , which is a schematic flowchart of a speech processing method provided by the speech processing system shown in Figure 1 . The speech processing method of the embodiments of the present application is mainly described from the server side. The speech processing method includes the following steps:
[0053] 201. Receive the target text sent by the user terminal.
[0054] Among them, the target text is the data information to be synthesized into speech, and specifically can be any text selected by the user from multiple texts. In the parent-child reading aloud scenario, the target text can be the text content of the story selected by the father or mother.
[0055] Specifically, the user terminal can be installed with a speech synthesis application. Through the application interface of the speech synthesis application, the user can submit the target text that needs to be synthesized into speech, and can also input the user's speech data. The user terminal sends the target text to the server.
[0056] 202. Call the speech synthesis model to process the target text, and generate the synthesized speech data corresponding to the target text. Among them, the speech synthesis model is trained according to the user's speech data, the text feature information of the speech data, and the user's identity feature information.
[0057] Among them, a personalized speech synthesis model that conforms to the characteristics such as the gender and voice color of different users can be trained for different users. Through this speech synthesis model, personalized speech synthesis services can be provided to the corresponding users. Among them, the speech synthesis model can be trained according to the user's speech data, the text feature information in the speech data, and the user's identity feature information. The user's identity feature information can refer to the user's gender, and this identity feature information can be obtained by the server through analyzing and processing the user's speech data, without the user setting it on the user terminal.
[0058] Specifically, after the server receives the target text submitted by the user terminal, it can call the speech synthesis model corresponding to the user to process the target text, and convert the content of the target text into speech data, that is, the synthesized speech data corresponding to the target text.
[0059] 203. Send the synthesized speech data corresponding to the target text to the user terminal.
[0060] Specifically, after obtaining the synthesized speech data corresponding to the target text, the server can send the synthesized speech data of the target text to the user terminal. After receiving it, the user terminal can respond to the user's play instruction and play the synthesized speech data.
[0061] In the embodiments of the present application, the server can receive the target text sent by the user terminal, call the speech synthesis model to process the target text, and generate the synthesized speech data corresponding to the target text. The speech synthesis model is trained based on the user's speech data, the text feature information of the speech data, and the user's identity feature information. Then, the server sends the synthesized speech data corresponding to the target text to the user terminal. It can be seen that when the user directly submits the text that needs to be synthesized into speech to the background, the background can quickly generate the corresponding synthesized speech data using the corresponding speech synthesis model, which simplifies the operation of the user in the speech synthesis service and at the same time ensures a good speech synthesis effect.
[0062] Please refer to Figure 3 , which is a schematic flowchart of another speech processing method provided by the speech processing system based on Figure 1 shown in the present application embodiment. The speech processing method of the present application embodiment is mainly described from the perspective of the server. The speech processing method includes the following steps:
[0063] 301. Receive the speech data sent by the user terminal, where the speech data includes at least one speech segment input by the user based on at least one reference text.
[0064] Among them, in order to train the user's personalized speech synthesis model, a certain amount of speech data needs to be provided by the user. The user can submit the speech data through the application interface of the speech synthesis application installed on the user terminal. When the user enters the speech data, the user terminal can output at least one reference text through the application interface of the speech synthesis application. The user reads the content of the reference text to enter the corresponding speech data. The speech data of each reference text entered by the user can be regarded as a speech segment. After the user finishes entering, the user terminal can obtain the speech data including at least one speech segment and send the speech data to the server.
[0065] 302. Obtain the user's identity feature information and the text feature information of the speech data according to the speech data.
[0066] Specifically, in order to train a speech synthesis model, it is necessary to determine the user's identity feature information (such as gender) and the text feature information corresponding to the speech data. The identity feature information can be used as auxiliary information to set the hyperparameters of the model. After receiving the speech data input by the user, the server can parse and process the speech data to obtain the user's identity feature information and the text feature information of the speech data. The text feature information can be understood as the text sequence contained in the speech data. The text sequence can be, for example, a phoneme sequence. For Chinese, the phoneme sequence is the initial and final consonant sequence. For example, the phoneme sequence of "ni hao" is "n i h ao".
[0067] In some feasible embodiments, since the user terminal does not need to detect environmental noise and does not require the user to record speech data in a quiet environment, but the user directly inputs speech data after opening the speech synthesis application, the user terminal directly submits the speech data input by the user to the server. The server can perform noise reduction processing on the speech data, thereby saving the steps of the user terminal for detecting environmental noise. To ensure the accuracy of extracting the identity feature information and the text feature information, after receiving the user's speech data, the server can perform noise reduction processing on the speech data to obtain the noise-reduced speech data, and perform speech recognition processing on the noise-reduced speech data to obtain the text feature information of the speech data, and determine the user's identity feature information according to the noise-reduced speech data. The identity feature information can include gender information.
[0068] In some feasible embodiments, the specific implementation method for the server to perform noise reduction processing on the speech data can be referred to Figure 4a , which mainly includes: performing Fourier transform processing FFT on the speech data x(n) to obtain the spectrum Y(w) of the speech data and the noise spectrum D(w). The noise spectrum D(w) can be estimated from the first n frames of silent data of the speech data; determining the target spectrum according to the spectrum Y(w) of the speech data and the noise spectrum D(w). For example, subtracting the noise spectrum D(w) from the spectrum Y(w) to obtain the target spectrum, obtaining the amplitude spectrum of the target spectrum, and performing inverse Fourier transform processing IFFT on the phase information of the amplitude spectrum and the spectrum of the speech data to obtain the noise-reduced speech data y(n). Among them, Figure 4a the noise reduction algorithm shown is the spectral subtraction method. The spectral subtraction method is a noise reduction algorithm based on digital signal processing. The Wiener filtering method in the noise reduction algorithms based on digital signal processing can also be used. Of course, a speech noise reduction algorithm based on machine learning can also be used. The embodiments of the present application do not make limitations.
[0069] In some feasible embodiments, the specific implementation method for the server to extract the text feature information of the speech data can be referred to Figure 4b, mainly including: extracting the acoustic feature information (such as MFCC spectrum features) of the denoised speech data, calling the speech recognition model in the model library to decode the acoustic feature information, obtaining the text sequence of the speech data, and using the text sequence as the text feature information of the speech data.
[0070] In some feasible implementations, since the user terminal directly sends the voice data to the server after obtaining the voice segment input by the user, the user terminal does not need to check the consistency between the voice segment and the corresponding reference text, such as whether the voice segment and the corresponding reference text are completely matched, the server needs to check the consistency between the voice input by the user and the reference text, and can delete the voice segment with large differences, thereby realizing the error correction function of the voice. Among them, the text feature information of the voice data includes the text sequence of each voice segment in at least one voice segment, and the server can obtain the matching degree between the text sequence of each voice segment and the corresponding reference text. If the matching degree corresponding to the target voice segment is less than or equal to the preset matching degree threshold, the preset matching degree threshold can be, for example, 90%, then the text sequence of the target voice segment is deleted from the text feature information of the voice data. It can be seen that in this application, the front end (i.e., the user terminal) does not need to ensure that the content of the recording and the reference text is completely consistent, nor does it need to be repeatedly detected at the front end, but directly submits the voice data entered by the user to the server, and only one recording operation is required for each reference text, thereby improving the recording efficiency of the user.
[0071] In some feasible implementations, the speech recognition model specifically includes an acoustic model and a language model. The server calls the speech recognition model to decode the acoustic feature information. The specific implementation method of obtaining the text sequence of the speech data may include: calling the acoustic model to determine the matching probability between the acoustic feature information and the corresponding phonemes or characters, calling the language model to determine the occurrence probability of each text sequence, and determining the text sequence of the speech data based on the matching probability and the occurrence probability.
[0072] Among them, the acoustic model and language model required by the speech recognition system can be pre-trained using the speech data of a large number of speakers, wherein the acoustic model learns the probability of acoustic features corresponding to phonemes or words, and the language model learns the probability of a certain word sequence occurring.
[0073] In some feasible implementations, since the user terminal does not need to set gender information, which saves the user's operation process in the recording process, the server needs to identify the user's gender based on the user's voice data and use it in the training of the speech synthesis model. The specific implementation method of the server determining the user's identity feature information based on the noise-reduced voice data can be found in Figure 4c, mainly including: in the application stage or the test stage of the identity discrimination model, the server can extract the acoustic feature information of the denoised speech data, and use the first identity discrimination model (such as the male gender model) and the second identity discrimination model (such as the female gender model) to score the acoustic feature information respectively to obtain the scoring results. According to the scoring results, the identity feature information of the user, such as gender, can be determined.
[0074] In some feasible implementation manners, the scoring results include the first score corresponding to the first identity discrimination model and the second score corresponding to the second identity discrimination model. The server can determine the highest score among the first score and the second score, and determine the target identity discrimination model corresponding to the highest score from the first identity discrimination model and the second identity discrimination model. The target identity discrimination model can be the first identity discrimination model or the second identity discrimination model. Then, the identity feature information of the user can be determined according to the gender type (i.e., male or female) corresponding to the target identity discrimination model.
[0075] Among them, when the user has multiple recordings, that is, the speech data of the user includes multiple speech segments, the first identity discrimination model and the second identity discrimination model can be used to score the acoustic feature information of each speech segment respectively. According to the scoring results of each speech segment, the gender with the most occurrences is taken as the final determination result of the user.
[0076] In some feasible implementation manners, in the training stage of the identity discrimination model, the server can obtain a training sample set. The training sample set includes the speech data of a large number of users of different gender types. Specifically, the number of male users and the number of female users can be equal. First, the universal background model (i.e., the UBM model) can be trained using the acoustic feature information of the speech data of each user in the training sample set. Then, the Gaussian mixture model (GMM model) can be obtained by adaptively training the universal background model using data of different gender types. For example, the first identity discrimination model (such as the male gender model) can be obtained by training the universal background model using the acoustic feature information of the speech data of the users of the first gender type (such as male users) in the training sample set. The first identity discrimination model is used to score the possibility that a user is male. The second identity discrimination model (such as the female gender model) can be obtained by training the universal background model using the acoustic feature information of the speech data of the users of the second gender type (such as female users) in the training sample set. The second identity discrimination model is used to score the possibility that a user is female. Thus, a model that can accurately identify the gender of the user can be trained.
[0077] 303. Train the speech synthesis model corresponding to the user by using the text feature information and the identity feature information.
[0078] Specifically, after obtaining the text feature information of the user's voice data and the user's identity feature information, the server can extract the spectral features of the voice data (including fundamental frequency and spectrum), set the hyperparameters of the voice synthesis model using the user's identity feature information, and use the spectral features as supervision information to train the voice synthesis model. The loss function can adopt the Mean Squared Error (MSE) loss. Input the text feature information of the voice data into the voice synthesis model to obtain the predicted synthesized voice data. Calculate the loss value of the loss function based on the spectral features of the predicted synthesized voice data and the spectral features of the extracted voice data. After training for a set number of iterations, the training can be stopped.
[0079] Among them, the server uses the text feature information of the voice data, the user's identity feature information, and the spectral features of the voice data to train and obtain the voice synthesis model, which specifically means: using the text feature information of the voice data, the user's identity feature information, and the spectral features of the voice data to perform fine-tuning training (finetune) on the pre-trained basic synthesis model. After training for a set number of iterations according to experience, the training can be stopped, and the voice synthesis model can be obtained. The basic synthesis model can be a basic model with a certain voice synthesis ability trained based on a large amount of voice synthesis data. By fine-tuning the basic synthesis model with the voice data of a specific user, a voice synthesis model that can synthesize voices matching the characteristics such as the gender and timbre of the user can be obtained.
[0080] 304. Receive the target text sent by the user terminal.
[0081] 305. Invoke the voice synthesis model to process the target text and generate the synthesized voice data corresponding to the target text.
[0082] 306. Send the synthesized voice data corresponding to the target text to the user terminal.
[0083] Among them, for the specific implementation of steps 304 to 306, reference can be made to the relevant descriptions of steps 201 to 203 in the foregoing embodiments, which will not be elaborated here.
[0084] In some feasible embodiments, as Figure 4d shown, it is a schematic diagram of the overall implementation process of a voice processing provided by an embodiment of the present application. It includes: a recording stage, a training stage, and a synthesis stage.
[0085] (1) Recording stage: The user starts the personalized speech synthesis program. For each recorded speech, it is determined whether all the text has been recorded. If not, the next speech is continued to be recorded until all the text is recorded, and the recording result (i.e., the user's speech data) is uploaded. It can be seen that operations such as environmental noise detection, text-recording matching check, and gender information input are omitted in the front-end recording link.
[0086] (2) Training stage: The speech is denoised, and then text annotation (i.e., text feature information) is obtained by dictating the denoised speech, and gender information is obtained by determining the gender of the denoised speech. The speech synthesis model is trained by combining the denoised speech, text annotation, and gender information. After the training is completed, it can be released and put online.
[0087] (3) Synthesis stage: The front-end sends a request text to the background. The background calls the trained personalized speech synthesis model to generate the synthesized speech data corresponding to the request text and sends it to the front-end. After receiving the synthesized speech data, the front-end can play it.
[0088] In the embodiments of the present application, the identity feature information of the user and the text feature information of the speech data can be obtained according to the speech data input by the user. The speech synthesis model corresponding to the user is trained by using the text feature information and the identity feature information. When the target text sent by the user terminal is received, the speech synthesis model is called to process the target text to generate the synthesized speech data corresponding to the target text. It can be seen that the user directly submits the text that needs to be synthesized into speech to the background, and the background can quickly generate the corresponding synthesized speech data by using the corresponding speech synthesis model, which simplifies the operation of the user in the speech synthesis service while ensuring a good speech synthesis effect. When designing the speech synthesis interaction system, a light front-end and heavy-back-end interaction method is adopted, and the operation of the front-end is greatly simplified during the user's use of the system. Specifically, in the front-end recording stage, operations such as environmental noise detection, text-recording matching check, and gender information input are reduced. At the same time, in the background training stage, technical means such as speech denoising, speech dictation, and gender classification are introduced accordingly to repair the recording quality and obtain gender auxiliary information, ensuring the good effect of the final speech synthesis model.
[0089] Please refer to Figure 5 , which is a schematic flowchart of another speech processing method provided by the speech processing system shown in Figure 1 of the embodiments of the present application. The speech processing method of the embodiments of the present application is mainly described from the perspective of the user terminal. The speech processing method includes the following steps:
[0090] 501. In response to the user's application start instruction, display the speech input interface of the speech synthesis application, where the speech input interface includes one or more of a recording progress indication area, a reference text display area, and a recording control operation area.
[0091] After the user starts the speech synthesis application, the user terminal can directly display the speech input interface without performing operations such as environmental noise detection. For example, Figure 6a As shown, the speech input interface may include a recording progress indication area 61, a reference text display area 62, and a recording control operation area 63. The recording progress indication area 61 is used to indicate whether the current recording is completed. The reference text display area 62 is used to display text content. The recording control operation area 63 can provide control commands such as "preview", "re-record", and "next".
[0092] 502. Obtain the speech data input by the user through the speech input interface. The speech data includes at least one speech segment input by the user based on at least one reference text.
[0093] Among them, the user can input the corresponding speech data by reading the specific content of the reference text. The user inputs the corresponding speech segment for each reference text. After the user finishes inputting, the user terminal can obtain the speech data including at least one speech segment.
[0094] 503. Send the speech data to the server so that the server can train a speech synthesis model based on the speech data, the text feature information of the speech data, and the identity feature information of the user.
[0095] In some feasible implementation manners, for example, Figure 6b As shown, after the user terminal sends the user's speech data to the server, it can display a prompt message in the speech synthesis application, such as "Congratulations on completing the recording. The model is preparing for training. Please come back to experience after 20 minutes", and can also set to receive a notification after the training is completed.
[0096] In some feasible implementation manners, during the user's recording process, the user terminal can sequentially display at least one reference text in the reference text display area according to a preset recording order, obtain the speech segment input by the user for each displayed reference text, and determine the speech data input by the user through the speech input interface according to the speech segment input by the user for each displayed reference text.
[0097] In some feasible embodiments, after the user terminal receives the notification message indicating that the voice synthesis model training is completed, it displays a content selection interface of the voice synthesis application, obtains the target text selected by the user through the content selection interface, and sends the target text to the server, so that the server can call the voice synthesis model to process the target text and generate the synthesized voice data corresponding to the target text. The user terminal receives the synthesized voice data corresponding to the target text sent by the server and plays the synthesized voice data corresponding to the target text.
[0098] In the embodiments of the present application, during recording, the user can directly input voice data without having to perform cumbersome operations such as environmental noise detection, gender information entry, and consistency verification between the recording and the text. Moreover, the user can directly submit the text that needs to be synthesized into voice to the background, and the background can then quickly generate the corresponding synthesized voice data using the corresponding voice synthesis model. While simplifying the user's operations in the voice synthesis service, it also ensures a good voice synthesis effect. Through the interaction method of light foreground and heavy background, the foreground operation of the user using the system is greatly simplified. During the foreground recording stage, cumbersome operations such as environmental noise detection, text-recording matching check, and gender information input are reduced.
[0099] Please refer to Figure 7 , which is a schematic structural diagram of a voice processing device according to an embodiment of the present application. The device includes:
[0100] A receiving module 701, configured to receive the target text sent by the user terminal.
[0101] A processing module 702, configured to call a voice synthesis model to process the target text and generate the synthesized voice data corresponding to the target text, where the voice synthesis model is trained based on the user's voice data, text feature information of the voice data, and identity feature information of the user.
[0102] A sending module 703, configured to send the synthesized voice data corresponding to the target text to the user terminal.
[0103] Optionally, the receiving module 701 is further configured to receive the voice data sent by the user terminal, where the voice data includes at least one voice segment input by the user based on at least one reference text.
[0104] The processing module 702 is further configured to obtain the identity feature information of the user and the text feature information of the voice data according to the voice data.
[0105] The processing module 702 is further configured to train the voice synthesis model corresponding to the user by using the text feature information and the identity feature information.
[0106] Optionally, the processing module 702 is specifically configured to:
[0107] Perform noise reduction processing on the voice data to obtain the noise-reduced voice data.
[0108] Perform speech recognition processing on the noise-reduced voice data to obtain the text feature information of the voice data.
[0109] Determine the identity feature information of the user according to the noise-reduced voice data, where the identity feature information includes gender information.
[0110] Optionally, the processing module 702 is specifically configured to:
[0111] Extract the acoustic feature information of the noise-reduced voice data.
[0112] Call a speech recognition model to perform decoding processing on the acoustic feature information to obtain the text sequence of the voice data.
[0113] Use the text sequence as the text feature information of the voice data.
[0114] Optionally, the text feature information of the voice data includes the text sequence of each voice segment in the at least one voice segment, and the processing module 702 is further configured to:
[0115] Obtain the matching degree between the text sequence of each voice segment and the corresponding reference text.
[0116] If the matching degree corresponding to the target voice segment is less than or equal to the preset matching degree threshold, delete the text sequence of the target voice segment from the text feature information of the voice data.
[0117] Optionally, the speech recognition model includes an acoustic model and a language model, and the processing module 702 is specifically configured to:
[0118] Call the acoustic model to determine the matching probability between the acoustic feature information and the corresponding phoneme or character.
[0119] Call the language model to determine the occurrence probability of each text sequence.
[0120] Determine the text sequence of the voice data according to the matching probability and the occurrence probability.
[0121] Optionally, the processing module 702 is specifically configured to:
[0122] Extract the acoustic feature information of the noise-reduced voice data.
[0123] Use the first identity discrimination model and the second identity discrimination model to score the acoustic feature information respectively to obtain a scoring result.
[0124] Determine the identity feature information of the user according to the scoring result.
[0125] Optionally, the scoring result includes a first score corresponding to the first identity discrimination model and a second score corresponding to the second identity discrimination model. The processing module 702 is specifically configured to:
[0126] Determine the highest score among the first score and the second score.
[0127] Determine the target identity discrimination model corresponding to the highest score from the first identity discrimination model and the second identity discrimination model.
[0128] Determine the identity feature information of the user according to the gender type corresponding to the target identity discrimination model.
[0129] Optionally, the device further includes: an acquisition module 704, where:
[0130] The acquisition module 704 is configured to acquire a training sample set, and the training sample set includes voice data of multiple users of different gender types.
[0131] The processing module 702 is further configured to train a general background model by using the acoustic feature information of the voice data of each user in the training sample set.
[0132] The processing module 702 is further configured to train the general background model by using the acoustic feature information of the voice data of the users of the first gender type in the training sample set to obtain a first identity discrimination model.
[0133] The processing module 702 is further configured to train the general background model by using the acoustic feature information of the voice data of the users of the second gender type in the training sample set to obtain a second identity discrimination model.
[0134] Optionally, the processing module 702 is specifically configured to:
[0135] Perform Fourier transform processing on the voice data to obtain the spectrum and noise spectrum of the voice data.
[0136] Determine the target spectrum according to the spectrum and noise spectrum of the voice data.
[0137] Obtain the amplitude spectrum of the target spectrum, and perform inverse Fourier transform processing on the phase information of the amplitude spectrum and the spectrum of the voice data to obtain the noise-reduced voice data.
[0138] It should be noted that the functions of the functional modules of the voice processing device according to the embodiments of the present application can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can refer to the relevant descriptions of the above method embodiments and will not be elaborated here.
[0139] Please refer to Figure 8 , which is a schematic structural diagram of another voice processing device according to an embodiment of the present application. The device includes:
[0140] A display module 801, configured to display a voice input interface of a voice synthesis application in response to a user's application start instruction. The voice input interface includes one or more of a recording progress indication area, a reference text display area, and a recording control operation area.
[0141] An acquisition module 802, configured to acquire voice data input by the user through the voice input interface. The voice data includes at least one voice segment input by the user based on at least one reference text.
[0142] A sending module 803, configured to send the voice data to a server, so that the server trains a voice synthesis model according to the voice data, text feature information of the voice data, and identity feature information of the user.
[0143] Optionally, the voice input interface includes the reference text display area. The acquisition module 802 is specifically configured to:
[0144] Sequentially display at least one reference text in the reference text display area through the display module 801 according to a preset recording order.
[0145] Acquire the voice segments input by the user for each displayed reference text.
[0146] Determine the voice data input by the user through the voice input interface according to the voice segments input by the user for each displayed reference text.
[0147] Optionally, the device further includes: a receiving module 804 and a playing module 805, where:
[0148] The display module 801 is further configured to display a content selection interface of the voice synthesis application after receiving a notification message sent by the server indicating that the training of the voice synthesis model is completed.
[0149] The acquisition module 802 is further configured to acquire a target text selected by the user through the content selection interface.
[0150] The sending module 803 is further configured to send the target text to the server, so that the server calls a speech synthesis model to process the target text and generate synthesized speech data corresponding to the target text.
[0151] The receiving module 804 is configured to receive the synthesized speech data corresponding to the target text sent by the server.
[0152] The playing module 805 is configured to play the synthesized speech data corresponding to the target text.
[0153] It should be noted that the functions of the functional modules of the speech processing device in the embodiments of the present application can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can refer to the relevant descriptions of the above method embodiments and will not be elaborated here.
[0154] Please refer to Figure 9 , which is a schematic structural diagram of a server in an embodiment of the present application. The server in the embodiment of the present application includes structures such as a power supply module, and includes a processor 901, a storage device 902, and a network interface 903. Data can be exchanged between the processor 901, the storage device 902, and the network interface 903.
[0155] The storage device 902 may include a volatile memory, such as a random-access memory (RAM); the storage device 902 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the storage device 902 may further include a combination of the above types of memories.
[0156] The processor 901 may be a central processing unit (CPU). In one embodiment, the processor 901 may also be a Graphics Processing Unit (GPU). The processor 901 may also be a combination of a CPU and a GPU. In one embodiment, the storage device 902 is used to store program instructions. The processor 901 may call the program instructions to perform the following operations:
[0157] Receive the target text sent by the user terminal.
[0158] Call the speech synthesis model to process the target text and generate the synthesized speech data corresponding to the target text, where the speech synthesis model is trained based on the user's speech data, the text feature information of the speech data, and the user's identity feature information.
[0159] Send the synthesized speech data corresponding to the target text to the user terminal.
[0160] Optionally, the processor 901 is further configured to:
[0161] Receive the speech data sent by the user terminal, where the speech data includes at least one speech segment input by the user based on at least one reference text.
[0162] Obtain the user's identity feature information and the text feature information of the speech data according to the speech data.
[0163] Train the speech synthesis model corresponding to the user by using the text feature information and the identity feature information.
[0164] Optionally, the processor 901 is specifically configured to:
[0165] Perform noise reduction processing on the speech data to obtain the denoised speech data.
[0166] Perform speech recognition processing on the denoised speech data to obtain the text feature information of the speech data.
[0167] Determine the user's identity feature information according to the denoised speech data, where the identity feature information includes gender information.
[0168] Optionally, the processor 901 is specifically configured to:
[0169] Extract the acoustic feature information of the denoised speech data.
[0170] Call the speech recognition model to perform decoding processing on the acoustic feature information to obtain the text sequence of the speech data.
[0171] Use the text sequence as the text feature information of the speech data.
[0172] Optionally, the text feature information of the speech data includes the text sequence of each speech segment in the at least one speech segment, and the processor 901 is further configured to:
[0173] Obtain the matching degree between the text sequence of each speech segment and the corresponding reference text.
[0174] If the matching degree corresponding to the target voice segment is less than or equal to the preset matching degree threshold, the text sequence of the target voice segment is deleted from the text feature information of the voice data.
[0175] Optionally, the voice recognition model includes an acoustic model and a language model. The processor 901 is specifically configured to:
[0176] Call the acoustic model to determine the matching probability between the acoustic feature information and the corresponding phoneme or character.
[0177] Call the language model to determine the occurrence probability of each text sequence.
[0178] Determine the text sequence of the voice data according to the matching probability and the occurrence probability.
[0179] Optionally, the processor 901 is specifically configured to:
[0180] Extract the acoustic feature information of the noise-reduced voice data.
[0181] Use the first identity discrimination model and the second identity discrimination model to score the acoustic feature information respectively to obtain a scoring result.
[0182] Determine the identity feature information of the user according to the scoring result.
[0183] Optionally, the scoring result includes a first score corresponding to the first identity discrimination model and a second score corresponding to the second identity discrimination model. The processor 901 is specifically configured to:
[0184] Determine the highest score among the first score and the second score.
[0185] Determine the target identity discrimination model corresponding to the highest score from the first identity discrimination model and the second identity discrimination model.
[0186] Determine the identity feature information of the user according to the gender type corresponding to the target identity discrimination model.
[0187] Optionally, the processor 901 is further configured to:
[0188] Obtain a training sample set, where the training sample set includes voice data of multiple users of different gender types.
[0189] Use the acoustic feature information of the voice data of each user in the training sample set to train a universal background model.
[0190] The general background model is trained using the acoustic feature information of the speech data of the user of the first gender type in the training sample set to obtain a first identity discrimination model.
[0191] The general background model is trained using the acoustic feature information of the speech data of the second gender type user in the training sample set to obtain a second identity discrimination model.
[0192] Optionally, the processor 901 is specifically configured to:
[0193] Perform Fourier transform processing on the speech data to obtain the frequency spectrum and noise spectrum of the speech data.
[0194] A target spectrum is determined according to the spectrum of the speech data and the noise spectrum.
[0195] The amplitude spectrum of the target spectrum is obtained, and the amplitude spectrum and the phase information of the spectrum of the speech data are subjected to inverse Fourier transform processing to obtain the denoised speech data.
[0196] In a specific implementation, the processor 901, the storage device 902, and the network interface 903 described in the embodiments of the present application can execute the embodiments of the present application. Figure 2 or Figure 3 The implementation method described in the relevant embodiments of the provided speech processing method can also be implemented in the embodiments of the present application. Figure 7 The implementation methods described in the relevant embodiments of the provided speech processing device will not be repeated here.
[0197] See also Figure 10 , is a schematic diagram of the structure of a user terminal according to an embodiment of the present invention, wherein the user terminal according to the embodiment of the present invention includes a power supply module and other structures, and includes a processor 1001, a storage device 1002, a display device 1003, and a communication device 1004. The processor 1001, the storage device 1002, the display device 1003, and the communication device 1004 can exchange data.
[0198] The storage device 1002 may include a volatile memory, such as a random-access memory (RAM); the storage device 1002 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the storage device 1002 may also include a combination of the above-mentioned types of memory.
[0199] The processor 1001 may be a central processing unit (CPU). In one embodiment, the processor 1001 may also be a Graphics Processing Unit (GPU). The processor 1001 may also be a combination of a CPU and a GPU. In one embodiment, the storage device 1002 is used to store program instructions. The processor 1001 may call the program instructions to perform the following operations:
[0200] In response to a user's application startup instruction, display a voice input interface for the voice synthesis application, where the voice input interface includes one or more of a recording progress indication area, a reference text display area, and a recording control operation area.
[0201] Obtain voice data input by the user through the voice input interface, where the voice data includes at least one voice segment input by the user based on at least one reference text.
[0202] Send the voice data to the server so that the server trains a voice synthesis model based on the voice data, text feature information of the voice data, and identity feature information of the user.
[0203] Optionally, the voice input interface includes the reference text display area, and the processor 1001 is specifically configured to:
[0204] Sequentially display at least one reference text in the reference text display area in a preset recording order.
[0205] Obtain the voice segments input by the user for each displayed reference text.
[0206] Determine the voice data input by the user through the voice input interface according to the voice segments input by the user for each displayed reference text.
[0207] Optionally, the processor 1001 is further configured to:
[0208] After receiving the notification message indicating that the training of the voice synthesis model is completed sent by the server, display a content selection interface for the voice synthesis application.
[0209] Obtain the target text selected by the user through the content selection interface.
[0210] Send the target text to the server so that the server calls the voice synthesis model to process the target text and generate synthesized voice data corresponding to the target text.
[0211] Receive the synthesized voice data corresponding to the target text sent by the server.
[0212] Play the synthesized voice data corresponding to the target text.
[0213] In specific implementation, the processor 1001, storage device 1002, display device 1003, and communication device 1004 described in the embodiments of the present invention can execute the implementation manners described in the relevant embodiments of the voice processing method provided by the embodiments of the present invention, and can also execute the implementation manners described in the relevant embodiments of the voice processing device provided by the embodiments of the present application, which will not be elaborated herein. Figure 5 The technical solution of the present application can be embodied in the form of a software product in essence, or in part that contributes to the prior art, or all or part of the technical solution. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc., specifically the processor in the computer device) to execute all or part of the steps of the methods in the various embodiments of the present application. Among them, the aforementioned storage medium can include: USB flash drive, mobile hard disk, magnetic disk, optical disk, read-only memory (abbreviation: ROM), or random access memory (abbreviation: RAM), etc., various media that can store program codes. Figure 8 As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
[0214] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
[0215] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A voice processing method, characterized in that, the method includes: Receiving a target text sent by a user terminal; Invoking a voice synthesis model corresponding to the user to process the target text, generating synthesized voice data corresponding to the target text; Sending the synthesized voice data corresponding to the target text to the user terminal; wherein, the training process of the voice synthesis model corresponding to the user includes: Receiving voice data sent by the user terminal, the voice data including at least one voice segment input by the user based on at least one reference text; Obtaining the identity feature information of the user and the text feature information of the voice data according to the voice data; wherein, the text feature information of the voice data includes the text sequence of each voice segment in the at least one voice segment; Obtaining the matching degree between the text sequence of each voice segment and the corresponding reference text; if the matching degree corresponding to the target voice segment is less than or equal to a preset matching degree threshold, deleting the text sequence of the target voice segment from the text feature information of the voice data; Extracting the spectral feature of the voice data, and setting hyperparameters of the voice synthesis model by using the identity feature information of the user; Training the voice synthesis model by using the text feature information of the voice data, the identity feature information of the user, and the spectral feature, obtaining the voice synthesis model corresponding to the user, wherein the spectral feature is used as supervision information for training the voice synthesis model.
2. The method according to claim 1, characterized in that, the obtaining the identity feature information of the user and the text feature information of the voice data according to the voice data includes: Performing noise reduction processing on the voice data to obtain denoised voice data; Performing voice recognition processing on the denoised voice data to obtain the text feature information of the voice data; Determining the identity feature information of the user according to the denoised voice data, the identity feature information including gender information.
3. The method according to claim 2, characterized in that, the performing voice recognition processing on the denoised voice data to obtain the text feature information of the voice data includes: Extracting acoustic feature information of the denoised voice data; Invoking a voice recognition model to perform decoding processing on the acoustic feature information to obtain the text sequence of the voice data; Taking the text sequence as the text feature information of the voice data.
4. The method according to claim 3, characterized in that, the voice recognition model includes an acoustic model and a language model, and the invoking the voice recognition model to perform decoding processing on the acoustic feature information to obtain the text sequence of the voice data includes: Invoking the acoustic model to determine the matching probability between the acoustic feature information and the corresponding phoneme or character; Invoking the language model to determine the occurrence probability of each text sequence; Determining the text sequence of the voice data according to the matching probability and the occurrence probability.
5. The method according to claim 2, characterized in that, Determining the identity feature information of the user according to the noise-reduced speech data includes: Extracting the acoustic feature information of the noise-reduced speech data; Respectively performing a scoring process on the acoustic feature information by using a first identity discrimination model and a second identity discrimination model to obtain a scoring result; Determining the identity feature information of the user according to the scoring result.
6. The method according to claim 5, wherein, the scoring result includes a first score corresponding to the first identity discrimination model and a second score corresponding to the second identity discrimination model, and determining the identity feature information of the user according to the scoring result includes: Determining the highest score among the first score and the second score; Determining a target identity discrimination model corresponding to the highest score from the first identity discrimination model and the second identity discrimination model; Determining the identity feature information of the user according to the gender type corresponding to the target identity discrimination model.
7. The method according to claim 5 or 6, wherein, the method further includes: Obtaining a training sample set, where the training sample set includes speech data of multiple users of different gender types; Training a universal background model by using the acoustic feature information of the speech data of each user in the training sample set; Training the universal background model by using the acoustic feature information of the speech data of users of the first gender type in the training sample set to obtain a first identity discrimination model; Training the universal background model by using the acoustic feature information of the speech data of users of the second gender type in the training sample set to obtain a second identity discrimination model.
8. The method according to claim 3, wherein, performing noise reduction processing on the speech data to obtain noise-reduced speech data includes: Performing Fourier transform processing on the speech data to obtain the spectrum and noise spectrum of the speech data; Determining a target spectrum according to the spectrum and noise spectrum of the speech data; Obtaining the amplitude spectrum of the target spectrum, and performing inverse Fourier transform processing on the amplitude spectrum and the phase information of the spectrum of the speech data to obtain noise-reduced speech data.
9. A speech processing method, wherein, the method includes: In response to a user's application start instruction, displaying a speech input interface of a speech synthesis application, where the speech input interface includes one or more of a recording progress indication area, a reference text display area, and a recording control operation area; Obtaining the speech data input by the user through the speech input interface, where the speech data includes at least one speech segment input by the user based on at least one reference text; Sending the speech data to a server so that the server trains a speech synthesis model corresponding to the user according to the speech data, the text feature information of the speech data, and the identity feature information of the user; wherein, the training process of the speech synthesis model corresponding to the user includes: Obtain the identity feature information of the user and the text feature information of the voice data according to the voice data; wherein, the text feature information of the voice data includes the text sequence of each voice segment in the at least one voice segment; Obtain the matching degree between the text sequence of each voice segment and the corresponding reference text; if the matching degree corresponding to the target voice segment is less than or equal to the preset matching degree threshold, delete the text sequence of the target voice segment from the text feature information of the voice data; Extract the spectral feature of the voice data, and set the hyperparameters of the voice synthesis model by using the identity feature information of the user; Train the voice synthesis model by using the text feature information of the voice data, the identity feature information of the user, and the spectral feature to obtain the voice synthesis model corresponding to the user, wherein the spectral feature is used as the supervision information for training the voice synthesis model.
10. The method according to claim 9, wherein, the voice input interface includes the reference text display area, and obtaining the voice data input by the user through the voice input interface includes: Sequentially display at least one reference text in the reference text display area according to a preset recording order; Obtain the voice segment input by the user for each displayed reference text; Determine the voice data input by the user through the voice input interface according to the voice segment input by the user for each displayed reference text.
11. The method according to claim 9 or 10, wherein, the method further includes: After receiving the notification message that the training of the voice synthesis model corresponding to the user is completed sent by the server, display the content selection interface of the voice synthesis application; Obtain the target text selected by the user through the content selection interface; Send the target text to the server so that the server calls the voice synthesis model corresponding to the user to process the target text and generate the synthesized voice data corresponding to the target text; Receive the synthesized voice data corresponding to the target text sent by the server, and play the synthesized voice data corresponding to the target text.
12. A voice processing device, wherein, the device includes: A receiving module, configured to receive the target text sent by the user terminal; A processing module, configured to call the voice synthesis model corresponding to the user to process the target text and generate the synthesized voice data corresponding to the target text; A sending module, configured to send the synthesized voice data corresponding to the target text to the user terminal; wherein, the training process of the voice synthesis model corresponding to the user includes: Receive the voice data sent by the user terminal, where the voice data includes at least one voice segment input by the user based on at least one reference text; Obtain the identity feature information of the user and the text feature information of the voice data according to the voice data; wherein, the text feature information of the voice data includes the text sequence of each voice segment in the at least one voice segment; Obtain the matching degree between the text sequence of each voice segment and the corresponding reference text; if the matching degree corresponding to the target voice segment is less than or equal to the preset matching degree threshold, delete the text sequence of the target voice segment from the text feature information of the voice data; Extract the spectral features of the voice data, and set the hyperparameters of the voice synthesis model using the identity feature information of the user; Train the voice synthesis model using the text feature information of the voice data, the identity feature information of the user, and the spectral features to obtain the voice synthesis model corresponding to the user, where the spectral features are used as the supervision information for training the voice synthesis model.
13. A voice processing device, Characterized in that, The device includes: A display module, configured to display a voice input interface of a voice synthesis application in response to a user's application start instruction, where the voice input interface includes one or more of a recording progress indication area, a reference text display area, and a recording control operation area; An acquisition module, configured to acquire voice data input by the user through the voice input interface, where the voice data includes at least one voice segment input by the user based on at least one reference text; A sending module, configured to send the voice data to a server so that the server trains a voice synthesis model corresponding to the user according to the voice data, the text feature information of the voice data, and the identity feature information of the user; Wherein, the training process of the voice synthesis model corresponding to the user includes: Obtain the identity feature information of the user and the text feature information of the voice data according to the voice data; wherein, the text feature information of the voice data includes the text sequence of each voice segment in the at least one voice segment; Obtain the matching degree between the text sequence of each voice segment and the corresponding reference text; if the matching degree corresponding to the target voice segment is less than or equal to the preset matching degree threshold, delete the text sequence of the target voice segment from the text feature information of the voice data; Extract the spectral features of the voice data, and set the hyperparameters of the voice synthesis model using the identity feature information of the user; Train the voice synthesis model using the text feature information of the voice data, the identity feature information of the user, and the spectral features to obtain the voice synthesis model corresponding to the user, where the spectral features are used as the supervision information for training the voice synthesis model.
14. A computer-readable storage medium, Characterized in that, The computer storage medium stores a computer program, the computer program includes program instructions, and the program instructions are executed by a processor to execute the voice processing method according to any one of claims 1-8, or the voice processing method according to any one of claims 9-11.
15. A computer program product, Characterized in that, The computer program product includes a computer program which, when executed by a processor, implements the speech processing method according to any one of claims 1-8, or the speech processing method according to any one of claims 9-11.
Citation Information
Patent Citations
Voice authentication and speech recognition system and method
CN104185868A
Speech recognition method and related product
CN111161739A
Electronic device and operating method thereof
US20210134269A1
Voice processing device, voice processing method, and program storage medium
WO2020049687A1