A voice processing method and related device

By predicting and generating target edited speech with similar pitch features, the problem of poor listening experience in existing speech editing methods is solved, thus improving the user experience of singing editing.

CN114882862BActive Publication Date: 2026-01-20HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210468926.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2026-01-20
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

Existing speech editing methods result in poor sound quality after correction when there are few speech segments in the database. This is especially true in singing editing, where the timbre and rhythm are unnatural, making it difficult to maintain the coherence and naturalness of the speech.

Method used

By acquiring the pitch features of the original speech and the target text information, the pitch features of the second text are predicted, and a neural network is used to generate the target edited speech, ensuring that the pitch features are similar, thereby improving the listening experience.

Benefits of technology

It achieves similarity in pitch characteristics between the edited and unedited voices, improving the user experience and making the edited voice sound similar to the original voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882862B_ABST
    Figure CN114882862B_ABST
Patent Text Reader

Abstract

A voice processing method applied to the field of song editing, the method comprising: obtaining an original voice and a second text; predicting a second pitch feature of the second text according to a first pitch feature of non-editing voice in the original voice and information of a target text; obtaining a first voice feature corresponding to the second text through a neural network according to the second pitch feature and the second text; and generating a target editing voice corresponding to the second text according to the first voice feature. The application predicts the pitch feature of the second text (to-be-edited text), generates the first voice feature of the second text according to the pitch feature, and generates the target editing voice corresponding to the second text based on the first voice feature, so that the pitch features of the voices before and after song editing are similar, and the hearing of the target editing voice is similar to that of the original voice.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of artificial intelligence, and in particular, to a speech processing method and related equipment. BACKGROUND

[0002] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is the design principle and implementation method of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making. The research in the field of artificial intelligence includes robots, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, AI basic theory, etc.

[0003] At present, speech editing has very important practical significance. For example, in the scene of user recording songs (such as a cappella), etc., some contents in the speech will often be wrong due to slip of the tongue. In this case, speech editing can help users quickly correct the wrong contents in the original song, and generate the corrected speech. The commonly used speech editing method is to pre-construct a database containing a large number of speech segments, obtain the segment of the pronunciation unit from the database, and replace the wrong segment in the original speech with the segment, and then generate the corrected speech.

[0004] However, the above-mentioned speech editing method depends on the diversity of the speech segments in the database, and in the case that the speech segments in the database are less, the corrected speech (such as the user's song) will have poor listening experience. SUMMARY

[0005] Embodiments of the present application provide a speech processing method and related equipment, which can realize that the listening experience of the edited song is similar to that of the original speech, and improve the user experience.

[0006] In a first aspect, the present application provides a speech processing method, which can be applied to scenarios such as user recording a short video, a teacher recording a teaching speech, etc. The method can be executed by a speech processing device or a component (such as a processor, a chip, or a chip system, etc.) of the speech processing device. The speech processing device can be a terminal device or a cloud device. The method comprises: obtaining an original speech and a second text, the second text being a text other than a first text in a target text, the original text corresponding to the original speech and the target text both comprising the first text, and the first text corresponding to a non-edited speech in the original speech; predicting a second pitch feature of the second text according to a first pitch feature of the non-edited speech and information of the target text; obtaining a first speech feature corresponding to the second text by a neural network according to the second pitch feature and the second text; and generating a target edited speech corresponding to the second text according to the first speech feature. The present application predicts the pitch feature of the second text (to-be-edited text), generates the first speech feature of the second text according to the pitch feature, and generates the target edited speech corresponding to the second text based on the first speech feature, so that the pitch features of the speech before and after the song editing are similar, and the listening experience of the target edited speech is similar to that of the original speech.

[0007] In addition, there are various ways to obtain the second text. The second text can be directly obtained. Alternatively, position information (which can also be understood as marking information, used to indicate the position of the second text in the target text) can be obtained first, and then the second text can be obtained according to the position and the target text. The position information is used to indicate the position of the second text in the target text. Alternatively, the target text and the original text (or the target text and the original speech, and the original speech is recognized to obtain the original text) can be obtained, and then the second text can be determined based on the original text and the target text.

[0008] In a possible implementation, the target edited speech corresponding to the second text is generated based on the second speech feature, comprising: generating the target edited speech based on the second speech feature by a vocoder.

[0009] In this possible implementation, the second speech feature is converted into the target edited speech by the vocoder, so that the target edited speech has similar speech features to the original speech, and the listening experience of the user is improved.

[0010] In a possible implementation, the content of the original speech is a user's song, for example, a speech recorded when the user sings.

[0011] In a possible implementation, the obtaining the original speech and the second text comprises: receiving the original speech and the second text sent by the terminal device; and the method further comprises: sending the target edited speech to the terminal device, the target edited speech being used by the terminal device to generate the target speech corresponding to the target text. It can also be understood as an interactive scenario, in which the cloud device performs complex calculation operations, the terminal device performs simple splicing operations, the original speech and the second text are obtained from the terminal device, the cloud device generates the target edited speech, and then sends the target edited speech to the terminal device, and finally the terminal device splices to obtain the target speech.

[0012] In the possible implementation, in the case where the speech processing device is the cloud device, on the one hand, the cloud device can generate the target edited speech through interaction with the terminal device and return the target edited speech to the terminal device, so as to reduce the computing power and storage space of the terminal device. On the other hand, the target edited speech corresponding to the modified text can be generated according to the speech features of the non-edited region in the original speech, and then the target speech corresponding to the target text is generated together with the non-edited speech.

[0013] Optionally, in a possible implementation of the first aspect, the step of obtaining the original speech and the second text comprises: receiving the original speech and the target text sent by the terminal device; and the method further comprises: generating the target speech corresponding to the target text based on the non-edited speech and the target edited speech, and sending the target speech to the terminal device.

[0014] In the possible implementation, the original speech and the target text sent by the terminal device are received, the non-edited speech is obtained, the second speech features corresponding to the second text are generated according to the first speech features of the non-edited speech, the target edited speech is obtained according to the vocoder, and the target speech is generated by splicing the target edited speech and the non-edited speech. In other words, the processing process is performed in the speech processing device, and the result is returned to the terminal device. The cloud device can generate the target speech through complex calculation and return the target speech to the terminal device, so as to reduce the computing power and storage space of the terminal device.

[0015] In a possible implementation, the second speech features include: the first pitch feature of the non-edited speech, information of the target text, and the second speech features of the non-edited speech; and the second speech features carry at least one of the following information: part or all of the speech frames of the non-edited speech; the voiceprint features of the non-edited speech; the timbre features of the non-edited speech; the prosody features of the non-edited speech; and the rhythm features of the non-edited speech.

[0016] The first speech feature can be the same as or similar to a rhythm, tone, signal-to-noise ratio, or the like of the second speech feature. The rhythm can reflect an emotional state or speaking form of a speaker, and generally refers to a tone, pitch, stress, pause, or rhythm.

[0017] In a possible implementation, the second speech feature carries a voiceprint feature of the original speech. The voiceprint feature can be obtained directly or by identifying the original speech.

[0018] In this possible implementation, on one hand, the voiceprint feature of the original speech is introduced, so that the first speech feature generated subsequently also carries the voiceprint feature of the original speech, and the similarity between the target edited speech and the original speech is improved. On the other hand, in the case where the number of speakers (or users) is multiple, the introduction of the voiceprint feature can improve the similarity between the subsequently predicted speech feature and the voiceprint of the speaker of the original speech.

[0019] In a possible implementation, the information of the target text includes:

[0020] Text embeddings of each phoneme in the target text.

[0021] In a possible implementation, the target text is a text obtained by inserting the second text into the first text, or the target text is a text obtained by deleting a first part of text in the first text, and the second text is a text adjacent to the first part of text.

[0022] The predicting the second pitch feature of the second text according to the first pitch feature of the non-edited speech and the information of the target text includes:

[0023] Fusing the first pitch feature of the non-edited speech and the information of the target text to obtain a first fusion result;

[0024] Inputting the first fusion result into a second neural network to obtain the second pitch feature of the second text.

[0025] In a possible implementation, the target text is obtained by replacing a second part of text in the first text with the second text.

[0026] The predicting the second pitch feature of the second text according to the first pitch feature of the non-edited speech and the information of the target text includes:

[0027] inputting a first pitch feature of the non-edited speech into a third neural network to obtain an initial pitch feature, the first initial pitch feature including a pitch of each frame in a plurality of frames;

[0028] inputting information of the target text into a fourth neural network to obtain a pronunciation feature of the second text, the pronunciation feature being used to indicate whether each frame in the plurality of frames included in the initial pitch feature pronounces;

[0029] fusing the initial pitch feature and the pronunciation feature to obtain a second pitch feature of the second text.

[0030] In a possible implementation, the method further includes:

[0031] predicting a frame number of each phoneme in the second text according to the frame number of each phoneme in the non-edited speech and the information of the target text.

[0032] In a possible implementation, the first pitch feature includes a pitch of each frame in a plurality of frames of the non-edited speech.

[0033] The second pitch feature includes a pitch of each frame in a plurality of frames of the target edited speech.

[0034] In a possible implementation, the predicting a frame number of each phoneme in the second text according to the frame number of each phoneme in the non-edited speech and the information of the target text includes:

[0035] the frame number of each phoneme in the non-edited speech, the information of the target text, and a second speech feature of the non-edited speech.

[0036] In a possible implementation, the above steps further include: obtaining a position of the second text in the target text; and splicing the target edited speech and the non-edited speech based on the position to obtain a target speech corresponding to the target text. It can also be understood that the target edited speech is used to replace the edited speech in the original speech, and the edited speech is the speech in the original speech except the non-edited speech.

[0037] In this possible implementation, the target edited speech and the non-edited speech can be spliced according to the position of the second text in the target text. If the first text is all the overlapping texts in the original text and the target text, the speech of the required text (i.e., the target text) can be generated without changing the non-edited speech in the original speech.

[0038] Optionally, in a possible implementation of the first aspect, the step of determining the non-edited speech based on the target text, the original text and the original speech can be: determining a first text based on the target text and the original text; and determining the non-edited speech based on the first text, the original text and the original speech.

[0039] In this possible implementation, the non-edited speech of the first text in the original speech is determined by comparing the original text and the original speech, which facilitates the generation of the first speech feature.

[0040] Optionally, in a possible implementation of the first aspect, the step of determining the first text based on the target text and the original text can include: determining an overlapping text based on the target text and the original text; displaying the overlapping text to a user; and determining the first text from the overlapping text in response to a second operation of the user.

[0041] In a second aspect, the present application provides a speech processing apparatus, which includes:

[0042] an obtaining module configured to obtain an original speech and a second text, the second text being a text other than a first text in a target text, the target text and an original text corresponding to the original speech both including the first text, and a speech corresponding to the first text in the original speech being a non-edited speech;

[0043] a pitch predicting module configured to predict a second pitch feature of the second text according to a first pitch feature of the non-edited speech and information of the target text;

[0044] a generating module configured to obtain a first speech feature corresponding to the second text by a neural network according to the second pitch feature and the second text;

[0045] generate a target edited speech corresponding to the second text according to the first speech feature.

[0046] In a possible implementation, the content of the original speech is a user's singing voice.

[0047] In a possible implementation, the first pitch feature of the non-edited speech and the second text can include:

[0048] the first pitch feature of the non-edited speech, the information of the target text and a second speech feature of the non-edited speech; the second speech feature carrying at least one of the following information:

[0049] part of speech frames or all speech frames of the non-edited speech;

[0050] a voiceprint feature of the non-edited speech;

[0051] a timbre feature of the non-edited speech;

[0052] a prosody feature of the non-edited speech; and

[0053] a rhythm feature of the non-edited speech.

[0054] In a possible implementation, the information of the target text includes: text embeddings of respective phonemes in the target text.

[0055] In a possible implementation, the target text is a text obtained by inserting the second text into the first text; or, the target text is a text obtained by deleting a first part of text of the first text, and the second text is text adjacent to the first part of text.

[0056] The pitch prediction module is specifically configured to:

[0057] fuse the first pitch feature of the non-edited speech and the information of the target text to obtain a first fusion result;

[0058] input the first fusion result into a second neural network to obtain a second pitch feature of the second text.

[0059] In a possible implementation, the target text is obtained by replacing a second part of text in the first text with the second text.

[0060] The pitch prediction module is specifically configured to:

[0061] input the first pitch feature of the non-edited speech into a third neural network to obtain an initial pitch feature, the first initial pitch feature including pitches of each frame in a plurality of frames;

[0062] input the information of the target text into a fourth neural network to obtain a pronunciation feature of the second text, the pronunciation feature being used to indicate whether each frame in the plurality of frames included in the initial pitch feature is pronounced;

[0063] fuse the initial pitch feature and the pronunciation feature to obtain a second pitch feature of the second text.

[0064] In a possible implementation, the apparatus further includes:

[0065] a duration prediction module configured to predict, according to the frame numbers of respective phonemes in the non-edited speech and the information of the target text, frame numbers of respective phonemes in the second text.

[0066] In a possible implementation, the first pitch feature includes a pitch feature of each frame in the plurality of frames of the non-edited speech.

[0067] The second pitch feature includes a pitch feature of each frame in the plurality of frames of the target edited speech.

[0068] In a possible implementation, the duration prediction module is specifically configured to:

[0069] According to the number of frames of each phoneme in the non-edited speech, the information of the target text, and the second speech feature of the non-edited speech.

[0070] In a possible implementation, the obtaining module is further configured to:

[0071] Obtain a position of the second text in the target text.

[0072] The generating module is further configured to splice the target edited speech and the non-edited speech based on the position to obtain the target speech corresponding to the target text.

[0073] The third aspect of the present application provides a speech processing device, which executes the method in the first aspect or any possible implementation manner of the first aspect.

[0074] The fourth aspect of the present application provides a speech processing device, which includes a processor and a memory, the memory is used to store programs or instructions, when the programs or instructions are executed by the processor, the speech processing device realizes the method in the first aspect or any possible implementation manner of the first aspect.

[0075] The fifth aspect of the present application provides a computer readable medium, which stores computer programs or instructions, when the computer programs or instructions are run on a computer, the computer executes the method in the first aspect or any possible implementation manner of the first aspect.

[0076] The sixth aspect of the present application provides a computer program product, which executes the method in the first aspect or any possible implementation manner of the first aspect when executed on a computer. BRIEF DESCRIPTION OF DRAWINGS

[0077] Figure 1 A structural schematic diagram of a system architecture provided by the present application;

[0078] Figure 2 A structural schematic diagram of a convolutional neural network provided by the present application;

[0079] Figure 3 Another convolutional neural network structure diagram provided by the present application is shown in the following figure:

[0080] Figure 4 A chip hardware structure diagram provided by the present application is shown in the following figure:

[0081] Figure 5 A schematic flow chart of a neural network training method provided by the present application is shown in the following figure:

[0082] Figure 6 A neural network structure diagram provided by the present application is shown in the following figure:

[0083] Figure 7a A flow chart of a speech processing method provided by the present application is shown in the following figure:

[0084] Figure 7b A duration prediction diagram provided by the present application is shown in the following figure:

[0085] Figure 7c A pitch prediction diagram provided by the present application is shown in the following figure:

[0086] Figure 7d A pitch prediction diagram provided by the present application is shown in the following figure:

[0087] Figures 8-10 Several display interface diagrams of a speech processing device provided by the present application are shown in the following figures:

[0088] Figure 11 A bidirectional decoder structure diagram provided by the present application is shown in the following figure:

[0089] Figure 12 Another display interface diagram of a speech processing device provided by the present application is shown in the following figure:

[0090] Figure 13 Another flow chart of a speech processing method provided by the present application is shown in the following figure:

[0091] Figures 14-16 Several structure diagrams of a speech processing device provided by the present application are shown in the following figures. DETAILED DESCRIPTION

[0092] The embodiments of the present application provide a speech processing method and related device, which can realize similar listening experience of edited speech and original speech, and improve user experience.

[0093] The technical solutions in the embodiments of the present application will be described below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0094] For the convenience of understanding, the related terms and concepts mainly involved in the embodiments of the present application are introduced below.

[0095] 1. Neural network

[0096] The neural network can be composed of neural units, and the neural unit can refer to an operation unit with X s and intercept 1 as input. The output of the operation unit can be:

[0097]

[0098] wherein s = 1, 2, … n, n is a natural number greater than 1, W s is the weight of X s , b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. The neural network is a network formed by connecting many single neural units as described above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.

[0099] 2. Deep neural network

[0100] Deep neural network (DNN), also known as multi-layer neural network, can be understood as a neural network with many hidden layers. Here, “many” has no special measurement standard. From the position of DNN according to different layers, the neural network inside the DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the number of layers in between is the hidden layer. The layers are fully connected, that is, any neuron in the i-th layer is connected to any neuron in the i+1-th layer. Of course, the deep neural network can also not include the hidden layer, which is not limited here.

[0101] The work of each layer in the deep neural network can be expressed by a mathematical expression Description: The working of each layer in a deep neural network from a physical perspective can be understood as a transformation from an input space (a set of input vectors) to an output space (a set of output vectors) by five operations on the input space, which are: 1, dimensionality increase / decrease; 2, scaling up / down; 3, rotation; 4, translation; 5, "warping". The operations 1, 2, 3 are done by , the operation 4 is done by , and the operation 5 is done by α(). Here, "space" is used because the objects being classified are not single things, but a class of things, and the space refers to the set of all individuals of the class. Here, W is a weight vector, each value in the vector represents the weight value of a neuron in the layer. The vector W determines the spatial transformation from the input space to the output space, i.e., the weight W of each layer controls how the space is transformed. The purpose of training a deep neural network is to ultimately obtain the weight matrix of all layers of the trained neural network (the weight matrix formed by the vectors W of many layers). Therefore, the training process of a neural network is essentially learning the way to control the spatial transformation, more specifically, learning the weight matrix.

[0102] 3. Convolutional neural network

[0103] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. The convolutional neural network includes a feature extractor composed of convolutional layers and subsampling layers. The feature extractor can be regarded as a filter, and the convolution process can be regarded as making the same trainable filter convolve with an input image or a convolutional feature plane. A convolutional layer refers to a layer of neurons in a convolutional neural network that performs convolution processing on an input signal. In the convolutional layer of the convolutional neural network, a neuron can only be connected to part of the adjacent layer neurons. A convolutional layer usually contains several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights are the convolution kernel. Shared weights can be understood as the way of extracting image information being independent of the location. The underlying principle is that the statistical information of a part of the image is the same as that of other parts, i.e., the image information learned in one part can also be used in another part. Therefore, the same learned image information can be used for all locations on the image. In the same convolutional layer, multiple convolution kernels can be used to extract different image information. Generally, the more the number of convolution kernels, the more image information the convolution operation reflects.

[0104] The convolution kernel can be initialized in the form of a matrix of random size, and the convolution kernel can obtain reasonable weights through learning in the training process of the convolution neural network. In addition, the direct benefit of sharing weights is to reduce the connections between layers of the convolution neural network, while reducing the risk of overfitting. The separation network, the identification network, the detection network, the depth estimation network and the like in the embodiments of the present application can all be CNNs.

[0105] 4. Recurrent neural network (RNN)

[0106] In a traditional neural network model, layers are fully connected, and nodes between each layer are unconnected. However, such a general neural network cannot solve many problems. For example, predicting the next word of a sentence, because the words in a sentence are not independent, and generally the previous words need to be used. The recurrent neural network (RNN) means that the current output of a sequence is also related to the previous output. The specific form is that the network memorizes the previous information in the internal state of the network and applies it to the calculation of the current output.

[0107] 5. Loss function

[0108] In the process of training a deep neural network, because it is hoped that the output of the deep neural network is as close as possible to the value that is really wanted to be predicted, the weight vector of each layer of the neural network can be updated according to the difference between the predicted value of the current network and the target value that is really wanted to be predicted (of course, there is usually an initialization process before the first update, that is, the parameters of each layer of the deep neural network are pre-configured), for example, if the predicted value of the network is too high, the weight vector is adjusted to make it predict lower, and the adjustment is continuously made until the neural network can predict the target value that is really wanted to be predicted. Therefore, it is necessary to define in advance "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function, which is an important equation for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, and then the training of the deep neural network becomes a process of trying to minimize the loss.

[0109] 6. Text to speech

[0110] Text to speech (TTS) is a program or software system that converts text into speech.

[0111] 7. Vocoder

[0112] A vocoder is a voice signal processing module or software that can generate a sound waveform from acoustic features.

[0113] 8、Pitch

[0114] Pitch can also be referred to as fundamental frequency. When a sound-emitting body emits sound due to vibration, the sound can generally be decomposed into many simple sinusoidal waves, that is, all natural sounds are basically composed of many sinusoidal waves with different frequencies, among which the sinusoidal wave with the lowest frequency is the fundamental frequency (which can be represented by F0), and the sinusoidal waves with higher frequencies are overtones.

[0115] 9、Prosody

[0116] In the field of speech synthesis, prosody generally refers to features that control intonation, pitch, stress, pause, and rhythm. Prosody can reflect the emotional state or speaking form of the speaker.

[0117] 10、Phone

[0118] Phone: is the smallest unit of speech divided according to the natural properties of speech, analyzed according to the pronunciation movements in a syllable, one movement constitutes a phone. Phones are divided into two categories: vowels and consonants. For example, the Chinese syllable a (e.g., one tone: ah) has only one phone, ai (e.g., four tones: love) has two phones, and dai (e.g., one tone: dull) has three phones.

[0119] 11、Embedding

[0120] Embedding can also be referred to as "word embedding", "vectorization", "vector mapping", "embedding", etc. Formally, a word vector is a dense vector that represents an object.

[0121] 12、Speech feature

[0122] Speech feature: converting a processed speech signal into a concise and logical representation, which is more discriminative and reliable than the actual signal. After obtaining a speech signal, speech features can be extracted from the speech signal. Among them, the extraction method is usually to extract a multi-dimensional feature vector for each speech signal. There are many parameterization representation methods for speech signals, such as perceptual linear prediction (PLP), linear predictive coding (LPC), and mel frequency cepstrum coefficient (MFCC).

[0123] 13、Transformer layer

[0124] The neural network comprises an embedding layer and at least one transformer layer, the at least one transformer layer can be N transformer layers (N is an integer greater than 0), wherein each transformer layer comprises, in sequence, an attention layer, an add&norm layer, a feed forward layer and an add&norm layer. In the embedding layer, the current input is embedded to obtain a plurality of feature vectors; in the attention layer, P input vectors are obtained from the previous layer of the first transformer layer, taking any first input vector in the P input vectors as the center, based on the correlation between each input vector in the preset attention window range and the first input vector, an intermediate vector corresponding to the first input vector is obtained, and thus P intermediate vectors corresponding to the P input vectors are determined; in the pooling layer, the P intermediate vectors are merged into Q output vectors, and the plurality of output vectors obtained by the last transformer layer in the transformer layer are used as the feature representation of the current input.

[0125] Next, the above steps will be specifically introduced in combination with specific examples.

[0126] Firstly, in the embedding layer, the current input is embedded to obtain a plurality of feature vectors.

[0127] The embedding layer can be referred to as the input embedding layer. The current input can be a text input, such as a piece of text or a sentence. The text can be in Chinese, English, or other languages. After obtaining the current input, the embedding layer can perform embedding processing on each word in the current input to obtain the feature vectors of each word. In some embodiments, the embedding layer includes an input embedding layer and a positional encoding layer. In the input embedding layer, word embedding processing can be performed on each word in the current input to obtain the word embedding vectors of each word. In the positional encoding layer, the position of each word in the current input can be obtained, and then position vectors can be generated for the positions of each word. In some examples, the position of each word can be the absolute position of each word in the current input. Taking the current input "When should I pay back Huabei" as an example, the position of "When" can be represented as the first position, the position of "should" can be represented as the second position, and so on. In some examples, the position of each word can be the relative position between each word. Still taking the current input "When should I pay back Huabei" as an example, the position of "When" can be represented as before "should", the position of "should" can be represented as after "When" and before "pay", and so on. When the word embedding vectors and position vectors of each word in the current input are obtained, the position vectors of each word and the corresponding word embedding vectors can be combined to obtain the feature vectors of each word, that is, multiple feature vectors corresponding to the current input are obtained. The multiple feature vectors can be represented as an embedding matrix with a preset dimension. It can be set that the number of feature vectors in the multiple feature vectors is M, and the preset dimension is H-dimensional, then the multiple feature vectors can be represented as an M×H embedding matrix.

[0128] 14. Attention mechanism

[0129] The attention mechanism mimics the internal process of biological observation behavior, that is, a mechanism that aligns internal experience and external sensations to increase the observation fineness of some regions, and can quickly screen out high-value information from a large amount of information using limited attention resources. The attention mechanism can quickly extract the important features of sparse data, so it is widely used in natural language processing tasks, especially machine translation. The self-attention mechanism is an improvement of the attention mechanism, which reduces the dependence on external information and is better at capturing the internal correlation of data or features. The essential idea of the attention mechanism can be rewritten as the following formula:

[0130] Wherein, Lx=||Source|| represents the length of Source, and the formula means that the elements in Source are imagined to be composed of a series of data pairs. When a certain element Query in the target Target is given, the similarity or correlation between Query and each Key is calculated to obtain the weight coefficient of the Value corresponding to each Key, and then the weighted sum of Value is obtained, that is, the final Attention value is obtained. Therefore, the essence of the Attention mechanism is to weight and sum the Value values of the elements in Source, and Query and Key are used to calculate the weight coefficient of the corresponding Value. Conceptually, Attention can be understood as selectively filtering out a small amount of important information from a large amount of information and focusing on these important information, ignoring a large amount of unimportant information. The focusing process is reflected in the calculation of the weight coefficient. The greater the weight, the more focused on the corresponding Value value, that is, the weight represents the importance of the information, and the Value is the corresponding information. The self-attention mechanism can be understood as intra attention. The Attention mechanism occurs between the element Query in the Target and all elements in the Source. The self-attention mechanism refers to the Attention mechanism occurring between elements in the Source or between elements in the Target. It can also be understood as the attention mechanism of the special case of Target=Source. The specific calculation process is the same, only the calculation object changes.

[0131] At present, the scene of voice editing is more and more, for example, the scene of song editing is for users to record songs (such as a cappella) and the like. In order to repair the error content in the original voice caused by slip of tongue, voice editing is usually used. The current voice editing method is to obtain a voice segment from a database and replace the error content with the voice segment, thereby generating a corrected voice.

[0132] However, this method relies too much on the voice segment stored in the database. If the voice segment is quite different from the original voice in timbre, prosody, signal-to-noise ratio, etc., it will cause the corrected vocal sound to be incoherent and the prosody to be unnatural, resulting in poor listening experience of the corrected voice. And although the scene of song editing is very similar to that of voice editing, unlike the smooth voice of speaking voice, the song data changes more in the dimensions of pronunciation duration, sound energy and pitch, and the existing voice editing technology is difficult to be directly applied to song editing.

[0133] To solve the above problems, the application provides a voice editing method. In the singing voice editing, the pitch feature will affect the hearing of the target edited voice and the original voice. The application predicts the pitch feature of the second text (to-be-edited text), generates the first voice feature of the second text according to the pitch feature, and generates the target edited voice corresponding to the second text based on the first voice feature, so that the pitch features of the voices before and after the singing voice editing are similar, and the hearing of the target edited voice is similar to the hearing of the original voice.

[0134] The technical solutions in the embodiments of the application will be described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.

[0135] First, the system architecture provided by the embodiments of the application is introduced.

[0136] Referring to the accompanying Figure 1 , the embodiments of the application provide a system architecture 10. As shown in the system architecture 10, the data acquisition device 16 is used to acquire training data, and the training data in the embodiments of the application includes training voice and training text corresponding to the training voice. The training data is stored in the database 13, and the target model / rule 101 is obtained by the training device 12 based on the training data maintained in the database 13. How the training device 12 obtains the target model / rule 101 based on the training data will be described in detail below. The target model / rule 101 can be used to implement the voice processing method provided by the embodiments of the application, that is, the text is input into the target model / rule 101 after relevant preprocessing, and the voice feature of the text can be obtained. The target model / rule 101 in the embodiments of the application can be a neural network. It should be noted that in actual application, the training data maintained in the database 13 may not all come from the acquisition of the data acquisition device 16, but may also be received from other devices. In addition, it should be noted that the training device 12 may not completely train the target model / rule 101 based on the training data maintained in the database 13, but may also obtain the training data from the cloud or other places for model training. The above description should not be regarded as a limitation of the embodiments of the application.

[0137] The target model / rule 101 trained by the training device 12 can be applied to different systems or devices, such as the execution device 11 shown in the figure. The execution device 11 can be a terminal, such as a mobile phone terminal, a tablet computer, a notebook computer, AR / VR, a vehicle-mounted terminal, etc., and can also be a server or a cloud, etc. In the accompanying Figure 1 , the execution device 11 is used to execute the voice processing method provided by the embodiments of the application. The execution device 11 can be a terminal, such as a mobile phone terminal, a tablet computer, a notebook computer, AR / VR, a vehicle-mounted terminal, etc., and can also be a server or a cloud, etc. In the accompanyingFigure 1 In some embodiments, the execution device 11 is configured with an I / O interface 112 for data interaction with external devices, and the user can input data to the I / O interface 112 through the client device 14, wherein the input data in the embodiments of the present application can include the second speech feature, the target text and the mark information, and the input data can also include the second speech feature and the second text. In addition, the input data can be input by the user, or uploaded by the user through other devices, or can come from a database, and the specific embodiments are not limited herein.

[0138] If the input data includes the second speech feature, the target text and the mark information, the pre-processing module 113 is configured to pre-process the target text and the mark information received by the I / O interface 112, and in the embodiments of the present application, the pre-processing module 113 can be configured to determine the target editing text in the target text based on the target text and the mark information. If the input data includes the second speech feature and the second text, the pre-processing module 113 is configured to pre-process the target text and the mark information received by the I / O interface 112, for example, to convert the target text into phonemes and the like.

[0139] During the pre-processing of the input data by the execution device 11, or during the processing of the calculation module 111 of the execution device 11, the execution device 11 can call the data, code and the like in the data storage system 15 for corresponding processing, or store the data, instructions and the like obtained by the corresponding processing in the data storage system 15.

[0140] Finally, the I / O interface 112 returns the processing result, such as the first speech feature obtained above, to the client device 14, thereby providing the user.

[0141] It is worth noting that the training device 12 can generate corresponding target models / rules 101 based on different training data for different targets or different tasks, and the corresponding target models / rules 101 can be used to achieve the above targets or complete the above tasks, thereby providing the required results for the user or providing the input for subsequent other processing.

[0142] In the attached Figure 1In the scenario shown, the user can manually provide input data, which can be done through the interface provided by I / O interface 112. Alternatively, the client device 14 can automatically send input data to I / O interface 112. If user authorization is required for the client device 14 to automatically send input data, the user can set the corresponding permissions in the client device 14. The user can view the output results of the execution device 11 on the client device 14, which can be presented in various forms such as display, sound, or animation. The client device 14 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 13. Alternatively, data can be collected directly from the I / O interface 112 without going through the client device 14, using the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 13.

[0143] It is worth noting that, attached Figure 1 This is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in the attached diagram... Figure 1 In this case, the data storage system 15 is an external memory relative to the execution device 11. In other cases, the data storage system 15 may also be placed in the execution device 11.

[0144] like Figure 1 As shown, the target model / rule 101 is trained by the training device 12. In this embodiment, the target model / rule 101 can be a neural network. Specifically, in the network provided in this embodiment, the neural network can be a recurrent neural network, a long short-term memory network, etc. The prediction network can be a convolutional neural network, a recurrent neural network, etc.

[0145] Optionally, the neural network and prediction network in the embodiments of this application can be two separate networks, or they can be a multi-task neural network, in which one task is to output duration, one task is to predict pitch features, and another task is to output speech features.

[0146] Since CNN is a very common type of neural network, the following will combine... Figure 2The structure of the CNN is described in detail. As described above, the CNN is a deep neural network with a convolutional structure, and is a deep learning architecture that learns at multiple levels of abstraction using machine learning algorithms. As a deep learning architecture, the CNN is a feed-forward artificial neural network in which each neuron can respond to an image input.

[0147] As shown in Figure 2 , the CNN 100 can include an input layer 110, a convolutional / pooling layer 120, and a neural network layer 130, in which the pooling layer is optional.

[0148] The convolutional / pooling layer 120:

[0149] The convolutional layer:

[0150] As shown in Figure 2 , the convolutional / pooling layer 120 can include layers 121-126, in one implementation, 121 is a convolutional layer, 122 is a pooling layer, 123 is a convolutional layer, 124 is a pooling layer, 125 is a convolutional layer, and 126 is a pooling layer; in another implementation, 121 and 122 are convolutional layers, 123 is a pooling layer, 124 and 125 are convolutional layers, and 126 is a pooling layer. That is, the output of the convolutional layer can be used as the input of the subsequent pooling layer, or as the input of another convolutional layer to continue the convolution operation.

[0151] Taking the convolution layer 121 as an example, the convolution layer 121 can include a plurality of convolution operators, also known as kernels, which are equivalent to filters for extracting specific information from an input image matrix in image processing. The convolution operator can be essentially a weight matrix, which is usually predefined. During the convolution operation on the image, the weight matrix is usually processed on the input image along the horizontal direction one pixel after another (or two pixels after another, depending on the value of the stride), thereby completing the work of extracting specific features from the image. The size of the weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as that of the input image, and the weight matrix extends to the entire depth of the input image during the convolution operation. Therefore, convolution with a single weight matrix will produce a single-depth convolution output, but most cases do not use a single weight matrix, but apply multiple weight matrices of the same dimension. The output of each weight matrix is stacked to form the depth dimension of the convolution image. Different weight matrices can be used to extract different features in the image, such as a weight matrix for extracting image edge information, another weight matrix for extracting specific colors of the image, and another weight matrix for blurring unwanted noise in the image. The multiple weight matrices are of the same dimension, and the feature maps extracted by the multiple weight matrices of the same dimension are also of the same dimension. The multiple extracted feature maps of the same dimension are combined to form the output of the convolution operation.

[0152] The weight values in these weight matrices need to be obtained through a large amount of training in actual applications. The weight values obtained through training form each weight matrix that can extract information from the input image, thereby helping the convolutional neural network 100 to make correct predictions.

[0153] When the convolutional neural network 100 has multiple convolution layers, the initial convolution layer (such as 121) often extracts more general features, which can also be referred to as low-level features. As the depth of the convolutional neural network 100 increases, the features extracted by the later convolution layers (such as 126) become more and more complex, such as high-level semantic features and the like. The higher the semantic features, the more suitable they are for the problem to be solved.

[0154] Pooling layer:

[0155] Since it is often necessary to reduce the number of training parameters, a pooling layer is often periodically introduced after the convolution layer, that is, as shown in FIG. 1, the pooling layer 122 is introduced after the convolution layer 121. Figure 2Layers 121-126 in example 120 can be a convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. In image processing, the sole purpose of pooling layers is to reduce the spatial size of the image. Pooling layers can include average pooling and / or max pooling operators to sample the input image to obtain a smaller image size. Average pooling calculates the average value of pixel values ​​within a specific range. Max pooling takes the pixel with the largest value within a specific range as the result of max pooling. Furthermore, just as the size of the weight matrix in a convolutional layer should be related to the image size, the operators in a pooling layer should also be related to the image size. The size of the output image after pooling can be smaller than the size of the input image of the pooling layer. Each pixel in the output image represents the average or maximum value of the corresponding sub-region of the input image of the pooling layer.

[0156] Neural network layer 130:

[0157] After processing by the convolutional / pooling layers 120, the convolutional neural network 100 is still insufficient to output the required information. As mentioned earlier, the convolutional / pooling layers 120 only extract features and reduce the parameters introduced by the input image. However, to generate the final output information (the required class information or other relevant information), the convolutional neural network 100 needs to utilize neural network layers 130 to generate one or more outputs representing the required number of classes. Therefore, neural network layers 130 may include multiple hidden layers (such as...). Figure 2 As shown in 131, 132 to 13n) and output layer 140, the parameters contained in these multi-layer hidden layers can be pre-trained based on relevant training data for specific task types, such as image recognition, image classification, image super-resolution reconstruction, etc.

[0158] After the multiple hidden layers in neural network layer 130, the final layer of the entire convolutional neural network 100 is the output layer 140. This output layer 140 has a loss function similar to classification cross-entropy, specifically used to calculate the prediction error. Once the entire convolutional neural network 100 has undergone forward propagation (e.g., ...), the loss function is applied. Figure 2 The propagation from 110 to 140 is completed (forward propagation), and the reverse propagation (such as...) Figure 2 The propagation from 140 to 110 (backpropagation) will begin to update the weight values ​​and biases of the layers mentioned above, in order to reduce the loss of the convolutional neural network 100 and the error between the output of the convolutional neural network 100 through the output layer and the ideal result.

[0159] It should be noted that, as Figure 2The convolutional neural network 100 shown is merely an example of a convolutional neural network. In specific applications, convolutional neural networks can also exist in the form of other network models, such as... Figure 3 The multiple convolutional / pooling layers shown are run in parallel, and the extracted features are all input into the full neural network layer 130 for processing.

[0160] The following describes a chip hardware structure provided by an embodiment of this application.

[0161] Figure 4 A chip hardware structure provided in this application embodiment includes a neural network processor 40. This chip can be configured as follows: Figure 1 The execution device 110 shown is used to perform the calculations of the calculation module 111. This chip can also be located in, for example... Figure 1 The training device 120 shown is used to complete the training work of the training device 120 and output the target model / rule 101. For example... Figure 2 The algorithms for each layer in the convolutional neural network shown can all be implemented in, for example... Figure 4 This is achieved in the chip shown.

[0162] The neural network processor 40 can be any processor suitable for large-scale XOR operations, such as a neural network processing unit (NPU), tensor processing unit (TPU), or graphics processing unit (GPU). Taking an NPU as an example: the neural network processor NPU40 is mounted as a coprocessor on the main central processing unit (CPU) (host CPU), and tasks are assigned by the main CPU. The core of the NPU is the arithmetic circuit 403, and the controller 404 controls the arithmetic circuit 403 to retrieve data from the memory (weight memory or input memory) and perform operations.

[0163] In some implementations, the arithmetic circuit 403 internally includes multiple process engines (PEs). In some implementations, the arithmetic circuit 403 is a two-dimensional pulsating array. The arithmetic circuit 403 can also be a one-dimensional pulsating array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 403 is a general-purpose matrix processor.

[0164] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The operation circuit fetches the data of matrix B from the weight memory 402 and caches it in each PE of the operation circuit. The operation circuit fetches the data of matrix A from the input memory 401 and performs matrix operation with matrix B to obtain partial results or final results of the matrix, which are stored in the accumulators 408.

[0165] The vector computation unit 407 can perform further processing on the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector computation unit 407 can be used for network computation of non-convolution / non-FC layers in a neural network, such as pooling, batch normalization, local response normalization, etc.

[0166] In some implementations, the vector computation unit 407 can store the processed output vector to the unified buffer 406. For example, the vector computation unit 407 can apply a non-linear function to the output of the operation circuit 403, e.g., a vector of accumulated values, to generate activation values. In some implementations, the vector computation unit 407 generates normalized values, merged values, or both. In some implementations, the processed output vector can be used as activation input to the operation circuit 403, e.g., for use in a subsequent layer in a neural network.

[0167] The unified memory 406 is used to store input data and output data.

[0168] The weight data is transferred from the external memory to the input memory 401 and / or the unified memory 406, from the external memory to the weight memory 402, and from the unified memory 506 to the external memory by the direct memory access controller (DMAC) 405.

[0169] The bus interface unit (BIU) 410 is used to interact between the main CPU, the DMAC, and the instruction fetch memory 409 through a bus.

[0170] The instruction fetch memory 409 connected to the controller 404 is used to store instructions used by the controller 404.

[0171] The controller 404 is used to invoke the instructions cached in the instruction fetch memory 409 to control the working process of the operation accelerator.

[0172] Generally, the unified memory 406, the input memory 401, the weight memory 402, and the instruction memory 409 are on-chip memories, and the external memory is a memory external to the NPU, which can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM), or other readable and writable memories.

[0173] wherein, Figure 2 or Figure 3 The operations of each layer in the convolutional neural network shown in the figure can be performed by the operation circuit 403 or the vector calculation unit 407.

[0174] First, the application scenario to which the speech processing method provided in the embodiments of the present application is applied is described. The speech processing method can be applied to a scenario that requires modification of speech content, for example, a scenario in which a user records a short video, a teacher records a teaching speech, etc. The speech processing method can be applied to, for example, a smart voice assistant on a mobile phone, a computer, a wearable terminal that can produce sound, an application program, software or speech processing device with a speech editing function, such as a smart speaker.

[0175] The speech processing device is a terminal device or a cloud device for serving a user. The terminal device can include a head mount display (HMD), which can be a combination of a virtual reality (VR) box and a terminal, a VR all-in-one machine, a personal computer (PC), an augmented reality (AR) device, a mixed reality (MR) device, etc. The terminal device can also include a cellular phone, a smart phone, a personal digital assistant (PDA), a tablet computer, a laptop computer, a personal computer (PC), a vehicle-mounted terminal, etc., without limitation.

[0176] The training method of the neural network, the prediction network, and the speech processing method of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0177] The neural network and the prediction network in the embodiments of the present application can be two separate networks, or can be a multi-task neural network, in which one task is to output the duration and the other task is to output the speech feature.

[0178] Secondly, in combination with Figure 5 The training method of the neural network in the embodiments of the present application is described in detail. Figure 5 The training method shown can be executed by a neural network training device. The neural network training device can be a cloud service device, or a terminal device, for example, a computer, a server, or other devices with sufficient computing power to execute the training method of the neural network, or a system composed of a cloud service device and a terminal device. Exemplarily, the training method can be executed by Figure 1 The training device 120 in the cloud service device 100, Figure 4 The neural network processor 40 in the cloud service device 100.

[0179] Alternatively, the training method can be processed by a CPU, or jointly processed by a CPU and a GPU, or not using a GPU, but using other processors suitable for neural network calculation, which is not limited in the present application.

[0180] Figure 5 The training method shown includes steps 501 and 502. The steps 501 and 502 are described in detail below.

[0181] First, the training process of the prediction network is described briefly. The prediction network in the embodiments of the present application can be a transformer network, an RNN, a CNN, etc., which is not limited here. In the training phase, the input of the prediction network is the vector of the training text, and the output is the duration, pitch feature or speech feature of each phoneme in the training text. Then, the difference between the duration, pitch feature or speech feature of each phoneme in the training text output by the prediction network and the actual duration, actual pitch feature or actual speech feature of the training speech corresponding to the training text is continuously reduced, and thus the trained prediction network is obtained.

[0182] Step 501, obtaining training data.

[0183] The training data in the embodiments of the present application includes training speech, or includes training speech and training text corresponding to the training speech. If the training data does not include training text, the training text can be obtained by recognizing the training speech.

[0184] Alternatively, if the number of speakers (or users) is multiple, in order to ensure the correctness of the subsequent predicted speech feature, the training speech feature in the training data can further include a user identifier, or include a voiceprint feature of the training speech, or include a vector for identifying the voiceprint feature of the training speech.

[0185] Optionally, the training data can further include start and end time length information of each phoneme in the training speech.

[0186] The training data in the embodiments of the present application can be obtained by directly recording the sound of the sound object, or by inputting audio information and video information by the user, or by receiving the sending of the collection device. In actual application, there are other ways to obtain training data, and the specific way of obtaining training data is not limited here.

[0187] In step 502, the training data is used as the input of the neural network, and the neural network is trained with the value of the loss function being less than a threshold value, to obtain a trained neural network.

[0188] Optionally, the training data can be preprocessed, for example, if the training data includes training speech, the training text can be obtained by recognizing the training speech, and the training text is input into the neural network in the form of phonemes.

[0189] In the training process, the entire training text can be regarded as a target editing text and input, and the neural network is trained with the value of the loss function being reduced, that is, the difference between the speech features output by the neural network and the actual speech features corresponding to the training speech is continuously reduced. The training process can be understood as a prediction task. The loss function can be understood as a loss function corresponding to the prediction task.

[0190] The neural network in the embodiments of the present application can be an attention mechanism model, for example: transformer, tacotron2, etc. The attention mechanism model includes an encoder-decoder, and the structure of the encoder or the decoder can be a recurrent neural network, a long short-term memory (LSTM), etc.

[0191] The neural network in the embodiments of the present application includes an encoder and a decoder, and the structure type of the encoder and the decoder can be RNN, LSTM, etc., which is not limited here. The function of the encoder is to encode the training text into a text vector (a vector represented by phonemes, each input corresponds to a vector), and the function of the decoder is to obtain the speech features corresponding to the text according to the text vector. In the training process, the decoder calculates each step with the real speech features corresponding to the previous step as a condition.

[0192] Further, in order to ensure the coherence of the front and rear speech, a prediction network can be used to correct the speech duration corresponding to the text vector. That is, the text vector can be up-sampled according to the duration of each phoneme in the training speech (which can also be understood as expanding the frame number of the vector) to obtain a vector corresponding to the frame number. The role of the decoder is to obtain the speech features corresponding to the text according to the above vector corresponding to the frame number.

[0193] Optionally, the decoder described above can be a one-way decoder or a two-way decoder (i.e., two directions in parallel), and the specific implementation is not limited here. The two directions refer to the direction of the training text, which can also be understood as the direction of the vector corresponding to the training text, and can also be understood as the forward order or reverse order of the training text. One direction is that one side of the training text points to the other side of the training text, and the other direction is that the other side of the training text points to the one side of the training text.

[0194] For example, if the training text is: "Have you had lunch?", the first direction or forward order can be from "Zhong" to "Mei", and the second direction or reverse order can be from "Mei" to "Zhong".

[0195] If the decoder is a two-way decoder, the decoders of the two directions (or forward and reverse orders) are trained in parallel, and each calculates independently in the training process, and there is no result dependency. Of course, if the prediction network and the neural network are a multi-task network, the prediction network can be referred to as a prediction module, and the decoder can correct the speech features output by the neural network according to the real duration information corresponding to the training text.

[0196] For example, taking song editing as an example, the input during model training can be the original song audio, the corresponding lyrics text (represented in phonemes), the duration information of each phoneme in the original audio obtained according to the original song audio, the singer's voiceprint feature, frame-level Pitch information, etc., which can be obtained through pre-trained other models or tools (such as song lyrics alignment tools, Singer voiceprint extraction tools, and Pitch extraction algorithms, etc.). The output can be a trained acoustic model, and the training target is to minimize the error between the predicted song features and the song speech features.

[0197] In the data preparation of the training sample, the corresponding training data sample can be constructed based on the song synthesis training data set, respectively simulating the "insertion, deletion and replacement" operation scenarios.

[0198] Training process:

[0199] Stage 1: First, use the ground-truth lyrics and audio, as well as Pitch and duration data, to train a song synthesis model, thereby obtaining a trained text encoding module and an audio feature decoding module;

[0200] Stage2: Fix the text encoding module and the audio feature decoding module, and use the simulated editing operation training data set to train the duration normalization module and the pitch prediction module;

[0201] Stage3: End-to-end training, using all the training data to fine-tune the entire model.

[0202] The architecture of the neural network in the embodiments of the present application can refer to Figure 6 . Among them, the neural network includes an encoder and a decoder. Optionally, the neural network can also include a prediction module and an up-sampling module. The prediction module is specifically used to implement the function of the prediction network described above, and the up-sampling module is specifically used to implement the process of up-sampling the text vector according to the duration of each phoneme in the training speech, which will not be described here.

[0203] It should be noted that the training process can also use other training methods instead of the foregoing training methods, which are not limited here.

[0204] The speech processing method of the embodiments of the present application will be described in detail below in conjunction with the drawings.

[0205] First, the speech processing method provided by the embodiments of the present application can be applied to a replacement scenario, an insertion scenario or a deletion scenario. The above scenarios can be understood as replacing, inserting, deleting, etc. the original speech corresponding to the original text to obtain a target speech, realizing that the target speech and the original speech are similar in hearing and / or improving the fluency of the target speech. Among them, the original speech can be considered as including the speech to be modified, and the target speech is the speech obtained after the user wants to modify the original speech.

[0206] In order to facilitate understanding, the following describes several examples of the above scenarios:

[0207] I. For the replacement scenario.

[0208] The original text is "Today Shenzhen weather is very good", and the target text is "Today Guangzhou weather is very good". Among them, the overlapping text is "Today weather is very good". The non-overlapping text in the original text is "Shenzhen", and the non-overlapping text in the target text is "Guangzhou". The target text includes a first text and a second text, and the first text is the overlapping text or part of the overlapping text. The second text is the text in the target text except the first text. For example: if the first text is "Today weather is very good", then the second text is "Guangzhou". If the first text is "Today weather is very good", then the second text is "Guangzhou".

[0209] II. For the insertion scenario.

[0210] The original text is "Today the weather in Shenzhen is very good", and the target text is "Today the weather in Shenzhen is very good". The overlapping text is "Today the weather in Shenzhen is very good". The non-overlapping text in the target text is "morning". In order to realize the coherence before and after the target voice, the insertion scene can be regarded as a replacement scene of replacing "today Shenzhen" in the original voice with "today morning Shenzhen". That is, the first text is "today the weather is very good", and the second text is "today morning Shenzhen".

[0211] III. For deletion scene.

[0212] The original text is "Today the weather in Shenzhen is very good", and the target text is "Today the weather is very good". The overlapping text is "Today the weather is very good". The non-overlapping text in the original text is "Shenzhen". In order to realize the coherence before and after the target voice, the deletion scene can be regarded as a replacement scene of replacing "today Shenzhen" in the original voice with "today". That is, the first text is "today the weather is very good", and the second text is "today".

[0213] Optionally, the above several scenes are only examples, and in actual application, there are other scenes, which are not limited here.

[0214] Since the above deletion scene and insertion scene can be replaced by the replacement scene, the speech processing method provided by the embodiments of the present application will be described below by taking the replacement scene as an example. The speech processing method provided by the embodiments of the present application can be executed by a terminal device or a cloud device alone, or can be completed by the terminal device and the cloud device together, which will be described as follows:

[0215] Embodiment one: the terminal device or the cloud device executes the speech processing method alone.

[0216] Please refer to Figure 7a , an embodiment of the speech processing method provided by the embodiments of the present application, which can be executed by a speech processing device or a component (such as a processor, a chip, or a chip system) of the speech processing device. The speech processing device can be a terminal device or a cloud device, and the embodiment includes steps 701 to 704.

[0217] Step 701, obtaining the original voice and the second text.

[0218] In the embodiments of the present application, the speech processing device can directly obtain the original voice, the original text and the second text. Alternatively, the original voice and the second text can be obtained first, and the original text corresponding to the original voice is obtained by recognizing the original voice. The second text is the text in the target text except the first text, and the original text and the target text contain the first text. The first text can be understood as part or all of the overlapping text in the original text and the target text.

[0219] In a possible implementation, the content of the original speech is a user's singing voice, for example, a voice recorded when the user sings a cappella.

[0220] In the embodiments of the present application, the voice processing device can acquire the second text in multiple ways, which are described as follows.

[0221] Firstly, the voice processing device can acquire the second text directly through input of other devices or the user.

[0222] Secondly, the voice processing device acquires the target text, and obtains the overlapping text according to the original text corresponding to the original speech and the target text, and then determines the second text according to the overlapping text. Specifically, the original text and the characters in the target text can be compared one by one or input into a comparison model to determine the overlapping text and / or non-overlapping text of the original text and the target text. Then, the first text is determined according to the overlapping text. The first text can be the overlapping text, or part of the overlapping text.

[0223] In the embodiments of the present application, the voice processing device can determine the overlapping text as the first text, or determine the first text in the overlapping text according to a preset rule, or determine the first text in the overlapping text according to the user's operation. The preset rule can be to remove N characters in the overlapping content to obtain the first text, where N is a positive integer.

[0224] It can be understood that the above two ways are only examples, and in actual application, there are other ways to acquire the second text, which are not limited here.

[0225] In addition, the voice processing device can align the original text with the original speech to determine the start and end positions of each phoneme in the original text in the original speech, and can obtain the duration of each phoneme in the original text. Then, the phonemes corresponding to the first text are acquired, that is, the speech (i.e., non-edited speech) corresponding to the first text in the original speech is acquired.

[0226] Optionally, the voice processing device can align the original text with the original speech in the following ways: using a forced alignment method, such as a montreal forced aligner (MFA), a neural network with alignment function, and the like, which are not limited here.

[0227] Optionally, after the voice processing device obtains the original voice and the original text, the voice processing device can display a user interface to the user, the user interface comprising the original voice and the original text. Further, the user performs a first operation on the original text through the user interface, and the voice processing device determines the target text in response to the first operation of the user. The first operation can be understood as an edit of the original text by the user, and the edit can be the aforementioned replacement, insertion, or deletion, etc.

[0228] For example, the above-mentioned replacement scenario is continued. The original text is "Today, the weather in Shenzhen is very good", and the target text is "Today, the weather in Guangzhou is very good". For example, the voice processing device is a mobile phone. After the voice processing device obtains the original text and the original voice, the voice processing device displays an interface as shown in FIG. 9A to the user, the interface comprising the original text and the original voice. As shown in FIG. 9B, the user can perform a first operation 901 on the original text, such as modifying "Shenzhen" to "Guangzhou", etc. The first operation is described by way of example. Figure 8 Figure 9

[0229] Optionally, after the voice processing device determines the overlapping text of the original text and the target text, the voice processing device displays the overlapping text to the user, and then determines the first text from the overlapping text according to a second operation of the user, and further determines the second text. The second operation can be a click, drag, slide, etc. operation, which is not limited here.

[0230] For example, the above-mentioned example is continued. The second text is "Guangzhou", the first text is "Today, the weather is very good", and the non-editing voice is the voice of the first text in the original voice. Assuming that one character corresponds to 2 frames, the original voice corresponding to the original text includes 16 frames, and the non-editing voice corresponds to the 1st frame to the 4th frame and the 9th frame to the 16th frame in the original voice. It can be understood that in actual applications, the correspondence between the character and the voice frame is not necessarily 1:2 as in the above-mentioned example, and the above-mentioned example is only for the convenience of understanding the non-editing area. The number of frames corresponding to the original text is not limited here. After the target text is determined, the voice processing device can display an interface as shown in FIG. 9C, the interface comprising the second text, the target text, the non-editing voice, and the editing voice in the original voice. The second text is "Guangzhou", the target text is "Today, the weather in Guangzhou is very good", the non-editing voice is the voice corresponding to "Today, the weather is very good", and the editing voice is the voice corresponding to "Shenzhen". It can also be understood that as the user edits the target text, the voice processing device determines the non-editing voice in the original voice based on the target text, the original text, and the original voice. Figure 10

[0231] ​​​Optionally, the voice processing device receives an editing request sent by the user, which includes the original voice and the second text. Optionally, the editing request further includes the original text and / or the speaker identifier. Of course, the editing request may also include the original voice and the target text.

[0232] Step 702, predict the second pitch feature of the second text according to the first pitch feature of the non-edited voice and the information of the target text.

[0233] In a possible implementation, the information of the target text includes: the text embedding of each phoneme in the target text.

[0234] In a possible implementation, according to the target text, the text embedding of each phoneme in the target text can be obtained through a text encoding module (Text Encoder). For example, the target text can be converted into a corresponding phoneme sequence (such as the phoneme corresponding to "爱怎么可以不问对错" is the sequence of the initials and finals of its pinyin), and then input into the Text Encoder to be converted into the corresponding text embedding in units of phonemes. The network structure of the Text Encoder can be exemplarily the Tacotron2 model.

[0235] In a possible implementation, the number of frames (which can also be called the duration) of each phoneme in the non-edited voice can be obtained, and the number of frames of each phoneme in the second text can be predicted according to the number of frames of each phoneme in the non-edited voice and the information of the target text.

[0236] In a possible implementation, the neural network used to predict the number of frames of each phoneme in the second text can be as Figure 7b shown (for example, it can be a duration prediction model based on a mask mechanism that fuses the original true duration), which takes the output of the TextEncoder, the original true duration (Reference Duration, that is, the duration of each phoneme in the first text), and the corresponding mask as inputs and predicts the duration (that is, the number of frames in the corresponding audio) of each phoneme to be edited (that is, each phoneme in the second text).

[0237] In a possible implementation, after obtaining the number of frames of each phoneme in the target text (including the first text and the second text), according to the predicted duration of each phoneme, each text embedding can be upsampled to obtain an embedding result corresponding to the number of frames (exemplarily, if the predicted duration of the phoneme ai is 10 frames, then the text embedding corresponding to ai can be copied N times, N is a positive number greater than 1, for example, N is 10).

[0238] It should be understood that, in the scenario of song editing, the song itself will follow a certain score, and the score will define the duration and pitch of each word. Therefore, in song editing, for the non-editing area (non-editing voice), the corresponding duration and pitch information do not need to be predicted, and the accurate real value can be directly obtained and used.

[0239] Next, an example of duration prediction for the second text is given:

[0240] Referring to Figure 7b , Reference Durations are the real durations of each phoneme in the original song audio, and the dashed box is the to-be-predicted duration of each phoneme in the second text (since it is unknown at this time, 0 can be used instead); and Edit Mask is used to mark the phonemes to be predicted (where Mask=0 indicates that prediction is required); Embedding Layer fuses Referencedurations and edit Mask (for example, fusion can be performed by performing an inner product operation), and the result is then accumulated with Text Embedding and Singer Embedding (extracted voiceprint features). Among them, 1 FFT Block can be a Transformer block, and for example, 4 (i.e., N=4) FFT blocks can be used; finally, the model predicts the duration of the phoneme corresponding to Mask=0, and outputs it together with the duration of other non-editing phonemes.

[0241] In one possible implementation, the predicted duration of each phoneme in the second text can be used for upsampling of each input in the pitch feature prediction, for example, the input for pitch feature prediction can include text embedding, and each text embedding before upsampling corresponds to one phoneme, and the text embedding after upsampling includes the number of text embeddings corresponding to the number of frames of the phoneme.

[0242] In one possible implementation, a second voice feature of the non-editing voice can also be obtained according to the non-editing voice. The second voice feature can carry at least one of the following information: part or all of the voice frames of the non-editing voice; the voiceprint feature of the non-editing voice; the timbre feature of the non-editing voice; the prosody feature of the non-editing voice; and the rhythm feature of the non-editing voice.

[0243] The voice feature in the embodiments of the present application can be used to represent the characteristics of the voice (for example: tone, rhythm, emotion or rhythm, etc.), and the voice feature has various forms of representation, which can be a voice frame, a sequence, a vector, etc., and the specific forms are not limited here. In addition, the voice feature in the embodiments of the present application can be a parameter extracted from the above-mentioned forms of representation by the aforementioned PLP, LPC, MFCC, etc.

[0244] Optionally, at least one voice frame is selected from the non-edited voice as the second voice feature. Further, the first voice feature is more combined with the context of the second voice feature. The text corresponding to the at least one voice frame can be the text adjacent to the second text in the first text.

[0245] Optionally, the non-edited voice is encoded by an encoding model to obtain a target sequence, and the target sequence is taken as the second voice feature. The encoding model can be CNN, RNN, etc., and the specific forms are not limited here.

[0246] In addition, the second voice feature can also carry the voiceprint feature of the original voice. The voiceprint feature can be obtained directly or by recognizing the original voice. On the one hand, by introducing the voiceprint feature of the original voice, the first voice feature generated subsequently also carries the voiceprint feature of the original voice, thereby improving the similarity between the target edited voice and the original voice. On the other hand, in the case where the number of speakers (or users) is multiple, introducing the voiceprint feature can improve the voice feature predicted subsequently to be more similar to the voiceprint of the speaker of the original voice.

[0247] Optionally, the voice processing device can also obtain the speaker identifier of the original voice, so as to match the voice corresponding to the corresponding speaker when the number of speakers is multiple, thereby improving the similarity between the target edited voice and the original voice.

[0248] Hereinafter, only the voice frame is taken as an example (or understood as the voice feature is obtained according to the voice frame) for description. For example, continuing the above example, at least one of the first frame to the fourth frame and the ninth frame to the sixteenth frame in the original voice is selected as the second voice feature.

[0249] For example, the second voice feature is a mel-frequency cepstrum feature.

[0250] In a possible implementation, the second voice feature can be expressed in the form of a vector. In a possible implementation, the duration of each phoneme in the predicted second text can be used for upsampling of each input in the pitch feature prediction, for example, the input for the pitch feature prediction can include the second voice feature, each vector before upsampling corresponds to a phoneme, and the text embedding after upsampling includes a vector corresponding to the number of frames of the phoneme.

[0251] In one possible implementation, the second pitch feature of the second text can be predicted based on the first pitch feature of the unedited speech and information from the target text.

[0252] In one possible implementation, the first pitch feature of the unedited speech can be obtained using existing pitch extraction algorithms, which are not limited in this application.

[0253] In one possible implementation, the second pitch feature of the second text can be predicted using a neural network based on the first pitch feature of the unedited speech, information about the target text, and the second speech feature of the unedited speech.

[0254] The following describes how to predict the second pitch feature of the second text based on the first pitch feature of the unedited speech and the information of the target text:

[0255] In one possible implementation, the target text is the text obtained by inserting the second text into the first text; or, the target text is the text obtained by deleting a first part of the first text, and the second text is the text adjacent to the first part of the first text; the first pitch feature of the unedited speech and the information of the target text can be fused to obtain a first fusion result; the first fusion result is input into a second neural network to obtain the second pitch feature of the second text.

[0256] For insertion and deletion operations: Use Figure 7c The model shown predicts frame-level pitch features for the target edited phonemes. The pitch prediction model for insertion and deletion operations can have the same model structure. Figure 7b The settings are the same or similar, the only difference being that the input at this time is at the frame level. Figure 7b The input is the pitch value extracted from the actual singing voice at the phoneme level. Figure 7b The input is the duration information, where the pitch of the area to be edited is marked by a dashed box, and its corresponding EditMask flag is set to 0.

[0257] In a possible implementation, the target text is obtained by replacing a second part of text in the first text with the second text; the first pitch feature of the non-editing speech can be input into a third neural network to obtain an initial pitch feature, the first initial pitch feature including a pitch of each frame in a plurality of frames; information of the target text can be input into a fourth neural network to obtain a pronunciation feature of the second text, the pronunciation feature being used to indicate whether each frame in the plurality of frames included in the initial pitch feature pronounces; and the initial pitch feature and the pronunciation feature can be fused to obtain a second pitch feature of the second text.

[0258] For the replacement operation (here, the replacement operation only represents a case where the number of characters of the new editing text is consistent with the number of characters of the replaced text, and if the numbers are inconsistent, the replacement operation is decomposed into two editing operations of deleting and then inserting). Since the replaced text can be quite different in pronunciation, a model shown in FIG. 6 is used to predict a new pitch to ensure the coherence of the singing before and after the replacement. Figure 7d

[0259] The pitch prediction model for the replacement operation. Frame-level voiced / unvoiced (V / UV) prediction can be introduced to help the prediction of the pitch. For example, the design of the V / UV predictor and the F0 predictor module can refer to the F0 predictor in Fastspeech2.

[0260] In a possible implementation, the input first pitch feature can include a pitch feature of each frame in a plurality of frames of the non-editing speech; and correspondingly, the output second pitch feature can include a pitch feature of each frame in a plurality of frames of the target editing speech.

[0261] In step 703, a first speech feature corresponding to the second text is obtained by a neural network according to the second pitch feature and the second text.

[0262] In a possible implementation, the second pitch feature and the second text (for example, a text embedding of the second text) can be fused (for example, added) and the fusion result can be input into a neural network to obtain the first speech feature corresponding to the second text. The first speech feature corresponding to the second text can be a mel-frequency spectrum feature.

[0263] In a possible implementation, the first pitch feature of the non-editing speech, information of the target text, and the second speech feature of the non-editing speech can be used to obtain the first speech feature corresponding to the second text. The description of the second speech feature can refer to the description of the second speech feature in the above embodiments, which is not repeated here.​

[0264] In a possible implementation, after the second speech feature is acquired, the first speech feature corresponding to the second text can be obtained based on the second speech feature and the second text by using a neural network. The neural network can include an encoder and a decoder. The second text is input into the encoder to obtain a first vector corresponding to the second text, and the first vector is decoded by the decoder based on the second speech feature to obtain the first speech feature. The second speech feature can be the same as or similar to the prosody, timbre and / or signal-to-noise ratio of the first speech feature. The prosody can reflect the emotional state or speaking form of the speaker, and the prosody generally refers to features such as intonation, pitch, stress, pause and rhythm.

[0265] Optionally, an attention mechanism can be introduced between the encoder and the decoder to adjust the correspondence between the input and the output.

[0266] Optionally, the target text in which the second text is located can be introduced in the encoding process of the encoder, so that the first vector of the generated second text refers to the target text, and the second text described by the first vector is more accurate. That is, the first speech feature corresponding to the second text can be obtained based on the second speech feature, the target text and the marking information by using a neural network. Specifically, the target text and the marking information can be input into the encoder to obtain the first vector corresponding to the second text, and the first vector is decoded by the decoder based on the second speech feature to obtain the first speech feature. The marking information is used to mark the second text in the target text.

[0267] The decoder in the embodiment of the present application can be a one-way decoder or a two-way decoder, which will be described below.

[0268] First, the decoder is a one-way decoder.

[0269] The decoder calculates the speech frame obtained by the first vector or the second vector from the first direction of the target text based on the second speech feature as the first speech feature. The first direction is a direction from one side of the target text to the other side of the target text. In addition, the first direction can be understood as the forward order or the reverse order of the target text (for details, please refer to the description of the forward order and the reverse order in the foregoing Figure 5 embodiment).

[0270] Optionally, the second speech feature and the first vector are input into the decoder to obtain the first speech feature. Or the second speech feature and the second vector are input into the decoder to obtain the first speech feature.

[0271] Second, if the second text is in the middle region of the target text, the decoder can be a two-way decoder (which can also be understood as that the encoder includes a first encoder and a second encoder).

[0272] The second text above is in the middle region of the target text, which can be understood as the second text not being at both ends of the target text.

[0273] There are various cases of the bidirectional decoder in the embodiments of the present application, which are described as follows:

[0274] 1. The first speech feature output from the first direction of the bidirectional decoder is the speech feature corresponding to the second text, and the fourth speech feature output from the second direction of the bidirectional decoder is the speech feature corresponding to the second text.

[0275] In this case, it can be understood that the complete speech feature corresponding to the second text can be obtained through the left and right sides (i.e. the forward order and the reverse order) respectively, and the first speech feature is obtained according to the two speech features.

[0276] The first decoder calculates the first vector or the second vector from the first direction of the target text based on the second speech feature to obtain the first speech feature (hereinafter referred to as LR) of the second text. The second decoder calculates the first vector or the second vector from the second direction of the target text based on the second speech feature to obtain the fourth speech feature (hereinafter referred to as RL) of the second text. And the first speech feature is generated according to the first speech feature and the fourth speech feature. Wherein, the first direction is a direction from one side of the target text to the other side of the target text, and the second direction is opposite to the first direction (or it can be understood that the second direction is a direction from the other side of the target text to one side of the target text). The first direction can be the forward order described above, and the second direction can be the reverse order described above.

[0277] For the bidirectional decoder, when the first encoder decodes the first frame of the first vector or the second vector in the first direction, it can decode the N frames of LR as a condition by using the speech frames adjacent to one side (also referred to as the left side) of the second text in the non-edited speech. When the second encoder decodes the first frame of the first vector or the second vector in the second direction, it can decode the N frames of RL as a condition by using the speech frames adjacent to the other side (also referred to as the right side) of the second text in the non-edited speech. Optionally, the structure of the bidirectional decoder can refer to Figure 11 After obtaining the N frames of LR and the N frames of RL, the frames with a difference less than a threshold value in LR and RL can be used as transition frames (position m, m < n, ), or the frame with the smallest difference in LR and RL can be used as the transition frame. Then the N frames of the first speech feature can include the first m frames in LR and the last n-m frames in RL, or the N frames of the first speech feature include the first n-m frames in LR and the last m frames in RL. Wherein, the difference between LR and RL can be understood as the distance between the vector and the vector. In addition, if the speaker identifier is obtained in the foregoing step 701, the first vector or the second vector in this step can also include a third vector for identifying the speaker. It can also be understood that the third vector is used to identify the voiceprint feature of the original speech.

[0278] For example, continuing the above example, assuming that the first encoder obtains the LR frames corresponding to "Guangzhou" including LR1, LR2, LR3, LR4, the second encoder obtains the RL frames corresponding to "Guangzhou" including RL1, RL2, RL3, RL4, and the difference between LR2 and RL2 is the smallest, then LR1, LR2, RL3, RL4 or LR1, RL2, RL3, RL4 are taken as the first speech feature.

[0279] 2、The first speech feature output from the first direction by the bidirectional decoder is the speech feature corresponding to the third text in the second text, and the fourth speech feature output from the second direction by the bidirectional decoder is the speech feature corresponding to the fourth text in the second text.

[0280] In this case, it can be understood that the partial speech feature corresponding to the second text can be obtained through the left and right sides (i.e. in the forward order and the reverse order), and the complete first speech feature is obtained according to the two partial speech features. That is, a part of the speech feature is taken from the forward direction, another part of the speech feature is taken from the reverse direction, and the whole speech feature is obtained by splicing the part of the speech feature and the other part of the speech feature.

[0281] For example, continuing the above example, assuming that the first encoder obtains the LR frames corresponding to the third text ("Guang") including LR1 and LR2, and the second encoder obtains the RL frames corresponding to the fourth text ("zhou") including RL3 and RL4, then the first speech feature is obtained by splicing LR1, LR2, RL3, RL4.

[0282] It can be understood that the above two ways are only examples, and in actual application, there are other ways to obtain the first speech feature, which is not limited here.

[0283] Step 704, generating a target edited speech corresponding to the second text according to the first speech feature.

[0284] In one possible implementation, after obtaining the first speech feature, the first speech feature can be converted into a target edited speech corresponding to the second text according to a vocoder. The vocoder can be a traditional vocoder (such as Griffin-lim algorithm), or a neural network vocoder (such as Melgan or Hifigan pre-trained using audio training data), etc., which is not limited here.

[0285] For example, continuing the above example, the target edited speech corresponding to "Guangzhou" is as shown in Figure 12 .

[0286] Step 705, obtaining the position of the second text in the target text. This step is optional.

[0287] Optionally, if the original speech and the second text are obtained in step 701, the position of the second text in the target text is obtained.

[0288] Optionally, if the target text is obtained in step 701, the start and end positions of each phoneme in the original speech in the original text can be determined by the alignment technology in the aforementioned step 701, and the position of the second text in the target text is determined according to the start and end positions of each phoneme.

[0289] In step 706, the target speech corresponding to the target text is generated by splicing the target edited speech and the non-edited speech based on the position. This step is optional.

[0290] The position used for splicing the non-edited speech and the target edited speech in the embodiment of the application can be the position of the second text in the target text, the position of the first text in the target text, the position of the non-edited speech in the original speech, or the position of the edited speech in the original speech.

[0291] Optionally, after the position of the second text in the target text is obtained, the start and end positions of each phoneme in the original speech in the original text can be determined by the alignment technology in the aforementioned step 701, and the position of the non-edited speech or the edited speech in the original speech is determined according to the position of the first text in the original text. Then, the target speech is obtained by splicing the target edited speech and the non-edited speech based on the position by the speech processing device. That is, the target speech corresponding to the second text is obtained by replacing the edited region in the original speech.

[0292] For example, continuing the above example, the non-edited speech corresponds to the first frame to the fourth frame and the ninth frame to the sixteenth frame in the original speech. The target edited speech is LR1, LR2, RL3, RL4, or LR1, RL2, RL3, RL4. Splicing the target edited speech and the non-edited speech can be understood as replacing the fifth frame to the eighth frame in the original speech with the obtained four frames, and then obtaining the target speech. That is, the speech corresponding to "Guangzhou" is replaced with the speech corresponding to "Shenzhen" in the original speech, and then the target speech corresponding to "Today, the weather in Guangzhou is very good" is obtained. The target speech corresponding to "Today, the weather in Guangzhou is very good" is shown in Figure 12 .

[0293] Optionally, the speech processing device plays the target edited speech or the target speech after obtaining the target edited speech or the target speech.

[0294] In a possible implementation manner, the speech processing method provided in the embodiments of the present application includes steps 701 to 704. In another possible implementation manner, the speech processing method provided in the embodiments of the present application includes steps 701 to 705. In another possible implementation manner, the speech processing method provided in the embodiments of the present application includes steps 701 to 706. In addition, in the embodiments of the present application Figure 7a The various steps shown are not limited in the time sequence. For example, step 705 in the above method can also be after step 704, before step 701, or jointly executed with step 701.

[0295] The embodiments of the present application provide a speech processing method, including: obtaining an original speech and a second text, the second text being a text other than a first text in a target text, the target text and an original text corresponding to the original speech both including the first text, the first text corresponding to a non-edited speech in the original speech; predicting a second pitch feature of the second text according to a first pitch feature of the non-edited speech and information of the target text; obtaining a first speech feature corresponding to the second text through a neural network according to the second pitch feature and the second text; and generating a target edited speech corresponding to the second text according to the first speech feature. The present application predicts the pitch feature of the second text (to-be-edited text), generates the first speech feature of the second text according to the pitch feature, and generates the target edited speech corresponding to the second text based on the first speech feature, so that the pitch features of the speech before and after the singing voice editing are similar, and the hearing sensation of the target edited speech is similar to that of the original speech.

[0296] Next, the speech processing method in the embodiments of the present application is introduced in combination with an example:

[0297] Taking the singing voice editing scenario as an example, the original to-be-edited singing voice W (wherein the speech content is S "Love can not ask for right or wrong") and the following three different target speeches are taken as examples respectively:

[0298] The editing request Q1 is that the target speech is W1 (the speech content corresponds to the text T1 "Love how can not ask for right or wrong"),

[0299] The editing request Q2 is that the target speech is W2 (the speech content corresponds to the text T2 "Love does not ask for right or wrong"),

[0300] The editing request Q3 is that the target speech is W3 (the speech content corresponds to the text T2 "Love how does not ask for right or wrong")

[0301] Step S1: receiving a "speech editing" request of a user;

[0302] The request includes at least original to-be-edited voice W, original lyrics text S, target text T (T1 or T2 or T3) and the like, and the pre-operation includes: comparing the original text S and the target text to determine the editing type of the current editing request, that is, for Q1, Q2 and Q3, it can be determined that they are insertion, deletion and replacement operations respectively; extracting the audio features and the pitch features of each frame from W; extracting the Singer embedding from W through a voiceprint model; converting S and the target text T* into phoneme representation, such as T2, whose phoneme sequence is [ai4 b u2 w en4 d ui4 c cuo4]; extracting the time length (i.e., the frame number) corresponding to each phoneme in S according to W and S; determining the Mask region according to the operation type, for Q1, which is an insertion operation (inserting the word “how”), the target Mask phoneme is the phoneme corresponding to “how”, i.e., the final target text phoneme of Q1 is [ai4 z en3 m e5 k e2 yi3 b u2 w en4 d ui4 c cuo4]; (where the red part represents the Masked phoneme), for Q2, which is a deletion operation (deleting the word “can”), the target phoneme is the phoneme of the word adjacent to “can” in S; i.e., the final target text phoneme of Q2 is [ai4 b u2 w en4 d ui4 c cuo4]; (where the red part represents the Masked phoneme); for Q3, which is a replacement operation (replacing “can” with “how”), so the target text phoneme is [ai4 z en3 m e5 b u2 w en4 d ui4 c cuo4]; (where the red part represents the Masked phoneme);

[0303] Step S2: The target text phoneme obtained in S1 generates text features, i.e., phoneme-level text embedding, through a text encoding module;

[0304] Step S3: The time length information of each phoneme in the target text is predicted through a time length normalization module; this step can be completed through the following sub-steps:

[0305] Generate a Mask vector and a reference time length vector according to the Mask label of the phoneme: for a non-Mask phoneme, its reference time length is the real time length extracted in step S1, otherwise it is set to 0; for a non-Mask phoneme, the corresponding position in the Mask vector is set to 1, otherwise it is set to 0;

[0306] Take the Text Embedding, Singer Embedding, reference time length vector and Mask vector as inputs, and use the time length prediction module shown in Figure 2-2 to predict the time length corresponding to the Mask phoneme

[0307] According to the duration corresponding to each phoneme, the Embedding of each phoneme is up-sampled (i.e. if the duration of phoneme A is 10, the Embedding of A is copied 10 times), thereby generating Frame-level Text Embedding;

[0308] Step S4: predicting the Pitch value of each frame through the Pitch prediction module, which can be completed through the following sub-steps:

[0309] For Q1 and Q2, the pitch of the frame corresponding to the Mask phoneme is predicted using the model shown in Figure 2-3:

[0310] Wherein, for non-Mask phonemes, the reference pitch is the real Pitch extracted in S1, and it is marked as 1 at the position corresponding to the Mask vector; for Mask phonemes, the pitch on the corresponding frame is set to 0, and the Mask is set to 0; the Frame-level pitch of the Mask phoneme is predicted.

[0311] For the replacement operation Q3, the Frame-level Pitch of the Mask phoneme is predicted using the model shown in Figure 2-4;

[0312] Step S5: adding the Frame-Level text Embedding and the Pitch together to input into the audio feature decoding module, and predicting the audio feature frame corresponding to the new Mask phoneme.

[0313] It should be understood that if multiple editing operations are involved in one editing request, the above-mentioned process can be used one by one according to the processing order from left to right. On the other hand, one replacement operation can also be realized by two operations of "deleting first and then inserting".

[0314] The above describes the speech processing method implemented by the terminal device or the cloud device alone, and the following describes the speech processing method implemented by the terminal device and the cloud device together.

[0315] Embodiment two: the terminal device and the cloud device implement the speech processing method together.

[0316] Please refer to Figure 13 An embodiment of the speech processing method provided by the present application can be implemented by the terminal device and the cloud device together, or by the components (such as processors, chips, or chip systems, etc.) of the terminal device and the components (such as processors, chips, or chip systems, etc.) of the cloud device, and the embodiment includes steps 1301 to 1306.

[0317] At step 1301, the terminal device acquires the original speech and the second text.

[0318] The step 1301 performed by the terminal device in this embodiment is similar to the step 1301 in the foregoing Figure 7a The step 701 performed by the speech processing device in the embodiment shown in the table is similar, and will not be described here again.

[0319] At step 1302, the terminal device sends the original speech and the second text to the cloud device.

[0320] After the terminal device acquires the original speech and the second text, the terminal device can send the original speech and the second text to the cloud device.

[0321] Optionally, if the terminal device acquires the original speech and the target text at step 1301, the terminal device sends the original speech and the target text to the cloud device.

[0322] At step 1303, the cloud device acquires the non-edited speech based on the original speech and the second text.

[0323] The step 1303 performed by the cloud device in this embodiment is similar to the step 1303 in the foregoing Figure 7a The description of determining the non-edited speech in the step 701 performed by the speech processing device in the embodiment shown in the table is similar, and will not be described here again.

[0324] At step 1304, the cloud device acquires the second pitch feature of the second text based on the first pitch feature of the non-edited speech and the information of the target text.

[0325] The step 1303 performed by the cloud device in this embodiment is similar to the step 1303 in the foregoing Figure 7a The description of determining the non-edited speech in the step 702 performed by the speech processing device in the embodiment shown in the table is similar, and will not be described here again.

[0326] At step 1305, the cloud device obtains the first speech feature corresponding to the second text by a neural network based on the second pitch feature and the second text.

[0327] At step 1306, the cloud device generates the target edited speech corresponding to the second text based on the first speech feature.

[0328] The steps 1304 to 1306 performed by the cloud device in this embodiment are similar to the steps 702 to 704 in the foregoing Figure 7a The steps 702 to 704 performed by the speech processing device in the embodiment shown in the table are similar, and will not be described here again.

[0329] At step 1307, the cloud device sends the target edited speech to the terminal device. This step is optional.

[0330] Optionally, after the cloud device acquires the target edited speech, the cloud device can send the target edited speech to the terminal device.

[0331] In step 1308, the terminal device or the cloud device acquires the position of the second text in the target text. This step is optional.

[0332] In step 1309, the terminal device or the cloud device splices the target edited speech and the non-edited speech based on the position to generate the target speech corresponding to the target text. This step is optional. This step is optional.

[0333] The steps 1308 and 1309 in this embodiment are similar to the steps 1308 and 1309 in the foregoing Figure 7a The steps 705 to 706 performed by the speech processing device in the illustrated embodiment are similar, and thus are not described herein again. The steps 1308 and 1309 in this embodiment can be performed by the terminal device or the cloud device.

[0334] In step 1310, the cloud device sends the target speech to the terminal device. This step is optional.

[0335] Optionally, if the steps 1308 and 1309 are performed by the cloud device, the cloud device sends the target speech to the terminal device after acquiring the target speech. If the steps 1308 and 1309 are performed by the terminal device, this step can not be performed.

[0336] Optionally, the terminal device plays the target edited speech or the target speech after acquiring the target edited speech or the target speech.

[0337] In one possible implementation, the speech processing method provided by the embodiments of the present application can include that the cloud device generates the target edited speech and sends the target edited speech to the terminal device, that is, the method includes the steps 1301 to 1307. In another possible implementation, the speech processing method provided by the embodiments of the present application can include that the cloud device generates the target edited speech and generates the target speech according to the target edited speech and the non-edited speech, and sends the target speech to the terminal device. That is, the method includes the steps 1301 to 1306, the steps 1308 to 1310. In another possible implementation, the speech processing method provided by the embodiments of the present application can include that the cloud device generates the target edited speech and sends the target edited speech to the terminal device. The terminal device generates the target speech according to the target edited speech and the non-edited speech. That is, the method includes the steps 1301 to 1309.

[0338] In an embodiment of the present application, on the one hand, the cloud device can perform complex calculations to obtain the target edited voice or the target voice through the interaction between the cloud device and the terminal device, and return the target edited voice or the target voice to the terminal device, so as to reduce the computing power and storage space of the terminal device. On the other hand, the target edited voice corresponding to the modified text can be generated according to the voice features of the non-edited region in the original voice, and then the target voice corresponding to the target text is generated with the non-edited voice. On the other hand, the user can modify the text in the original text to obtain the target edited voice corresponding to the modified text (i.e., the second text). The editing experience of the user based on text editing is improved. On the other hand, the non-edited voice is not modified when generating the target voice, and the pitch features of the target edited voice are similar to the pitch features of the non-edited voice, so that the user can hardly hear the difference between the original voice and the target voice in voice features when listening to the original voice and the target voice.

[0339] The above describes the voice processing method in the embodiment of the present application, and the voice processing device in the embodiment of the present application is described below. Please refer to Figure 14 One embodiment of the voice processing device in the embodiment of the present application includes:

[0340] The acquisition module 1401 is configured to acquire an original voice and a second text, wherein the second text is a text in a target text except for a first text, the target text and an original text corresponding to the original voice both include the first text, and a voice corresponding to the first text in the original voice is a non-edited voice.

[0341] Specific description of the acquisition module 1401 can be referred to the description of step 701 in the above embodiment, which is not repeated here.

[0342] The pitch prediction module 1402 is configured to predict a second pitch feature of the second text according to a first pitch feature of the non-edited voice and information of the target text.

[0343] Specific description of the pitch prediction module 1402 can be referred to the description of step 702 in the above embodiment, which is not repeated here.

[0344] The generation module 1403 is configured to obtain a first voice feature corresponding to the second text by a neural network according to the second pitch feature and the second text.

[0345] According to the first voice feature, a target edited voice corresponding to the second text is generated.

[0346] Specific description of the generation module 1403 can be referred to the description of steps 703 and 704 in the above embodiment, which is not repeated here.

[0347] In a possible implementation, the content of the original speech is a user's singing.

[0348] In a possible implementation, the first pitch feature of the non-edited speech and the information of the target text include:

[0349] the first pitch feature of the non-edited speech, the information of the target text, and the second speech feature of the non-edited speech; the second speech feature carries at least one of the following information:

[0350] part or all of speech frames of the non-edited speech;

[0351] a voiceprint feature of the non-edited speech;

[0352] a timbre feature of the non-edited speech;

[0353] a prosody feature of the non-edited speech; and

[0354] a rhythm feature of the non-edited speech.

[0355] In a possible implementation, the information of the target text includes: text embeddings of each phoneme in the target text.

[0356] In a possible implementation, the target text is a text obtained by inserting the second text into the first text; or the target text is a text obtained by deleting a first part of text in the first text, and the second text is text adjacent to the first part of text.

[0357] The pitch prediction module is specifically configured to:

[0358] fuse the first pitch feature of the non-edited speech and the information of the target text to obtain a first fusion result;

[0359] input the first fusion result into a second neural network to obtain a second pitch feature of the second text.

[0360] In a possible implementation, the target text is obtained by replacing a second part of text in the first text with the second text.

[0361] The pitch prediction module is specifically configured to:

[0362] input the first pitch feature of the non-edited speech into a third neural network to obtain an initial pitch feature, and the first initial pitch feature includes a pitch of each frame in a plurality of frames.

[0363] inputting the information of the target text into a fourth neural network to obtain a pronunciation feature of the second text, the pronunciation feature being used to indicate whether each frame included in the initial pitch feature is pronounced or not.

[0364] fusing the initial pitch feature and the pronunciation feature to obtain a second pitch feature of the second text.

[0365] In a possible implementation, the apparatus further includes:

[0366] a duration prediction module configured to predict a frame number of each phoneme in the second text according to the frame number of each phoneme in the non-edited speech and the information of the target text.

[0367] In a possible implementation, the first pitch feature includes a pitch feature of each frame in the multiple frames of the non-edited speech.

[0368] The second pitch feature includes a pitch feature of each frame in the multiple frames of the target edited speech.

[0369] In a possible implementation, the duration prediction module is specifically configured to:

[0370] according to the frame number of each phoneme in the non-edited speech, the information of the target text, and a second speech feature of the non-edited speech.

[0371] In a possible implementation, the obtaining module is further configured to:

[0372] obtain a position of the second text in the target text.

[0373] The generating module is further configured to splice the target edited speech and the non-edited speech based on the position to obtain a target speech corresponding to the target text.

[0374] For example, the speech processing device can be a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), a vehicle-mounted computer, or any terminal device. Figure 15 For example, the speech processing device can be a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), a vehicle-mounted computer, or any terminal device.

[0375] Figure 15 The figure shows a block diagram of part of the structure of a mobile phone related to the speech processing device provided by the embodiments of the present application. For example, the speech processing device can be a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), a vehicle-mounted computer, or any terminal device. Figure 15The mobile phone includes radio frequency (RF) circuit 1510, memory 1520, input unit 1530, display unit 1540, sensor 1550, audio circuit 1560, wireless fidelity (WiFi) module 1570, processor 1580, and power supply 1590, etc. Those skilled in the art can understand that Figure 15 The mobile phone structure shown in the figure does not constitute a limitation on the mobile phone, and can include more or less components than the figure, or combine certain components, or different component arrangements.

[0376] The following will be described in detail Figure 15 The various components of the mobile phone will be specifically introduced:

[0377] RF circuit 1510 can be used for receiving and sending signals in the process of information or call, especially, receiving the downlink information of the base station and processing by processor 1580; in addition, sending the uplink data to the base station. Usually, RF circuit 1510 includes but is not limited to antenna, at least one amplifier, transceiver, coupler, low noise amplifier (LNA), duplexer, etc. In addition, RF circuit 1510 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to global system for mobile communication (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), long term evolution (LTE), email, short message service (SMS), etc.

[0378] The memory 1520 can be used to store software programs and modules, and the processor 1580 can execute various function applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1520. The memory 1520 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), and the like; and the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), and the like. In addition, the memory 1520 can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.

[0379] The input unit 1530 can be used to receive inputted digital or character information, and to generate key signal input related to the user settings and function control of the mobile phone. Specifically, the input unit 1530 can include a touch panel 1531 and other input devices 1532. The touch panel 1531, also called a touch screen, can collect the touch operation of a user thereon or nearby (such as the operation of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 1531), and drive the corresponding connection device according to the pre-set program. Optionally, the touch panel 1531 can include two parts of a touch detection device and a touch controller. The touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, and converts it into touch coordinates, and then sends it to the processor 1580, and can also receive the command from the processor 1580 and execute it. In addition, the touch panel 1531 can be implemented in various types such as a resistive type, a capacitive type, an infrared type, and a surface acoustic wave type. In addition to the touch panel 1531, the input unit 1530 can also include other input devices 1532. Specifically, the other input devices 1532 can include one or more of a physical keyboard, a function key (such as a volume control key, an on-off key, etc.), a trackball, a mouse, a joystick, and the like.

[0380] The display unit 1540 can be used to display information input by a user or information provided to the user as well as various menus of the phone. The display unit 1540 can include a display panel 1541, which can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like, as an option. Further, a touch panel 1531 can cover the display panel 1541, and when the touch panel 1531 detects a touch operation thereon or nearby, it transmits to the processor 1580 to determine the type of touch event, and then the processor 1580 provides a corresponding visual output on the display panel 1541 according to the type of touch event. Although in the above description, the touch panel 1531 and the display panel 1541 are implemented as two independent components to realize the input and output functions of the phone, in some embodiments, the touch panel 1531 and the display panel 1541 can be integrated to realize the input and output functions of the phone. Figure 15

[0381] The phone can also include at least one sensor 1550, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor can include an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 1541 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1541 and / or the backlight when the phone is moved to the ear. As one of the motion sensors, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes), and when at rest, it can detect the magnitude and direction of gravity, which can be used for applications that identify the posture of the phone (such as landscape / portrait screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), and the like. As for other sensors that the phone can also be configured, such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, and the like, they will not be described here.

[0382] The audio circuit 1560, the speaker 1561, and the microphone 1562 can provide an audio interface between the user and the phone. The audio circuit 1560 can convert received audio data into an electrical signal, transmit it to the speaker 1561, and convert it into a sound signal output by the speaker 1561; on the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560 and converted into audio data, which is then output to the processor 1580 for processing, and then transmitted to another phone, for example, via the RF circuit 1510, or output to the memory 1520 for further processing.

[0383] ​WiFi is a short-range wireless transmission technology. The WiFi module 1570 can help users send and receive emails, browse web pages, and access streaming media, etc. It provides users with wireless broadband Internet access. Figure 15 The WiFi module 1570 is shown, but it is understood that it is not a necessary component of the mobile phone.

[0384] The processor 1580 is the control center of the mobile phone. It connects all parts of the mobile phone through various interfaces and lines, executes various functions of the mobile phone and processes data by running or executing software programs and / or modules stored in the memory 1520 and calling data stored in the memory 1520, thereby overall monitoring the mobile phone. Optionally, the processor 1580 can include one or more processing units; preferably, the processor 1580 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application program, etc., and the modem processor mainly processes wireless communication. It is understood that the above-mentioned modem processor can also be integrated into the processor 1580.

[0385] The mobile phone also includes a power supply 1590 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 1580 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management, etc. through the power management system.

[0386] Although not shown, the mobile phone can also include a camera, a Bluetooth module, etc., which will not be described here.

[0387] In the embodiments of the present application, the processor 1580 included in the terminal device can execute the functions of the voice processing device in the foregoing Figure 7a embodiments, or execute the functions of the terminal device in the foregoing Figure 13 embodiments, which will not be described here.

[0388] Referring to Figure 16 , another structural schematic diagram of a voice processing device is provided. The voice processing device can be a cloud device. The cloud device can include a processor 1601, a memory 1602, and a communication interface 1603. The processor 1601, the memory 1602, and the communication interface 1603 are interconnected through lines. The memory 1602 stores program instructions and data.

[0389] The memory 1602 stores the program instructions and data corresponding to the steps executed by the voice processing device in the foregoing Figure 7a corresponding embodiments. Or the memory 1602 stores the program instructions and data corresponding to the steps executed by the cloud device in the foregoing Figure 13 corresponding embodiments.

[0390] The processor 1601 is configured to perform the steps shown in any of the embodiments described above. Figure 7a The processor 1601 is configured to perform the steps shown in any of the embodiments described above. Figure 13 The processor 1601 is configured to perform the steps shown in any of the embodiments described above.

[0391] The communication interface 1603 can be configured to receive and send data, and perform the steps related to obtaining, sending, and receiving in any of the embodiments described above. Figure 7a Or Figure 13 The communication interface 1603 can be configured to receive and send data, and perform the steps related to obtaining, sending, and receiving in any of the embodiments described above.

[0392] In an implementation manner, the cloud device can include more or fewer components, and the present application is only illustrative and not limited. Figure 16 More or fewer components, and the present application is only illustrative and not limited.

[0393] In the several embodiments provided by the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative, and the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0394] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0395] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized by software, hardware, firmware or any combination thereof, in whole or in part.

[0396] When the units are implemented by using software, the units can be totally or partially realized in a form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are totally or partially generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatuses. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center through wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (SSD)), etc.

[0397] The terms "first", "second", and the like in the description and in the claims of the present application and above-described drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attribute used in the description of the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or apparatus including a series of units does not necessarily limit to those units, but can include other units not clearly listed or inherent to these processes, methods, products or apparatuses.

Claims

1. A voice processing method, characterized by, The method comprises: obtaining original speech and second text, the second text being text other than the first text in target text, the original text corresponding to the original speech including the first text, the speech corresponding to the first text in the original speech being non-edited speech, and the content of the original speech being a user's singing voice; predicting a second pitch feature of the second text according to a first pitch feature of the non-edited speech and information of the target text; obtaining a first speech feature corresponding to the second text through a neural network according to the second pitch feature and the second text; generating target edited speech corresponding to the second text according to the first speech feature; in a case where the target text is obtained by replacing a second part of text in the first text with the second text, the predicting a second pitch feature of the second text according to a first pitch feature of the non-edited speech and information of the target text comprises: inputting the first pitch feature of the non-edited speech into a third neural network to obtain an initial pitch feature, the initial pitch feature including a pitch of each frame in a plurality of frames; inputting the information of the target text into a fourth neural network to obtain a pronunciation feature of the second text, the pronunciation feature being used to indicate whether each frame in the plurality of frames included in the initial pitch feature pronounces; fusing the initial pitch feature and the pronunciation feature to obtain the second pitch feature of the second text.

2. The method of claim 1, wherein, The predicting a second pitch feature of the second text according to a first pitch feature of the non-edited speech and information of the target text comprises: according to the first pitch feature of the non-edited speech, the information of the target text, and a second speech feature of the non-edited speech, the second speech feature carrying at least one of the following information: part or all of speech frames of the non-edited speech; a voiceprint feature of the non-edited speech; a timbre feature of the non-edited speech; a prosody feature of the non-edited speech; and a rhythm feature of the non-edited speech.

3. The method of claim 1, wherein, The information of the target text comprises: text embedding of each phoneme in the target text.

4. The method according to any one of claims 1 to 3, characterized in that, The target text is text obtained by inserting the second text into the first text, or the target text is text obtained by deleting a first part of text of the first text, the second text being text adjacent to the first part of text; The predicting a second pitch feature of the second text according to a first pitch feature of the non-edited speech and information of the target text comprises: fusing the first pitch feature of the non-edited speech and the information of the target text to obtain a first fusion result; inputting the first fusion result into a second neural network to obtain the second pitch feature of the second text.

5. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: According to the frame number of each phoneme in the non-edited speech and information of the target text, frame numbers of each phoneme in the second text are predicted.

6. The method according to any one of claims 1 to 3, characterized in that, The first pitch feature includes a pitch feature of each frame in multiple frames of the non-edited speech. The second pitch feature includes a pitch feature of each frame in multiple frames of the target edited speech.

7. The method of claim 5, wherein, The information of the target text includes a text embedding of each phoneme in the target text. According to the frame number of each phoneme in the non-edited speech, information of the target text, and a second speech feature of the non-edited speech.

8. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtaining a position of the second text in the target text; Splicing the target edited speech and the non-edited speech based on the position to obtain a target speech corresponding to the target text.

9. A speech processing device, characterized by The device includes: An obtaining module is configured to obtain an original speech and a second text, the second text being a text other than a first text in a target text, the target text and an original text corresponding to the original speech both including the first text, a speech corresponding to the first text in the original speech being a non-edited speech, and a content of the original speech being a user's singing voice. A pitch prediction module is configured to predict a second pitch feature of the second text according to a first pitch feature of the non-edited speech and information of the target text. A generating module is configured to obtain, by using a neural network, a first speech feature corresponding to the second text according to the second pitch feature and the second text. According to the first speech feature, a target edited speech corresponding to the second text is generated. In a case where the target text is obtained by replacing a second part of text in the first text with the second text, the pitch prediction module is specifically configured to: input the first pitch feature of the non-edited speech into a third neural network to obtain an initial pitch feature, the initial pitch feature including a pitch of each frame in multiple frames; input information of the target text into a fourth neural network to obtain a pronunciation feature of the second text, the pronunciation feature being used to indicate whether each frame in the multiple frames included in the initial pitch feature pronounces; fuse the initial pitch feature and the pronunciation feature to obtain the second pitch feature of the second text.

10. The apparatus of claim 9, wherein, According to the first pitch feature of the non-edited speech and the second text, the following is included: According to the first pitch feature of the non-edited speech, information of the target text, and a second speech feature of the non-edited speech, the second speech feature carrying at least one of the following information: part or all of speech frames of the non-edited speech; a voiceprint feature of the non-edited speech; a timbre feature of the non-edited speech; a prosody feature of the non-edited speech; and a rhythm feature of the non-edited speech.

11. The apparatus of claim 9, wherein, The information of the target text includes a text embedding of each phoneme in the target text.

12. The apparatus of any one of claims 9 to 11, wherein, The target text is a text obtained by inserting the second text into the first text, or the target text is a text obtained by deleting a first part of the first text, and the second text is text adjacent to the first part of the first text. The pitch prediction module is specifically configured to: fuse the first pitch feature of the non-edited speech and information of the target text to obtain a first fusion result; input the first fusion result into a second neural network to obtain a second pitch feature of the second text.

13. The apparatus of any one of claims 9 to 11, wherein, The apparatus further includes: a duration prediction module configured to predict a frame number of each phoneme in the second text according to the frame number of each phoneme in the non-edited speech and information of the target text.

14. The apparatus of any one of claims 9 to 11, wherein, The first pitch feature includes a pitch feature of each frame in multiple frames of the non-edited speech. The second pitch feature includes a pitch feature of each frame in multiple frames of the target edited speech.

15. The apparatus of claim 13, wherein, The duration prediction module is specifically configured to: predict the frame number of each phoneme in the second text according to the frame number of each phoneme in the non-edited speech, the information of the target text, and a second speech feature of the non-edited speech.

16. The apparatus of any one of claims 9 to 11, wherein, The obtaining module is further configured to: obtain a position of the second text in the target text; The generation module is further configured to splice the target edited speech and the non-edited speech based on the position to obtain a target speech corresponding to the target text.

17. A speech processing device, characterized by includes: a processor coupled with a memory, the memory being configured to store programs or instructions, when the programs or instructions are executed by the processor, the processor is caused to execute the method in any one of claims 1 to 8.

18. The apparatus of claim 17, wherein, The apparatus further includes: an input unit configured to receive a second text; an output unit configured to play a target edited speech corresponding to the second text or a target speech corresponding to a target text.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, when the instructions are executed on a computer, the computer is caused to execute the method in any one of claims 1 to 8.

20. A computer program product, characterised in that, The computer program product, when executed on a computer, causes the computer to execute the method in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voice processing method and related equipment

    CN113421547A

  • Speech synthesis model, model training method and speech synthesis method

    CN113920977A