An artificial intelligence-based sound conversion method, device, equipment and medium

By extracting pitch features from the fundamental frequency of the source speech and extracting timbre and emotion features from the Mel spectrum of the target speech, and then performing frame expansion and fusion, the problem of poor speech conversion effect in the past has been solved, and a better speech conversion effect has been achieved.

CN119559956BActive Publication Date: 2026-05-01PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2024-11-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing speech conversion methods have poor conversion results when transferring one person's voice to another person's speech.

Method used

Pitch features are extracted by obtaining the fundamental frequency of the source speaker, and timbre and emotion features are extracted by obtaining the Mel spectrum information of the target speaker. The number of frames of timbre and emotion features is expanded to be equal to that of pitch features before fusion. The converted speech is obtained by combining the source speech content with the spectrum reconstruction.

Benefits of technology

It improves the speech conversion effect, making the converted speech better reflect the voice characteristics of the target speaker.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559956B_ABST
    Figure CN119559956B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, and more particularly to a voice conversion method, device and equipment based on artificial intelligence and a medium. The method comprises the following steps: obtaining a pitch feature of a source speaker, determining the frame number of the pitch feature, obtaining a timbre feature and an emotion feature of a target speaker, expanding the frame number of the timbre feature and the emotion feature to be equal to the frame number of the pitch feature, aligning and fusing the pitch feature, the expanded timbre feature and the expanded emotion feature to obtain a first fusion feature, extracting source speech content, fusing the source speech content and the first fusion feature to obtain a second fusion feature, and obtaining converted speech according to the second fusion feature. In the present application, the emotion information and the pitch information corresponding to the target speaker are fused into the content information of the source speaker in the process of voice conversion, so that the converted target voice can better reflect the voice of the target speaker, thereby improving the voice conversion effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a sound conversion method, apparatus, device, and medium based on artificial intelligence. Background Technology

[0002] Speech conversion is a speech technology that preserves the content information of the source speaker's voice and converts it into the voice of the target speaker. This technology has a wide range of applications, such as "voice-changing bowties." Furthermore, the development of speech conversion technology is of great significance to fields such as personalized speech synthesis, voiceprint recognition, and voiceprint security. Current speech conversion methods unemotionally transfer one person's voice to another's speech content, resulting in poor conversion quality. Therefore, improving the conversion effect is a pressing issue that needs to be addressed. Summary of the Invention

[0003] In view of this, embodiments of this application provide a voice conversion method, apparatus, device, and medium based on artificial intelligence to solve the problem of poor voice quality after conversion during the voice conversion process.

[0004] In a first aspect, embodiments of this application provide an artificial intelligence-based voice conversion method, the voice conversion method comprising:

[0005] Obtain the source speech of the source speaker, extract the fundamental frequency from the source speech, extract the pitch features from the fundamental frequency to obtain the pitch features, and determine the frame number of the pitch features;

[0006] The target speech of the target speaker is obtained, the Mel spectrum information in the target speech is extracted to obtain the target Mel spectrum, the timbre features are extracted from the target Mel spectrum to obtain the timbre features, and the emotional features are extracted from the target Mel spectrum to obtain the emotional features.

[0007] The number of frames for both the timbre feature and the emotional feature is expanded to be equal to the number of frames for the pitch feature, resulting in expanded timbre features and expanded emotional features.

[0008] The pitch feature, the expanded timbre feature, and the expanded emotional feature are aligned and then fused to obtain the first fused feature;

[0009] The source speech is extracted to obtain the source speech content. The source speech content is then fused with the first fusion feature to obtain the second fusion feature. The second fusion feature is then spectrally reconstructed to obtain the reconstructed Mel spectrum. Based on the reconstructed Mel spectrum, the converted speech is obtained.

[0010] Secondly, embodiments of this application provide an artificial intelligence-based voice conversion device, the voice conversion device comprising:

[0011] The first acquisition module is used to acquire the source speech of the source speaker, extract the fundamental frequency from the source speech, extract the pitch features from the fundamental frequency to obtain the pitch features, and determine the frame number of the pitch features;

[0012] The second acquisition module is used to acquire the target speech of the target speaker, extract the Mel spectrum information in the target speech to obtain the target Mel spectrum, extract timbre features from the target Mel spectrum to obtain timbre features, and extract emotional features from the target Mel spectrum to obtain emotional features.

[0013] An expansion module is used to expand the number of frames of both the timbre feature and the emotional feature to be equal to the number of frames of the pitch feature, so as to obtain expanded timbre features and expanded emotional features.

[0014] The fusion module is used to align and fuse the pitch feature, the expanded timbre feature, and the expanded emotional feature to obtain a first fused feature;

[0015] The reconstruction module is used to extract the speech content from the source speech to obtain the source speech content, fuse the source speech content with the first fusion feature to obtain the second fusion feature, reconstruct the spectrum of the second fusion feature to obtain the reconstructed Mel spectrum, and obtain the converted speech based on the reconstructed Mel spectrum.

[0016] Thirdly, embodiments of this application provide a terminal device, the terminal device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the artificial intelligence-based voice conversion method as described in the first aspect.

[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the artificial intelligence-based voice conversion method as described in the first aspect.

[0018] The advantages of this invention compared to the prior art are:

[0019] The process involves: acquiring the source speaker's speech, extracting the fundamental frequency (FFM) from the source speech, extracting pitch features from the FFM, determining the number of frames for the pitch features, acquiring the target speaker's speech, extracting Mel spectrum information from the target speech, obtaining the target Mel spectrum, extracting timbre features from the target Mel spectrum, extracting emotion features from the target Mel spectrum, expanding the number of frames for both timbre and emotion features to be equal to the number of frames for the pitch features, obtaining expanded timbre and expanded emotion features, aligning and fusing the pitch features, obtaining the first fused feature, extracting the speech content from the source speech, fusing the source speech content with the first fused feature, obtaining the second fused feature, reconstructing the spectrum of the second fused feature, obtaining the reconstructed Mel spectrum, and obtaining the converted speech based on the reconstructed Mel spectrum. In this application, during the speech conversion process, the emotional information and pitch information of the target speaker are integrated into the content information of the source speaker, so that the converted target speech can better reflect the voice of the target speaker, thereby improving the speech conversion effect. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of an application environment for an artificial intelligence-based voice conversion method provided in an embodiment of this application;

[0022] Figure 2 This is a flowchart illustrating an artificial intelligence-based voice conversion method provided in an embodiment of this application;

[0023] Figure 3 This is a schematic diagram of the structure of an artificial intelligence-based voice conversion device provided in an embodiment of this application;

[0024] Figure 4 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0027] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0028] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0029] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0030] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0031] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0032] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0033] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0034] This application provides an artificial intelligence-based voice conversion method that can be applied to, for example... Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, smart TVs, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0035] See Figure 2 This is a flowchart illustrating an artificial intelligence-based voice conversion method provided in an embodiment of this application. The aforementioned artificial intelligence-based voice conversion method is applied to the aforementioned server. Figure 2 As shown, this AI-based voice conversion method may include the following steps:

[0036] S201: Obtain the source speech of the source speaker, extract the fundamental frequency from the source speech, extract the pitch features from the fundamental frequency, obtain the pitch features, and determine the number of frames for the pitch features.

[0037] In step S201, the source speech to be converted is obtained, the pitch features in the source speech are extracted, and the number of frames of the pitch features is determined so as to expand the timbre and emotional features of the target speaker according to the number of frames of the pitch features.

[0038] In this embodiment, the source speech of the speaker can be acquired in a voice interaction system, and the fundamental frequency (FFM) can be extracted from the source speech to obtain a continuous FFM curve. Specifically, the source speech is initially extracted, and its spectrum is extracted. Harmonic fringes on the spectrum are used to correct the harmonics and half-frequency components generated during the extraction process. For parts where the fundamental frequency cannot be extracted, it is extracted using a fundamental frequency extraction method. Then, spline functions are used to interpolate the positions where the fundamental frequency is missing, thereby obtaining a continuous FFM curve. Here, the initial extraction of the source speech can be performed using any non-frequency domain algorithm, such as Praat's autocorrelation method, the AMDF algorithm, the YIN algorithm, and fundamental frequency recognition algorithms based on statistical models. Praat is a software name, and AMDF stands for Average Magnitude Difference Function. Based on the corresponding fundamental frequency curve, features are extracted from the fundamental frequency to obtain the corresponding pitch features.

[0039] When determining the number of frames for pitch features, the continuous fundamental frequency curve can be divided into equal-length segments to obtain the corresponding number of frames. For example, a segmentation step size can be set, and the number of frames for pitch features can be determined based on the duration of the continuous fundamental frequency curve and the set step size.

[0040] Optionally, pitch features are extracted from the fundamental frequency to obtain pitch features, including:

[0041] Use a preset vocoder to extract the fundamental frequency from the source speech;

[0042] The fundamental frequency is logarithmically normalized to obtain normalized features, which are then used as pitch features.

[0043] In this embodiment, the vocoder can be a world vocoder. After the source speech is input to the world vocoder, the world vocoder can use the DIO algorithm to analyze the source speech, thereby obtaining the fundamental frequency information corresponding to the source speech. The fundamental frequency is processed to obtain the corresponding pitch features. During processing, the logarithm of the fundamental frequency is taken and then normalized to obtain normalized features. The normalized features are determined as pitch features.

[0044] It should be noted that the vocoder in the embodiments of this application can also be a neural network vocoder, including but not limited to autoregressive neural network vocoders such as waveRNN vocoders, and Gann network-based neural network vocoders such as melGan vocoders and hifiGan vocoders.

[0045] S202: Obtain the target speech of the target speaker, extract the Mel spectrum information in the target speech to obtain the target Mel spectrum, extract the timbre features from the target Mel spectrum to obtain the timbre features, and extract the emotion features from the target Mel spectrum to obtain the emotion features.

[0046] In step S202, the target speech of the target speaker is the speech information that the source speaker needs to convert. Mel spectrum information is extracted from the target speech, and the spectrum information is scaled to a coordinate system similar to human ear perception features to facilitate better analysis and processing. Timbre features and emotional features are extracted from the target Mel spectrum to obtain the corresponding timbre features and emotional features.

[0047] In this embodiment, a web crawler can be written to selectively crawl data after setting up a data source, thereby obtaining the target speech data. The data source can be various types of online platforms, social media, or specific audio databases, etc. The target speech data can be the target speaker's music, speeches, chat conversations, etc. Target speech data can also be obtained through other methods, and is not limited to these.

[0048] When extracting Mel-frequency spectrum information from target speech, the target speech is first converted into its corresponding Fourier transform spectrum, and then a preset function is used to convert it into a Mel-frequency spectrum that better matches human hearing. This transforms one-dimensional, difficult-to-process time-series signals into easily processed and more information-rich two-dimensional frequency domain data. Mel-frequency coefficients in Mel-frequency cepstrum are a set of key coefficients used to construct the Mel-frequency cepstrum. From segments of the target speech, a spectrum sufficient to represent the target speech can be obtained, and the Mel-frequency cepstrum coefficients are the spectrum derived from this spectrum (i.e., the spectrum of the spectrum). Unlike ordinary spectra, the most distinctive feature of Mel-frequency cepstrum is that the frequency bands on the Mel-frequency spectrum are uniformly distributed across the Mel scale. In other words, compared to the generally seen linear spectrum representation, such frequency bands are closer to the non-linear human auditory system. For example, Mel-frequency spectrum is often used in audio compression technology.

[0049] Phonographic features are extracted from the target Mel spectrum. When extracting timbre features, a timbre encoder is used to extract the corresponding timbre features. The timbre encoder is a source domain private encoder used to extract private features of the Mel spectrum. The timbre features are used to characterize the speaker's identity. By inputting the Mel spectrum into the timbre encoder, the corresponding timbre features can be extracted from the Mel spectrum.

[0050] Emotional features are extracted from the target Mel spectrum. An emotion encoder is used to extract these features; for example, the encoder can consist of two convolutional neural network layers and two bidirectional long short-term memory (LSTM) layers. The convolutional kernels of the two convolutional neural network layers are 7×7 and 20×7, respectively. Following the convolutional layers are batch normalization layers, ReLU nonlinear activation layers, and max-pooling layers with kernel sizes of 2×2 and 1×5, respectively. The convolutional operation yields a 74×128-dimensional intermediate emotional representation sequence M = [m1, m2, ..., mn, ..., mN], where mn is the feature vector at the nth position. Emotion-related features are extracted from the FBank acoustic features using the two convolutional neural network layers and used as input to the LSTM layer, outputting the corresponding emotional features. In the emotion encoder, a two-layer bidirectional LSTM network is used to model the temporal relationship of the input intermediate sequence features M. The latent vector representations of the two-layer bidirectional long short-term memory network are derived from the forward and reverse long short-term memory networks, respectively. Each long short-term memory network has 128 hidden layer nodes. By using nonlinear activation, the output sequence of the final latent vector N time steps can be obtained, which together constitute the sentiment feature.

[0051] It should be noted that before using the timbre encoder and emotion encoder, they need to be trained. The training process may include: obtaining the initial timbre encoder and initial emotion encoder; obtaining sample training data, which includes the corresponding speaker's speech, pitch features, and speech content; using the initial timbre encoder to extract the timbre features corresponding to the speaker's speech; using the initial emotion encoder to extract the emotion features corresponding to the speaker's speech; and calculating the mutual information loss based on the speaker's pitch features, speech content, timbre features, and emotion features, using the following formula:

[0052]

[0053]

[0054]

[0055] in, This represents the mutual information loss corresponding to the timbre encoder. This refers to the mutual information loss corresponding to the emotion encoder. For timbre characteristics, Voice content, Pitch characteristics, As an emotional characteristic, For the mutual information between timbre features and speech content, This refers to the mutual information between timbre and pitch features. For the mutual information between timbre features and emotional features, Mutual information between emotional features and speech content For mutual information between emotional features and pitch features, For mutual information between emotional features and timbre features, The total loss is used for training. When the total loss reaches its maximum value, training is stopped, resulting in a trained timbre encoder and an emotion encoder. The trained timbre encoder is used to extract timbre features from the Mel spectrum information in the target speech, and the trained emotion encoder is used to extract emotion features from the Mel spectrum information in the target speech, and the corresponding emotion features are obtained.

[0056] S203: Expand the number of frames for both timbre and emotion features to be equal to the number of frames for pitch features, thus obtaining expanded timbre and emotion features.

[0057] In step S203, in order to better integrate timbre features and emotional features into the speech content of the source speech, the timbre features and emotional features are expanded to obtain corresponding features with the same number of pitch feature frames as the source speech.

[0058] In this embodiment, to increase the number of frames for timbre features and emotional features, the number of frames for timbre features and emotional features is expanded to be equal to the number of frames for pitch features, resulting in expanded timbre features and expanded emotional features. In this embodiment, a preset adapter can be used to expand the timbre features and emotional features during the expansion process.

[0059] Optionally, the number of frames for both timbre and emotion features is expanded to be equal to the number of frames for pitch features, resulting in expanded timbre and emotion features, including:

[0060] The timbre features are copied and expanded to obtain expanded timbre features with the same number of frames as the pitch features;

[0061] The emotional features are copied and expanded to obtain expanded emotional features with the same number of frames as the pitch features.

[0062] In this embodiment, the timbre feature is expanded by copying to obtain an expanded timbre feature with the same number of frames as the pitch feature. Similarly, the emotion feature is expanded by copying to obtain an expanded emotion feature with the same number of frames as the pitch feature. By expanding the features through copying, the original information of the corresponding timbre and emotion features is preserved.

[0063] S204: Align the pitch features, the expanded timbre features, and the expanded emotional features and then fuse them to obtain the first fused feature.

[0064] In step S204, the pitch features, expanded timbre features, and expanded emotion features are fused together to integrate the emotional and timbre information of the target speaker into the content of the source speech, thereby improving the speech conversion effect.

[0065] In this embodiment, the pitch feature, the expanded timbre feature, and the expanded emotion feature are first aligned. The number of frames of the expanded timbre feature, the expanded emotion feature, and the pitch feature are equal. Each frame of the pitch feature, the expanded timbre feature, and the expanded emotion feature is aligned and then fused. During fusion, the pitch feature, the expanded timbre feature, and the expanded emotion feature of the corresponding frame can be added and fused together, or the pitch feature, the expanded timbre feature, and the expanded emotion feature can be aligned and then spliced ​​together to obtain the fused feature.

[0066] S205: Extract the speech content from the source speech to obtain the source speech content, fuse the source speech content with the first fusion feature to obtain the second fusion feature, reconstruct the spectrum of the second fusion feature to obtain the reconstructed Mel spectrum, and obtain the converted speech based on the reconstructed Mel spectrum.

[0067] In step S205, the pitch features, expanded timbre features, and expanded emotion features are fused with the speech content in the source speech to obtain complete speech information features. The complete speech information features are then used for speech reconstruction to obtain the reconstructed Mel spectrum, so as to synthesize the corresponding converted speech based on the reconstructed Mel spectrum.

[0068] In this embodiment, the speech content is extracted from the source speech to obtain the source speech content, which represents the text information of the source speech. The source speech content is then superimposed and fused with a first fusion feature to obtain a second fusion feature. The second fusion feature is then spectrally reconstructed to obtain a reconstructed Mel spectrum. During spectral reconstruction, a decoder is used. The decoder is first trained to obtain a trained decoder, which is then used to decode and reconstruct the second fusion feature.

[0069] The decoder training process includes: acquiring sample training data, which contains the real Mel spectrum of the corresponding sample training data and the original second fusion features of the corresponding sample training data. The original second fusion features represent the timbre features, expanded emotion features, expanded pitch features, and features after being superimposed and fused with the speech content in the corresponding sample training data; acquiring an initial decoder and a discriminator in a preset adversarial network, training the initial decoder as the generator in the preset adversarial network, and using the discriminator to judge the decoding results obtained by the initial decoder decoding the original fusion features, obtaining a discrimination result. When the discrimination result meets preset conditions, the training of the initial decoder ends, and a trained decoder is obtained. Specifically, the discriminator's judgment of the decoding results obtained by the initial decoder decoding the original fusion features includes: using the initial decoder to decode the original fusion features to obtain the decoding results of the corresponding sample training data; classifying the decoding results of the corresponding sample training data based on the real Mel spectrum, obtaining classification probability values, and determining the classification probability values ​​as the discrimination results of the decoding results. When the classification probability value is greater than a preset threshold, it is considered that the initial decoder, which acts as the generator, can reconstruct and decode into a Mel spectrum that is close to the real Mel spectrum.

[0070] It should be noted that when training the initial decoder, the same speech can be used as the sample training data. The emotion features and timbre features, as well as the speech content and pitch features, can be extracted from the same speech. When performing inference, the emotion features and timbre features use the extracted features from the target speech of the target speaker, while the pitch features and speech content use the features from the source speech of the source speaker.

[0071] The reconstructed Mel spectrum is synthesized into the corresponding converted speech using a vocoder. In this embodiment, the vocoder can also be a neural network vocoder, including but not limited to autoregressive neural network vocoders (such as waveRNN vocoders) and Gann network-based neural network vocoders (such as melGan vocoders and hifiGan vocoders).

[0072] Optionally, the source speech is extracted to obtain the source speech content, including:

[0073] Extract Mel spectrum information from the source speech information to obtain the source Mel spectrum;

[0074] By performing a dimension-shifting process on the source Mel spectrum, the hidden state features corresponding to the source Mel spectrum are obtained.

[0075] The hidden state features are extracted using a pre-defined residual network to obtain the corresponding source speech content.

[0076] In this embodiment, Mel spectrum information is extracted from the source speech information to obtain the source Mel spectrum. During extraction, the source speech is first converted into a corresponding Fourier transform spectrum, and then a preset function is used to convert it into a Mel spectrum that better matches human hearing. This allows one-dimensional, difficult-to-process time-series signals to be transformed into easily processed and more information-rich two-dimensional frequency domain data. The Mel spectrum coefficients are a set of key coefficients used to construct the Mel spectrum. From a segment of the source speech, a set of cepstrums sufficient to represent the source speech can be obtained, and the Mel spectrum coefficients are the spectrum derived from this cepstrum (i.e., the spectrum of the spectrum). Unlike ordinary spectra, the most distinctive feature of the Mel spectrum is that the frequency bands on the Mel spectrum are uniformly distributed across the Mel scale. Therefore, the Mel spectrum of the corresponding source speech is extracted, and processing is performed based on the Mel spectrum.

[0077] It should be noted that the source speech information is subjected to spectral calculation using Short-Time Fourier Transform (SFT) to obtain a spectrogram. Specifically, the source speech information is processed by signal framing and windowing to obtain multiple speech segments. A SFT is performed on each speech segment to convert its time-domain features into frequency-domain features. Finally, the frequency-domain features of each frame are stacked along the time dimension to obtain the spectrogram. Each speech segment has a frame length of 20ms, a frame shift of 10ms, and 512 points in the SFT. The spectrogram is filtered using a 64-dimensional Mel-frequency filter bank. First, a logarithmic operation is performed on the spectrogram to obtain the logarithmic spectrum. Then, an inverse Fourier transform is performed on the logarithmic spectrum to obtain the source Mel-frequency spectrum. Feature extraction is performed on the source Mel-frequency spectrum to obtain Mel-frequency coefficients. The feature dimension of the source Mel-frequency coefficients is T*64, where T is the number of frames in the source speech information.

[0078] The source Mel spectrum is subjected to dimension-shifting processing to obtain the corresponding hidden state features. This is achieved through a convolutional network (CNN) containing at least one convolutional layer with a kernel size of 3 and 512 channels. This CNN adjusts the feature dimension of the source Mel spectrum to 512, filtering out irrelevant information and obtaining the Mel spectrum hidden state, i.e., the hidden state features. A pre-defined residual network is then used to extract features from these hidden state features, yielding the corresponding source speech content. This residual network consists of four identical layers. The hidden state features are the input to the first layer, the output of the first layer is the input to the second layer, and so on. Finally, the fourth layer outputs the speech features, providing the corresponding source speech content.

[0079] Optionally, the hidden state features are extracted using a pre-defined residual network to obtain the corresponding source speech content, including:

[0080] The latent state features are convolved using a residual network to obtain convolutional feature vectors.

[0081] The hidden state features are activated using the activation function of the residual network to obtain the activated feature vector;

[0082] The convolutional features and activation feature vectors are summed according to preset weight parameters to obtain target features, and the source speech content is determined based on the target features.

[0083] In this embodiment, a residual network is used to convolve the hidden state features to obtain a convolutional feature vector. The residual network consists of four layers, each with the same network structure. The hidden state features are the input to the first layer, the output of the first layer is the input to the second layer, and so on, until the fourth layer outputs the speech features. The convolutional kernels of each layer are 3, 7, 11, and 15, respectively. Each layer includes two convolutional layers (convolutional layer A and convolutional layer B) and one activation layer. The convolutional layer A of the residual network convolves the hidden state features to capture their spatial characteristics, resulting in a convolutional feature vector. The size of convolutional layer A is 1*1. The activation layer of the residual network includes two activation functions to activate the hidden state features, thus obtaining an activation feature vector. The target activation feature vector can be obtained by multiplying the activation vectors output by each activation function and then convolving the vector obtained by the multiplication by convolutional layer B. The size of convolutional layer A is 1*1. The preset weight parameters can be set according to actual business needs and are not limited. The convolutional features and activation feature vectors are summed according to preset weight parameters to obtain speech features.

[0084] It should be noted that the activation layer of the residual network includes two activation functions. The first activation function activates the hidden state features, resulting in a first activation feature vector. The second activation function activates the hidden state features, resulting in a second activation feature vector. The first and second activation feature vectors are then multiplied to obtain the target activation feature vector. The first activation function is the tanh function, which transforms the element values ​​of the input hidden state features to between -1 and 1, resulting in the first activation feature vector. The second activation function is the sigmoid function, which transforms the element values ​​of the input hidden state features to between 0 and 1, resulting in the second activation feature vector.

[0085] The first and second activation feature vectors are subjected to matrix dot product to obtain the dot product vector. Then, the dot product vector is subjected to one-dimensional convolution to obtain the target activation feature vector.

[0086] In this embodiment, the Mel-spectral features corresponding to the source speech information are extracted to obtain Mel-spectral hidden states of variable length. The Mel-spectral hidden states are then activated to obtain speech features, thereby enriching the semantic information of the speech features and improving the accuracy of speech content recognition.

[0087] The process involves: acquiring the source speaker's speech, extracting the fundamental frequency (FFM) from the source speech, extracting pitch features from the FFM, determining the number of frames for the pitch features, acquiring the target speaker's speech, extracting Mel spectrum information from the target speech, obtaining the target Mel spectrum, extracting timbre features from the target Mel spectrum, extracting emotion features from the target Mel spectrum, expanding the number of frames for both timbre and emotion features to be equal to the number of frames for the pitch features, obtaining expanded timbre and expanded emotion features, aligning and fusing the pitch features, obtaining the first fused feature, extracting the speech content from the source speech, fusing the source speech content with the first fused feature, obtaining the second fused feature, reconstructing the spectrum of the second fused feature, obtaining the reconstructed Mel spectrum, and obtaining the converted speech based on the reconstructed Mel spectrum. In this application, during the speech conversion process, the emotional information and pitch information of the target speaker are integrated into the content information of the source speaker, so that the converted target speech can better reflect the voice of the target speaker, thereby improving the speech conversion effect.

[0088] See Figure 3 , Figure 3 A structural block diagram of an AI-based voice conversion device provided in an embodiment of this application is shown. This voice conversion device is applied to the aforementioned server. For ease of explanation, only the parts relevant to the embodiments of this application are shown.

[0089] See Figure 3 The sound conversion device 30 includes: a first acquisition module 31, a second acquisition module 32, an expansion module 33, a fusion module 34, and a reconstruction module 35.

[0090] The first acquisition module 31 is used to acquire the source speech of the source speaker, extract the fundamental frequency from the source speech, extract the pitch features from the fundamental frequency, obtain the pitch features, and determine the number of frames of the pitch features.

[0091] The second acquisition module 32 is used to acquire the target speech of the target speaker, extract the Mel spectrum information in the target speech to obtain the target Mel spectrum, extract the timbre features from the target Mel spectrum to obtain the timbre features, and extract the emotional features from the target Mel spectrum to obtain the emotional features.

[0092] The expansion module 33 is used to expand the number of frames of both timbre features and emotional features to be equal to the number of frames of pitch features, so as to obtain expanded timbre features and expanded emotional features.

[0093] The fusion module 34 is used to align and fuse the pitch features, the expanded timbre features, and the expanded emotional features to obtain the first fusion feature.

[0094] The reconstruction module 35 is used to extract the speech content from the source speech, obtain the source speech content, fuse the source speech content with the first fusion feature to obtain the second fusion feature, reconstruct the spectrum of the second fusion feature to obtain the reconstructed Mel spectrum, and obtain the converted speech based on the reconstructed Mel spectrum.

[0095] Optionally, the aforementioned refactoring module 35 includes:

[0096] The extraction unit is used to extract Mel spectrum information from the source speech information to obtain the source Mel spectrum;

[0097] The variable-dimensional unit is used to perform variable-dimensional processing on the source Mel spectrum to obtain the hidden state features corresponding to the source Mel spectrum.

[0098] The obtained unit is used to extract features from the hidden state features through a preset residual network to obtain the corresponding source speech content.

[0099] Optionally, the obtained unit includes:

[0100] The convolutional subunit is used to perform convolution processing on the hidden state features through the residual network to obtain the convolutional feature vector.

[0101] The activation subunit is used to activate the hidden state features through the activation function of the residual network to obtain the activation feature vector.

[0102] The summation subunit is used to sum the convolutional features and activation feature vectors according to preset weight parameters to obtain target features, and to determine the source speech content based on the target features.

[0103] Optionally, the first acquisition module 31 mentioned above includes:

[0104] The fundamental frequency extraction unit is used to extract the fundamental frequency from the source speech using a preset vocoder.

[0105] The normalization unit is used to normalize the fundamental frequency by taking its logarithm to obtain the normalized feature, which is then used as the pitch feature.

[0106] Optionally, the aforementioned expansion module 33 includes:

[0107] The timbre features are copied and expanded to obtain expanded timbre features with the same number of frames as the pitch features.

[0108] The emotional features are copied and expanded to obtain expanded emotional features with the same number of frames as the pitch features.

[0109] Optionally, the aforementioned refactoring module 35 includes:

[0110] The segmentation unit is used to divide the source speech content into frames, resulting in a segmented source speech content with the same number of frames as the pitch features.

[0111] The fusion unit is used to align the segmented source speech content with the first fusion feature and then fuse them to obtain the second fusion feature.

[0112] Figure 4 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 4 As shown, the terminal device of this embodiment includes: at least one processor ( Figure 4 Only one is shown in the diagram), a memory, and a computer program stored in the memory and capable of running on at least one processor, wherein the processor executes the computer program to implement the steps in any of the above embodiments of the AI-based voice conversion method.

[0113] The terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 4 This is merely an example of a terminal device and does not constitute a limitation on the terminal device. A terminal device may include more or fewer components than shown in the figure, or a combination of certain components, or different components, such as network interfaces, displays, and input devices.

[0114] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0115] The memory includes readable storage media, internal memory, etc., wherein the internal memory can be the main memory of the terminal device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard disk of the terminal device, or in some embodiments, it can be the external storage device of the terminal device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device. Furthermore, the memory can also include both internal storage units and external storage devices of the terminal device. The memory is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of computer programs. The memory can also be used to temporarily store data that has been output or will be output.

[0116] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0117] The implementation of all or part of the processes in the methods of the above embodiments can also be accomplished by a computer program product. When the computer program product is run on a terminal device, the terminal device executes the steps in the above method embodiments.

[0118] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0119] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0120] In the embodiments provided in this application, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0121] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0122] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A voice conversion method based on artificial intelligence, characterized in that, The sound conversion method includes: Obtain the source speech of the source speaker, extract the fundamental frequency from the source speech, extract the pitch features from the fundamental frequency to obtain the pitch features, and determine the frame number of the pitch features; The process involves acquiring the target speech of the target speaker, extracting Mel-spectral information from the target speech to obtain the target Mel-spectrum, extracting timbre features from the target Mel-spectrum to obtain timbre features, and extracting emotional features from the target Mel-spectrum to obtain emotional features. Specifically, a timbre encoder is used to extract the corresponding timbre features, and an emotional encoder is used to extract the corresponding emotional features. Before using the timbre encoder and emotional encoder, they need to be trained. The training process includes: Obtain the initial timbre encoder and initial emotion encoder, and acquire sample training data. The sample training data includes the corresponding speaker's speech, pitch features, and speech content. Use the initial timbre encoder to extract the timbre features corresponding to the speaker's speech, and use the initial emotion encoder to extract the emotion features corresponding to the speaker's speech. Based on the speaker's pitch features, speech content, timbre features, and emotion features, calculate the mutual information loss using the following formula: in, This represents the mutual information loss corresponding to the timbre encoder. This refers to the mutual information loss corresponding to the emotion encoder. For timbre characteristics, Voice content, Pitch characteristics, As an emotional characteristic, For the mutual information between timbre features and speech content, This refers to the mutual information between timbre and pitch features. For the mutual information between timbre features and emotional features, Mutual information between emotional features and speech content For mutual information between emotional features and pitch features, For mutual information between emotional features and timbre features, The initial timbre encoder and initial emotion encoder are trained based on the mutual information loss to obtain the trained timbre encoder and emotion encoder. The number of frames for both the timbre feature and the emotional feature is expanded to be equal to the number of frames for the pitch feature, resulting in expanded timbre features and expanded emotional features. The pitch feature, the expanded timbre feature, and the expanded emotional feature are aligned and then fused to obtain the first fused feature; The source speech is extracted to obtain the source speech content. The source speech content is then fused with the first fusion feature to obtain the second fusion feature. The second fusion feature is then spectrally reconstructed to obtain the reconstructed Mel spectrum. Based on the reconstructed Mel spectrum, the converted speech is obtained.

2. The sound conversion method as described in claim 1, characterized in that, The step of extracting the spoken content from the source speech to obtain the source speech content includes: Extract the Mel spectrum information from the source speech information to obtain the source Mel spectrum; The source Mel spectrum is subjected to dimensionality-changing processing to obtain the hidden state features corresponding to the source Mel spectrum; The hidden state features are extracted using a pre-defined residual network to obtain the corresponding source speech content.

3. The sound conversion method as described in claim 2, characterized in that, The step of extracting features from the hidden state features using a preset residual network to obtain the corresponding source speech content includes: The hidden state features are convolved using the residual network to obtain the convolutional feature vector; The hidden state features are activated using the activation function of the residual network to obtain the activated feature vector; The convolutional features and the activation feature vector are summed according to preset weight parameters to obtain target features, and the source speech content is determined based on the target features.

4. The sound conversion method as described in claim 1, characterized in that, The step of extracting pitch features from the fundamental frequency to obtain pitch features includes: Extract the fundamental frequency from the source speech using a preset vocoder; The fundamental frequency is logarithmically normalized to obtain a normalized feature, which is then determined as the pitch feature.

5. The sound conversion method as described in claim 1, characterized in that, The step of expanding the frame count of both the timbre feature and the emotional feature to be equal to the frame count of the pitch feature, to obtain the expanded timbre feature and the expanded emotional feature, includes: The timbre feature is copied and expanded to obtain an expanded timbre feature with the same number of frames as the pitch feature; The emotional feature is copied and expanded to obtain an expanded emotional feature with the same number of frames as the pitch feature.

6. The sound conversion method as described in claim 1, characterized in that, The step of fusing the source speech content with the first fusion feature to obtain the second fusion feature includes: The source speech content is divided into frames to obtain a number of frames equal to the pitch feature. The segmented source speech content is aligned with the first fusion feature and then fused to obtain the second fusion feature.

7. A voice conversion device based on artificial intelligence, characterized in that, The sound conversion device includes: The first acquisition module is used to acquire the source speech of the source speaker, extract the fundamental frequency from the source speech, extract the pitch features from the fundamental frequency to obtain the pitch features, and determine the frame number of the pitch features; The second acquisition module is used to acquire the target speech of the target speaker, extract the Mel spectrum information from the target speech to obtain the target Mel spectrum, extract timbre features from the target Mel spectrum to obtain timbre features, and extract emotion features from the target Mel spectrum to obtain emotion features. Specifically, a timbre encoder is used to extract the corresponding timbre features, and an emotion encoder is used to extract the corresponding emotion features. Before using the timbre encoder and emotion encoder, they need to be trained. The training process includes: Obtain the initial timbre encoder and initial emotion encoder, and acquire sample training data. The sample training data includes the corresponding speaker's speech, pitch features, and speech content. Use the initial timbre encoder to extract the timbre features corresponding to the speaker's speech, and use the initial emotion encoder to extract the emotion features corresponding to the speaker's speech. Based on the speaker's pitch features, speech content, timbre features, and emotion features, calculate the mutual information loss using the following formula: in, This represents the mutual information loss corresponding to the timbre encoder. This refers to the mutual information loss corresponding to the emotion encoder. For timbre characteristics, Voice content, Pitch characteristics, As an emotional characteristic, For the mutual information between timbre features and speech content, This refers to the mutual information between timbre and pitch features. For the mutual information between timbre features and emotional features, Mutual information between emotional features and speech content For mutual information between emotional features and pitch features, For mutual information between emotional features and timbre features, The initial timbre encoder and initial emotion encoder are trained based on the mutual information loss to obtain the trained timbre encoder and emotion encoder. An expansion module is used to expand the number of frames of both the timbre feature and the emotional feature to be equal to the number of frames of the pitch feature, so as to obtain expanded timbre features and expanded emotional features. The fusion module is used to align and fuse the pitch feature, the expanded timbre feature, and the expanded emotional feature to obtain a first fusion feature; The reconstruction module is used to extract the speech content from the source speech to obtain the source speech content, fuse the source speech content with the first fusion feature to obtain the second fusion feature, reconstruct the spectrum of the second fusion feature to obtain the reconstructed Mel spectrum, and obtain the converted speech based on the reconstructed Mel spectrum.

8. The sound conversion device as described in claim 7, characterized in that, The reconstruction module includes: The extraction unit is used to extract Mel spectrum information from the source speech information to obtain the source Mel spectrum; A variable-dimensional unit is used to perform variable-dimensional processing on the source Mel spectrum to obtain the hidden state features corresponding to the source Mel spectrum. The unit is used to extract features from the hidden state features through a preset residual network to obtain the corresponding source speech content.

9. A terminal device, characterized in that, The terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the sound conversion method as described in any one of claims 1 to 6.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the sound conversion method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Tone conversion method and system focusing on audio feature extraction and separation

    CN118379984A