Timbre mixing method and apparatus, audio processing method and apparatus, and electronic device and storage medium

By performing dimensionality reduction and dimensionality upgrading processing on the tone feature sequence, combined with clustering analysis and weighted summing, the problems of mixed tone instability and noise are solved, and a more natural and stable audio output is achieved, which is suitable for speech synthesis, speech conversion and singing synthesis.

WO2025139400A1PCT designated stage expired Publication Date: 2025-07-03BEIJING XIYU JIZHI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/130857
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-11-08
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the existing audio processing technology, there are problems such as instability, unnaturalness and high background noise when mixing timbres. Especially in the fields of speech synthesis and speech conversion, traditional linear mixing methods are difficult to effectively control the stability and naturalness of the timbre.

Method used

The dimensionality reduction and dimensionality upscaling technology of tone feature sequences is used to perform dimensionality reduction processing through principal component analysis (PCA), and the timbre characteristics are fusion combined with clustering analysis and weighted summing to generate a hybrid timbre sequence suitable for acoustic neural networks.

Benefits of technology

Improves the stability and nature of the mixed tone, reduces background noise, achieves a more natural and stable audio output, and is compatible with Fastspeed and other audio processing models, improving processing speed and output quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130857_03072025_PF_FP_ABST
    Figure CN2024130857_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a timbre mixing method and apparatus, an audio processing method and apparatus, and an electronic device and a storage medium. The timbre mixing method in the embodiments of the present application comprises: acquiring a plurality of timbre feature sequences; performing dimensionality reduction on the plurality of timbre feature sequences, so as to obtain a plurality of dimensionality-reduced timbre sequences; fusing the plurality of dimensionality-reduced timbre sequences, so as to obtain a fused dimensionality-reduced timbre sequence; and performing dimensionality expansion on the fused dimensionality-reduced timbre sequence, so as to obtain a mixed timbre sequence used for controlling the timbre of an audio output from an acoustic neural network. The timbre mixing method and the audio processing method in the present application can retain key timbre features and effectively eliminate background noise and other interference information during timbre mixing, thereby handling the defects in conventional crude linear timbre mixing, and realizing the output of an audio which has a mixed timbre and is clearer, more natural, more stable and more realistic.
Need to check novelty before this filing date? Find Prior Art

Description

Tone mixing method and device, audio processing method and device, electronic device, and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application number 2023118645081 filed with the Chinese Patent Office on December 29, 2023, entitled "Tone mixing method and device, audio processing method and device, electronic device, storage medium", the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of audio processing technology, and in particular to a timbre mixing method and device, an audio processing method and device, and related electronic equipment and storage media. Background Art

[0004] In related fields, a variety of audio processing technologies are under development, including but not limited to text-to-speech (TTS) and voice conversion (VC). These technologies have a wide range of applications, such as in speech synthesis, voice assistants, and audiobooks. Among these audio processing technologies, especially in the field of speech synthesis, there is a desire to create diverse sound effects and provide personalized audio output.

[0005] The description of the background technology is intended to help understand the relevant technology in the relevant field, and does not mean that the background technology content is admitted to be prior art.

[0006] Summary of the Invention

[0007] Therefore, the embodiments of the present application aim to provide a timbre mixing method and device, an audio processing method and device, an electronic device, and a storage medium, thereby proposing an excellent mixed timbre control solution for audio processing technologies, including but not limited to speech synthesis, speech conversion, singing synthesis, etc., and optionally solving or improving at least one of the problems of mixed audio, such as unstable timbre of speech or singing, unnatural timbre, poor sound quality, and large background noise.

[0008] In a first aspect, an embodiment of the present application provides a timbre mixing method, the timbre mixing method comprising:

[0009] Acquire multiple timbre feature sequences;

[0010] Performing dimensionality reduction on the plurality of timbre feature sequences to obtain a plurality of reduced-dimensionality timbre sequences;

[0011] fusing the plurality of reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence; and

[0012] The fused reduced-dimensionality timbre sequence is dimensionally increased to obtain a mixed timbre sequence for controlling the audio timbre output by the acoustic neural network.

[0013] In some embodiments of the present application, the dimensionality reduction of the multiple timbre feature sequences to obtain multiple reduced-dimensional timbre sequences includes:

[0014] Performing dimensionality reduction on the plurality of timbre feature sequences using PCA transformation, and determining principal component coefficients and means corresponding to the dimensionality reduction using principal component analysis;

[0015] The step of increasing the dimension of the fused reduced-dimensionality timbre sequence to obtain a mixed timbre sequence for controlling the audio timbre output by the acoustic neural network includes:

[0016] Based on the determined principal component coefficients and mean values, the mixed timbre sequence is dimensionally increased by using an inverse PCA transformation.

[0017] In some embodiments of the present application, fusing the multiple reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence includes:

[0018] The plurality of dimension-reduced timbre sequences are weightedly summed to obtain the fused dimension-reduced timbre sequence.

[0019] In some embodiments of the present application, obtaining multiple timbre feature sequences includes:

[0020] Get multiple initial timbre feature sequences:

[0021] performing clustering processing on the multiple initial timbre feature sequences;

[0022] The plurality of timbre feature sequences are generated based on cluster centers of the plurality of initial timbre feature sequences.

[0023] In some embodiments of the present application, the multiple timbre feature sequences are subjected to dimensionality reduction to obtain multiple reduced-dimensionality timbre sequences, including:

[0024] Based on the number of clusters of the multiple initial timbre feature sequences, the dimension reduction dimensions of the multiple timbre feature sequences are determined.

[0025] In a second aspect, an embodiment of the present application provides an audio processing method, the audio processing method comprising:

[0026] Acquire a first control sequence for controlling target audio content;

[0027] Acquire multiple timbre feature sequences;

[0028] Performing dimensionality reduction on the plurality of timbre feature sequences to obtain a plurality of reduced-dimensionality timbre sequences;

[0029] fusing the multiple reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence;

[0030] Performing dimension increase on the fused reduced-dimensionality timbre sequence to obtain a second control sequence for controlling the target audio timbre; and

[0031] The first control sequence and the second control sequence are input into a given acoustic neural network to obtain target audio.

[0032] In some embodiments of the present application, reducing the dimensionality of the multiple timbre feature sequences to obtain multiple reduced-dimensional timbre sequences includes: reducing the dimensionality of the multiple timbre feature sequences using PCA transformation, and determining the principal component coefficients and means corresponding to the principal component analysis method;

[0033] The step of increasing the dimension of the fused reduced-dimensionality timbre sequence to obtain a mixed timbre sequence for controlling the audio timbre output by the acoustic neural network includes:

[0034] Based on the determined principal component coefficients and mean values, the mixed timbre sequence is dimensionally increased by using an inverse PCA transformation.

[0035] In some embodiments of the present application, fusing the multiple reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence includes:

[0036] The plurality of dimension-reduced timbre sequences are weightedly summed to obtain the fused dimension-reduced timbre sequence.

[0037] In some embodiments of the present application, obtaining multiple timbre feature sequences includes:

[0038] Get multiple initial timbre feature sequences:

[0039] performing clustering processing on the multiple initial timbre feature sequences;

[0040] The plurality of timbre feature sequences are generated based on cluster centers of the plurality of initial timbre feature sequences.

[0041] In some embodiments of the present application, the multiple timbre feature sequences are subjected to dimensionality reduction to obtain multiple reduced-dimensionality timbre sequences, including:

[0042] Based on the number of clusters of the multiple initial timbre feature sequences, the dimension reduction dimensions of the multiple timbre feature sequences are determined.

[0043] In some embodiments of the present application, the audio processing method is a text-to-speech (TTS) method, wherein the first control sequence is a text sequence for synthesizing a target speech, and the acoustic neural network is a speech synthesis model.

[0044] In some embodiments of the present application, the audio processing method is a voice conversion (VC) method, the first control sequence is a speech sequence to be converted for conversion into a target speech, and the acoustic neural network is a speech conversion model.

[0045] In some embodiments of the present application, the acquiring of multiple timbre feature sequences includes: extracting timbre features from the speech sequence to be converted to obtain a first timbre feature sequence, and acquiring one or more pre-provided second timbre feature sequences;

[0046] The step of reducing the dimension of the plurality of timbre feature sequences to obtain a plurality of reduced-dimensional timbre sequences comprises: reducing the dimension of the first timbre feature sequence to obtain a first reduced-dimensional timbre sequence, and reducing the dimension of the one or more second timbre feature sequences to obtain one or more second reduced-dimensional timbre sequences;

[0047] The fusion of the plurality of reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence comprises: weighted summing the first reduced-dimensionality timbre sequence and the one or more second reduced-dimensionality timbre sequences to obtain the fused reduced-dimensionality timbre sequence.

[0048] In some embodiments of the present application, the weight of the first reduced-dimensionality timbre sequence is greater than the total weight of the one or more second reduced-dimensionality timbre sequences.

[0049] In some embodiments of the present application, the audio processing method is a singing synthesis method, the first control sequence is a lyrics text sequence for synthesizing a target singing voice, and the acoustic neural network is a singing synthesis model.

[0050] The audio processing method further includes: acquiring a third control sequence for controlling the melody of the target singing voice, wherein the third control sequence includes one or more melody features of the song corresponding to the lyrics text sequence;

[0051] The step of inputting the first control sequence and the second control sequence into a given acoustic neural network to obtain target audio includes: inputting the first control sequence, the second control sequence and the third control sequence into the singing synthesis model to obtain the target singing.

[0052] In a third aspect, an embodiment of the present application provides a timbre mixing device, the timbre mixing device comprising:

[0053] an acquisition unit configured to acquire a plurality of timbre feature sequences;

[0054] a dimensionality reduction unit configured to reduce the dimensionality of the plurality of timbre feature sequences to obtain a plurality of reduced-dimensionality timbre sequences;

[0055] a fusion unit configured to fuse the plurality of reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence; and

[0056] The dimension increasing unit is configured to increase the dimension of the fused reduced-dimensionality timbre sequence to obtain a mixed timbre sequence for controlling the audio timbre output by the acoustic neural network.

[0057] In a fourth aspect, an embodiment of the present application provides an audio processing device, the audio processing device comprising:

[0058] a first acquiring unit configured to acquire a first control sequence for controlling target audio content;

[0059] a second acquiring unit configured to acquire a plurality of timbre feature sequences;

[0060] a dimensionality reduction unit configured to reduce the dimensionality of the plurality of timbre feature sequences to obtain a plurality of reduced-dimensionality timbre sequences;

[0061] a fusion unit configured to fuse the plurality of reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence;

[0062] a dimension-increasing unit configured to perform dimension-increasing on the fused reduced-dimensionality timbre sequence to obtain a second control sequence for controlling the target audio timbre; and

[0063] An input unit is configured to input the first control sequence and the second control sequence into a given acoustic neural network to obtain target audio.

[0064] In some embodiments of the present application, the audio processing device includes a speech synthesis device, wherein the first control sequence is a text sequence for synthesizing a target speech, and the acoustic neural network is a speech synthesis model.

[0065] In some embodiments of the present application, the audio processing device includes a speech conversion device, wherein the first control sequence is a speech sequence to be converted into a target speech, and the acoustic neural network is a speech conversion model.

[0066] In some embodiments of the present application, the audio processing device includes a singing synthesis device, wherein the first control sequence is a lyrics text sequence for synthesizing a target singing voice, and the acoustic neural network is a singing synthesis model.

[0067] The audio processing device further includes: a third acquisition unit configured to acquire a third control sequence for controlling the melody of the target singing voice, wherein the third control sequence includes one or more melody features of the song corresponding to the lyrics text sequence;

[0068] The input unit is configured to input the first control sequence, the second control sequence and the third control sequence into the singing synthesis model to obtain the target singing.

[0069] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, wherein when the program is executed by a processor, it implements the method of any embodiment of the present application.

[0070] In a sixth aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory storing a computer program, wherein the processor is configured to execute any method of the embodiment of the present application when running the computer program.

[0071] The inventors note that timbre control has been proposed to help create diverse sound effects and provide personalized audio output. The inventors also note that current timbre control solutions rarely involve controlling audio output through timbre mixing. At the same time, the inventors have discovered that simple timbre feature mixing can result in unstable or unnatural timbre after mixing. These issues limit the application of timbre mixing in audio processing, including but not limited to speech synthesis, voice conversion, and singing synthesis.

[0072] The technical solution of the present application provides a timbre mixing method and related audio processing methods. By reducing the dimensionality of multiple high-dimensional timbre feature sequences and fusing them into low-dimensional space, a fused dimensionality-reduced timbre sequence is obtained that effectively eliminates noise and interference information while retaining multiple timbre attributes. By increasing the dimensionality of the fused dimensionality-reduced timbre sequence, a mixed timbre sequence is generated that is ultimately used to control the output of the acoustic neural network. Inputting it into the acoustic model can produce natural, stable, and realistic audio output. In addition, the timbre mixing method and related audio processing methods of the present application are also compatible with the Fastspeech model and other existing audio processing models, improving processing speed, ensuring the stability and naturalness of the output, and are suitable for a variety of audio output requirements.

[0073] Other optional features and technical effects of the embodiments of the present application are partially described below, and partially can be understood by reading this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use. It should be noted that the drawings described below only cover some embodiments of the present application. Ordinary technicians in this field can derive other drawings based on these drawings without having to engage in creative work. The purpose of these drawings is to better illustrate the technical details to facilitate understanding of the implementation methods of the present application, including:

[0075] FIG1 is a schematic flow chart showing a speech synthesis method of the prior art;

[0076] FIG2 is a schematic flow chart showing a timbre mixing method according to an embodiment of the present application;

[0077] FIG3 is a schematic flow chart showing a timbre mixing method according to an embodiment of the present application;

[0078] FIG4 is a schematic distribution diagram showing the color gathering results of the timbre mixing method according to an embodiment of the present application;

[0079] FIG5 is a schematic flow chart of an audio processing method according to an embodiment of the present application;

[0080] FIG6 shows a schematic structural diagram of a text-to-speech (TTS) method for implementing the audio processing solution in accordance with an embodiment of the present application;

[0081] FIG7 shows a structural diagram of an exemplary Fastspeech-based acoustic model architecture for implementing an audio processing method according to an embodiment of the present application;

[0082] FIG8 shows a schematic structural diagram for implementing a voice conversion (VC) method combined with an audio processing solution according to an embodiment of the present application;

[0083] FIG9 shows a schematic structural diagram for implementing a song generation method combined with the audio processing solution of an embodiment of the present application;

[0084] FIG10 shows an exemplary structural diagram of a timbre mixing device according to an embodiment of the present application;

[0085] FIG11 shows an exemplary structural diagram of an audio processing device according to an embodiment of the present application; and

[0086] FIG12 is a schematic diagram of an exemplary structure of an electronic device that implements a method according to an embodiment of the present application. DETAILED DESCRIPTION

[0087] The following detailed description of exemplary embodiments of the present application is provided, and illustrations of the exemplary embodiments are shown in the accompanying drawings. When referring to the drawings, unless otherwise specified, identical numbers or symbols in different figures represent identical or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Instead, they are merely examples of some aspects of the apparatus and methods covered by one or more embodiments of this specification, as detailed in the claims of this application.

[0088] As used in this specification, the term "including" and its variations are intended to be broadly inclusive, meaning "including but not limited to" the listed items. Unless otherwise stated, the term "or" means "and / or," the term "based on" means relying on, or at least partially relying on, the terms "an example embodiment" and "an embodiment" refer to at least one example embodiment, and the term "another embodiment" refers to at least one different embodiment. The terms "first," "second," and so on may refer to different or the same items. Other explicit and implicit definitions may be included below.

[0089] Currently, traditional audio processing technologies such as text-to-speech (TTS) and voice conversion (Voice Conversion) are widely used. As shown in Figure 1, some conventional TTS methods known to the inventors generally include two main parts: a first acoustic model that generates a Mel spectrum from an input text sequence containing text information and a timbre feature sequence; and a second acoustic model that converts the Mel spectrum into the final target audio waveform.

[0090] However, few existing solutions involve inputting mixed timbre information obtained by mixing two or more timbres. The inventors have discovered that when obtaining mixed timbre information, using only a crude method of linearly mixing multiple timbres can result in instability, unnaturalness, and high background noise in the generated audio, such as speech.

[0091] In this regard, the embodiments of the present application provide a timbre mixing method and apparatus, an audio processing method and apparatus that apply the timbre mixing scheme of the embodiments of the present application, related electronic equipment, and storage media. The timbre mixing scheme of the embodiments of the present application introduces dimensionality increase and dimensionality reduction techniques for timbre sequences, which greatly improves the stability and naturalness of mixed timbre audio such as that generated by a speech synthesis model, and reduces the presence of irrelevant noise. In addition, the timbre mixing scheme of the embodiments of the application can also be widely applied to other audio processing technologies based on artificial intelligence neural networks, especially machine learning, such as speech conversion technology and singing synthesis technology, and further effects can be produced by combining it with these audio processing technologies.

[0092] The methods and devices may be implemented by one or more computers. In some embodiments, the implementation of these devices may be accomplished by software, hardware, or a combination of software and hardware. In some embodiments, the electronic device or computer may be implemented by the computers described herein or other electronic devices capable of performing the corresponding functions.

[0093] In the embodiment of the present application, as shown in FIG2 , a timbre mixing method is provided, which may include the following steps:

[0094] S210: Acquire multiple timbre feature sequences.

[0095] In the embodiment of the present application, the timbre refers to the specific characteristics of the sound, including the frequency distribution, spectral characteristics, resonance characteristics and vibration mode of the sound waves. Each speaker's voice has a unique timbre, which enables people to distinguish different speakers.

[0096] In an embodiment of the present application, the multiple timbre feature sequences are typically in the form of high-dimensional features, such as in the form of embedding vectors (512 or 1024 dimensions). In some embodiments of the present application, for example, in step S210 above, multiple timbre embedding vectors containing timbre information may be obtained. In some embodiments of the present invention, the multiple initial timbre feature sequences may be timbre feature sequences for speakers of different genders, ages, and emotional states, without limitation.

[0097] In some embodiments of the present application, the dimensionality reduction process described below may be directly performed on the acquired or generated multiple timbre feature sequences (such as embedding vectors).

[0098] In other embodiments of the present application, clustering processing can be advantageously performed on at least part, and preferably all, of the multiple timbre feature sequences (such as embedding vectors), and the subsequent dimensionality reduction processing can be assisted based on the results of the clustering processing.

[0099] As shown in FIG3 , the step of obtaining a plurality of timbre feature sequences may further include:

[0100] S211: Acquire multiple initial timbre feature sequences.

[0101] Similarly, the multiple initial timbre feature sequences are usually in the form of high-dimensional features, such as 512-dimensional or 1024-dimensional Embedding vectors.

[0102] S212: performing clustering processing on the multiple initial timbre feature sequences.

[0103] In the embodiments of the present application, clustering refers to a method of grouping objects in a dataset. Clustering divides the data into multiple groups or "clusters" so that data points within the same cluster are similar to each other, while data points in different clusters are relatively dissimilar. Cluster centroids are generated from this, representing the center point or representative feature of each cluster. These cluster centroids can be used to represent the common features of all data points in the cluster. The generated cluster centroids can be used for further analysis, feature extraction, or model training.

[0104] In some embodiments of the present application, for example, in the above step S212, clustering processing may be performed on the multiple initial timbre feature sequences to generate cluster centers corresponding to the multiple initial timbre feature sequences.

[0105] In some embodiments of the present application, the clustering processing may include K-means clustering processing, hierarchical clustering processing, Gaussian mixture model processing, etc., which are not limited here.

[0106] S213: Generate the multiple timbre feature sequences based on the cluster centers of the multiple initial timbre feature sequences.

[0107] In the embodiment of the present invention, each cluster center of the plurality of initial timbre feature sequences represents a typical timbre feature of all timbre feature sequences in the cluster.

[0108] In an embodiment of the present application, for example, in the above step S213 , a plurality of timbre feature sequences having typical timbre features of the plurality of initial timbre feature sequences may be generated based on the cluster centers of the plurality of initial timbre feature sequences.

[0109] S220: Perform dimensionality reduction on the multiple timbre feature sequences to obtain multiple reduced-dimensional timbre sequences.

[0110] In an embodiment of the present application, dimensionality reduction implements the process of converting a high-dimensional feature sequence into one or more low-dimensional feature sequences. This process will reduce the complexity of the high-dimensional feature sequence, remove redundant information, and improve computational efficiency, while retaining the main features of the original data as much as possible.

[0111] In an embodiment of the present application, for example, in the above-mentioned step S220, the multiple timbre feature sequences are usually high-dimensional features, such as timbre embedding vectors (512 dimensions or 1024 dimensions), which, in addition to timbre information, also contain some background noise or ambient sound information.

[0112] As mentioned above, in a preferred embodiment of the present application, the number of clusters can be used to assist in determining the dimension reduction. In a specific embodiment of the present application, the above step S220 can further include the following steps: determining the dimension reduction of the multiple timbre feature sequences based on the number of clusters of the multiple initial timbre feature sequences.

[0113] In an embodiment of the present application, the number of clusters refers to the total number of clusters formed by the multiple initial timbre feature sequences during the clustering process, and in the graphical representation, these clusters may be displayed as different point groups, each point group represents a cluster, and the dimensionality reduction dimension can be set according to the number of clusters. For example, in one embodiment, the dimensionality reduction dimension can be determined roughly based on the number of clusters or roughly based on the number of clusters plus a reasonable margin. In an exemplary embodiment, for example, in the above-mentioned step S221, the number of dimensions of the dimensionality reduction dimension of the multiple timbre feature sequences can be set to 1 to 2 times the number of clusters of the multiple initial timbre feature sequences. In a specific embodiment of the present application, in the above-mentioned step of determining the dimensionality reduction dimension of multiple timbre feature sequences based on the number of clusters of multiple initial timbre feature sequences, the dimension after dimensionality reduction can be determined to be 100. In a specific example, for example, the number of clusters is determined to be 76 and the dimensionality reduction dimension is 100.

[0114] In some embodiments of the present application, in the above step S220, principal component analysis (PCA) can be used to reduce the dimension of the multiple timbre feature sequences, and determine the principal component coefficients and means corresponding to the dimensionality reduction using the principal component analysis method.

[0115] In an embodiment of the present application, principal component analysis (PCA) is a statistical method for data dimensionality reduction and feature extraction, and its core idea is to map the original high-dimensional feature space (n-dimensional) to a new low-dimensional space (k-dimensional), where k is less than n. In the new low-dimensional space (k-dimensional), each dimension is a new, mutually orthogonal feature, which is also called a principal component. Principal component analysis (PCA) can find a set of mutually orthogonal coordinate axes in an orderly manner. The selection of these coordinate axes is closely related to the variance distribution of the original data. First, PCA selects the first new coordinate axis, which is the direction with the largest variance in the original data. Next, it selects the second new coordinate axis, which is located in a plane orthogonal to the first coordinate axis to maximize the variance. Subsequently, principal component analysis (PCA) selects the third new coordinate axis, which is orthogonal to the first two coordinate axes and maximizes the variance in this plane, and so on, until k new coordinate axes are obtained.

[0116] Through this process, most of the data variance is contained in the first k axes, while the variance in subsequent axes is almost negligible. Therefore, the remaining axes can be safely discarded, retaining only the first k axes, which contain the vast majority of the data variance information. Thus, principal component analysis (PCA) reduces the dimensionality of the data features, retaining the most informative feature dimensions while ignoring those with little variance.

[0117] For example, in the above step S220, when the principal component analysis method is used to reduce the dimensionality of the plurality of timbre feature sequences, it is preferred to further determine the principal component coefficients and means corresponding to the dimensionality reduction of the principal component analysis method.

[0118] In some embodiments, data standardization can be performed when the principal component analysis method is processed, including but not limited to calculating the mean (mu) and standard deviation of each timbre feature sequence by mean zeroing and standard deviation normalization to ensure that all features are on the same scale. The covariance matrix can also be calculated for the standardized multiple timbre feature data sets, and the covariance matrix represents the correlation between different features. In addition, the covariance matrix can also be subjected to eigenvalue decomposition (eigenvalue decomposition) to obtain eigenvalues ​​and corresponding eigenvectors. Moreover, the principal component (coeff) can be determined as follows: the eigenvectors are sorted by the size of the eigenvalues, and the first k eigenvectors are selected as the principal component, where k is the number of dimensions after the expected dimensionality reduction. Further, the selected first k eigenvectors are used to project the original multiple timbre feature sequences into a new k-dimensional space to obtain the principal component score (score) of each timbre feature sequence on each principal component, thereby generating multiple timbre feature sequences after dimensionality reduction.

[0119] In an embodiment of the present invention, using a universal PCA for dimensionality reduction means using a unified principal component (coeff) and mean (mu), so that the multiple timbre feature sequences obtained after dimensionality reduction will be suitable for mixing with each other in the dimensionality reduction space, for example, suitable for fusion processing in the subsequent step S230.

[0120] In one example of the present application, when PCA is used to reduce the high-dimensional timbre embedding (1024 dimensions) to 100 dimensions to obtain a low-dimensional timbre embedding_pca (100 dimensions), clustering processing is performed on the timbre embedding_pca (100 dimensions), and embeddings with similar timbre can still be clustered, as shown in Figure 4. This shows that directly reducing the dimensionality of multiple high-dimensional timbre feature sequences to the 100-dimensional embedding_pca discards useless features (background noise, ambient sound, etc.) while still preserving most of the required timbre features.

[0121] In the embodiment of the present application, in addition to using principal component analysis, independent component analysis (ICA), linear discriminant analysis (LDA), non-negative matrix factorization (NMF) and t-distributed stochastic neighbor embedding (t-SNE) and other methods can also be used to reduce the dimension of multiple timbre feature sequences, without limitation here.

[0122] S230: Fusing multiple reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence.

[0123] In the embodiment of the present application, the fused reduced-dimensionality timbre sequence obtained by fusion includes the timbre information of the multiple reduced-dimensionality timbre sequences.

[0124] In some embodiments of the present application, for example, in the above step S230, multiple reduced-dimensionality timbre sequences may be fused to obtain a fused reduced-dimensionality timbre sequence containing the timbre information of the multiple reduced-dimensionality timbre sequences.

[0125] In some embodiments of the present application, for example, in the above step S230, the plurality of reduced-dimensionality timbre sequences may be weighted and summed to obtain the fused reduced-dimensionality timbre sequence.

[0126] In some embodiments of the present application, for example, in the above step S230, a weighted summation may be performed on the timbre embeddings of multiple timbre timbres after dimensionality reduction to obtain a fused timbre embedding.

[0127] In a specific embodiment of the present application, for example, the three timbres Embedding_pca_A, Embedding_pca_B and Embedding_pca_C obtained after PCA dimensionality reduction can be weighted and summed according to the following formula to obtain the fused timbre Embedding_pca_X: Embedding_pca_X = a*Embedding_pca_A+b*Embedding_pca_B+c*Embedding_pca_C

[0128] Where a, b, and c are weight coefficients of the corresponding timbres, and a + b + c = 1. The weight coefficients determine the contribution of each reduced timbre embedding in the final embedding_pca_x. Different weight settings can produce different fusion timbres.

[0129] The above shows the mixing of three timbres, but the number of mixed timbres can be greater than or equal to two, and the specific number is not limited.

[0130] S240: Dimensionality-increasing the fused reduced-dimensionality timbre sequence to obtain a mixed timbre sequence for controlling the audio timbre output by the acoustic neural network.

[0131] In an embodiment of the present application, for example, in the above-mentioned step S240, the fused reduced-dimensionality timbre sequence can be upgraded to obtain a mixed timbre sequence whose dimension is the same as the dimension of the original speech synthesis model for controlling the audio timbre output by the acoustic neural network.

[0132] In some embodiments of the present application, such as in step S240, the mixed timbre sequence may be increased in dimension using an inverse PCA transform based on the principal component coefficients and means determined during the PCA dimensionality reduction process. For example, the mixed timbre sequence may be increased in dimension using an inverse PCA transform based on the principal component (coeff) and mean (mu) in step S220. This ensures consistency in the dimensionality increase process, thereby obtaining a mixed timbre sequence suitable for controlling the audio timbre output by the acoustic neural network.

[0133] Thus, the technical solution of the present application provides a timbre mixing method. By reducing the dimensionality of multiple high-dimensional timbre feature sequences and fusing them into low-dimensional space, a fused dimensionality-reduced timbre sequence is obtained that effectively eliminates noise and interference information while retaining multiple timbre attributes. By increasing the dimensionality of the fused dimensionality-reduced timbre sequence, a mixed timbre sequence is generated that is ultimately used to control the output of the acoustic neural network. Inputting it into the acoustic model can produce natural, stable, and realistic audio output. The technical solution of the present application provides a timbre mixing method that effectively achieves the precise extraction of timbre attributes from multiple timbre feature sequences by optionally combining a clustering method before dimensionality reduction.

[0134] As previously mentioned, the timbre mixing scheme of the application embodiments can also be widely applied to other audio processing technologies based on artificial intelligence neural networks, especially machine learning, such as voice conversion technology and singing synthesis technology, and can produce further benefits by combining with these audio processing technologies. The following describes the application of the timbre mixing scheme in audio processing technologies based on artificial intelligence neural networks in conjunction with specific embodiments.

[0135] In some embodiments of the present application, as shown in FIG5 , an audio processing method is further provided, which may include the following steps:

[0136] S510: Acquire a first control sequence for controlling target audio content.

[0137] S520: Acquire multiple timbre feature sequences.

[0138] S530: Perform dimensionality reduction on the multiple timbre feature sequences to obtain multiple reduced-dimensional timbre sequences.

[0139] S540: Fusing multiple reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence.

[0140] S550: Perform dimension increase on the fused reduced-dimensionality timbre sequence to obtain a second control sequence for controlling the target audio timbre.

[0141] S560: Input the first control sequence and the second control sequence into a given acoustic neural network to obtain target audio.

[0142] In some embodiments of the present application, the above-mentioned timbre mixing method, such as the timbre mixing method shown in Figure 2 and its steps S210 to S240, sub-steps such as steps S211 to S213 or features can be combined with the audio processing method of the embodiment of the present application in a non-contradictory manner, and vice versa, which will not be elaborated here.

[0143] In the embodiments of the present application, the above audio processing method can be applied to a variety of application scenarios.

[0144] In some embodiments of the present application, the audio processing method may be a text-to-speech (TTS) method.

[0145] FIG6 shows an exemplary structural diagram of a speech synthesis method incorporating a timbre mixing solution according to an embodiment of the present application.

[0146] In the embodiment shown in Figure 6, for example, in the above-mentioned step S510, when the audio processing method is applied to speech synthesis technology, the first control sequence for controlling the target audio content can be a text sequence for synthesizing the target speech, and the acoustic neural network is a speech synthesis model.

[0147] In the embodiment shown in FIG6 , the text sequence may be subjected to front-end processing. For example, steps such as text normalization, word segmentation, syntactic analysis, and phoneme conversion may be performed on the text sequence to convert the text sequence into a format suitable for acoustic processing, such as similarly performing embedding processing.

[0148] In the embodiment shown in FIG. 6 , for example, in steps S520 to S550 , the timbre mixing scheme described in the embodiment of the present application may be used to perform timbre mixing processing on multiple timbre feature sequences, such as timbre feature sequences 1 to N.

[0149] As shown in Figure 6, the processed first control sequence (such as a text sequence) and the second control sequence (a timbre control sequence processed with timbre mixing) will be input into a given acoustic neural network, here a text-to-speech (TTS) model. In some embodiments, the TTS model may include a first acoustic model at the front end. In some embodiments of the present application, it may be based on the Fastspeech architecture, as described below. After processing by the first acoustic model, a characteristic mel-spectrogram of the sound will be generated. In the model, the characteristic mel-spectrogram of the sound generated by combining the two control information can serve as a key intermediate representation in speech synthesis, capturing the key characteristics of the speech to be generated. However, the present application does not intend to limit the intermediate representation of speech synthesis. The characteristic mel-spectrogram will be input into a second acoustic model, which may be, for example, a vocoder. In embodiments of the present application, the second acoustic model, such as a vocoder, may include, but is not limited to, WaveNet, WaveGlow, and WaveRNN, etc., without limitation. The vocoder can convert the mel-spectrogram into a final audio waveform, outputting a natural and stable target audio containing mixed timbre.

[0150] The audio processing method of the embodiment of the present application can also be implemented in conjunction with the Fastspeech architecture. As previously described, the first acoustic model can have a Fastspeech architecture for speech synthesis, such as shown in Figure 7. In the embodiment of the present application, the Fastspeech architecture is a new feedforward network that improves performance in many aspects compared to traditional TTS models. Fastspeech can quickly generate stable, controllable, and high-quality mel-spectrograms.

[0151] In some embodiments of the present application, as shown in FIG7 , the Fastspeech architecture model may mainly include a feed-forward transformer block (FFT Block Feed-Forward Transformer Block), a length regulator (Length Regulator), and a linear layer (Linear Layer):

[0152] a. Feed-Forward Transformer Block (FFT Block Feed-Forward Transformer Block): It includes a self-attention module and a 1D convolution module, which can convert the input phoneme sequence into a Mel-spectrogram sequence to prepare for the subsequent speech synthesis step.

[0153] b. Length Regulator: Because the length of a phoneme sequence is smaller than the length of a mel-spectrogram sequence, one phoneme corresponds to multiple mel-spectrograms. The number of mel-spectrograms aligned with a phoneme is called the phoneme duration. The length regulator expands the hidden sequence of phonemes based on their duration to match the length of the mel-spectrogram sequence. We can proportionally increase or decrease the phoneme duration to adjust the speed of sound.

[0154] c. Linear Layer: The linear layer is a basic neural network layer used to implement linear transformations. In deep learning and neural networks, it is often called a fully connected layer or dense layer. It is used to convert the network's intermediate representation into the final output format, the mel-spectrogram.

[0155] In some embodiments of the present application, as shown in FIG. 7 , an example of applying the audio processing method of the present application to a speech synthesis model of the Fastspeech architecture to output synthesized speech with mixed timbre will be described in conjunction with step S560, wherein:

[0156] As shown in Figure 7, the processed phoneme sequence and the mixed timbre sequence (for example, in the form of embedding) will be input to the feedforward transformer block. In this embodiment of the application, unlike the conventional Fastspeech architecture that only inputs the phoneme sequence, the phoneme sequence and the mixed timbre sequence are input. Optionally, the phoneme sequence and the mixed timbre sequence can be spliced ​​together, for example, the phoneme sequence is placed before the timbre sequence.

[0157] The length-adjusted fused phonemes and mixed timbre embeddings are fed into Fastspeech's Feed-Forward Transformer (FFT) Blocks, for example, using M FFT Block modules. These modules use a self-attention mechanism and a 1D convolutional network to process the fused phoneme embeddings, capturing the relationship and contextual information between the fused phonemes and the mixed timbre embeddings.

[0158] The obtained fused phoneme and mixed timbre Embedding sequence is input into the length regulator to adjust the length of the Embedding.

[0159] The Embedding sequence processed by the length adjuster will be sent to the positional encoding module for positional encoding to add further position information to maintain the order information of the Embedding sequence.

[0160] Then, a second FFT Block process is performed. The Embedding sequence after the second position encoding is sent to M FFT Block modules again to further capture the relationship and contextual information between the fused phonemes and the mixed timbre Embedding.

[0161] The fused phoneme embedding sequence processed as above will be input into the linear layer, which will implement linear transformation and convert the fused phoneme embedding sequence representation into the final sound feature mel-spectrogram.

[0162] Then, as described above and shown in FIG6 , the mel-spectrogram may be processed by a second acoustic model, such as a vocoder, to obtain synthesized speech.

[0163] In some embodiments of the present application, the audio processing method may be a voice conversion technology.

[0164] In some embodiments of the present application, for example, in the above-mentioned step S510, when the audio processing method is applied to speech conversion technology, the first control sequence may be a speech sequence to be converted into a target speech, and the acoustic neural network may be a speech conversion (VC) model.

[0165] In some embodiments of the present application, for example, in steps S520 to S550, the timbre mixing scheme described in the embodiments of the present application may be used to perform timbre mixing processing on multiple timbre feature sequences.

[0166] In a preferred embodiment of the present application, it is also conceivable to extract the timbre of the speech to be converted for mixing when performing speech conversion.

[0167] In the embodiment shown in FIG8 , when the audio processing method is applied to speech conversion technology, the above step S510 may further include the following steps:

[0168] S511: Extracting timbre features from the speech sequence to be converted to obtain a first timbre feature sequence.

[0169] In an embodiment of the present application, the speech sequence to be converted may be preprocessed by noise reduction, pre-emphasis, and framing, and then the preprocessed speech sequence to be converted (original speech) may be subjected to feature extraction and feature selection to obtain a first timbre feature sequence containing the timbre features of the speech sequence to be converted. Any suitable method may be used to extract the speech feature sequence, which will not be described in detail herein.

[0170] S512: Acquire one or more pre-provided second timbre feature sequences.

[0171] In a preferred embodiment, the second timbre feature sequence may be, for example, a preferred voice that can be used to beautify and optimize the original voice.

[0172] Preferably, the timbre of the original speech can be used as the main factor. Accordingly, as shown in FIG9 , the above step S530 may further include the following steps: reducing the dimension of the first timbre feature sequence to obtain a first reduced-dimensional timbre sequence; reducing the dimension of the one or more second timbre feature sequences to obtain one or more second reduced-dimensional timbre sequences. Accordingly, the above step S540 may further include the following steps: weighted summing the first reduced-dimensional timbre sequence and the one or more second reduced-dimensional timbre sequences to obtain a fused reduced-dimensional timbre sequence. In such an embodiment, the weight of the first reduced-dimensional timbre sequence is greater than the total weight of the one or more second reduced-dimensional timbre sequences.

[0173] For example, in a specific example, the fusion weight formula may be set as follows: Embedding_pca_X = a*Embedding_pca_A + b*Embedding_pca_B + c*Embedding_pca_C

[0174] Where a, b, and c are weight coefficients of the first reduced-dimensionality timbre sequence Embedding_pca_A and the two second reduced-dimensionality timbre sequences Embedding_pca_B and Embedding_pca_C, respectively, and a+b+c=1, and a>(b+c). This ultimately generates the fused reduced-dimensionality timbre sequence Embedding_pca_x.

[0175] Those skilled in the art will understand that the audio processing method of the embodiment of the present application, here the speech conversion method, can be implemented based on any suitable speech conversion model, as long as it can be combined with the timbre mixing method described in the embodiment of the present application, including but not limited to any speech conversion model based on the self-attention mechanism, and the present application does not limit it here. In an optional embodiment, the speech conversion model can be appropriately modified to be suitable for inputting the speech sequence to be converted and the mixed timbre sequence (for example, in the form of embedding).

[0176] For example, as shown in FIG8 , the speech conversion model may include a feature extraction module for extracting features of the speech sequence to be converted and the mixed timbre sequence, converting the features into an intermediate representation, such as a mel-spectrogram, through an acoustic model, and mapping the source speaker's features to the target speaker's features (including mixed timbre features) through feature conversion. The target speech waveform is then reconstructed and generated, for example, through a vocoder, such as WaveNet, WaveGlow, and WaveRNN, etc., which are not limited here. The vocoder can convert the mel-spectrogram into a final audio waveform, outputting a target audio with a mixed timbre that retains the speech content of the speech sequence to be converted.

[0177] In some embodiments of the present application, the audio processing method may be a singing synthesis method.

[0178] In some embodiments of the present application, for example, in the above-mentioned step S510, when the audio processing method may be a singing synthesis method, the first control sequence may be a lyrics text sequence for synthesizing the target singing, and the acoustic neural network may be a singing synthesis model.

[0179] In some embodiments of the present application, for example, in steps S520 to S550, the timbre mixing scheme described in the embodiments of the present application may be used to perform timbre mixing processing on multiple timbre feature sequences.

[0180] In some embodiments of the present application, such as the above-mentioned step S510, the following steps may be further included: obtaining a third control sequence for controlling the target singing melody, wherein the third control sequence includes one or more melody features of the song corresponding to the lyrics text sequence, such as a pitch sequence obtained by pitch feature extraction, as shown in Figure 9.

[0181] In some embodiments of the present application, the melody feature may be the pitch information of the song corresponding to the lyrics text sequence, specifically, the fundamental frequency (F0) information of the song. The change of the fundamental frequency constitutes the main line of the melody in the song.

[0182] Accordingly, the above step S560 will include: inputting the first control sequence (text sequence), the second control sequence (mixed timbre sequence) and the third control sequence (melody sequence) into the singing synthesis model to obtain the target singing.

[0183] Those skilled in the art will understand that the audio processing method of the embodiment of the present application, here the singing synthesis method, can be implemented based on any suitable singing synthesis model, as long as it can be combined with the timbre mixing method described in the embodiment of the present application. Preferably, the singing synthesis model is not limited to the singing synthesis model based on the extension of the speech synthesis model and / or any speech conversion model based on the self-attention mechanism, and the present application does not limit this. In an optional embodiment, the singing synthesis model can be appropriately modified to be suitable for inputting the speech sequence to be converted and the mixed timbre sequence (for example, in the form of embedding).

[0184] For example, as shown in Figure 9, the singing synthesis model may include a first acoustic model for extracting the features of the speech sequence to be converted, the mixed timbre sequence, and the pitch sequence, and converting them into an intermediate representation such as a Mel-spectrogram through the acoustic model, and generating a target singing waveform through feature conversion.

[0185] The timbre mixing method and audio processing method of the embodiments of the present application realize the precise extraction of the timbre attributes of the timbre feature sequence through cluster analysis, and by reducing the dimensionality of the obtained multiple high-dimensional timbre feature sequences to low-dimensional space for fusion, the obtained fused reduced-dimensional timbre sequence can not only retain multiple timbre attributes, but also remove noise and interference information. The fused reduced-dimensional timbre sequence is finally upgraded to obtain a mixed timbre sequence for controlling the output of the acoustic neural network. The mixed timbre sequence is used in the acoustic model to realize natural, stable and realistic audio output controlled by the mixed timbre sequence. The timbre mixing method and audio processing method of the present application can also be combined with Fastspeech and a variety of existing audio processing models to realize fast, stable and natural mixed timbre audio output.

[0186] In some embodiments of the present application, as shown in FIG10 , a timbre mixing device 1000 is further provided. The timbre mixing device 1000 includes:

[0187] An acquisition unit 1010 is configured to acquire a plurality of timbre feature sequences;

[0188] A dimension reduction unit 1020 is configured to reduce the dimension of the plurality of timbre feature sequences to obtain a plurality of reduced-dimensional timbre sequences;

[0189] A fusion unit 1030 is configured to fuse the multiple reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence; and

[0190] The dimension increasing unit 1040 is configured to increase the dimension of the fused reduced-dimensionality timbre sequence to obtain a mixed timbre sequence for controlling the audio timbre output by the acoustic neural network.

[0191] In some embodiments of the present application, as shown in FIG11 , an audio processing device 1100 is further provided. The audio processing device 1100 includes:

[0192] A first acquiring unit 1110 is configured to acquire a first control sequence for controlling target audio content;

[0193] The second acquisition unit 1120 is configured to acquire a plurality of timbre feature sequences;

[0194] A dimension reduction unit 1130 is configured to reduce the dimension of the plurality of timbre feature sequences to obtain a plurality of reduced-dimensional timbre sequences;

[0195] a fusion unit 1140 configured to fuse the plurality of reduced-dimensionality timbre sequences to obtain a fused reduced-dimensionality timbre sequence;

[0196] A dimension increasing unit 1150 is configured to increase the dimension of the fused reduced-dimensionality timbre sequence to obtain a second control sequence for controlling the target audio timbre; and

[0197] The input unit 1160 is configured to input the first control sequence and the second control sequence into a given acoustic neural network to obtain target audio.

[0198] The method and its steps, sub-steps and features in the embodiments of the present application can be combined with the device of the embodiments of the present application in a non-contradictory manner, and the device and its components, modules, units and features in the embodiments of the present application can be combined with the method of the embodiments of the present application in a non-contradictory manner.

[0199] In some embodiments of the present application, an electronic device is provided, which includes a processor and a memory storing a computer program, and the processor is configured to implement any method of the embodiments of the present application when running the computer program.

[0200] FIG12 shows a schematic diagram of an electronic device 1200 that can be used to implement the method or realize the embodiment of the present application. In some embodiments, the number of electronic devices may be more or less than the number shown. In some embodiments, a single or multiple electronic devices may be used for implementation. In some embodiments, cloud-based or distributed electronic devices may also be used for implementation.

[0201] As shown in Figure 12, the electronic device 1200 includes a processor 1210 and a memory 1220. The processor is used to execute programs stored in the memory, which can implement the methods, steps or functions described in the above embodiments when executed by a computer. The processor 1210 may include various types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. The processor 1210 and the memory 1220 are interconnected via a bus 1230. An input / output (I / O) interface and the like can also be connected to the bus 1230.

[0202] Although not shown in the figure, an embodiment of the present application also provides a storage medium that stores computer programs that are configured to execute speech synthesis model training and speech synthesis methods related to any embodiment at runtime.

[0203] It should be understood that all operations in the method described above are merely exemplary, and the present disclosure is not limited to any operation in the method or the order of these operations, but should cover all other equivalent transformations under the same or similar concept.

[0204] It should also be understood that all modules in the above-described device can be implemented in various ways. These modules can be implemented as hardware, software, or a combination thereof. In addition, any module in these modules can be further divided into submodules or combined together in function.

[0205] Processor has been described in conjunction with various devices and methods.These processors can be implemented using electronic hardware, computer software or its arbitrary combination.Whether these processors are implemented as hardware or software will depend on specific application and the overall design constraint imposed on the system.As an example, the processor provided in this disclosure, any part of the processor or any combination of processors can be implemented as microprocessor, microcontroller, digital signal processor (DSP), field programmable gate array (FPGA), programmable logic device (PLD), state machine, gate logic, discrete hardware circuit and other suitable processing components configured for performing the various functions described in this disclosure.The function of the processor provided in this disclosure, any part of the processor or any combination of processors can be implemented as software performed by microprocessor, microcontroller, DSP or other suitable platform.

[0206] Software should be broadly considered to mean instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, running threads, processes, functions, etc. Software can reside in a computer-readable medium. A computer-readable medium can include, for example, a memory, which can be, for example, a magnetic storage device (e.g., a hard disk, a floppy disk, a magnetic stripe), an optical disk, a smart card, a flash memory device, a random access memory (RAM), a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, or a removable disk. Although the memory is shown as being separate from the processor in various aspects provided in the present disclosure, the memory can also be located inside the processor (e.g., a cache or register).

[0207] The above description is provided to enable any person skilled in the art to implement the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents of the elements of the various aspects described in this disclosure that are known or to be known to those skilled in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Industrial Applicability

[0208] Through the above-mentioned timbre mixing method and related audio processing methods, the dimensionality reduction processing and low-dimensional space fusion of multiple high-dimensional timbre feature sequences are carried out to obtain a fused dimensionality reduction timbre sequence that effectively eliminates noise and interference information while retaining multiple timbre attributes. By increasing the dimensionality of the fused dimensionality reduction timbre sequence, a mixed timbre sequence is generated to ultimately control the output of the acoustic neural network. Inputting it into the acoustic model can produce natural, stable and realistic audio output. In addition, the timbre mixing method and related audio processing method of the present application are also compatible with the Fastspeech model and other existing audio processing models, which improves the processing speed, ensures the stability and naturalness of the output, and is suitable for a variety of audio output requirements.

Claims

1. A timbre mixing method, characterized in that, The tone mixing method includes: Obtaining a plurality of tone feature sequences; Reducing the dimension of the plurality of tone feature sequences to obtain a plurality of dimension-reduced tone sequences; Fusing the plurality of dimension-reduced tone sequences to obtain a fused dimension-reduced tone sequence; and Increasing the dimension of the fused dimension-reduced tone sequence to obtain a mixed tone sequence for controlling the audio tone output by the acoustic neural network.

2. The tone mixing method according to claim 1, wherein The reducing the dimension of the plurality of tone feature sequences to obtain a plurality of dimension-reduced tone sequences includes: Reducing the dimension of the plurality of tone feature sequences by using PCA transformation, and determining the principal component coefficients and means corresponding to the dimension reduction by the principal component analysis method; The increasing the dimension of the fused dimension-reduced tone sequence to obtain a mixed tone sequence for controlling the audio tone output by the acoustic neural network includes: Based on the determined principal component coefficients and means, increasing the dimension of the mixed tone sequence by using inverse PCA transformation.

3. The tone mixing method according to claim 1 or 2, characterized in that The fusing the plurality of dimension-reduced tone sequences to obtain a fused dimension-reduced tone sequence includes: Performing weighted summation on the plurality of dimension-reduced tone sequences to obtain the fused dimension-reduced tone sequence.

4. The tone mixing method according to claim 1, wherein The obtaining a plurality of tone feature sequences includes: Obtaining a plurality of initial tone feature sequences: Performing clustering processing on the plurality of initial tone feature sequences; Generating the plurality of tone feature sequences based on the clustering centers of the plurality of initial tone feature sequences.

5. The tone mixing method according to claim 4, wherein The reducing the dimension of the plurality of tone feature sequences to obtain a plurality of dimension-reduced tone sequences includes: Determining the dimension reduction dimension of the plurality of tone feature sequences based on the clustering number of the plurality of initial tone feature sequences.

6. An audio processing method, characterized in that, The audio processing method includes: Obtaining a first control sequence for controlling target audio content; Obtaining a plurality of tone feature sequences; Reducing the dimension of the plurality of tone feature sequences to obtain a plurality of dimension-reduced tone sequences; Fusing the plurality of dimension-reduced tone sequences to obtain a fused dimension-reduced tone sequence; Increasing the dimension of the fused dimension-reduced tone sequence to obtain a second control sequence for controlling the target audio tone; and Inputting the first control sequence and the second control sequence into a given acoustic neural network to obtain the target audio.

7. The audio processing method according to claim 6, characterized in that, The reducing the dimension of the plurality of tone feature sequences to obtain a plurality of dimension-reduced tone sequences includes: reducing the dimension of the plurality of tone feature sequences by using PCA transformation, and determining the principal component coefficients and means corresponding to the dimension reduction by the principal component analysis method; The increasing the dimension of the fused dimension-reduced tone sequence to obtain a mixed tone sequence for controlling the audio tone output by the acoustic neural network includes: Based on the determined principal component coefficients and means, increasing the dimension of the mixed tone sequence by using inverse PCA transformation.

8. The audio processing method according to claim 6 or 7, characterized in that The fusing the plurality of dimension-reduced tone sequences to obtain a fused dimension-reduced tone sequence includes: Performing weighted summation on the plurality of dimension-reduced tone sequences to obtain the fused dimension-reduced tone sequence.

9. The audio processing method according to claim 6, wherein The obtaining a plurality of tone feature sequences includes: Obtaining a plurality of initial tone feature sequences: Performing clustering processing on the plurality of initial tone feature sequences; Generating the plurality of tone feature sequences based on the clustering centers of the plurality of initial tone feature sequences.

10. The audio processing method according to claim 9, wherein Perform dimensionality reduction on the multiple timbre feature sequences to obtain multiple dimensionality-reduced timbre sequences, including: Determine the dimensionality reduction dimension of the multiple timbre feature sequences based on the clustering number of the multiple initial timbre feature sequences.

11. The audio processing method according to any one of claims 6 to 10, characterized in that, The audio processing method is a text-to-speech (TTS) method. Among them, the first control sequence is a text sequence for synthesizing target speech, and the acoustic neural network is a speech synthesis model.

12. The audio processing method according to any one of claims 6 to 10, characterized in that, The audio processing method is a voice conversion (VC) method. The first control sequence is a speech sequence to be converted into target speech, and the acoustic neural network is a voice conversion model.

13. The audio processing method according to claim 12, wherein The obtaining of the multiple timbre feature sequences includes: extracting timbre features from the speech sequence to be converted to obtain a first timbre feature sequence, and obtaining one or more second timbre feature sequences provided in advance; The performing of dimensionality reduction on the multiple timbre feature sequences to obtain multiple dimensionality-reduced timbre sequences includes: performing dimensionality reduction on the first timbre feature sequence to obtain a first dimensionality-reduced timbre sequence, and performing dimensionality reduction on the one or more second timbre feature sequences to obtain one or more second dimensionality-reduced timbre sequences; The fusing of the multiple dimensionality-reduced timbre sequences to obtain a fused dimensionality-reduced timbre sequence includes: performing weighted summation on the first dimensionality-reduced timbre sequence and the one or more second dimensionality-reduced timbre sequences to obtain the fused dimensionality-reduced timbre sequence.

14. The audio processing method according to claim 13, wherein Wherein the weight of the first dimensionality-reduced timbre sequence is greater than the total weight of the one or more second dimensionality-reduced timbre sequences.

15. The audio processing method according to any one of claims 6 to 10, characterized in that, The audio processing method is a singing synthesis method. The first control sequence is a lyrics text sequence for synthesizing target singing, and the acoustic neural network is a singing synthesis model. The audio processing method further includes: obtaining a third control sequence for controlling the melody of the target singing, where the third control sequence includes one or more melody features of the song corresponding to the lyrics text sequence; The inputting of the first control sequence and the second control sequence into a given acoustic neural network to obtain a target audio includes: inputting the first control sequence, the second control sequence, and the third control sequence into the singing synthesis model to obtain the target singing.

16. A timbre mixing device, characterized in that, The timbre mixing device includes: An obtaining unit configured to obtain multiple timbre feature sequences; A dimensionality reduction unit configured to perform dimensionality reduction on the multiple timbre feature sequences to obtain multiple dimensionality-reduced timbre sequences; A fusing unit configured to fuse the multiple dimensionality-reduced timbre sequences to obtain a fused dimensionality-reduced timbre sequence; and A dimensionality increase unit configured to perform dimensionality increase on the fused dimensionality-reduced timbre sequence to obtain a mixed timbre sequence for controlling the audio timbre output by the acoustic neural network.

17. An audio processing device, characterized in that, The audio processing device includes: A first obtaining unit configured to obtain a first control sequence for controlling target audio content; A second obtaining unit configured to obtain multiple timbre feature sequences; A dimensionality reduction unit configured to perform dimensionality reduction on the multiple timbre feature sequences to obtain multiple dimensionality-reduced timbre sequences; A fusing unit configured to fuse the multiple dimensionality-reduced timbre sequences to obtain a fused dimensionality-reduced timbre sequence; A dimension-raising unit configured to raise the dimension of the fused and dimension-reduced timbre sequence to obtain a second control sequence for controlling the target audio timbre; and An input unit configured to input the first control sequence and the second control sequence into a given acoustic neural network to obtain the target audio.

18. The audio processing device according to claim 17, wherein The audio processing device includes a speech synthesis device, wherein the first control sequence is a text sequence for synthesizing the target speech, and the acoustic neural network is a speech synthesis model.

19. The audio processing device according to claim 17, wherein The audio processing device includes a voice conversion device, wherein the first control sequence is a voice sequence to be converted for converting into the target voice, and the acoustic neural network is a voice conversion model.

20. The audio processing device according to claim 17, wherein The audio processing device includes a singing synthesis device, wherein the first control sequence is a lyric text sequence for synthesizing the target singing, and the acoustic neural network is a singing synthesis model. The audio processing device further includes: a third acquisition unit configured to acquire a third control sequence for controlling the target singing melody, wherein the third control sequence includes one or more melody features of the song corresponding to the lyric text sequence; wherein the input unit is configured to input the first control sequence, the second control sequence, and the third control sequence into the singing synthesis model to obtain the target singing.

21. An electronic device, characterized in that, Comprising a processor and a memory storing a computer program, the processor being configured to implement the method according to any one of claims 1 to 15 when running the computer program.

22. A storage medium, characterized in that, The storage medium stores a computer program, the computer program being configured to implement the method according to any one of claims 1 to 15 when being run.

Citation Information

Patent Citations

  • Speech synthesis method and device, equipment, storage medium and program product

    CN114242033A

  • Singing synthesis method, device and equipment and storage medium

    CN115457923A

  • Many-to-many real-time voice inflexion method and device and storage medium

    CN116959422A

  • Timbre mixing method and device, audio processing method and device, electronic equipment and storage medium

    CN117975933A

  • Speech synthesis method and apparatus, electronic device, and readable storage medium

    WO2023045954A1