Multi-scene vocoder training method and device based on value zero decomposition, and audio generation method and device

By using a multi-scenario vocoder training method based on value-zero decomposition, and generating Mel spectrum samples for different scenarios using value space and null space models, the problem of computational resource consumption in existing vocoders under multiple scenarios is solved, achieving savings in computational resources and improvement in audio prediction accuracy.

CN120998225AActive Publication Date: 2025-11-21INST OF ACOUSTICS CHINESE ACAD OF SCI
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510768261.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-11-21
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Existing vocoders can only adapt to one Mel spectrum configuration, which means that different vocoders for different scenarios need to be trained when generating audio in multiple scenarios, resulting in a large consumption of computing resources.

Method used

By using a multi-scenario vocoder training method based on value-zero decomposition, Mel spectrum samples with different parameter configurations are generated using value space and null space models. The vocoder model parameters are then adjusted by adjusting the loss value to adapt to Mel filter configurations in various scenarios, thus avoiding the need to train a vocoder model for each parameter configuration separately.

Benefits of technology

The vocoder model was adapted to various Mel filter configurations in different scenarios, saving computational resources and improving the accuracy of audio prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998225A_ABST
    Figure CN120998225A_ABST
Patent Text Reader

Abstract

The invention provides a value zero decomposition-based multi-scene vocoder training method and device. The method comprises the following steps of: obtaining a plurality of audio samples; determining a Mel spectrum sample of each audio sample; the Mel-spectrum samples and the Mel filters corresponding to the Mel-spectrum samples are input into vocoder models, the vocoder models are trained, audio prediction results corresponding to the Mel-spectrum samples are obtained, the vocoder models comprise a value space model and a null space model, and the value space model and the null space model correspond to each Mel-spectrum sample. The vocoder model generates a first feature through a value space model according to the Mel spectrum sample and a Mel filter corresponding to the Mel spectrum sample, and the vocoder model performs linear amplitude domain reconstruction and phase information recovery on the first feature through a null space model to obtain an audio prediction result; determining a loss value between the audio prediction result and the audio sample; and according to the loss value, adjusting parameters of the vocoder model to obtain a trained vocoder model. Thus, the vocoder model can adapt to various different scenes, vocoders in different scenes do not need to be trained, and computing resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, and in particular to a training method of a multi-scene vocoder based on value zero decomposition, an audio generation method and device. BACKGROUND

[0002] A vocoder is used for speech synthesis, audio generation, music generation, and speech enhancement tasks. In speech synthesis, audio generation, and music generation tasks, when a user gives a textual description, a generative model is used to estimate the corresponding Mel spectrogram of the audio. In the speech enhancement task, the Mel spectrogram of the noisy speech is extracted, and then the Mel spectrogram of the clean speech is recovered through a speech enhancement algorithm. The vocoder is responsible for converting it into the corresponding time-domain waveform. Although existing deep learning-based neural vocoders can effectively realize the conversion of Mel spectrogram to waveform.

[0003] However, the current vocoder, after being trained, can only adapt to one configuration of the Mel spectrogram. When audio in multiple scenes needs to be generated, vocoders in different scenes need to be trained to adapt to the configuration of the Mel spectrogram in different scenes. In this case, a large amount of computing resources is required, resulting in a large occupation of computing resources. SUMMARY

[0004] The present application provides a training method of a multi-scene vocoder based on value zero decomposition, an audio generation method and device, which can realize that the vocoder model can adapt to multiple different scenes, thereby saving computing resources without training vocoders in different scenes.

[0005] In a first aspect, the present application provides a training method of a multi-scene vocoder based on value zero decomposition, the method comprising:

[0006] obtaining a plurality of audio samples and label information of each audio sample in the plurality of audio samples;

[0007] determining a Mel spectrogram sample of each audio sample, wherein the Mel spectrogram samples of different audio samples are generated by Mel filters having different parameter configurations, and the Mel filters having different parameter configurations adapt to different scenes;

[0008] inputting the Mel spectrogram samples and the Mel filters corresponding to the Mel spectrogram samples into a vocoder model, training the vocoder model, and obtaining an audio prediction result corresponding to each Mel spectrogram sample, wherein the vocoder model comprises a value space model and a zero space model, the vocoder model generates a first feature from the Mel spectrogram samples and the Mel filters corresponding to the Mel spectrogram samples through the value space model, and the vocoder model performs linear amplitude domain reconstruction and phase information recovery on the first feature through the zero space model to obtain the audio prediction result;

[0009] determine a loss value between the audio prediction result and the audio sample;

[0010] adjust parameters of the vocoder model according to the loss value, to obtain a trained vocoder model.

[0011] According to the above scheme, the mel filter with different parameter configurations is used to generate the mel spectrum sample of the audio sample in different scenes. When training the vocoder model, the mel filter and the mel spectrum sample in different scenes are used, so that the trained vocoder model can adapt to the parameter configuration of different mel filters. Therefore, it is not necessary to train the vocoder model for each parameter configuration of the mel filter, thereby saving the computing resources. In the process of training the vocoder model, in order to enable the vocoder model to adapt to the mel filter in different scenes (i.e. different parameter configurations), the value space model is used in the vocoder model to process the mel spectrum sample and the mel filter, so as to realize the conversion of the mel spectrum sample from the mel domain to the linear domain, so that the zero space model can continue to perform the subsequent linear amplitude domain reconstruction and phase information recovery, to obtain the audio prediction result. According to the loss value between the audio prediction result and the audio sample, the parameters of the vocoder model are adjusted, so as to obtain the trained vocoder.

[0012] In a possible implementation, the mel spectrum sample of each audio sample is determined, including:

[0013] According to the pre-configured band interval and the frequency interval, the mel filter with different parameter configurations is generated;

[0014] According to the audio sample, the target mel filter corresponding to the audio sample and the target parameter configuration corresponding to the target mel filter are determined;

[0015] According to the target mel filter and the target parameter configuration, the mel spectrum sample of the audio sample is generated.

[0016] In this way, the mel spectrum sample of the audio sample can be generated in different scenes.

[0017] In a possible implementation, the first feature satisfies the following formula:

[0018]

[0019] wherein, denotes the first feature, A + denotes the pseudo-inverse matrix of the mel filter, A + has a size of |X mel |denotes the mel spectrum sample.

[0020] In this way, the mel-spectrum sample is converted from the mel domain to the linear domain, the mel-spectrum sample of the audio sample in different scenes is used to train the vocoder model, and thus the trained vocoder model can be adapted to the audio in different scenes.

[0021] In a possible implementation, the vocoder model performs linear amplitude domain reconstruction and phase information recovery on the first feature through the null space model to obtain an audio prediction result, including:

[0022] The vocoder model compresses and encodes the spectral feature in the first feature in the order from low frequency to high frequency through the null space model to obtain an encoded feature.

[0023] The encoded feature is modeled to obtain a target audio.

[0024] The target audio is subjected to amplitude reconstruction and phase reconstruction to obtain the audio prediction result.

[0025] In this way, the target audio is subjected to amplitude reconstruction and phase reconstruction through the null space model, so that the missing spectral details and phase information of the mel-spectrum sample can be recovered, and the accuracy of the audio prediction result is improved.

[0026] In a possible implementation, the audio prediction result satisfies the following formula:

[0027]

[0028] wherein, represents an amplitude estimation of the first feature, represents a phase estimation of the first feature,

[0029] represents a representation function of the null space model, and I represents a unit matrix.

[0030] In this way, the target audio is subjected to amplitude reconstruction and phase reconstruction through the null space model, so that the missing spectral details and phase information of the mel-spectrum sample can be recovered, and the accuracy of the audio prediction result is improved.

[0031] In a possible implementation, the loss value includes at least one of the following:

[0032] a logarithmic amplitude loss value, a phase loss value, a real-imaginary part loss value, a mel-spectrum loss value, and a consistency loss value.

[0033] The logarithmic amplitude loss value satisfies the following formula:

[0034]

[0035] wherein, represents the value of the audio prediction result at frequency index f and time frame index t, X f,t represents the value of the label information at frequency index f and time frame index t, F and T represent the size of the frequency dimension and the number of time frames respectively, log(.) represents the logarithm operation, and ‖.‖ 2 represents the L2 norm;

[0036] phase loss value satisfies the following formula:

[0037]

[0038] wherein, "*" represents a two-dimensional convolution operation, f AW (·) represents the inverse closed function, round(.) represents the rounding operation, and ‖.‖ 1 represents the L1 norm;

[0039] real-imaginary loss value satisfies the following formula:

[0040]

[0041] wherein, and represent the real part and the imaginary part operation on a complex number respectively.

[0042] mel-spectrogram loss value satisfies the following formula:

[0043]

[0044] wherein, F mel is the number of compressed mel-bands in the frequency dimension, represents the mel-spectrogram of the audio prediction result at frequency index f and time frame index t, represents the mel-spectrogram of the label information at frequency index f and time frame index t;

[0045] consistency loss value satisfies the following formula:

[0046]

[0047] wherein, represents the consistency estimation of .

[0048] In a second aspect, the present application provides an audio generation method, which comprises:

[0049] obtaining a mel-spectrogram;

[0050] obtaining a mel-filter of the mel-spectrogram;

[0051] inputting the mel-spectrogram and the mel-filter into the vocoder model provided in the first aspect to generate an audio.

[0052] According to the above scheme, the mel spectrum samples of the audio samples in different scenes are generated based on the mel filters with different parameter configurations. When training the vocoder model, the mel filters and the mel spectrum samples in different scenes are used, so that the trained vocoder model can adapt to the parameter configurations of different mel filters, and it is not necessary to train the vocoder model for each parameter configuration of the mel filter, thereby saving computing resources. In the process of training the vocoder model, in order to enable the vocoder model to adapt to the mel filters in different scenes (i.e. different parameter configurations), the value space model in the vocoder model is used to process the mel spectrum samples and the mel filters, so as to realize the conversion of the mel spectrum samples from the mel domain to the linear domain, so that the zero space model continues to perform subsequent linear amplitude domain reconstruction and phase information recovery to obtain an audio prediction result. According to the loss value between the audio prediction result and the audio sample, the parameters of the vocoder model are adjusted, so as to obtain the trained vocoder. Through the above process, the vocoder trained can generate audio corresponding to the mel spectrum in different scenes, thereby improving the audio generation efficiency in different scenes.

[0053] In a third aspect, the present application provides a device for training a multi-scene vocoder based on value zero decomposition, which comprises:

[0054] The acquisition module is configured to acquire a plurality of audio samples and label information of each audio sample in the plurality of audio samples.

[0055] The first determination module is configured to determine a mel spectrum sample of each audio sample, and the mel spectrum samples of different audio samples are generated by mel filters with different parameter configurations.

[0056] The training module is configured to input the mel spectrum samples and the mel filters corresponding to the mel spectrum samples into a vocoder model, train the vocoder model, and obtain an audio prediction result corresponding to each mel spectrum sample. The vocoder model comprises a value space model and a zero space model. The vocoder model generates a first feature according to the mel spectrum samples and the mel filters corresponding to the mel spectrum samples through the value space model. The vocoder model performs linear amplitude domain reconstruction and phase information recovery on the first feature through the zero space model to obtain the audio prediction result.

[0057] The second determination module is configured to determine a loss value between the audio prediction result and the audio sample.

[0058] The adjustment module is configured to adjust the parameters of the vocoder model according to the loss value, and obtain a trained vocoder model.

[0059] According to the above scheme, the mel spectrum samples of the audio samples in different scenes are generated based on the mel filters with different parameter configurations. When training the vocoder model, the mel filters and the mel spectrum samples in different scenes are used, so that the trained vocoder model can adapt to the parameter configurations of different mel filters, and it is not necessary to train the vocoder model for each parameter configuration of the mel filter, thereby saving computing resources. In the process of training the vocoder model, in order to enable the vocoder model to adapt to the mel filters in different scenes (i.e., different parameter configurations), the value space model is used in the vocoder model to process the mel spectrum samples and the mel filters, so as to realize the conversion of the mel spectrum samples from the mel domain to the linear domain, so that the zero space model continues to perform subsequent linear amplitude domain reconstruction and phase information recovery to obtain an audio prediction result. According to the loss value between the audio prediction result and the audio sample, the parameters of the vocoder model are adjusted, so as to obtain the trained vocoder.

[0060] In a possible implementation, the first determining module is configured to:

[0061] generate the mel filters with different parameter configurations according to the preconfigured band interval and the frequency interval;

[0062] determine, according to the audio sample, a target mel filter corresponding to the audio sample and a target parameter configuration corresponding to the target mel filter;

[0063] generate a mel spectrum sample of the audio sample according to the target mel filter and the target parameter configuration.

[0064] In this way, the mel spectrum sample of the audio sample can be generated in different scenes.

[0065] In a possible implementation, the first feature satisfies the following formula:

[0066]

[0067] wherein, denotes the first feature, A + denotes a pseudo-inverse matrix of the mel filter, A + has a size of |X mel denotes the mel spectrum sample.

[0068] In this way, the mel spectrum sample is converted from the mel domain to the linear domain, the mel sample of the audio sample in different scenes is used to train the vocoder model, so that the trained vocoder model can adapt to the audio in different scenes.

[0069] In a possible implementation, the training module is configured to:

[0070] The vocoder model compresses and encodes the spectral features in the first features in order from low frequency to high frequency through the null space model to obtain encoded features;

[0071] The target audio is modeled to obtain an audio prediction result.

[0072] The target audio is subjected to amplitude reconstruction and phase reconstruction to obtain an audio prediction result.

[0073] In this way, the target audio is subjected to amplitude reconstruction and phase reconstruction through the null space model, so that spectral details and phase information missing in the mel-spectral samples can be recovered, and the accuracy of the audio prediction result is improved.

[0074] In a possible implementation, the audio prediction result satisfies the following formula:

[0075]

[0076] wherein, represents an amplitude estimate of the first features, represents a phase estimate of the first features,

[0077] represents a representation function of the null space model, I represents a unit matrix, and iSTFT(.) represents an inverse transform of a short-time Fourier transform.

[0078] In this way, the target audio is subjected to amplitude reconstruction and phase reconstruction through the null space model, so that spectral details and phase information missing in the mel-spectral samples can be recovered, and the accuracy of the audio prediction result is improved.

[0079] In a possible implementation, the loss value includes at least one of the following:

[0080] a log-amplitude loss value, a phase loss value, a real-imaginary loss value, a mel-spectral loss value, and a consistency loss value.

[0081] The log-amplitude loss value satisfies the following formula:

[0082]

[0083] wherein, represents a value of the audio prediction result at a frequency index f and a time frame index t, X f,t represents a value of the label information at the frequency index f and the time frame index t, F and T represent a frequency dimension size and a time frame number respectively, log(.) represents a logarithm operation, and ‖.‖2 represents an L2 norm;

[0084] The phase loss value satisfies the following formula:

[0085]

[0086] wherein, "*" represents a two-dimensional convolution operation, f AW (·) represents an inverse closed function, round(.) represents an integer operation, and ||.||1represents an L1norm;

[0087] The real-imaginary part loss value satisfies the following formula:

[0088]

[0089] wherein, and respectively represent the real part and the imaginary part operation on a complex number.

[0090] The mel-spectrogram loss value satisfies the following formula:

[0091]

[0092] wherein, F mel is the number of mel-bands after compression in the frequency dimension, represents a mel-spectrogram of an audio prediction result at a frequency index f and a time frame index t, represents a mel-spectrogram of label information at a frequency index f and a time frame index t;

[0093] The consistency loss value satisfies the following formula:

[0094]

[0095] wherein, represents a consistency estimation on .

[0096] In a fourth aspect, the present application provides an audio generation device, which comprises:

[0097] A first obtaining module is configured to obtain a mel-spectrogram.

[0098] A second obtaining module is configured to obtain a mel-filter of the mel-spectrogram.

[0099] A generating module is configured to input the mel-spectrogram and the mel-filter into a vocoder model as provided in the first aspect, and generate an audio.

[0100] In a fifth aspect, the present application provides a computing device, which comprises a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the method as provided in the first aspect or any possible implementation manner of the first aspect, or the method as provided in the second aspect is implemented.

[0101] In a sixth aspect, the present application provides a computer storage medium, having stored therein instructions which, when executed on a computer, cause the computer to perform the method provided in the first aspect or any possible implementation manner of the first aspect, or perform the method provided in the second aspect. BRIEF DESCRIPTION OF DRAWINGS

[0102] Figure 1 A flowchart of a training method of a multi-scenario vocoder based on value zero decomposition is shown;

[0103] Figure 2 A training flowchart of a vocoder model is shown;

[0104] Figure 3 A flowchart of another training method of a multi-scenario vocoder based on value zero decomposition is shown;

[0105] Figure 4 An exemplary flowchart of a phase loss function is shown;

[0106] Figure 5 A flowchart of an audio generation method is shown;

[0107] Figure 6 A structural diagram of a training device of a multi-scenario vocoder based on value zero decomposition is shown;

[0108] Figure 7 A structural diagram of an audio generation device is shown;

[0109] Figure 8 A structural diagram of a computing device is shown. DETAILED DESCRIPTION

[0110] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the drawings.

[0111] In the description of the embodiments of the present application, the words “exemplary”, “for example”, or “for instance” are used to mean serving as an example, instance, or illustration. Any embodiment or design solution described as “exemplary”, “for example”, or “for instance” in the embodiments of the present application should not be interpreted as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of the words “exemplary”, “for example”, or “for instance” is intended to present concepts in a specific way.

[0112] In the description of the embodiments of the present application, the term "and / or" is merely an association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent three cases of A alone, B alone, and A and B together. In addition, unless otherwise specified, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals.

[0113] In addition, the terms "first" and "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features. The terms "include", "contain", "have" and their variants mean "include but not limited to", unless otherwise specifically emphasized.

[0114] The vocoder is used for speech synthesis, audio generation, music generation, and speech enhancement, etc. In speech synthesis, audio generation, and music generation, etc., when a user gives a textual description, the audio corresponding Mel spectrogram is estimated by a generative model; in the speech enhancement task, the Mel spectrogram of the noisy speech is extracted, and then the Mel spectrogram of the clean speech is recovered by the speech enhancement algorithm. The vocoder is responsible for converting it into the corresponding time-domain waveform. Although the existing deep learning-based neural vocoder can effectively realize the conversion of Mel spectrogram to waveform.

[0115] However, the current vocoder can only adapt to one configuration of Mel spectrogram after training. When audio in multiple scenarios needs to be generated, vocoders in different scenarios need to be trained to adapt to the configuration of Mel spectrogram in different scenarios. In such a case, a large amount of computing resources is required, resulting in a large amount of occupation of computing resources.

[0116] Based on this, the embodiment of the present application provides a training method and device of a multi-scene vocoder based on value zero decomposition. The mel filter based on different parameter configurations is used to generate the mel spectrum sample of the audio sample in different scenes. When training the vocoder model, the mel filter and the mel spectrum sample in different scenes are used, so that the trained vocoder model can adapt to the parameter configuration of different mel filters, and the vocoder model does not need to be trained separately for each parameter configuration of the mel filter, thereby saving the computing resources. In the process of training the vocoder model, in order to enable the vocoder model to adapt to the mel filter in different scenes (i.e. different parameter configurations), the value space model in the vocoder model is used to process the mel spectrum sample and the mel filter, so as to realize the conversion of the mel spectrum sample from the mel domain to the linear domain, so that the zero space model continues to perform the subsequent linear amplitude domain reconstruction and phase information recovery to obtain the audio prediction result. According to the loss value between the audio prediction result and the audio sample, the parameters of the vocoder model are adjusted, so as to obtain the trained vocoder.

[0117] Figure 1 The training method of a multi-scene vocoder based on value zero decomposition is provided by the embodiment of the present application. The training method of a multi-scene vocoder based on value zero decomposition provided by the embodiment of the present application can be applied to a device or platform with computing function and storage function, such as a computing device (such as a server, etc.), a computing device cluster (such as a cloud computing platform, etc.). As shown in Figure 1 The training method of a multi-scene vocoder based on value zero decomposition provided by the embodiment of the present application includes the following steps S101 to S106.

[0118] S101, obtaining a plurality of audio samples.

[0119] The audio sample can be a sample generated in different scenes.

[0120] In some embodiments, S101 includes the following process: first, generating mel filters with different parameter configurations according to the pre-configured band number interval and frequency interval; then, determining the target mel filter corresponding to the audio sample and the target parameter configuration corresponding to the target mel filter according to the audio sample; finally, generating the mel spectrum sample of the audio sample according to the target mel filter and the target parameter configuration. In this way, the mel spectrum sample of the audio sample can be generated in different scenes.

[0121] Exemplarily, as shown in Figure 2 In S101, the interval of the mel band number and the highest calculation frequency is selected, the sampling configuration is performed based on the interval of the mel band number and the highest calculation frequency, thereby obtaining a plurality of mel filter matrices, and the plurality of mel filter matrices are stored, so that a plurality of parameter configurations of the mel filter can be obtained. For example, it is assumed that the band number interval of the mel filter is wherein and respectively represent the minimum and maximum values of the number of bands, and the maximum frequency interval is wherein and respectively represent the minimum and maximum values thereof, and the minimum calculation frequency value f min is usually set to 0. When the interval is set, the interval of the number of bands is set to δ n , and the interval of the maximum frequency is set to δ f . In this case, M = M n M f parameter configurations of the Meller filter can be obtained, wherein M n and M f respectively represent the maximum number of samples in the interval of the number of bands and the maximum frequency, and the calculation method satisfies the following formula (1):

[0122]

[0123] wherein, represents the floor symbol. For example, and are respectively set to 64 and 128, and for an audio generation scene with a sampling rate f s = 24 kHz, δ and δ are respectively set to 8 kHz and 12 kHz, and it is verified through practice that such a setting is sufficient to cover most actual application scenarios. δ n and δ f are respectively set to 4 and 100 Hz. For the M parameter configurations, the Meller filter matrix corresponding to each configuration is calculated and stored, and the parameter configurations of different Meller filter matrices are different. For example, if the Meller filter matrix set is {A1,..., A M}, then the parameter configuration matrix set corresponding to the matrix set is {C1,..., C M}, wherein C m and A m respectively represent the mth parameter configuration and the corresponding Meller filter matrix, and m is a natural number, m = 1, 2,..., M.

[0124] S102, determine the Meller spectrum sample of each audio sample, and the Meller spectrum samples of different audio samples are generated by Meller filters with different parameter configurations. Different parameter configurations of the Meller filter are adapted to different scenes.

[0125] In the embodiments of the present application, for an audio sample x(n), the audio sample x(n) is transformed into a time-frequency domain by a short-time Fourier transform (STFT), and a time-frequency domain representation of the audio sample x(n) is where F and T represent the size of the frequency and time dimensions, respectively. For a mel filter with a parameter configuration of C m , a mel spectrum sample of the audio sample x(n) compressed by the mel filter satisfies the following formula (2):

[0126] |X mel |=A|X| (2)

[0127] where and represent a mel spectrum and an amplitude spectrum, respectively. Here, a configuration subscript (.) is ignored for the sake of brevity. m F mel is the number of mel bands after compression in the frequency dimension. For example, as shown in FIG. 2, in the process of model training, when an audio sample is obtained, a mel filter is sampled from a stored mel filter matrix, and the audio sample is processed based on the mel filter to obtain a mel spectrum sample. Figure 2

[0128] For example, the configuration of the short-time Fourier transform is as follows: a Hanning window is used as a window function, the size of each frame window is set to 1024, the frame shift point number is 256, the Fourier transform (FFT) point number is 1024, and the feature size F in the frequency dimension is 513 according to the conjugate symmetry. In network training, the mel spectrum sample |X mel | and the mel filter A are used as inputs of the network, and x(n) is used as a training target, that is, x(n) is label information of the mel spectrum sample.

[0129] S103, input the mel spectrum sample and the mel filter corresponding to the mel spectrum sample into the vocoder model, train the vocoder model, and obtain an audio prediction result corresponding to each mel spectrum sample, wherein the vocoder model includes a value space model and a zero space model, the vocoder model generates a first feature according to the mel spectrum sample and the mel filter corresponding to the mel spectrum sample through the value space model, and the vocoder model restores the linear amplitude domain and the phase information of the first feature through the zero space model to obtain the audio prediction result.

[0130] ​The vocoder model compresses and encodes the spectral features in the first features in an order from low frequency to high frequency through the null space model to obtain encoded features, models the encoded features to obtain target audio, and performs amplitude reconstruction and phase reconstruction on the target audio to obtain an audio prediction result. In this way, the amplitude reconstruction and the phase reconstruction are performed on the target audio through the null space model, so that spectral details and phase information of missing mel-spectrum samples can be recovered, and the accuracy of the audio prediction result is improved.

[0131] In some embodiments, the vocoder model is established based on a sub-band structure neural network (RNDVoC).

[0132] In the embodiments of the present application, as shown in Figure 3 , the vocoder model includes two parts, namely a range space model and a null space model. In the training process, the range space model aims to convert the input mel-spectrum from the mel domain to the linear domain by using the inverse transform of the mel filter A. Since the inverse transform is difficult to perfectly reconstruct, a pseudo-inverse form is actually used, that is, A + , for example, the pseudo-inverse that can be used is the Moore-Penrose pseudo-inverse. Therefore, the operation process of the range space model can be represented as the following formula (3):

[0133]

[0134] , wherein, is the output of the range space model, that is, the first features, the pseudo-inverse transform matrix A + has a size of |X mel | represents the mel-spectrum sample.

[0135] In this way, the mel-spectrum sample is converted from the mel domain to the linear domain, the mel-spectrum sample of the audio sample in different scenes is trained on the vocoder model, so that the trained vocoder model can adapt to the audio in different scenes.

[0136] The null space model aims to reconstruct the missing spectral details in the linear domain. It is worth noting that since the mel-spectrum loses the phase information, the null space model also needs to additionally recover the phase information, and the recovery process of the phase information satisfies the following formula (4):

[0137]

[0138] , wherein, is the amplitude estimation of the first features, a phase estimate representing the first feature,

[0139] is a representation function of the null space model. Following the value zero decomposition theory, the audio prediction result satisfies the following formula (5):

[0140]

[0141] As Figure 3 shown, the null space model mainly includes three parts, which are the subband encoding module, the subband decoding module and the dual-path modeling module. The subband encoding module utilizes the prior characteristic that the amount of information of the spectrum is different in different frequency intervals, and compresses and encodes the spectral features from low to high frequencies, from fine to coarse. Assuming that the number of compressed subbands is N, the amplitude of the subband index n in the linear domain is where F n is the frequency size of the subband, and the feature encoding process of the interval spectrum satisfies the following formula (6):

[0142]

[0143] where LN(.) represents layer normalization, and Conv1d(.) represents one-dimensional convolution. The N encoded features are spliced to obtain a three-dimensional representation where C represents the channel dimension size after encoding. Under the sampling rate of 24 kHz, the frequency interval division of the spectrum compression is as follows: the 0-3 kHz interval is merged into a subband every 250 Hz, the 3-7 kHz interval is merged into a subband every 500 Hz, the 7-10 kHz interval is merged into a subband every 1 kHz, and the interval above 10 kHz is merged into a subband, so N = 12 + 8 + 3 + 1 = 24.

[0144] After the subband encoding module, the output F in is input into the dual-path modeling module for target information modeling. The dual-path modeling module is composed of B dual-path blocks, and each dual-path block includes a cross-subband module and P subband modules. For the cross-subband module, given the input feature with a size of NxCxT, first perform a transpose operation, then use a group one-dimensional convolution with a convolution kernel size of J and a group number of H to model along the subband dimension; then, use a one-dimensional point convolution to compress the channel number to one quarter of the original, model the global subband features on the compressed features using a linear layer, and then use a one-dimensional point convolution to restore the channel number; finally, use the group one-dimensional convolution with the convolution kernel size of J and the group number of H again to model in the subband dimension. For the subband module, first perform a transpose operation, then use a group one-dimensional convolution with a convolution kernel size of J ′The deep one-dimensional convolution along the time axis is modeled, followed by a one-dimensional point convolution and a global response normalization layer, and finally output by a one-dimensional point convolution. The hyperparameter description and setting of the specific network are as shown in Table 1, where C = 256, B = 6, P = 2, J = 3, H = 8, J ′ = 7, and the parameter amount under this configuration is 9.48M, and the complexity is 24.98GMACs / 5s.

[0145] Table 1

[0146]

[0147] After passing through the double-path modeling module, the output is denoted as and is sent to the subband decoding module. The subband decoding module involves two branches, which are used to reconstruct the amplitude and phase information respectively. Taking the amplitude branch as an example, given the input feature O, the feature of the subband index n is which is passed through the decoding layer, and then the amplitude of the subband is recovered through the exponential operation exp(.), and then the N amplitudes are spliced along the frequency dimension to estimate the amplitude spectrum, which satisfies the following formula (7):

[0148]

[0149] where GELU(.) represents Gaussian Error Linear Unit, and Concat(.) represents splicing operation. In the phase branch, similar decoding operations are used as in the amplitude branch, except that the phase is estimated through the Atan2(.) operation at the end.

[0150] During training, the batch size of the network is 16, the generator iterates a total of one million steps, and the AdamW optimizer is used for training, where {β1, β2} are set to 0.8 and 0.99 respectively, the learning rate is initialized to 0.0002, and the exponential decay strategy is used to adjust the learning rate, and the decay rate is set to 0.999.

[0151] In S104, a loss value between the audio prediction result and the audio sample is determined.

[0152] The loss function can determine the gap between the audio prediction result and the audio sample, i.e., the loss value. In order to enable the vocoder model after training to accurately generate audio in different scenes, the loss function of the present embodiment includes two parts, namely a reconstruction loss function and an adversarial loss function. The reconstruction loss function includes a logarithmic amplitude loss function, a phase loss function, a real-imaginary part loss function, a mel loss function, and a consistency loss function.

[0153] wherein the log-magnitude loss function is the Mean-Square Error (MSE) between the log-estimated magnitude and the target magnitude spectrum, and the log-magnitude loss function satisfies the following equation (8):

[0154]

[0155] wherein, represents the value of the audio prediction result at frequency index f and time frame index t, X f,t represents the value of the label information at frequency index f and time frame index t, F and T represent the size of frequency dimension and the number of time frames respectively, log(.) represents the logarithm operation, and ‖.‖2 represents the L2 norm;

[0156] The phase loss function can be an omnidirectional phase loss function, as shown in Figure 4 , a fixed-parameter convolution kernel is designed to calculate the difference between a certain phase and the 8 adjacent phases in the phase spectrum, wherein the 5th 3x3 convolution calculation is the instantaneous phase. Given the estimated phase and the target phase Φ, the omnidirectional phase difference after convolution is obtained by and respectively. Then the loss is calculated by an Anti-Wrapping function. The loss function satisfies the following equation (9):

[0157]

[0158] wherein “*” represents the two-dimensional convolution operation, and the Anti-Wrapping function f AW (x) is defined as wherein round(.) represents the rounding operation, and ‖.‖1 represents the L1 norm.

[0159] The real-imaginary part loss loss function is the Mean Absolute Error (MAE) between the estimated real-imaginary part and the target real-imaginary part, and the real-imaginary part loss loss function satisfies the following equation (10):

[0160]

[0161] wherein, and represent the real part and the imaginary part operation on a complex number respectively.

[0162] The mel loss function is the Mean Absolute Error between the estimated mel spectrum and the target mel spectrum, and the mel loss function satisfies the following equation (11):

[0163] ​​

[0164] The consistency loss loss function satisfies the following formula (12):

[0165]

[0166] wherein, is a consistency estimate of The reconstruction loss function satisfies the following formula (13):

[0167]

[0168] wherein, {λ a ,λ p ,λ ri ,λ mel ,λ c} are the weighting coefficients of each loss term. For example, these weighting coefficients can be set to {45, 100, 45, 45, 45}. The specific weighting coefficients can be set according to the actual situation, and the embodiments of the present application are not limited specifically.

[0169] For the adversarial loss function, two discriminators are adopted, which are a multi-period discriminator (MPD) and a multi-resolution spectrogram discriminator (MRSD). For the MPD, it includes 5 sub-discriminators, and the period values are set to {2, 3, 5, 7, 11} respectively; for the MRSD, it includes 3 sub-discriminators, and the (window length, frame shift, Fourier transform coefficient) are set to {(512, 128, 512), (1024, 256, 1024), (2048, 512, 2048)} respectively. The hinge is adopted as the adversarial loss in the embodiments of the present application, and the adversarial loss function of the discriminator part satisfies the following formula (14):

[0170]

[0171] wherein, D m (s) and respectively represent the outputs of the target and the estimated audio in the sub-discriminator, and M is the number of sub-discriminators. For the generator, the adversarial loss function satisfies the following formula (15):

[0172]

[0173] In addition, a feature matching loss function is also combined, and the feature matching loss function satisfies the following formula (16):

[0174]

[0175] wherein, represents the output feature of the e th layer of the m th sub-discriminator, and E represents the total number of layers of the sub-discriminator. In the embodiments of the present application, for and The weighting coefficients {λ D ,λ g ,λ fm} are set to 1.

[0176] S105, adjusting the parameters of the vocoder model according to the loss value to obtain a trained vocoder model.

[0177] According to the embodiments of the present application, based on the mel filter with different parameter configurations, the mel spectrum samples of the audio samples in different scenes are generated. When training the vocoder model, through the mel filter and the mel spectrum samples in different scenes, the vocoder model obtained by training can adapt to the parameter configurations of different mel filters, without training the vocoder model for each parameter configuration of the mel filter, saving the computing resources. In the process of training the vocoder model, in order to make the vocoder model adapt to the mel filter in different scenes (i.e. different parameter configurations), the value space model is used in the vocoder model to process the mel spectrum samples and the mel filter, so as to realize the conversion of the mel spectrum samples from the mel domain to the linear domain, so that the zero space model continues to perform the subsequent linear amplitude domain reconstruction and phase information recovery to obtain the audio prediction result. According to the loss value between the audio prediction result and the audio sample, the parameters of the vocoder model are adjusted, so as to obtain the trained vocoder.

[0178] Figure 5 is a flowchart of the audio generation method provided by the embodiments of the present application. As shown in Figure 5 the audio generation method provided by the embodiments of the present application includes the following steps S501 to S503.

[0179] S501, obtaining a mel spectrum.

[0180] A user can give a mel spectrum.

[0181] S502, obtaining a mel filter of the mel spectrum.

[0182] According to the mel spectrum, the mel filter corresponding to the mel spectrum can be obtained. For example, according to the mel spectrum, the mel filter matrix corresponding to the mel spectrum is matched from the mel filter matrix set generated by the embodiment corresponding to S101, that is, the mel filter. If the mel filter matrix set does not include the mel filter matrix corresponding to the mel spectrum, the mel filter matrix of the mel spectrum is generated.

[0183] S503, input the mel spectrum and the mel filter into a vocoder model to generate audio.

[0184] The mel spectrum and the mel filter are input into a vocoder model to generate audio. The vocoder is a vocoder model trained by the embodiment shown in the figure. Based on the vocoder model, audio in different scenes can be generated. Figure 1

[0185] Based on the same concept as the method embodiments of the present application, the embodiments of the present application also provide a training device for a multi-scene vocoder based on value zero decomposition. The training device for a multi-scene vocoder based on value zero decomposition includes several modules, each module being used to execute each step in the training method for a multi-scene vocoder based on value zero decomposition provided by the embodiments of the present application. The division of the modules is not limited here. Those skilled in the art can clearly understand that in actual applications, each step in the training method for a multi-scene vocoder based on value zero decomposition provided by the embodiments of the present application can be completed by different modules according to needs, that is, the internal structure of the device is divided into different modules to complete all or part of the functions described above. Each module in the embodiments can be integrated in a processing unit, or each unit can exist physically, or two or more modules can be integrated in a unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit. In addition, the specific names of the modules are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the modules in the above device can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0186] Figure 6 The training device for a multi-scene vocoder based on value zero decomposition provided by the embodiments of the present application is shown in the figure. The training device for a multi-scene vocoder based on value zero decomposition provided by the embodiments of the present application includes: Figure 6 The acquisition module 601 is used to acquire a plurality of audio samples and label information of each audio sample in the plurality of audio samples.

[0187] The first determination module 602 is used to determine a mel spectrum sample of each audio sample, and the mel spectrum samples of different audio samples are generated by mel filters with different parameter configurations.

[0188]

[0189] ​​The training module 603 is configured to input the mel spectrum sample and the mel filter corresponding to the mel spectrum sample into a vocoder model, train the vocoder model, and obtain an audio prediction result corresponding to each mel spectrum sample. The vocoder model includes a value space model and a zero space model. The vocoder model generates a first feature from the mel spectrum sample and the mel filter corresponding to the mel spectrum sample through the value space model. The vocoder model performs linear amplitude domain reconstruction and phase information recovery on the first feature through the zero space model to obtain the audio prediction result.

[0190] The second determination module 604 is configured to determine a loss value between the audio prediction result and the audio sample.

[0191] The adjustment module 605 is configured to adjust parameters of the vocoder model according to the loss value, and obtain a trained vocoder model.

[0192] According to the above scheme, the mel spectrum sample of the audio sample in different scenarios is generated based on the mel filter with different parameter configurations. When training the vocoder model, the mel filter and the mel spectrum sample in different scenarios are used, so that the trained vocoder model can adapt to different parameter configurations of the mel filter. Therefore, it is not necessary to train the vocoder model for each parameter configuration of the mel filter, and the computing resources are saved. In the process of training the vocoder model, in order to enable the vocoder model to adapt to the mel filter in different scenarios (i.e., different parameter configurations), the value space model is used in the vocoder model to process the mel spectrum sample and the mel filter, so as to convert the mel spectrum sample from the mel domain to the linear domain. Then, the zero space model continues to perform subsequent linear amplitude domain reconstruction and phase information recovery to obtain the audio prediction result. According to the loss value between the audio prediction result and the audio sample, the parameters of the vocoder model are adjusted, so as to obtain the trained vocoder.

[0193] In a possible implementation, the first determination module is configured to:

[0194] Generate the mel filter with different parameter configurations according to the preconfigured band interval and frequency interval;

[0195] Determine, according to the audio sample, a target mel filter corresponding to the audio sample and a target parameter configuration corresponding to the target mel filter;

[0196] Generate the mel spectrum sample of the audio sample according to the target mel filter and the target parameter configuration.

[0197] In this way, the mel spectrum sample of the audio sample can be generated in different scenarios.

[0198] In a possible implementation, the first feature satisfies the following formula:

[0199]

[0200] wherein, represents a first feature, A + represents a pseudo-inverse transform matrix of the mel filter, A + has a size of |X mel |represents a mel-spectral sample.

[0201] In this way, the mel-spectral sample is converted from the mel domain to the linear domain, the mel-spectral sample of the audio sample in different scenarios is trained on the vocoder model, so that the trained vocoder model can adapt to the audio in different scenarios.

[0202] In a possible implementation, the training module is configured to:

[0203] The vocoder model compresses and encodes the spectral feature in the first feature in order from low frequency to high frequency through the null space model to obtain an encoded feature.

[0204] The encoded feature is modeled to obtain a target audio.

[0205] The target audio is subjected to amplitude reconstruction and phase reconstruction to obtain an audio prediction result.

[0206] In this way, the target audio is subjected to amplitude reconstruction and phase reconstruction through the null space model, so that the missing spectral details and phase information of the mel-spectral sample can be recovered, and the accuracy of the audio prediction result is improved.

[0207] In a possible implementation, the audio prediction result satisfies the following formula:

[0208]

[0209] wherein, represents an amplitude estimate of the first feature, represents a phase estimate of the first feature,

[0210] represents a representation function of the null space model, I represents a unit matrix, and iSTFT(.) represents an inverse transform of a short-time Fourier transform.

[0211] In this way, the target audio is subjected to amplitude reconstruction and phase reconstruction through the null space model, so that the missing spectral details and phase information of the mel-spectral sample can be recovered, and the accuracy of the audio prediction result is improved.

[0212] In a possible implementation, the loss value includes at least one of the following:

[0213] log magnitude loss value, phase loss value, real-imaginary loss value, mel-spectrogram loss value, consistency loss value

[0214] wherein the log magnitude loss value satisfies the following equation:

[0215]

[0216] wherein, represents the value of the audio prediction result at frequency index f and time frame index t, X f,t represents the value of the label information at frequency index f and time frame index t, F and T represent the size of the frequency dimension and the number of time frames respectively, log(.) represents the logarithm operation, and ‖.‖2 represents the L2 norm;

[0217] phase loss value satisfies the following equation:

[0218]

[0219] wherein, “*” represents the two-dimensional convolution operation, f AW (·) represents the inverse closed function, round(.) represents the rounding operation, and ‖.‖1 represents the L1 norm;

[0220] real-imaginary loss value satisfies the following equation:

[0221]

[0222] wherein, and represent the real part and imaginary part operations on a complex number, respectively.

[0223] mel-spectrogram loss value satisfies the following equation:

[0224]

[0225] wherein, F mel is the number of mel bands after compression of the frequency dimension, represents the mel-spectrogram of the audio prediction result at frequency index f and time frame index t, represents the mel-spectrogram of the label information at frequency index f and time frame index t;

[0226] consistency loss value satisfies the following equation:

[0227]

[0228] wherein, represents the consistency estimation of .

[0229] Based on the same idea as the method embodiments of the present application, the embodiments of the present application also provide an audio generation device. The audio generation device comprises a plurality of modules, each module being configured to perform each step of the audio generation method provided by the embodiments of the present application. The division of the modules is not limited herein. Those skilled in the art can clearly understand that, in actual application, each step of the audio generation method provided by the embodiments of the present application can be assigned to be completed by different modules according to needs, i.e., the internal structure of the device is divided into different modules to complete all or part of the functions described above. Each module in the embodiments can be integrated in one processing unit, or each unit can exist physically independently, or two or more modules can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the modules are only for the purpose of mutual distinction, and do not limit the protection scope of the present application. The specific working process of the modules in the above device can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0230] Figure 7 is a training device of a multi-scene vocoder based on value zero decomposition provided by the embodiments of the present application. As shown in Figure 7 , the training device of the multi-scene vocoder based on value zero decomposition provided by the embodiments of the present application comprises:

[0231] The first acquisition module 701 is configured to acquire a mel spectrum.

[0232] The second acquisition module 702 is configured to acquire a mel filter of the mel spectrum.

[0233] The generation module 703 is configured to input the mel spectrum and the mel filter into the vocoder model provided by any of the foregoing embodiments to generate an audio.

[0234] Next, a computing device provided by the embodiments of the present application is introduced.

[0235] Figure 8 is a structural schematic diagram of a computer device provided by the embodiments of the present application. As shown in Figure 8 , the computer device provided by the embodiments of the present application can be used to implement the Tibetan multi-dialect language recognition and speech recognition joint model training method or the speech recognition method described in the foregoing method embodiments.

[0236] The computer device can include a processor 801 and a memory 802 having computer program instructions stored therein.

[0237] In particular, the processor 801 can include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to perform one or more of the embodiments of the present application.

[0238] The memory 802 can include mass storage for data or instructions. As an example and not by way of limitation, the memory 802 can include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc (e.g., a compact disc (CD) or a digital versatile disc (DVD)), a solid-state drive (SSD), a USB drive, or a combination of two or more of these. Where appropriate, the memory 802 can include removable or non-removable (or fixed) media, where appropriate. The memory 802 can be internal or external to the integrated gateway disaster recovery device. In particular embodiments, the memory 802 is non-volatile, solid-state memory.

[0239] The memory can include read-only memory (ROM), random-access memory (RAM), magnetic disk storage mediums, optical storage mediums, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software that, when executed (by one or more processors), is operable to perform any of the operations described with reference to the methods in this application.

[0240] The processor 801 implements any of the Tibetan multi-dialect language recognition and speech recognition joint model training methods or speech recognition methods in the above embodiments by reading and executing computer program instructions stored in the memory 802.

[0241] In one example, the electronic device can further include a communication interface 803 and a bus 810. As shown, the processor 801, the memory 802, and the communication interface 803 are connected through the bus 810 and complete communication with each other. Figure 8

[0242] The communication interface 803 is mainly used to realize the communication between the modules, devices, units, and / or equipment in the embodiments of the present application.

[0243] ​Bus 810 includes a hardware, software, or both that couples components of electronic device to each other. As an example and not by way of limitation, bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or another suitable bus or a combination of two or more of these. Where appropriate, bus 810 can include one or more buses. Although this application describes and shows a particular bus, this application contemplates any suitable bus or interconnect.

[0244] In addition, the embodiments of the present application can provide a computer storage medium to implement, in combination with the above-mentioned embodiments. The computer storage medium has computer program instructions stored thereon; the computer program instructions are executed by a processor to implement any of the above-mentioned embodiments.

[0245] The functional blocks shown in the structural block diagrams described above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, functional cards, and the like. When implemented in software, the elements of the present application are program or code segments used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. The "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and the like. The code segments can be downloaded via a computer network such as the Internet, an intranet, and the like.

[0246] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above-mentioned steps, that is, the steps can be performed in the order mentioned in the embodiments, or in an order different from the embodiments, or several steps can be performed simultaneously.

[0247] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0248] The above description is only specific implementation of the present application. For the convenience and brevity of description, the specific working process of the above-described system, module and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described herein. It should be understood that the protection scope of the present application is not limited in this way. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present application, and these modifications or replacements should be covered within the protection scope of the present application.

Claims

1. A training method of a multi-scenario vocoder based on value zero decomposition, characterized in that, The method comprises: obtaining a plurality of audio samples; determining a mel-spectrum sample of each audio sample, the mel-spectrum samples of different audio samples being generated by mel filters with different parameter configurations, the mel filters with different parameter configurations being adapted to different scenes; inputting the mel-spectrum samples and the mel filters corresponding to the mel-spectrum samples into a vocoder model, training the vocoder model, and obtaining an audio prediction result corresponding to each mel-spectrum sample, wherein the vocoder model comprises a value space model and a null space model, the vocoder model generates a first feature from the mel-spectrum samples and the mel filters corresponding to the mel-spectrum samples through the value space model, and the vocoder model restores linear amplitude domain reconstruction and phase information of the first feature through the null space model to obtain the audio prediction result; determining a loss value between the audio prediction result and the audio sample; adjusting parameters of the vocoder model according to the loss value to obtain a trained vocoder model.

2. The method of claim 1, wherein, The method comprises: generating mel filters with different parameter configurations according to a pre-configured band interval and a frequency interval; determining a target mel filter corresponding to the audio sample and a target parameter configuration corresponding to the target mel filter according to the audio sample; generating a mel-spectrum sample of the audio sample according to the target mel filter and the target parameter configuration.

3. The method of claim 1, wherein, The first feature satisfies the following formula: wherein denotes a first feature, A + denotes a pseudo-inverse matrix of the mel filter, A + has a size of |X mel denotes a mel-spectral sample.

4. The method of claim 1, wherein, The vocoder model restores linear amplitude domain reconstruction and phase information of the first feature through the null space model to obtain the audio prediction result, comprising: The vocoder model compresses and encodes spectral features in the first feature in order from low frequency to high frequency through the null space model to obtain encoded features; modeling the encoded features to obtain a target audio; performing amplitude reconstruction and phase reconstruction on the target audio to obtain the audio prediction result.

5. The method of claim 4, wherein, The audio prediction result satisfies the following equation: wherein denotes an amplitude estimate of the first feature, denotes a phase estimate of the first feature, denotes a representation function of the null space model, I denotes the identity matrix.

6. The method according to any one of claims 1 to 4, characterized in that, The loss value comprises at least one of the following: a logarithmic amplitude loss value, a phase loss value, a real-imaginary part loss value, a mel-spectrum loss value, and a consistency loss value; The log amplitude loss value satisfies the following formula: wherein, represents the value of the audio prediction result at frequency index f and time frame index t, X f,t represents the value of the label information at frequency index f and time frame index t, F and T represent the size of frequency dimension and the number of time frames, respectively, log(.) represents the logarithm operation, and ‖.‖2 represents the L2 norm. The phase loss value satisfies the following equation: where "*" denotes a two-dimensional convolution operation, f AW (·) denotes an anti-closing function, round(.) denotes a rounding operation, and ||.||1denotes an L1 norm. The real-imaginary part loss value satisfies the following formula: wherein and respectively denote taking the real and imaginary parts of a complex number. The mel-spectrum loss value satisfies the following formula: wherein F mel is the frequency dimension after compression, represents the mel-spectrogram of the audio prediction result at frequency index f and time frame index t, represents the mel-spectrogram of the label information at frequency index f and time frame index t; The consistency loss value satisfies the following formula: wherein represents an estimate of the consistency of to the data.

7. An audio generating method, characterized by, The method comprises: obtaining a mel-spectrum; obtaining a mel filter of the mel-spectrum; inputting the mel-spectrum and the mel filter into the vocoder model of any one of claims 1-6 to generate an audio.

8. A training apparatus for a multi-scenario vocoder based on value zero factorization, characterized by, The device comprises: an obtaining module, configured to obtain a plurality of audio samples and label information of each audio sample in the plurality of audio samples; a first determining module, configured to determine a mel-spectrum sample of each audio sample, the mel-spectrum samples of different audio samples being generated by mel filters with different parameter configurations, the mel filters with different parameter configurations being adapted to different scenes; The training module is configured to input the mel spectrum sample and the mel filter corresponding to the mel spectrum sample into a vocoder model, train the vocoder model, and obtain an audio prediction result corresponding to each mel spectrum sample, wherein the vocoder model comprises a value space model and a zero space model, the vocoder model generates a first feature according to the mel spectrum sample and the mel filter corresponding to the mel spectrum sample through the value space model, and the vocoder model performs linear amplitude domain reconstruction and phase information recovery on the first feature through the zero space model to obtain the audio prediction result. The second determination module is configured to determine a loss value between the audio prediction result and the audio sample. The adjustment module is configured to adjust parameters of the vocoder model according to the loss value to obtain a trained vocoder model.

9. An audio generating apparatus, characterized by comprising: The device comprises: The first acquisition module is configured to acquire a mel spectrum. The second acquisition module is configured to acquire a mel filter of the mel spectrum. The generation module is configured to input the mel spectrum and the mel filter into the vocoder model in any one of claims 1-6 to generate audio.

10. A computing device, comprising: The computer program is stored in the memory and executable on the processor, and when the computer program is executed by the processor, the method in any one of claims 1-6 is implemented, or the method in claim 7 is implemented.

Citation Information

Patent Citations

  • Method and device for carrying out codebook processing on channel information

    CN103780289A

  • Speech corpus generation system training method

    CN114120973A

  • Noise reduction method of vocoder, vocoder, electronic equipment and storage medium

    CN114299970A

  • Coding and decoding scale factor information

    US20070016427A1

  • Methods and Systems for Voice Conversion

    US20160005403A1