Discrete audio feature generation method and device and audio data word segmentation device training method and device

By using a vector quantization module with multiple codebooks and a gating module to adaptively select codebook vectors in the audio word segmenter, the problems of audio reconstruction quality degradation and resource waste caused by traditional fixed codebooks are solved, and the audio generation quality and generalization ability are improved.

CN120748434AActive Publication Date: 2025-10-03BEIJING XIYU JIZHI TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511124479.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-03
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Traditional audio word segmenters adopt a fixed codebook strategy, which leads to differences in the distribution of acoustic features of different types of audio, resulting in a decrease in the quality of audio reconstruction and limiting its generalization ability in multimodal tasks.

Method used

A vector quantization module with multiple codebooks is used to adaptively select the target codebook vector as the discrete audio feature based on the acoustic feature vector of the initial audio data. The matching probability is determined by combining the gating module and audio statistics to select the most appropriate codebook vector.

Benefits of technology

It improves the audio generation quality, solves the problems of resource waste and insufficient quality caused by traditional fixed codebooks, and enhances the generalization ability of audio word segmenters in multimodal tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748434A_ABST
    Figure CN120748434A_ABST
Patent Text Reader

Abstract

The invention provides a discrete audio feature generation method and device and an audio data word segmentation device training method and device, the discrete audio feature generation method is realized based on an audio data word segmentation device, a vector quantization module comprises a plurality of codebooks, and each codebook comprises a plurality of codebook vectors; the generation method comprises the following steps: inputting initial audio data into the encoder to obtain an acoustic feature vector; and for each piece of initial audio data, matching from the vector quantization module based on the acoustic feature vector to obtain a target codebook vector, and taking the target codebook vector as a discrete audio feature corresponding to the initial audio data. Therefore, different target codebook vectors are adaptively selected as the discrete audio features according to the acoustic feature vector of the initial audio data, resources can be better balanced, meanwhile, the generation quality of audio generation by using the discrete audio features subsequently is improved, and the problem of resource waste or insufficient quality caused by a traditional fixed codebook is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to a method and device for generating discrete audio features, and a method and device for training an audio data word segmenter. Background Art

[0002] With the rapid development of deep learning technology, significant progress has been made in the field of audio signal processing. Audio word segmenters, as key components, are widely used in speech-related tasks. Traditional audio word segmenters are primarily based on vector quantization (VQ) technology. Its core goal is to achieve high-quality signal reconstruction by discretizing continuous audio signals into sequences of acoustic units.

[0003] In existing technologies, vector quantization is typically implemented using a codebook. Existing codebook designs typically employ a fixed strategy: the same set of codebooks is applied to different audio types (such as clean speech, noisy speech, and music). However, the acoustic feature distributions of different audio types vary significantly. This quantization approach degrades audio reconstruction quality, further limiting the generalization capabilities of audio word segmenters in multimodal tasks. Summary of the Invention

[0004] In view of this, the purpose of the present application is to provide a discrete audio feature generation method and device, an audio data word segmenter training method and device, the vector quantization module includes multiple codebooks, each codebook includes multiple codebook vectors, and different target codebook vectors are adaptively selected as discrete audio features according to the acoustic feature vectors of the initial audio data, which can better balance resources and improve the subsequent generation quality of audio using discrete audio features, thereby solving the problems of resource waste or insufficient quality caused by traditional fixed codebooks.

[0005] In a first aspect, an embodiment of the present application provides a method for generating discrete audio features. The discrete audio feature generation method is implemented based on an audio data word segmenter. The audio data word segmenter includes an encoder and a vector quantization module. The vector quantization module includes multiple codebooks, each codebook includes multiple codebook vectors. The discrete audio feature generation method includes: Inputting the initial audio data into the encoder to obtain an acoustic feature vector; For each piece of initial audio data, a target codebook vector is matched from the vector quantization module based on the acoustic feature vector as a discrete audio feature corresponding to the initial audio data.

[0006] Furthermore, an audio feature extraction module is further included between the encoder and the vector quantization module, and the discrete audio feature generation method further includes: Inputting the acoustic feature vector into the audio feature extraction module to obtain audio statistics corresponding to the initial audio data; Determining a target codebook that matches the initial audio data based on a first preset matching rule and the audio statistics; In the target codebook, a codebook vector matching the acoustic feature vector is determined based on a second preset matching rule as a discrete audio feature corresponding to the initial audio data.

[0007] Furthermore, the vector quantization module includes a gating module; The first preset matching rule includes: inputting the audio statistics into the gating module, the gating module determining the matching probability between each codebook and the acoustic feature vector according to the value of the audio statistics, and selecting the codebook with the highest matching probability among the multiple codebooks as the target codebook.

[0008] Furthermore, the determining, based on the second preset matching rule, a codebook vector matching the acoustic feature vector as the audio feature corresponding to the initial audio data includes: Calculating the spatial distance between the acoustic feature vector and all codebook vectors in the target codebook; A codebook vector with the smallest spatial distance among the multiple codebook vectors is obtained, and the codebook vector is used as a discrete audio feature corresponding to the initial audio data.

[0009] In a second aspect, an embodiment of the present application further provides a method for training an audio data word segmenter, the training method being used to train an audio data word segmenter, the audio data word segmenter being used to implement a discrete audio feature generation method, the training method comprising: Input the audio sample data into the encoder to obtain the acoustic sample feature vector; Based on a vector quantization module, the acoustic feature vector is converted into discrete audio sample features; Inputting the discrete audio sample features into a decoder to obtain reconstructed audio data; At least one loss function is established, and at least one parameter of the audio data segmenter and the decoder is iteratively trained based on the at least one loss function result.

[0010] Furthermore, establishing at least one loss function includes: Acquire text sample data corresponding to the audio sample data, convert the text sample data into a discrete text feature vector, and establish a first loss function between the discrete audio sample features and the discrete text feature vector; establishing a second loss function between the acoustic sample feature vector and the discrete audio sample feature; Establishing a third loss function and a fourth loss function between the audio sample data and the reconstructed audio data; Based on the weighted fusion strategy, the first loss function, the second loss function, the third loss function and the fourth loss function are fused to obtain a fusion loss function, and at least one parameter of the audio data segmenter and the decoder is iteratively trained based on the result of the fusion loss function.

[0011] Furthermore, the weighted fusion strategy includes: configuring a first weight, a second weight, a third weight, and a fourth weight for the first loss function, the second loss function, the third loss function, and the fourth loss function, respectively, where the first weight, the second weight, the third weight, and the fourth weight are learnable weight parameters; The iterative process of the learnable weight parameters includes: Obtain the loss change rate of each loss function per unit time, and dynamically adjust the corresponding weight based on the loss change rate corresponding to each loss function.

[0012] Furthermore, the training phase of each codebook in the vector quantization module includes: When training begins, the acoustic sample feature vector is allowed to be mapped to at least two codebook vectors in at least two codebooks; during the training process, the gating function is gradually adjusted so that the output vector of the gating function approaches the one-hot vector.

[0013] Furthermore, gradually adjusting the gating function so that the output vector of the gating function approaches the one-hot vector includes: The gating function includes a temperature parameter, and the temperature parameter is preset to a maximum value so that the matching probability of the acoustic sample feature vector and each codebook and each codebook vector approaches the same when training starts, so that all codebook vectors are optimized; The value of the temperature parameter is gradually reduced to make the matching probabilities of different codebook vectors different, until the temperature parameter approaches 0 and the output vector of the gating function approaches a one-hot vector.

[0014] In a third aspect, an embodiment of the present application further provides a discrete audio feature generation device, which is implemented based on an audio data word segmenter. The audio data word segmenter includes an encoder and a vector quantization module. The vector quantization module includes multiple codebooks, each codebook includes multiple codebook vectors, and the discrete audio feature generation device includes: an acoustic feature vector determination module, configured to input initial audio data into the encoder to obtain an acoustic feature vector; The discrete audio feature generation module is used to obtain a target codebook vector from the vector quantization module based on the acoustic feature vector for each piece of initial audio data as a discrete audio feature corresponding to the initial audio data.

[0015] In a fourth aspect, an embodiment of the present application further provides a training device for an audio data word segmenter, the training device being used to train the audio data word segmenter, the audio data word segmenter being used to implement a discrete audio feature generation method, the training device comprising: A sample feature vector determination module is used to input audio sample data into an encoder to obtain an acoustic sample feature vector; an audio sample feature determination module, configured to convert the acoustic feature vector into discrete audio sample features based on a vector quantization module; An audio reconstruction module, configured to input the discrete audio sample features into a decoder to obtain reconstructed audio data; The model training module is used to establish at least one loss function and iteratively train at least one parameter of the audio data segmenter and the decoder based on the result of at least one loss function.

[0016] The embodiments of the present application provide a discrete audio feature generation method and device, and an audio data word segmenter training method and device. The discrete audio feature generation method is implemented based on an audio data word segmenter. The audio data word segmenter includes an encoder and a vector quantization module. The vector quantization module includes multiple codebooks, and each codebook includes multiple codebook vectors. First, the initial audio data is input into the encoder to obtain an acoustic feature vector. Then, for each piece of initial audio data, a target codebook vector is matched from the vector quantization module based on the acoustic feature vector as the discrete audio feature corresponding to the initial audio data.

[0017] In this way, since the vector quantization module includes multiple codebooks, each codebook includes multiple codebook vectors, different target codebook vectors are adaptively selected as discrete audio features based on the acoustic feature vectors of the initial audio data, which can better balance resources and improve the subsequent generation quality of audio using discrete audio features, solving the problems of resource waste or insufficient quality caused by traditional fixed codebooks.

[0018] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 A flowchart of a method for generating discrete audio features provided in an embodiment of the present application; Figure 2 A flowchart of a method for training an audio data word segmenter provided in another embodiment of the present application; Figure 3 A schematic diagram of the structure of a discrete audio feature generation device provided in an embodiment of the present application; Figure 4 A schematic structural diagram of a training device for an audio data word segmenter provided in another embodiment of the present application; Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, each other embodiment obtained by those skilled in the art without making creative work falls within the scope of protection of the present application.

[0022] First, the application scenarios to which this application is applicable are introduced. This application can be applied in the field of audio processing technology.

[0023] With the rapid development of deep learning technology, significant progress has been made in the field of audio signal processing. Audio word segmenters, as key components, are widely used in speech-related tasks. Traditional audio word segmenters are primarily based on vector quantization (VQ) technology. Its core goal is to achieve high-quality signal reconstruction by discretizing continuous audio signals into sequences of acoustic units.

[0024] Research has found that, in existing technologies, vector quantization is typically implemented using a codebook. Existing codebook designs typically employ a fixed strategy: the same set of codebooks is applied to different audio types (such as clean speech, noisy speech, and music), despite the significant differences in the distribution of acoustic features across these different audio types. This quantization approach degrades audio reconstruction quality, further limiting the generalization capabilities of audio word segmenters in multimodal tasks.

[0025] Based on this, the embodiments of the present application provide a discrete audio feature generation method and device, and an audio data word segmenter training method and device, which adaptively select different target codebook vectors as discrete audio features based on the acoustic feature vectors of the initial audio data, which can better balance resources and improve the subsequent generation quality of audio using discrete audio features, solving the problems of resource waste or insufficient quality caused by traditional fixed codebooks.

[0026] See also Figure 1 , Figure 1 This is a flowchart of a discrete audio feature generation method provided in an embodiment of the present application. The discrete audio feature generation method provided in an embodiment of the present application is implemented based on an audio data word segmenter, wherein the audio data word segmenter includes an encoder and a vector quantization module, wherein the vector quantization module includes multiple codebooks, each codebook includes multiple codebook vectors, such as Figure 1 As shown in , the discrete audio feature generation method provided by the embodiment of the present application includes: S101: Input initial audio data into the encoder to obtain an acoustic feature vector.

[0027] Here, the encoder is used to extract the acoustic features of the audio data and convert the audio data into several continuous vectors.

[0028] Regarding step S101, in a specific implementation, initial audio data is obtained as input. Here, the initial audio data can be an audio signal or a Mel spectrum, which is not specifically limited in this application. The initial audio data is then input into an encoder to extract acoustic features and obtain an acoustic feature vector.

[0029] S102 : For each piece of initial audio data, obtain a target codebook vector from the vector quantization module based on the acoustic feature vector as a discrete audio feature corresponding to the initial audio data.

[0030] The vector quantization (VQ) module is a key component in deep learning, particularly in autoencoder frameworks and generative models, for discretizing continuous vector representations. Its core concept is to approximate the input vector using a finite, predefined "codebook." In the embodiments provided herein, the vector quantization module is used to map the continuous latent space output by the encoder into a discrete sequence of audio tokens. Each latent vector is matched with the most similar vector in the codebook and represented by its index (token ID), thereby forming discrete audio features. The vector quantization module includes multiple codebooks, each containing multiple codebook vectors. The core function of the codebook in the vector quantization module is to convert the continuous input feature vector into a discrete symbol (index). The number of codebook vectors in a codebook is pre-set. For example, if the number of codebooks is 3, the first codebook contains 1024 codebook vectors, the second codebook contains 4096 codebook vectors, and the third codebook contains 8192 codebook vectors. The codebook is equipped with a gated network that can select different codebooks to perform vector quantization based on the continuous feature vectors that characterize the dynamic range of the audio. Specifically, for audio with a smaller dynamic range (relatively flat or even silent), only a smaller codebook is required, such as a smaller codebook containing 1024 vectors; for audio with a larger dynamic range (such as the chorus of a song), a larger codebook is selected, such as a larger codebook containing 8192 vectors; for ordinary speech audio, a medium codebook of 4096 vectors is selected. It should be understood that the above codebook data is only an example. In fact, a larger or smaller number of codebooks can be trained based on model resources, etc. The examples of this application should not be understood as limiting the scope of protection of this application.

[0031] For the above-mentioned step S102, during the specific implementation, for each piece of initial audio data, the acoustic feature vector corresponding to the initial audio data is input into the vector quantization module, and the target codebook vector is matched from the multiple codebooks included in the vector quantization module based on the acoustic feature vector, and the target codebook vector is used as the discrete audio feature corresponding to the initial audio data.

[0032] As an optional embodiment, an audio feature extraction module is further included between the encoder and the vector quantization module. Furthermore, the discrete audio feature generation method provided in the embodiment of the present application further includes: I: Inputting the acoustic feature vector into the audio feature extraction module to obtain audio statistics corresponding to the initial audio data.

[0033] Here, an audio feature extraction module is introduced between the encoder and vector quantization modules. This module extracts statistical features related to the audio dynamic range from the encoder output as input to the gating module. This module can be implemented using a neural network (such as an MLP or multi-layer LSTM / GRU). The neural network processes the encoder's global features and outputs a dynamic feature vector.

[0034] Regarding step I above, during implementation, the encoder output is received, the acoustic feature vector is input into the audio feature extraction module, the statistical features of the current audio segment are calculated, and an audio statistic representing the "activity" of the audio is output. This audio statistic is then used in subsequent codebook selection decisions, enhancing the model's ability to understand the audio content. Statistical features, such as variance, standard deviation, and peak-to-valley difference, are used to indicate the degree of audio dispersion, and this is not a limitation of the present invention.

[0035] II: Determine a target codebook matching the initial audio data based on a first preset matching rule and the audio statistics.

[0036] Regarding the above step II, during specific implementation, a target codebook matching the initial audio data is determined based on the first preset matching rule and the audio statistics output by the audio feature extraction module.

[0037] According to an embodiment provided by the present application, the vector quantization module includes a gating module. The first preset matching rule includes: inputting the audio statistics into the gating module, the gating module determining the matching probability between each codebook and the acoustic feature vector based on the value of the audio statistics, and selecting the codebook with the highest matching probability among the multiple codebooks as the target codebook.

[0038] Here, the gating module can be composed of a multi-layer perceptron or a lightweight neural network based on the attention mechanism, and its output layer dimension is equal to the total number of codebooks. The gating module receives audio statistics that represent the dynamic range of the audio. The gating function in the gating module determines the matching probability of each codebook and the acoustic feature vector based on the data of the audio statistics, and outputs a result vector L, which is converted into the matching probability distribution P = {p1, p2, p3, ... p n}, n is the number of codebooks. The codebook with the highest matching probability among the multiple codebooks is then selected as the target codebook. Preferably, before obtaining the matching probability distribution for each codebook, the result vector L is normalized to obtain the matching probability distribution P. The normalization method can use a commonly used normalization algorithm such as softmax, which is not limited in this application.

[0039] III: In the target codebook, a codebook vector matching the acoustic feature vector is determined based on a second preset matching rule as a discrete audio feature corresponding to the initial audio data.

[0040] Regarding the above step III, in specific implementation, after the target codebook is determined in the above step II, a codebook vector matching the acoustic feature vector is determined in the target codebook based on the second preset matching rule to serve as a discrete audio feature corresponding to the initial audio data.

[0041] Specifically, with respect to the above step III, determining a codebook vector that matches the acoustic feature vector based on the second preset matching rule as the audio feature corresponding to the initial audio data includes: Calculating the spatial distance between the acoustic feature vector and all codebook vectors in the target codebook; obtaining a codebook vector with the smallest spatial distance among multiple codebook vectors, and using the codebook vector as a discrete audio feature corresponding to the initial audio data.

[0042] For the above two steps, in the specific implementation, in the target codebook, the spatial distance between the acoustic feature vector and each codebook vector in the target codebook is calculated, and the codebook vector with the smallest spatial distance among multiple codebook vectors is used as the discrete audio feature corresponding to the initial audio data.

[0043] The discrete audio feature generation method provided in the embodiment of the present application is implemented based on an audio data segmenter, and the audio data segmenter includes an encoder and a vector quantization module, and the vector quantization module includes multiple codebooks, and each codebook includes multiple codebook vectors; first, the initial audio data is input into the encoder to obtain an acoustic feature vector; then, for each piece of initial audio data, a target codebook vector is matched from the vector quantization module based on the acoustic feature vector to obtain a discrete audio feature corresponding to the initial audio data. In this way, since the vector quantization module includes multiple codebooks, and each codebook includes multiple codebook vectors, different target codebook vectors are adaptively selected as discrete audio features according to the acoustic feature vectors of the initial audio data, which can better balance resources and improve the quality of subsequent audio generation using discrete audio features, solving the problem of resource waste or insufficient quality caused by traditional fixed codebooks.

[0044] See also Figure 2 , Figure 2 This is a flowchart of a method for training an audio data word segmenter provided in another embodiment of the present application. The method for training an audio data word segmenter provided in the embodiment of the present application is used to train an audio data word segmenter, and the audio data word segmenter is used to implement a discrete audio feature generation method, such as Figure 2 As shown in , the training method provided in the embodiment of the present application includes: S201: Input audio sample data into an encoder to obtain an acoustic sample feature vector.

[0045] S202: Based on a vector quantization module, convert the acoustic feature vector into discrete audio sample features.

[0046] Regarding steps S201 and S202 above, in a specific implementation, audio sample data is obtained and input into an encoder to obtain an acoustic sample feature vector. The acoustic sample feature vector is input into a vector quantization module, which converts the acoustic feature vector into discrete audio sample features. The method for obtaining the acoustic sample feature vector and the method for obtaining the discrete audio sample features can be referred to the description of steps S101 and S102 above, and the same technical effects can be achieved, so this description is not repeated here.

[0047] S203: Input the discrete audio sample features into a decoder to obtain reconstructed audio data.

[0048] Here, the decoder is used to restore the discrete audio feature sequence to audio data or Mel spectrum, depending on whether the input to the encoder is audio data or Mel spectrum.

[0049] Regarding the above step S203, in a specific implementation, the discrete audio sample features output by the vector quantization module are input into the decoder for data restoration to obtain reconstructed audio data.

[0050] S204: Establish at least one loss function, and iteratively train at least one parameter of the audio data segmenter and the decoder based on the at least one loss function result.

[0051] For the above step S204, in the specific implementation, at least one loss function is established using audio sample data, acoustic sample feature vectors, discrete audio sample features and reconstructed audio data, and at least one parameter of the audio data segmenter and decoder is iteratively trained based on the result of at least one loss function.

[0052] As an optional embodiment, with respect to the above step S204, establishing at least one loss function includes: Step 2041: Acquire text sample data corresponding to the audio sample data, convert the text sample data into a discrete text feature vector, and establish a first loss function between the discrete audio sample features and the discrete text feature vector.

[0053] Here, the first loss function is the CTC loss (Connectionist Temporal Classification loss) function. The core of the CTC loss function is that the algorithm can solve the problem of inconsistent lengths between the input sequence and the output sequence and / or the two sequences are not strictly time-aligned. In this application, it is established between the discrete audio sample features predicted by the vector quantization module and the discrete text feature vector processed by the text encoder. It can measure the probability of obtaining text sample data from discrete audio sample features under all possible alignment methods, and is used to enhance the semantic ability of the overall model so that the content in the output audio feature sequence can better correspond to the text. Each codebook in the vector quantization module can be composed of a single-layer VQ or a multi-layer VQ. When multi-layer VQ is used, the input vector is decomposed into multiple parts during the quantization stage, and each part is quantized using a codebook. This method can significantly increase the representation capability while keeping the codebook size controllable.

[0054] Regarding step 2041, in a specific implementation, text sample data corresponding to the audio sample data is obtained and converted into a discrete text feature vector. The discrete audio sample features are compared with the discrete text feature vector, and a first loss function is established between the discrete audio sample features and the discrete text feature vector.

[0055] Step 2042: Establish a second loss function between the acoustic sample feature vector and the discrete audio sample feature.

[0056] Here, the second loss function is a codebook learning loss (VQ loss) function, which is used to measure the difference between the acoustic sample feature vector output by the encoder and the sample target codebook corresponding to the discrete audio sample feature.

[0057] Regarding the above step 2042, during specific implementation, a second loss function is established between the acoustic sample feature vector and the discrete audio sample feature.

[0058] Step 2043: Establish a third loss function and a fourth loss function between the audio sample data and the reconstructed audio data.

[0059] Here, the third loss function is the adversarial loss function of the GAN model. It introduces a discriminator network, built between the initial audio sample data and the reconstructed audio data reconstructed by the decoder. The discriminator measures which audio is real and which is reconstructed, ensuring that the reconstructed audio is as close as possible to the audio originally input to the encoder. The fourth loss function is the reconstruction loss function, also built between the initial audio sample data and the reconstructed audio data reconstructed by the decoder. It is used to measure the difference between the audio data reconstructed by the decoder and the audio sample data. This is a common loss type and calculation method in the model. Based on this reconstruction loss, the difference between the original input and the reconstructed output can be minimized, ensuring that the encoder, vector quantization module, and decoder can output higher-quality audio information during inference.

[0060] Regarding the above step 2043, during specific implementation, a third loss function and a fourth loss function are established between the audio sample data and the reconstructed audio data.

[0061] Step 2044: Based on the weighted fusion strategy, the first loss function, the second loss function, the third loss function and the fourth loss function are fused to obtain a fusion loss function, and at least one parameter of the audio data segmenter and the decoder is iteratively trained based on the result of the fusion loss function.

[0062] In the specific implementation of step 2044, the first, second, third, and fourth loss functions are fused together based on a weighted fusion strategy to obtain a fused loss function. Based on the result of the fused loss function, at least one parameter of the audio data word segmenter and decoder is iteratively trained. Backpropagation is performed based on the fused loss function to update the parameters of the audio data word segmenter and decoder.

[0063] Here, the weighted fusion strategy includes: configuring a first weight, a second weight, a third weight, and a fourth weight for the first loss function, the second loss function, the third loss function, and the fourth loss function, respectively.

[0064] Here, the first, second, third, and fourth weights are assigned to the first, second, third, and fourth loss functions, respectively. The fusion loss function is the weighted sum of the first, second, third, and fourth loss functions. The weights can, to a certain extent, represent the model's capabilities. For example, a larger weight in the CTC loss function indicates stronger semantic capabilities, a larger weight in the adversarial loss function indicates a higher acoustic performance, and so on.

[0065] Furthermore, the first weight, the second weight, the third weight, and the fourth weight are learnable weight parameters. The iterative process of the learnable weight parameters includes: obtaining the loss change rate of each loss function per unit time, and dynamically adjusting the corresponding weight based on the loss change rate corresponding to each loss function.

[0066] The weights of each loss function can be dynamically adjusted based on the model's performance and expected performance. For example, if the current audio task requires stronger semantic capabilities, the weight of the CTC loss can be increased accordingly. Specifically, the loss change rate per unit time of each loss function is obtained, and the corresponding weight is dynamically adjusted based on the corresponding loss change rate of each loss function.

[0067] Here, the unit time is preferably a fixed number of iterations, for example, 5 iterations as a unit time or 10 iterations as a unit time, etc., which is not specifically limited in this application. For each loss function, the loss change rate of the loss function is calculated by the following formula:

[0068] in, is the average loss function value in the current unit time, is the average loss function value in the previous unit time.

[0069] Dynamically adjusting weights based on the loss change rate includes: ,in, is the weight in the current unit time, is the weight in the previous unit time. When it is greater than 1, it indicates that the loss has increased. At this time, the value of the exponential term is greater than 1, and the weight will increase, will emphasize the task; when =1, indicating that the loss has not changed and the weight unchanged; when When it is less than 1, it means that the loss decreases. At this time, the value of the exponential term is less than 1, and the weight Will decrease. is a sensitivity parameter, if Increase, the weight adjustment will be greater, if If it is too large, the weight fluctuation may lead to unstable training. If the value is smaller, the adjustment range of the weight will be smaller, resulting in a slower adjustment of the weight. =0.5, and then gradually fine-tune according to the results of each round of training size.

[0070] Preferably, when dynamically adjusting the weights of the loss functions, the weights of each loss function are adjusted simultaneously. Preferably, after each adjustment of the weight of any loss function, the overall probability is normalized. In this way, by weighted fusion of multiple loss functions and dynamic adjustment of the weights of each loss function, it is possible to ensure that the capabilities represented by each loss function in the model are balanced, resolving the problem of the model only iterating parameters for one or two capabilities, resulting in weaknesses in the capabilities represented by other loss functions.

[0071] In order to reduce the training cost and difficulty, each codebook in the vector quantization module can also be trained independently. As an optional embodiment, the training phase of each codebook in the vector quantization module includes: When training begins, the acoustic sample feature vector is allowed to be mapped to at least two codebook vectors in at least two codebooks; during the training process, the gating function is gradually adjusted so that the output vector of the gating function approaches the one-hot vector.

[0072] Furthermore, gradually adjusting the gating function so that the output vector of the gating function approaches the one-hot vector includes: The gating function includes a temperature parameter, which is preset to a maximum value so that the matching probabilities of the acoustic sample feature vector and each codebook and each codebook vector are close to the same at the beginning of training, so that all codebook vectors are optimized; the value of the temperature parameter is gradually reduced so that the matching probabilities of different codebook vectors differ until the temperature parameter approaches 0, and the output vector of the gating function approaches a one-hot vector.

[0073] Here, the gating function is expressed by the following formula:

[0074] in, is the gating function, Represents the acoustic sample feature vector and the The similarity between codebook vectors, Represents the temperature parameter.

[0075] Here, the above gating function adds the normalization function , That is the temperature parameter. >0, The maximum value of T refers to the maximum value within its possible value range. For example, it is set to 2.0 at the beginning of training. By the end of training, the value of T can be a decimal close to 0, such as 0.001.

[0076] in, The larger the When it approaches positive infinity, the exponential value of each codebook vector approaches 1. Therefore, for all codebooks and codebook vectors, it approaches 1 / N. Therefore, the probability of all codebook vectors being selected is similar, which can ensure that all codebook vectors have the possibility of being trained.

[0077] when As the value of gradually decreases, the probabilities of the codebook vectors begin to differ until As the value approaches 0, the gate function outputs a vector that approaches the form [0, 0, 0, 1, 0, 0, 0…], with only one 1 and all other dimensions 0. This is known as a one-hot vector. 1 represents a selected codebook or selected codebook vector, and 0 represents an unselected codebook or unselected codebook vector. The temperature parameter can decay linearly or exponentially.

[0078] In addition to training each codebook vector using the VQ loss, an exponential moving average (EMA) update can also be used to indirectly update codebook vectors that have been selected too few times. This ensures that all codebook vectors in the vector quantization module are trained, preventing the selection of the same or similar codebook vectors for each training session, which could result in some codebook vectors not being trained.

[0079] The embodiment of the present application provides a training method for an audio data word segmenter, which is used to train an audio data word segmenter, and the audio data word segmenter is used to implement a discrete audio feature generation method; first, the audio sample data is input into an encoder to obtain an acoustic sample feature vector; then, based on a vector quantization module, the acoustic feature vector is converted into a discrete audio sample feature; the discrete audio sample feature is input into a decoder to obtain reconstructed audio data; finally, at least one loss function is established, and at least one parameter of the audio data word segmenter and decoder is iteratively trained based on the result of at least one loss function. By setting at least one loss function, it is possible to simultaneously take into account the learning capabilities of the model in different aspects, such as reconstruction capability, semantic capability, and acoustic performance capability, thereby improving the quality of subsequent audio generation.

[0080] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a discrete audio feature generation device provided in an embodiment of the present application. The discrete audio feature generation device is implemented based on an audio data word segmenter, and the audio data word segmenter includes an encoder and a vector quantization module. The vector quantization module includes multiple codebooks, each codebook includes multiple codebook vectors, such as Figure 3 As shown, the discrete audio feature generating device 300 includes: an acoustic feature vector determination module 301, configured to input initial audio data into the encoder to obtain an acoustic feature vector; The discrete audio feature generation module 302 is configured to obtain, for each piece of initial audio data, a target codebook vector from the vector quantization module based on the acoustic feature vector, as a discrete audio feature corresponding to the initial audio data.

[0081] Furthermore, an audio feature extraction module is further included between the encoder and the vector quantization module, and the discrete audio feature generation module 302 is further used to: Inputting the acoustic feature vector into the audio feature extraction module to obtain audio statistics corresponding to the initial audio data; Determining a target codebook that matches the initial audio data based on a first preset matching rule and the audio statistics; In the target codebook, a codebook vector matching the acoustic feature vector is determined based on a second preset matching rule as a discrete audio feature corresponding to the initial audio data.

[0082] Furthermore, the vector quantization module includes a gating module; The first preset matching rule includes: inputting the audio statistics into the gating module, the gating module determining the matching probability between each codebook and the acoustic feature vector according to the value of the audio statistics, and selecting the codebook with the highest matching probability among the multiple codebooks as the target codebook.

[0083] Furthermore, when the discrete audio feature generation module 302 is used to determine a codebook vector that matches the acoustic feature vector based on a second preset matching rule as the audio feature corresponding to the initial audio data, the discrete audio feature generation module 302 is further used to: Calculating the spatial distance between the acoustic feature vector and all codebook vectors in the target codebook; A codebook vector with the smallest spatial distance among the multiple codebook vectors is obtained, and the codebook vector is used as a discrete audio feature corresponding to the initial audio data.

[0084] See also Figure 4 , Figure 4 This is a structural diagram of a training device for an audio data word segmenter provided in another embodiment of the present application. The training device is used to train an audio data word segmenter, and the audio data word segmenter is used to implement a discrete audio feature generation method, such as Figure 4 As shown in FIG, the training device 400 includes: The sample feature vector determination module 401 is used to input the audio sample data into the encoder to obtain the acoustic sample feature vector; An audio sample feature determination module 402, configured to convert the acoustic feature vector into discrete audio sample features based on a vector quantization module; An audio reconstruction module 403 is configured to input the discrete audio sample features into a decoder to obtain reconstructed audio data; The model training module 404 is used to establish at least one loss function and iteratively train at least one parameter of the audio data segmenter and the decoder based on the at least one loss function result.

[0085] Furthermore, when the model training module 404 is used to establish at least one loss function, the model training module 404 is further used to: Acquire text sample data corresponding to the audio sample data, convert the text sample data into a discrete text feature vector, and establish a first loss function between the discrete audio sample features and the discrete text feature vector; establishing a second loss function between the acoustic sample feature vector and the discrete audio sample feature; Establishing a third loss function and a fourth loss function between the audio sample data and the reconstructed audio data; Based on the weighted fusion strategy, the first loss function, the second loss function, the third loss function and the fourth loss function are fused to obtain a fusion loss function, and at least one parameter of the audio data segmenter and the decoder is iteratively trained based on the result of the fusion loss function.

[0086] Furthermore, the weighted fusion strategy includes: configuring a first weight, a second weight, a third weight, and a fourth weight for the first loss function, the second loss function, the third loss function, and the fourth loss function, respectively, where the first weight, the second weight, the third weight, and the fourth weight are learnable weight parameters; The iterative process of the learnable weight parameters includes: Obtain the loss change rate of each loss function per unit time, and dynamically adjust the corresponding weight based on the loss change rate corresponding to each loss function.

[0087] Furthermore, the training device 400 further includes a vector quantization module training module, and the training phase of each codebook in the vector quantization module by the vector quantization module training module includes: When training begins, the acoustic sample feature vector is allowed to be mapped to at least two codebook vectors in at least two codebooks; during the training process, the gating function is gradually adjusted so that the output vector of the gating function approaches the one-hot vector.

[0088] Furthermore, when the vector quantization module training module is used to gradually adjust the gating function so that the output vector of the gating function approaches the one-hot vector, the vector quantization module training module is further used to: The gating function includes a temperature parameter, and the temperature parameter is preset to a maximum value so that the matching probability of the acoustic sample feature vector and each codebook and each codebook vector approaches the same when training starts, so that all codebook vectors are optimized; The value of the temperature parameter is gradually reduced to make the matching probabilities of different codebook vectors different, until the temperature parameter approaches 0 and the output vector of the gating function approaches a one-hot vector.

[0089] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 As shown in FIG, the electronic device 500 includes a processor 510, a memory 520 and a bus 530.

[0090] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 is running, the processor 510 communicates with the memory 520 via the bus 530. When the machine-readable instructions are executed by the processor 510, the above-mentioned Figure 1 as well as Figure 2 The steps of the discrete audio feature generation method and the audio data word segmenter training method in the method embodiment shown are specifically implemented in the method embodiment and will not be repeated here.

[0091] The embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the computer program can execute the above-mentioned Figure 1 as well as Figure 2 The steps of the discrete audio feature generation method and the audio data word segmenter training method in the method embodiment shown are specifically implemented in the method embodiment and will not be repeated here.

[0092] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0093] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. There may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.

[0094] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0095] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0096] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0097] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in the present application, or perform equivalent replacements for some of the technical features thereof. These modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for generating discrete audio features, characterized in that: The discrete audio feature generation method is implemented based on an audio data word segmenter, wherein the audio data word segmenter includes an encoder and a vector quantization module, wherein the vector quantization module includes multiple codebooks, each codebook includes multiple codebook vectors, and the discrete audio feature generation method includes: Inputting the initial audio data into the encoder to obtain an acoustic feature vector; For each piece of initial audio data, a target codebook vector is matched from the vector quantization module based on the acoustic feature vector as a discrete audio feature corresponding to the initial audio data.

2. The method for generating discrete audio features according to claim 1, wherein: An audio feature extraction module is also included between the encoder and the vector quantization module, and the discrete audio feature generation method further includes: Inputting the acoustic feature vector into the audio feature extraction module to obtain audio statistics corresponding to the initial audio data; Determining a target codebook that matches the initial audio data based on a first preset matching rule and the audio statistics; In the target codebook, a codebook vector matching the acoustic feature vector is determined based on a second preset matching rule as a discrete audio feature corresponding to the initial audio data.

3. The method for generating discrete audio features according to claim 2, wherein: The vector quantization module includes a gating module; The first preset matching rule includes: inputting the audio statistics into the gating module, the gating module determining the matching probability between each codebook and the acoustic feature vector according to the value of the audio statistics, and selecting the codebook with the highest matching probability among the multiple codebooks as the target codebook.

4. The method for generating discrete audio features according to claim 3, wherein: The determining, based on the second preset matching rule, a codebook vector matching the acoustic feature vector as the audio feature corresponding to the initial audio data includes: Calculating the spatial distance between the acoustic feature vector and all codebook vectors in the target codebook; A codebook vector with the smallest spatial distance among the multiple codebook vectors is obtained, and the codebook vector is used as a discrete audio feature corresponding to the initial audio data.

5. A method for training an audio data word segmenter, characterized in that: The training method is used to train an audio data word segmenter, and the audio data word segmenter is used to implement the discrete audio feature generation method according to any one of claims 1 to 4, and the training method includes: Input the audio sample data into the encoder to obtain the acoustic sample feature vector; Based on a vector quantization module, the acoustic feature vector is converted into discrete audio sample features; Inputting the discrete audio sample features into a decoder to obtain reconstructed audio data; At least one loss function is established, and at least one parameter of the audio data segmenter and the decoder is iteratively trained based on the at least one loss function result.

6. The training method according to claim 5, characterized in that The establishing of at least one loss function comprises: Acquire text sample data corresponding to the audio sample data, convert the text sample data into a discrete text feature vector, and establish a first loss function between the discrete audio sample features and the discrete text feature vector; establishing a second loss function between the acoustic sample feature vector and the discrete audio sample feature; Establishing a third loss function and a fourth loss function between the audio sample data and the reconstructed audio data; Based on the weighted fusion strategy, the first loss function, the second loss function, the third loss function and the fourth loss function are fused to obtain a fusion loss function, and at least one parameter of the audio data segmenter and the decoder is iteratively trained based on the result of the fusion loss function.

7. The training method according to claim 6, characterized in that The weighted fusion strategy includes: configuring a first weight, a second weight, a third weight, and a fourth weight for the first loss function, the second loss function, the third loss function, and the fourth loss function, respectively, where the first weight, the second weight, the third weight, and the fourth weight are learnable weight parameters; The iterative process of the learnable weight parameters includes: Obtain the loss change rate of each loss function per unit time, and dynamically adjust the corresponding weight based on the loss change rate corresponding to each loss function.

8. The training method according to claim 5, characterized in that The training phase of each codebook in the vector quantization module includes: When training begins, the acoustic sample feature vector is allowed to be mapped to at least two codebook vectors in at least two codebooks; during the training process, the gating function is gradually adjusted so that the output vector of the gating function approaches the one-hot vector.

9. The training method according to claim 8, characterized in that The step of gradually adjusting the gating function so that the output vector of the gating function approaches the one-hot vector includes: The gating function includes a temperature parameter, and the temperature parameter is preset to a maximum value so that the matching probability of the acoustic sample feature vector and each codebook and each codebook vector approaches the same when training starts, so that all codebook vectors are optimized; The value of the temperature parameter is gradually reduced to make the matching probabilities of different codebook vectors different, until the temperature parameter approaches 0 and the output vector of the gating function approaches a one-hot vector.

10. A discrete audio feature generation device, characterized in that: The discrete audio feature generation device is implemented based on an audio data word segmenter, the audio data word segmenter includes an encoder and a vector quantization module, the vector quantization module includes multiple codebooks, each codebook includes multiple codebook vectors, and the discrete audio feature generation device includes: an acoustic feature vector determination module, configured to input initial audio data into the encoder to obtain an acoustic feature vector; The discrete audio feature generation module is used to obtain a target codebook vector from the vector quantization module based on the acoustic feature vector for each piece of initial audio data as a discrete audio feature corresponding to the initial audio data.

11. A training device for an audio data word segmenter, characterized in that: The training device is used to train an audio data word segmenter, and the audio data word segmenter is used to implement the discrete audio feature generation method according to any one of claims 1 to 4, and the training device includes: A sample feature vector determination module is used to input audio sample data into an encoder to obtain an acoustic sample feature vector; an audio sample feature determination module, configured to convert the acoustic feature vector into discrete audio sample features based on a vector quantization module; An audio reconstruction module, configured to input the discrete audio sample features into a decoder to obtain reconstructed audio data; The model training module is used to establish at least one loss function and iteratively train at least one parameter of the audio data segmenter and the decoder based on the result of at least one loss function.

Citation Information

Patent Citations

  • Multiple-codebook coding parameter quantification method based on audio emergent event

    CN101587710A

  • Speech synthesis method and device, equipment and storage medium

    CN117727288A

  • Audio processing method and device, electronic equipment and storage medium

    CN118571238A

  • Data processing method and device and related equipment

    CN118690182A

  • Voice synthesis method and device based on gated attention mechanism, equipment and medium

    CN119314463A