Data transmission method and device, model training method and device, chip and terminal
Patent Information
- Application Number
- CN202311787476.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-12-22
AI Technical Summary
[0004]本发明的目的在于提供一种数据传输方法、模型训练方法、装置、芯片及终端,用以改善现有可变码率AI声码器模型的可选码率必须是log2N的倍数的问题
[0032]本发明的有益效果在于:所述目标码率不在所述可选码率范围内,可根据所述目标码率对所述编码向量量化器进行裁剪,以调整所述编码向量量化器的码本大小,得到裁剪后的前Nt个编码向量量化器以及包含所述目标码率的裁剪后可选码率范围,从而在不重新训练模型的情况下,定制化可选码率以满足不同应用场景下对语音编码和数据传输的多样化需求。
Smart Images

Figure CN117789701B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vocoder technology, and in particular to a data transmission method, a model training method, an apparatus, a chip, and a terminal. Background Technology
[0002] Vector quantization (VQ) is a data compression and encoding technique developed in the late 1970s. Its basic principle is to construct a codebook where any vector (extracted feature vector) can be approximated by the most similar vector in the codebook. During storage, only the index value corresponding to the codeword needs to be recorded to achieve compression. Vector quantization was first introduced into artificial intelligence (AI) models. The codebook can be trained on large-scale datasets and subsequently applied to AI speech encoders. Existing variable bitrate AI vocoder models require selectable bitrates that are multiples of log₂N, where N is the codebook size. To meet customized customer needs, the training model must be redesigned, consuming significant time and resources.
[0003] Therefore, it is necessary to propose a data transmission method, a model training method, a device, a chip, and a terminal to solve the above problems. Summary of the Invention
[0004] The purpose of this invention is to provide a data transmission method, a model training method, a device, a chip, and a terminal to improve the problem that the selectable bit rate of existing variable bit rate AI vocoder models must be a multiple of log2N.
[0005] In a first aspect, the present invention provides a voice data transmission method, the method comprising:
[0006] The speech data is acquired, and the speech data is encoded using an encoding model to obtain N. q N codeword index values, wherein the encoding model includes N q N encoding vector quantizers q It is a positive integer;
[0007] Determine the target bitrate based on the current network environment, and then determine N based on the target bitrate. q The top N are determined from the quantizers of the encoded vector. t N are the encoded vector quantizers, where N t For less than or equal to N q Positive integers;
[0008] If the target bitrate is within the selectable bitrate range of the coding model, then the first N... tIf the codeword index values of the aforementioned code vector quantizers are determined as the data to be processed, and the target code rate is not within the selectable code rate range, then the code vector quantizers are pruned according to the target code rate to obtain the first N pruned values. t A number of encoded vector quantizers and a cropped selectable bitrate range containing the target bitrate, the first N cropped bits... t The codeword index value of each encoded vector quantizer is determined as the data to be processed;
[0009] The data to be processed is processed to obtain a data packet, and the data packet is sent to the receiving end according to the target bit rate.
[0010] In one possible embodiment, the code vector quantizer is pruned according to the target code rate to obtain the pruned first N bits. t Each encoded vector quantizer and a cropped selectable bitrate range including the target bitrate include:
[0011] Based on the target bit rate and N t The value is calculated to obtain the second to Nth values. t The target codebook of the Nth coded vector quantizer, the first coded vector quantizer remains unchanged, and according to the second to the Nth coded vector quantizers... t The target codebook of each of the aforementioned encoded vector quantizers is correspondingly pruned from the second to the Nth. t The aforementioned encoding vector quantizers are used to obtain the second to Nth quantizations. t A pruned encoded vector quantizer and a pruned selectable bitrate range including the target bitrate.
[0012] In one possible embodiment, the speech data is encoded using an encoding model to obtain N. q Each codeword index value includes:
[0013] Feature data is obtained by extracting features from the speech data using the encoder of the encoding model, and the feature data is input into N. q N encoded vector quantizers yield N q Each codeword index value.
[0014] Secondly, the present invention provides a model training method, the method comprising:
[0015] The encoding and decoding models in the above embodiments are trained using a training sample set. The encoding model includes an encoder and a residual vector quantization encoding module, and the residual vector quantization encoding module includes N... q The decoding model includes N encoding vector quantizers, and the residual vector quantization decoding module includes a decoder. The residual vector quantization decoding module includes N... q N decoder vector quantizersq It is a positive integer;
[0016] The model parameters of the trained encoding model are frozen. The frozen encoding model is tested using the first test sample set. The frequency of each codeword index value obtained from the test is counted. The codewords obtained from the test are reordered according to the frequency from high to low to obtain a new codebook and update the codeword index value corresponding to the codeword.
[0017] Thirdly, the present invention provides a voice data transmission method applied at a receiving end, the method comprising:
[0018] Receive data packets sent from the sending end, parse the data packets to obtain parsed data, the data including the first N... t Each codeword index value;
[0019] The parsed data is decoded using a decoding model to obtain decoded speech data, wherein the decoding model includes N q N decoder vector quantizers q N is a positive integer. t For less than or equal to N q Positive integers.
[0020] In one possible embodiment, the parsed data is decoded using a decoding model to obtain decoded speech data, including:
[0021] Determine if it includes the first N t N of the codeword index values q N codeword index values, of which the first N t The codeword index value is a valid value, excluding the first N. t Codeword index values other than the stated codeword index values are invalid;
[0022] N q Invalid values in the codeword index values are assigned the value 0, resulting in the updated N. q Each codeword index value will be updated to N. q Input N codeword index values q The decoder vector quantizers obtained N q Each codeword is assigned an all-zero vector corresponding to the invalid index value. All codewords are summed to obtain an inverse quantized vector. The inverse quantized vector is then decoded by the decoder of the decoding model to obtain the decoded speech data.
[0023] Fourthly, the present invention provides a model training method, the method comprising:
[0024] The encoding model and the decoding model in the above embodiments are trained using a training sample set. The encoding model includes an encoder and a residual vector quantization encoding module, and the residual vector quantization encoding module includes N... q The decoding model includes N encoding vector quantizers, and the residual vector quantization decoding module includes a decoder. The residual vector quantization decoding module includes N... q N decoder vector quantizers q It is a positive integer;
[0025] The model parameters of the trained decoding model are frozen, and the frozen decoding model is tested using a second test sample set to evaluate the performance of the decoding model.
[0026] Fifthly, this invention also provides a voice data transmission device, which includes modules / units that execute any of the possible design methods described in the first aspect. These modules / units can be implemented in hardware or by hardware executing corresponding software.
[0027] Sixthly, this embodiment of the invention also provides a model training apparatus, which includes modules / units for executing any of the possible design methods described in the first aspect. These modules / units can be implemented in hardware or by hardware executing corresponding software.
[0028] In a seventh aspect, embodiments of the present invention also provide a voice data transmission device, which includes modules / units for executing any of the possible design methods described in the third aspect above. These modules / units can be implemented in hardware or by hardware executing corresponding software.
[0029] Eighthly, this invention also provides a model training apparatus, which includes modules / units for executing any of the possible design methods described in the third aspect above. These modules / units can be implemented in hardware or by hardware executing corresponding software.
[0030] In a ninth aspect, the present invention also provides a chip for use in an electronic device, the chip being used to execute any of the possible designs of the first to fourth aspects described above.
[0031] In a tenth aspect, embodiments of the present invention provide a terminal, including a processor and a memory. The memory stores one or more computer programs; when the one or more computer programs stored in the memory are executed by the processor, the terminal is able to implement any of the possible design methods described in the first to fourth aspects.
[0032] The beneficial effect of this invention is that: since the target bitrate is not within the selectable bitrate range, the code vector quantizer can be pruned according to the target bitrate to adjust the codebook size of the code vector quantizer and obtain the pruned first N bits. t Each encoding vector quantizer and a cropped optional bitrate range including the target bitrate allow for customization of the optional bitrate to meet the diverse needs of speech coding and data transmission in different application scenarios without retraining the model. Attached Figure Description
[0033] Figure 1 This is a flowchart illustrating the application of the voice data transmission method of the present invention at the sending end.
[0034] Figure 2 This is a schematic diagram illustrating the flow of the voice data transmission method of the present invention from the sending end to the receiving end.
[0035] Figure 3 This is a schematic diagram of the pruning code vector quantizer in the voice data transmission method of the present invention.
[0036] Figure 4 This is a schematic diagram illustrating the working principle of the encoding model in the voice data transmission method of this invention.
[0037] Figure 5 This is a schematic diagram of the model training method of the present invention in a first preferred embodiment.
[0038] Figure 6 This is a schematic diagram illustrating the application of the voice data transmission method of the present invention at the receiving end.
[0039] Figure 7 This is a schematic diagram illustrating the working principle of the decoding model in the voice data transmission method of this invention.
[0040] Figure 8 This is a flowchart illustrating the model training method of the present invention in a second preferred embodiment.
[0041] Figure 9 This is a schematic diagram of the voice data transmission device of the present invention applied to the transmitting end.
[0042] Figure 10 This is a schematic diagram of the process from the sending end to the receiving end of the voice data transmission device of the present invention.
[0043] Figure 11 This is a schematic diagram of the model training device of the present invention in a first specific embodiment.
[0044] Figure 12 This is a schematic diagram of the voice data transmission device of the present invention applied to the receiving end.
[0045] Figure 13This is a schematic diagram of the model training device of the present invention in a second specific embodiment.
[0046] Figure 14 This is a schematic diagram of the terminal of the present invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] The technical terms involved in this invention will be explained below.
[0049] Vocoder: A device or algorithm that converts speech signals into digital codes.
[0050] Bit rate: refers to the number of bits transmitted or processed per unit of time. In vocoders, bit rate represents the number of bits that need to be transmitted or stored per second, measured in bps or kbps.
[0051] Codebook: A codebook is a table used to map signals into discrete codewords. In a vocoder, the codebook contains a series of fixed-length codewords, each representing a possible speech feature.
[0052] Code book size: refers to the number of codewords in the code book.
[0053] AI (Artificial Intelligence) model: refers to an artificial intelligence model built by machine learning algorithms or deep learning neural networks.
[0054] Training graph: A computational graph used during training. It describes the structure of the model and the relationships between its parameters, defining how each parameter is updated during backpropagation. The training graph is used to adjust the model's parameters based on the input data, enabling the model to better fit the training data and improve its performance.
[0055] Model freezing refers to setting the parameters of a model to a non-trainable (non-modifiable) state after training. By freezing the model, it is ensured that the model parameters will not be modified again when performing inference or applying the model, thus maintaining the model's stability and consistency.
[0056] An inference graph is a computational graph used for prediction or inference in real-world applications. Unlike a training graph, an inference graph only contains the parts used to process input data and generate output results; it does not involve gradient calculations or parameter updates. Once the model is frozen, the inference graph can be used for efficient real-time inference without a complex training process.
[0057] Existing variable bitrate artificial intelligence (AI) vocoder models have the following problems:
[0058] (1) The selectable bitrate must be a multiple of log2N, where N is the codebook size. To meet the customized needs of customers, the training model must be redesigned, which will consume a lot of time and resources.
[0059] (2) Residual Vector Quantization (RVQ) progressive encoding obtains some codeword index values. There is no obvious correlation between the index values, so packet loss compensation algorithms cannot be directly applied.
[0060] (3) RVQ contains a variable loop, and the dynamic model is not conducive to edge deployment.
[0061] An embodiment of the present invention provides a voice data transmission method, see below. Figure 1 and Figure 2 The method includes:
[0062] S101: Acquire speech data, and encode the speech data using a coding model to obtain N. q There are N codeword index values, where the encoding model includes N q N encoding vector quantizers q It is a positive integer.
[0063] S102: Determine the target bitrate based on the current network environment, and then determine the bitrate from N based on the target bitrate. q Determine the top N required from the encoded vector quantizers. t N encoding vector quantizers, where N t For less than or equal to N q N is a positive integer. t It equals the ratio of the target bit rate to log2N, where N is the codebook value of the encoding vector quantizer.
[0064] S103: If the target bitrate is within the selectable bitrate range of the coding model, then the first N... t If the codeword index values of the quantizers are determined as the data to be processed, and the target code rate is not within the selectable code rate range, then the quantizers are pruned according to the target code rate to obtain the first N pruned codes. tEach encoded vector quantizer and a cropped selectable bitrate range containing the target bitrate, thereby adjusting the first N... t The storage space occupied by the codeword index values of each encoder vector quantizer, after pruning the first N... t The codeword index value of each encoded vector quantizer is determined as the data to be processed.
[0065] S104: Process the data to be processed to obtain data packets, and send the data packets to the receiving end through the channel or storage medium according to the target code rate.
[0066] In this embodiment, if the target bitrate is within the selectable bitrate range of the coding model, that is, the target bitrate is N t If the value is *log2N, then the output of the encoding model needs to be truncated, retaining only the first N bits. t The codeword index values of each code vector quantizer are transmitted. If the target code rate is not within the selectable code rate range, the code vector quantizer is pruned according to the target code rate by adjusting the codebook size of the code vector quantizer, adjusting the first N... t The storage value occupied by the codeword index value of each encoding vector quantizer, i.e. the amount of data transmitted or stored, is used to meet the requirements of the target bitrate, thus realizing the customization of the selectable bitrate without retraining the model.
[0067] In a preferred embodiment, see [link to previous document]. Figure 3 The code vector quantizer is pruned according to the target bit rate to obtain the first N pruned bits. t Each encoded vector quantizer and a cropped selectable bitrate range including the target bitrate, including: based on the target bitrate and N t The value is calculated to obtain the second to Nth values. t The target codebook for N code vector quantizers, the first code vector quantizer remains unchanged, and the second through Nth code vector quantizers are used as the basis for... t The target codebook of each encoded vector quantizer is correspondingly pruned from the second to the Nth. t One encoded vector quantizer is used to obtain the second to Nth quantizers. t Each pruned encoded vector quantizer and a pruned optional bitrate range containing the target bitrate. In this embodiment, the first encoded vector quantizer remains unchanged to retain the most representative vector in the entire codebook, which helps to obtain better decoding performance and more accurate quantization results.
[0068] For example, the number of encoding vector quantizers N q =3, codebook size N = 256 for each quantizer of the encoded vector. See also Figure 3 From the 0th to the Nth q-1 Each cell represents the first to the Nth cell. qThe storage value occupied by each codeword index value is as follows: the storage value occupied by the first codeword index value is log2(256) = 8 bits; the storage value occupied by the first and second codeword index values together is 2*log2(256) = 16 bits; the storage value occupied by the first to third codeword index values together is 3*log2(256) = 24 bits.
[0069] When 1 second of audio is divided into 50 frames, the selectable bitrates are 50*8=400bps, 50*16=800bps, and 50*24=1200bps, so the selectable bitrate range is {400bps, 800bps, 1200bps}.
[0070] After pruning the code vector quantizer, the storage space occupied by the codeword index value of the code vector quantizer becomes smaller. If the size of the first codebook remains unchanged at 256, the size of the second codebook is 128, and the size of the third codebook is 32, the storage space occupied by the first codeword index value is 8 bits; the storage space occupied by the first and second codeword index values is log2(256)+log2(128)=15 bits; the storage space occupied by the first to third codeword index values is log2(256)+log2(128)+log2(32)=20 bits. Similarly, with 50 frames of speech per second, the selectable bitrate range becomes {400bps, 750bps, 1000bps}.
[0071] In a preferred embodiment, see [link to previous document]. Figure 4 The speech data is encoded using a coding model to obtain N. q Each codeword index value includes: feature data obtained by feature extraction of speech data through the encoder of the coding model, and the feature data input to N. q N encoded vector quantizers yield N q Each codeword index value is used to combine N codeword index values using the concat operator. q Concatenate the codeword index values to obtain N q A sequence consisting of codeword index values.
[0072] In one specific embodiment, see Figure 2 The speech data is encoded using a coding model to obtain N. q Before the codeword index value, the process also includes: preprocessing the speech data to obtain preprocessed data, wherein the preprocessing includes at least one of converting the audio waveform into a spectrogram or a Mel spectrogram, removing noise, and suppressing echo.
[0073] The speech data is encoded using a coding model to obtain N. q Each codeword index value includes: N obtained by encoding the preprocessed data using an encoding model. q Each codeword index value.
[0074] In one specific embodiment, after preprocessing the voice data, the method further includes: normalizing the preprocessed data to the range [-1, 1] to obtain normalized data.
[0075] The speech data is encoded using a coding model to obtain N. q Each codeword index value includes: N obtained by encoding the normalized data using an encoding model. q Each codeword index value.
[0076] In a first preferred embodiment, the present invention provides a model training method, the method comprising:
[0077] S501: Train the encoding model and decoding model in the above embodiments using the training sample set. The encoding model includes an encoder and a Residual Vector Quantization (RVQ) encoding module. The Residual Vector Quantization encoding module includes N... q The encoding vector quantizer and decoding model include a residual vector quantization (RVQ) decoding module and a decoder. The residual vector quantization decoding module includes N... q N decoder vector quantizers q It is a positive integer.
[0078] Specifically, the training sample set includes large-scale speech data with different languages, speakers, speech rates, emotions, and background noise to cover as many application scenarios as possible. The speech data file format can be PCM, WAV, etc. The encoder is used to extract feature representations of the audio signal. The encoding and decoding models can employ recurrent neural networks (RNNs) such as Long Short-Term Memory (LSTM) or Gated Recurrent Units (GRUs), or convolutional neural networks (CNNs), or even model architectures such as Transformers.
[0079] During training, the encoder feeds the extracted features into N. q There are N codebook quantizers with a codebook size of N and a codebook dimension of D. In each training round, one of the N codebooks is randomly selected. t The value can be set at intervals to narrow down the range of options, for example, N q =8, during training, {2, 4, 6, 8} can be selected, or certain N values can be trained in a focused manner according to the target bitrate requirements. tValue. Since it is end-to-end training, the encoded vector is then immediately passed to N. q Decoding is performed in a decoder vector quantizer. Before decoding, some codeword index values are randomly discarded to simulate data loss scenarios and improve the robustness of the model. Finally, the data is fed into a decoder to recover the speech; the decoder is generally the inverse process of the encoder.
[0080] S502: Freeze the model parameters of the trained coding model, test the frozen coding model using the first test sample set, count the frequency of each codeword index value obtained from the test, reorder the codewords obtained from the test according to the frequency from high to low, obtain a new codebook, and update the codeword index value corresponding to the codeword.
[0081] The model freezing phase freezes the dynamic training graph into a fixed-shape inference graph and splits the training model into encoding and decoding models, freezing them separately. Different N t The value will result in different numbers of layers in the residual vector quantization encoding module, by using all N q The encoded vector quantizer is frozen (example N in the diagram). q =3), the codeword index values generated by each encoding vector quantizer are then concatenated using the concat operator. A codebook is actually a feature space composed of N D-dimensional vectors, which can be denoted as {q0, q1, ..., q...} N-1 The codeword index value corresponding to a codeword is the subscript i, which ranges from 0 to N-1. The smaller the codeword index value, the fewer bits are required for storage. Codewords are reordered according to frequency, placing codewords with higher frequency index values at the front, thus saving storage space. Even if the codebook is pruned, high-frequency codewords can be retained to maintain quantization loss. During the freezing phase, conditional masking is used to freeze the coding model, facilitating variable code rate implementation by directly pruning codeword index values.
[0082] In addition, during the model freezing phase, the encoding and decoding models can be optimized by performing parameter pruning, weight quantization, graph optimization, and other processes, depending on the actual deployment platform.
[0083] This invention provides a voice data transmission method applied at a receiving end, the method comprising:
[0084] S601: Receives data packets sent from the sender, parses the data packets to obtain parsed data, which includes the first N... t Each codeword index value. The code rate information and encoded data are parsed from the data packet; if insufficient to meet the input size of the decoding model, they are padded with -1.
[0085] S602: The parsed data is decoded using a decoding model to obtain decoded speech data. The decoding model includes N... qN decoder vector quantizers q N is a positive integer. t For less than or equal to N q Positive integers. The numerical range of the decoded speech data is between [-1, 1].
[0086] In a preferred embodiment, see [link to previous document]. Figure 7 The parsed data is decoded using a decoding model to obtain decoded speech data, including:
[0087] Determine if it includes the first N t N codeword index values q N codeword index values, of which the first N t Each codeword index value is a valid value, divided by the first N. t Codeword index values other than the specified codeword index value are invalid, and the invalid value is -1.
[0088] N q Invalid values in each codeword index are assigned the value 0, resulting in the updated N. q Each codeword index value will be updated to N. q Input N codeword index values q N decoder vector quantizers yield N q Each codeword is converted into N using the concat operator. q Concatenate the codeword index values to obtain N q A sequence of codeword index values is formed. The codewords corresponding to invalid index values are assigned to a vector of all zeros. All codewords are accumulated to obtain the dequantized vector. The dequantized vector is then decoded by the decoder of the decoding model to obtain the decoded speech data.
[0089] The decoding process can be broken down into two steps, such as Figure 7 As shown, N q =3. First, find the corresponding codeword by its index value. The rule requires the codeword index value to be no less than 0. Since invalid values are -1, they are subsequently assigned the value 0 to ensure the codeword index value is not less than 0. Second, sum up the codewords decoded by all the decoder vector quantizers. To obtain a fixed-shape inference graph, the model input for one frame of speech is still set to N. q N codeword index values, of which only the first N t The input is assigned a valid value and an invalid value is assigned -1. The model needs to determine if the input is equal to -1 and divides into two branches. One branch assigns 0 to invalid values instead of -1, leaving other invalid values unchanged. It then finds the corresponding codewords according to the new codeword index and concatenates all codewords together using the concat operator. The other branch assigns the codeword corresponding to the invalid value -1 as a vector of all zeros. Finally, the reducesum operator is used to sum the codewords to obtain the dequantized vector, which is then sent to the decoder to recover the speech.
[0090] In a preferred embodiment, see [link to previous document]. Figure 2 After parsing the data packets to obtain the parsed data, the process also includes: using forward error correction and packet loss compensation techniques to compensate for packet loss in the parsed data.
[0091] In one specific embodiment, see Figure 2 After decoding the parsed data using a decoding model to obtain decoded speech data, the process also includes post-processing the decoded speech data to obtain post-processed speech data. Post-processing includes at least one of inverse normalization and gain compensation to improve the quality of the speech signal.
[0092] In a second preferred embodiment, the present invention provides a model training method, see [link to previous embodiment]. Figure 8 The method includes:
[0093] S801: Train the encoding model and the decoding model in the above embodiments using the training sample set. The encoding model includes an encoder and a residual vector quantization encoding module. The residual vector quantization encoding module includes N... q The encoding vector quantizer and decoding model include a residual vector quantization decoding module and a decoder. The residual vector quantization decoding module includes N... q N decoder vector quantizers q It is a positive integer.
[0094] When training the decoding model, some codeword index values are randomly discarded before the residual vector quantization decoding module performs decoding to simulate data packet loss scenarios and improve the model's ability to withstand packet loss.
[0095] S802: Freeze the model parameters of the trained decoding model and test the frozen decoding model using the second test sample set to evaluate the performance of the decoding model.
[0096] In addition, this invention also proposes a voice data transmission device for use at the transmitting end, see [link to relevant documentation]. Figure 9 and Figure 10 The device includes: a speech encoding unit 901, used to acquire speech data and encode the speech data using an encoding model to obtain N. q There are N codeword index values, where the encoding model includes N q N encoding vector quantizers q The value is a positive integer; the determining unit 902 is used to determine the target bitrate based on the current network environment, and to determine the target bitrate from N. q The top N are determined from the encoded vector quantizers. t N encoding vector quantizers, where N t For less than or equal to N qA positive integer; the data packet processing unit 903, used to process the first N data packets if the target bit rate is within the selectable bit rate range of the coding model. t If the codeword index values of the quantizers are determined as the data to be processed, and the target code rate is not within the selectable code rate range, then the quantizers are pruned according to the target code rate to obtain the first N pruned codes. t Each encoded vector quantizer and a cropped selectable bitrate range containing the target bitrate will crop the first N bits. t The codeword index value of each encoding vector quantizer is determined as the data to be processed; the data packet processing unit 903 is also used to process the data to be processed to obtain data packets; the data packet sending unit 904 is used to send the data packets to the receiving end according to the target bit rate. All relevant content of each step involved in any embodiment of the above voice data transmission method can be referred to the functional description of the corresponding functional module, and will not be repeated here.
[0097] In a first specific embodiment, the present invention also proposes a model training device, see [link to relevant documentation]. Figure 11 The device includes: a first training unit 1101, used to train the encoding model and decoding model of any embodiment of the above-described speech data transmission method using a training sample set, wherein the encoding model includes an encoder and a residual vector quantization encoding module, and the residual vector quantization encoding module includes N q The encoding vector quantizer and decoding model include a residual vector quantization decoding module and a decoder. The residual vector quantization decoding module includes N... q N decoder vector quantizers q The variable is a positive integer; the first freezing unit 1102 is used to freeze the model parameters of the trained encoding model, test the frozen encoding model using the first test sample set, count the frequency of each codeword index value obtained from the test, and reorder the codewords obtained from the test according to the frequency from high to low to obtain a new codebook and update the codeword index value corresponding to the codeword. All relevant content of each step involved in any embodiment of the above voice data transmission method can be referred to the functional description of the corresponding functional module, and will not be repeated here.
[0098] This invention also proposes a voice data transmission device for use at the receiving end, see [link to relevant documentation]. Figure 10 and Figure 12 The device includes: a data packet parsing unit 1201, used to receive data packets sent from a sending end, parse the data packets to obtain parsed data, the data including the first N... t Each codeword index value; a speech decoding unit 1202, used to decode the parsed data through a decoding model to obtain decoded speech data, wherein the decoding model includes N... q N decoder vector quantizers q N is a positive integer.t For less than or equal to N q Positive integers. All relevant content regarding each step in any embodiment of the above voice data transmission method can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0099] In a second specific embodiment, the present invention also proposes a model training device, see [link to relevant documentation]. Figure 13 The device includes a second training unit 1301, used to train the encoding model and the decoding model of any embodiment of the above-described speech data transmission method using a training sample set. The encoding model includes an encoder and a residual vector quantization encoding module, the residual vector quantization encoding module including N... q The encoding vector quantizer and decoding model include a residual vector quantization decoding module and a decoder. The residual vector quantization decoding module includes N... q N decoder vector quantizers q The value is a positive integer; the second freezing unit 1302 is used to freeze the model parameters of the trained decoding model, and to test the frozen decoding model using the second test sample set to evaluate the performance of the decoding model. All relevant content of each step involved in any embodiment of the above voice data transmission method can be referred to in the functional description of the corresponding functional module, and will not be repeated here.
[0100] Compared with existing variable bitrate AI vocoder models, the present invention has the following advantages:
[0101] (1) Based on the needs of different application scenarios, such as the current network environment, voice transmission requirements, and device performance, the coding vector quantizer is pruned and the codebook size of the coding vector quantizer is adjusted to obtain a range of selectable bitrates after pruning that includes the target bitrate. Customizable selectable bitrates are achieved through codebook pruning, without the need to redesign the training model, thus avoiding a lot of time and resources.
[0102] (2) The present invention sorts codewords according to frequency and establishes the correlation between codeword index values, so that packet loss compensation algorithm can be directly applied.
[0103] (3) Randomly discard some codeword index values during the training phase to simulate data loss scenarios and improve the model's ability to resist packet loss.
[0104] (4) Currently, the residual vector quantization module contains a variable loop body, which can be set to different N values. t The dynamic model of prematurely interrupting the loop in RVQ is not conducive to edge deployment. However, this invention prunes the code vector quantizer according to the target bitrate, eliminating the need to prematurely interrupt the loop in RVQ, making it more user-friendly for edge deployment.
[0105] This invention also provides a chip for use in electronic devices, which executes the methods described in the above embodiments. From the chip's perspective, a dedicated hardware accelerator or neural network processor is required to execute the AI vocoder. The AI vocoder includes an encoding model and a decoding model. An AI vocoder integrated on a chip has the characteristics of high efficiency and real-time performance.
[0106] In other embodiments of the present invention, a terminal is disclosed, see [link to relevant documentation]. Figure 14 The terminal may include: one or more processors 1401; a memory 1402; a display 1403; one or more application programs (not shown); and one or more computer programs 1404. These devices can be connected via one or more communication buses 1405. The one or more computer programs 1404 are stored in the memory 1402 and configured to be executed by the one or more processors 1401. The one or more computer programs 1404 include instructions that can be used to perform actions such as... Figure 1 , Figure 5 , Figure 6 , Figure 8 , Figure 9 , Figures 11 to 13 And the steps in the corresponding embodiments.
[0107] From the perspective of the terminal, AI vocoders are usually installed as software libraries or applications (APPs) on terminal devices, such as smartphones, smart speakers, or smart headphones, and have a high degree of flexibility.
[0108] From the perspective of base stations, AI vocoders are not strictly limited in size and can make full use of deep learning technology to improve the quality and efficiency of voice communication and enhance the user experience of mobile communication networks.
[0109] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0110] In the various embodiments of this invention, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0111] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.
[0112] While embodiments of the invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations fall within the scope and spirit of the invention as defined in the claims. Furthermore, the invention described herein may have other embodiments and can be implemented or carried out in various ways. Unless otherwise defined, the technical or scientific terms used herein should be understood in their ordinary sense by one of ordinary skill in the art to which this invention pertains. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word encompasses the element or object listed following the word and its equivalents, but do not exclude other elements or objects.
Claims
1. A voice data transmission method, applied at a transmitting end, characterized in that, The method includes: The speech data is acquired, and then encoded using an encoding model to obtain... Each codeword index value, wherein the encoding model includes Each encoded vector quantizer It is a positive integer; Determine the target bitrate based on the current network environment, and then proceed from the target bitrate... The first one is determined in the encoded vector quantizer. The quantizers of the encoded vector, wherein... less than or equal to Positive integers; If the target bitrate is within the selectable bitrate range of the coding model, then the preceding bitrate will be... If the codeword index value of the quantizer is determined as the data to be processed, and the target code rate is not within the selectable code rate range, then the quantizer is pruned according to the target code rate to obtain the pruned data. Each encoded vector quantizer and a cropped selectable bitrate range containing the target bitrate, will quantize the cropped bitrate... The codeword index value of each encoded vector quantizer is determined as the data to be processed; The data to be processed is processed to obtain a data packet, and the data packet is sent to the receiving end according to the target bit rate.
2. The method according to claim 1, characterized in that, The quantizer of the encoded vector is pruned according to the target bitrate to obtain the pruned first... Each encoded vector quantizer and a cropped selectable bitrate range including the target bitrate include: Based on the target bit rate and The value is calculated to obtain the second to the third. The target codebook of the quantizers, the first quantizer remains unchanged, and according to the second to the third... The target codebook of the quantizer of the quantizer is correspondingly pruned from the second to the third. The first encoding vector quantizer is used to obtain the second to the third... A pruned encoded vector quantizer and a pruned selectable bitrate range including the target bitrate.
3. A model training method, characterized in that, The method includes: The encoding model and decoding model are trained using a training sample set, wherein the encoding model adopts the encoding model as described in any one of claims 1 or 2 and includes an encoder and a residual vector quantization encoding module, the residual vector quantization encoding module including... A coded vector quantizer, the decoding model includes a residual vector quantization decoding module and a decoder, the residual vector quantization decoding module includes... One decoder vector quantizer, It is a positive integer; The model parameters of the trained encoding model are frozen. The frozen encoding model is tested using the first test sample set. The frequency of each codeword index value obtained from the test is counted. The codewords obtained from the test are reordered according to the frequency from high to low to obtain a new codebook and update the codeword index value corresponding to the codeword.
4. A voice data transmission method, applied at a receiving end, characterized in that, The method includes: Receive data packets sent from the sending end, parse the data packets to obtain parsed data, the data including the preceding... Each codeword index value; The parsed data is decoded using a decoding model to obtain decoded speech data, wherein the decoding model includes... One decoder vector quantizer, It is a positive integer. less than or equal to Positive integers.
5. The method according to claim 4, characterized in that, The parsed data is decoded using a decoding model to obtain decoded speech data, including: Determine if included before The codeword index values Each codeword index value, of which the first... The codeword index value is a valid value, except for the first one. Codeword index values other than the stated codeword index values are invalid; Will Invalid values in the codeword index values are assigned the value of 0, resulting in the updated... Each codeword index value, the updated Input the codeword index value The decoder vector quantizers obtained For each codeword, the codeword corresponding to the invalid index value is assigned a vector of all zeros. All codewords are accumulated to obtain the dequantized vector. The dequantized vector is then decoded by the decoder of the decoding model to obtain the decoded speech data.
6. A model training method, characterized in that, The method includes: The encoding model and the decoding model as described in claim 4 or 5 are trained using a training sample set, wherein the encoding model includes an encoder and a residual vector quantization encoding module, and the residual vector quantization encoding module includes... A coded vector quantizer, the decoding model includes a residual vector quantization decoding module and a decoder, the residual vector quantization decoding module includes... One decoder vector quantizer, It is a positive integer; The model parameters of the trained decoding model are frozen, and the frozen decoding model is tested using a second test sample set to evaluate the performance of the decoding model.
7. A voice data transmission device, applied at a transmitting end, characterized in that, The device includes: A speech encoding unit is used to acquire the speech data and encode the speech data using an encoding model. Each codeword index value, wherein the encoding model includes Each encoded vector quantizer It is a positive integer; The determining unit is used to determine the target bitrate based on the current network environment, and to determine the target bitrate from... The first one is determined in the encoded vector quantizer. The quantizers of the encoded vectors, wherein... less than or equal to Positive integers; The data packet processing unit is used to process the data packet with the target bitrate within the selectable bitrate range of the encoding model. If the codeword index value of the quantizer is determined as the data to be processed, and the target code rate is not within the selectable code rate range, then the quantizer is pruned according to the target code rate to obtain the pruned first... Each encoded vector quantizer and a cropped selectable bitrate range containing the target bitrate, will quantize the cropped bitrate... The codeword index value of each encoded vector quantizer is determined as the data to be processed; The data packet processing unit is also used to process the data to be processed to obtain a data packet; A data packet sending unit is used to send the data packet to the receiving end according to the target bit rate.
8. A model training device, characterized in that, The device includes: The first training unit is used to train the encoding model and the decoding model as described in any one of claims 1 or 2 using a training sample set, wherein the encoding model includes an encoder and a residual vector quantization encoding module, and the residual vector quantization encoding module includes... A coded vector quantizer, the decoding model includes a residual vector quantization decoding module and a decoder, the residual vector quantization decoding module includes... One decoder vector quantizer, It is a positive integer; The first freezing unit is used to freeze the model parameters of the trained coding model, test the frozen coding model using the first test sample set, count the frequency of each codeword index value obtained from the test, reorder the codewords obtained from the test according to the frequency from high to low, obtain a new codebook, and update the codeword index value corresponding to the codeword.
9. A voice data transmission device, applied at a receiving end, characterized in that, The device includes: The data packet parsing unit is used to receive data packets sent from the sending end, parse the data packets to obtain parsed data, the data including the preceding data. Each codeword index value; A speech decoding unit is used to decode the parsed data using a decoding model to obtain decoded speech data, wherein the decoding model includes... One decoder vector quantizer, It is a positive integer. less than or equal to Positive integers.
10. A model training device, characterized in that, The device includes: The second training unit is used to train the encoding model and the decoding model as described in claim 4 or 5 using a training sample set, wherein the encoding model includes an encoder and a residual vector quantization encoding module, and the residual vector quantization encoding module includes... A coded vector quantizer, the decoding model includes a residual vector quantization decoding module and a decoder, the residual vector quantization decoding module includes... One decoder vector quantizer, It is a positive integer; The second freezing unit is used to freeze the model parameters of the trained decoding model and to test the frozen decoding model using the second test sample set to evaluate the performance of the decoding model.
11. A chip used in electronic devices, characterized in that, The chip is used to perform the method as described in any one of claims 1 to 6.
12. A terminal, characterized in that, include: Processor and memory, wherein the memory is used to store computer programs; The processor is configured to execute the computer program stored in the memory to cause the terminal to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method for video encoding or decoding based on orthogonal transform and vector quantization, and apparatus thereof
CN101009839A
Method and system for processing video data
CN101090495A