Method, device and equipment for training voice coding model and readable medium

By training the speech coding model by generating discrete features and probabilistic information, the problems of weak noise resistance and overfitting in unsupervised pre-training are solved, achieving efficient model training and improved speech recognition performance.

CN121122255APending Publication Date: 2025-12-12BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410749776.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing speech coding models suffer from weak noise resistance and overfitting during unsupervised pre-training, resulting in poor model training performance.

Method used

By using a first speech coding model to generate discrete features, generating label information based on the discrete features, and using a second speech coding model to generate probability information, the training loss is determined by combining the weight information, and the parameters of the second speech coding model are adjusted to improve the training stability and effectiveness of the model.

Benefits of technology

It reduces the amount of data required for training, improves the stability and capabilities of unsupervised model training, and enhances the performance of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122255A_ABST
    Figure CN121122255A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for training a voice coding model, equipment and a readable medium. The method comprises: processing a speech feature representation of a speech sample using a first speech coding model to generate a set of discrete features; generating label information based on the set of discrete features, the label information comprising a set of labels, the set of labels indicating clustering centers corresponding to the corresponding discrete features; processing an intermediate feature representation generated based on the voice feature representation by using a second voice coding model to generate probability information corresponding to the tag information; the training loss is determined based on the label information, the probability information and the weight information, and the weight information is determined based on the distance from each discrete feature to the corresponding clustering center; and adjusting parameters of the second speech coding model based on the training loss. Therefore, the data volume required by training can be reduced, the stability of the unsupervised model training process can be improved, the training effect of the model is improved, and the model capability is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computer technology, and more specifically, to methods, apparatus, devices, and computer-readable storage media for training speech coding models. Background Technology

[0002] With the development of internet technology, an increasing number of applications and platforms offer natural language processing (NLP) capabilities, bringing numerous conveniences to users. Applications and platforms with NLP functionality can provide NLP services based on trained machine learning models. Speech recognition is a crucial task within NLP. The aim is to improve the quality of NLP services by enhancing the capabilities of machine learning models. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for training a speech coding model is provided. The method includes: processing speech feature representations of speech samples using a first speech coding model to generate a set of discrete features; generating label information based on the set of discrete features, the label information including a set of labels indicating cluster centers corresponding to the respective discrete features; processing intermediate feature representations generated based on the speech feature representations using a second speech coding model to generate probability information corresponding to the label information; determining a training loss based on the label information, probability information, and weight information, the weight information being determined based on the distances from each discrete feature to its corresponding cluster center; and adjusting the parameters of the second speech coding model based on the training loss.

[0004] In a second aspect of this disclosure, an apparatus for training a speech coding model is provided. The apparatus includes: a discrete feature generation module configured to process speech feature representations of speech samples using a first speech coding model to generate a set of discrete features; a label information generation module configured to generate label information based on the set of discrete features, the label information including a set of labels indicating cluster centers corresponding to the respective discrete features; a probability information generation module configured to process intermediate feature representations generated based on the speech feature representations using a second speech coding model to generate probability information corresponding to the label information; a training loss determination module configured to determine a training loss based on the label information, probability information, and weight information, the weight information being determined based on the distances from each discrete feature to its corresponding cluster center; and a model parameter adjustment module configured to adjust the parameters of the second speech coding model based on the training loss.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.

[0008] It should be understood that the content described in this summary section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of various implementations of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1A A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0011] Figure 1B A schematic diagram illustrating an example of a model according to some embodiments of the present disclosure is shown;

[0012] Figure 2 A schematic diagram of an example architecture for training a speech coding model according to some embodiments of the present disclosure is shown;

[0013] Figure 3 Schematic diagrams illustrating examples of some embodiments according to this disclosure are shown;

[0014] Figure 4 A flowchart is shown illustrating a process for training a speech coding model according to some embodiments of the present disclosure;

[0015] Figure 5 A schematic structural block diagram of an apparatus for training a speech coding model according to certain embodiments of the present disclosure is shown; and

[0016] Figure 6A block diagram of a computing device in which one or more embodiments of the present disclosure may be implemented is shown. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0019] It should be noted that the acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0020] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0021] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.

[0022] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0023] It is understood that the above notification and user authorization acquisition process is merely illustrative and does not constitute a limitation on the embodiments of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the embodiments of this disclosure.

[0024] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0025] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0026] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating its parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual inputs based on the trained parameter values ​​to determine the corresponding output.

[0027] Figure 1A A schematic diagram of an example environment 100A in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1A As shown, environment 100A may include electronic device 110.

[0028] Electronic device 110 can, for example, convert speech 102 into a text sequence 112 that matches speech 102. That is, electronic device 110 can perform speech recognition on speech 102 to generate the corresponding text sequence 112. Here, speech 102 can be speech in any suitable language and of any duration. Electronic device 110 can perform speech recognition on speech 102 to generate a text sequence in its corresponding language. For example, electronic device 110 can recognize speech in English to generate a text sequence in English. Speech 102 can be speech locally generated by electronic device 110 or speech acquired by electronic device 110 in real time.

[0029] Electronic device 110 may utilize a trained model 120 (e.g., a machine learning model) to perform a speech recognition task. Model 120 may be a model native to electronic device 110 or a model installed on another electronic device 110 (e.g., installed on a remote device). Model 120 may include one or more models. If model 120 includes multiple models, these multiple models may include the same model or different models.

[0030] Electronic device 110 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, server devices, etc. Terminal devices may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof.

[0031] Server-side equipment can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server-side equipment may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0032] It should be understood that the structure and function of the various elements in environment 100A are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0033] Figure 1B A schematic diagram of example 100B of a model (e.g., model 120) according to some embodiments of the present disclosure is shown. Example 100B relates to encoding model 130, transformation model 140, and language model 160. Encoding model 130 encodes speech 102 into speech features. Transformation model 140 further transforms the speech features to a dimension that language model 160 can process to obtain transformed speech features 153. Transformed speech features 153 may also be referred to as speech embeddings or a set of speech tokens. The input to language model 160 includes contextual information 151, which may include the content of historical dialogues, a scene description of the speech, and any other information that may be helpful for speech recognition. The input to language model 160 may also include a cue 152, which can be used to instruct language model 160 to perform a speech recognition task. In some scenarios, cue 152 may also instruct other tasks, such as language recognition. Furthermore, the input to language model 160 may also include the transformed speech features 153.

[0034] Language Model 160 can be executed based on NTP (next token prediction), as shown in the diagram. <bos>It indicates the beginning of a sentence and is a marker. <eos>Indicates the end of a sentence, which is also a marker. Each time a prediction is made, the language model 160 can output a token (e.g., a Chinese character or a word). When predicting the next token, the previously generated tokens can be used as a basis for the language model 160 to predict the next token. For example, when the prediction of the token "day" is completed, the language model 160 can predict the token "day" based on the already generated token "today".

[0035] As briefly mentioned above, an application or platform with natural language processing capabilities can provide natural language processing services to users based on a trained machine learning model (e.g., a speech encoding model). Speech recognition tasks are important tasks in natural language processing tasks. With the continuous progress and popularization of artificial intelligence technology, people's expectations for the effect of speech recognition are also getting higher and higher. People expect to better train machine learning models to provide the model capabilities of machine learning models.

[0036] Traditionally, to improve the model capabilities, unsupervised pre-training is usually used to train speech encoding models. Training based on discrete labels is the most common type of method for unsupervised pre-training. Specifically, the speech segment information can be quantized into discrete labels, and the mapping relationship between the model speech segments and the discrete labels can be established, so that the speech encoding model can have the ability to model speech information. However, due to the information loss in the discrete labels themselves, this may lead to the defects of weak anti-noise ability and easy overfitting in the speech encoding model obtained by the method of pre-training based on discrete labels.

[0037] In view of this, embodiments of the present disclosure provide a method for training a speech encoding model. The method includes: processing the speech feature representation of a speech sample by a first speech encoding model to generate a set of discrete features. Generating label information based on the set of discrete features. The label information includes a set of labels, and the set of labels indicates the cluster centers corresponding to the respective discrete features. Processing the intermediate feature representation generated based on the speech feature representation by a second speech encoding model to generate probability information corresponding to the label information. Determining a training loss based on the label information, the probability information, and weight information. The weight information is determined based on the distances from the respective discrete features to the corresponding cluster centers. Adjusting the parameters of the second speech encoding model based on the training loss.

[0038] In this way, embodiments of the present disclosure can train a second speech encoding model through a trained speech encoding model, can train the second speech encoding model only with speech samples including a small amount of data, reduce the amount of data required for training, and can also improve the stability of the unsupervised model training process, improve the training effect of the model, and thus improve the model capabilities.

[0039] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0040] Figure 2 A schematic diagram of an example architecture 200 for training a speech coding model (e.g., coding model 130) according to some embodiments of the present disclosure is shown. The example architecture 200 may be implemented at an electronic device 110. For ease of discussion, the architecture 200 will be described with reference to the environment 100 of FIG1.

[0041] like Figure 2 As shown, architecture 200 includes a feature extraction unit 210, a feature generation unit 220, a label information generation unit 240, a masking unit 250, a probability information generation unit 260, and a loss determination unit 280. Feature extraction unit 210 can extract a speech feature representation 215 of speech sample 201. For example, feature extraction unit 210 can encode speech sample 201 to determine the speech feature representation 215 of speech sample 201. Speech feature representation 245 includes, for example, the spectral features of speech sample 201. Feature generation unit 220 can determine a set of discrete features corresponding to speech sample 201 based on speech feature representation 215. Specifically, feature generation unit 220 can determine multiple segments of speech sample 201. Multiple segments can have the same duration. For example, if speech sample 201 corresponds to 100 seconds, each segment can correspond to 10 seconds of speech.

[0042] The feature generation unit 220 can generate multiple segment feature representations corresponding to multiple segments of the speech sample 201 based on the speech feature representation 215. The feature generation unit 220 can process the multiple segment feature representations using the first speech coding model 230 to generate a set of discrete features 225 corresponding to the multiple segment feature representations. The set of discrete features may, for example, include the results output by the intermediate layers of the first speech coding model 230. The first speech coding model 230 can have any suitable model structure. The first speech coding model 230 includes a speech coding model determined based on a best-RQ (Best-RQ) random discrete label pre-training process.

[0043] For example, if multiple segments of speech sample 201 are (where t is the number of segments, and multiple segments have the same duration), the feature generation unit 220 can determine multiple segment feature representations corresponding to these multiple segments, and can use the first speech coding model 230 to process the multiple segment feature representations to generate a set of intermediate layer features corresponding to the multiple segment feature representations. Intermediate layer features F i This refers to the output of the intermediate layer of the first speech coding model 230. Similarly, the feature generation unit 220 can also generate other features corresponding to multiple segment features, and these other features can be combined with the intermediate layer features F. i Together they form a feature set F, which is also a set of discrete features 225.

[0044] The label information generation unit 240 can generate label information 245 based on a set of discrete features 225. The label information 245 may include a set of labels, which may, for example, indicate cluster centers corresponding to the respective discrete features. In some embodiments, the label information generation unit 240 can determine multiple cluster centers by clustering a set of discrete features 225. The label information generation unit 240 can cluster a set of discrete features 225 in any suitable manner. For example, the label information generation unit 240 can cluster a set of discrete features 225 based on the K-means algorithm to determine multiple cluster centers C = [C1, C2, ..., C...]. n ], where n is the number of cluster centers.

[0045] The label information generation unit 240 can, for example, determine a set of labels corresponding to a set of discrete features 225 based on the distances from a set of discrete features 225 to multiple cluster centers. For example, a set of discrete features 225 might be defined as F = [F1, F2, ..., F...]. t For example, the label information generation unit 240 can generate labels based on each discrete feature F. i To each cluster center C j The distance is calculated to obtain the distance vector D. i ={d ij }, where d ij This represents the distance from the i-th discrete feature to the j-th cluster center. In some embodiments, the label information generation unit 240 can also determine the closest cluster center among multiple cluster centers for each discrete feature in a set of discrete features 225, L = [l1, l2, ..., l]. t ].

[0046] Alternatively or additionally, in addition to automatically determining cluster centers, in some embodiments, the label information generation unit 240 may also determine a set of labels corresponding to a set of discrete features 225 based on the distances from a set of discrete features 225 to multiple preset cluster centers. The multiple preset cluster centers may be, for example, multiple cluster centers historically determined based on the above method. For example, they may be cluster centers determined in previous rounds of training. The multiple preset cluster centers may also be multiple cluster centers specified by the user.

[0047] Masking unit 250 can generate intermediate feature representation 255 based on speech feature representation 215. Specifically, masking unit 250 can apply a target mask to speech feature representation 215 to generate intermediate feature representation 255. The target mask indicates that the feature values ​​of one or more segments of speech feature representation 215 should be set to a predetermined value (e.g., 0).

[0048] The probability information generation unit 260 can process the intermediate feature representation 255 using the second speech coding model 270 to generate probability information 265 corresponding to the label information 245. The second speech coding model 270 can also have any suitable model structure. For example, the second speech coding model 270 can be viewed as a speech coding model determined by a multi-round training (Hubert) process based on K-means clustering of speech spectral features. The probability information 265 may include a set of probabilities.

[0049] The loss determination unit 280 can determine the training loss 285 based on label information 245, probability information 265, and weight information 290. The weight information 290 can be determined, for example, based on the distance from each discrete feature in a set of discrete features 225 to its corresponding cluster center. In some embodiments, the loss determination unit 280 can determine the cross-entropy of the label information 245 and the probability information 265. The loss determination unit 280 can also determine a set of weighting coefficients corresponding to a set of labels included in the label information 245 based on the weight information 290. The loss determination unit 280 can then determine the training loss 285 based on the cross-entropy and the set of weighting coefficients. Exemplarily, the loss determination unit 280 can determine the training loss 285 based on the following formula:

[0050] w = F.softmax(-D i ,dim=-1) (1)

[0051] loss=∑w×CrossEntropy(logit,L) (2)

[0052] Where w is the weighting coefficient, F is a set of discrete features, and D is a weighting coefficient. i For each discrete feature F i To each cluster center C j The distance vector is obtained by measuring the distance between two points. Since the distance vector is a two-dimensional vector, the first dimension is the number of frames in the 201 speech samples, and the second dimension is the number of cluster centers. `dim = -1` indicates that all cluster centers in each frame participate in the calculation, meaning the weighting coefficients are applied to each category of the cluster. Here, D... i The minus sign "-" indicates that the target weighting coefficient corresponding to the target label is negatively correlated with the distance from the target label to the corresponding target cluster center. loss is the training loss 285, CrossEntropy is the cross-entropy loss function, L is the class with the closest distance among multiple cluster centers for each discrete feature, and logit is the output of the last layer of the model.

[0053] The electronic device 110 can adjust the parameters of the second speech coding model 270 based on the training loss 285. It should be noted that during the training of the second speech coding model 270, the label information generation unit 240 can, for example, consist of the first speech coding model 230 and cluster centers, and can also be referred to as a discrete label generator. For example, when the label information generation unit 240 clusters a set of discrete features 225 based on the K-means algorithm to determine multiple cluster centers, the label information generation unit 240 can be referred to as a K-means discrete label generator T. In some embodiments, if the second speech coding model 270 and the first speech coding model 230 have the same structure, the electronic device 110 can also initialize the second speech coding model 270 with the parameters of the first speech coding model 230 to improve the convergence speed.

[0054] refer to Figure 3 , Figure 3 A schematic diagram of example 300 according to some embodiments of the present disclosure is shown. Figure 3 As shown, after acquiring the speech sample 201, the electronic device 110 can determine the spectral features 310 of the speech sample 201. The electronic device 110 can process the spectral features 310 using the first speech coding model 230 to generate a set of discrete features. The discrete label generator 340 can generate label information 330 based on the set of discrete features. The label information 330 includes a set of labels (e.g., A1 A2 A3…AN), which can indicate the cluster centers corresponding to the respective discrete features. The discrete label generator 340 can, for example, be composed of the first speech coding model 230 and the cluster centers corresponding to the set of discrete features.

[0055] Electronic device 110 can apply a target mask to spectral feature 310 to generate masked spectral feature 320. Masked spectral feature 320 can be considered as an intermediate feature representation generated based on spectral feature 310. Electronic device 110 can process masked spectral feature 320 using second speech coding model 270 to generate probability information corresponding to label information. Electronic device 110 can determine a set of weighting coefficients 350 corresponding to a set of labels based on weight information determined by the distance from each discrete feature to the corresponding cluster center. Weighting coefficients 350 can also include a set of weighting coefficients (e.g., B1 B2 B3…BN). For example, if a set of labels is (15 123 3…27), then a set of weighting coefficients can be (0.5 1.0 0.3…0.9). Electronic device 110 can determine the training loss based on the cross-entropy of the determined label information 330 and probability information, and the weighting coefficients 350, and adjust the parameters of the second speech coding model 270 based on the training loss.

[0056] In some embodiments, after the second speech coding model has been trained, the electronic device 110 can further use the trained second speech coding model to process the target speech to generate a target feature representation of the target speech, and use the speech decoding model to process the target feature representation to generate a speech recognition result of the target speech. The speech decoding model can also have any suitable model structure, which may include, for example, a language model.

[0057] In summary, the embodiments of this disclosure can train a second speech coding model using a trained speech coding model, and can train the second speech coding model using only speech samples including a small amount of data, which reduces the amount of data required for training, improves the stability of the unsupervised model training process, improves the training effect of the model, and thus improves the model capability.

[0058] Figure 4 A flowchart of a process 400 for training a speech coding model according to some embodiments of the present disclosure is shown. Process 400 may be implemented at an electronic device 110.

[0059] In box 410, electronic device 110 uses a first speech coding model to process the speech feature representation of a speech sample to generate a set of discrete features.

[0060] In box 420, electronic device 110 generates label information based on a set of discrete features. The label information includes a set of labels that indicate the cluster centers corresponding to the respective discrete features.

[0061] In box 430, electronic device 110 uses a second speech coding model to process intermediate feature representations generated based on speech feature representations to generate probability information corresponding to label information.

[0062] In box 440, electronic device 110 determines the training loss based on label information, probability information, and weight information. The weight information is determined based on the distance from each discrete feature to the corresponding cluster center.

[0063] In box 450, electronic device 110 adjusts the parameters of the second speech coding model based on the training loss.

[0064] In some embodiments, determining the training loss based on label information, probability information, and weight information includes: determining the cross-entropy of label information and probability information; determining a set of weighting coefficients corresponding to a set of labels based on weight information; and determining the training loss based on the cross-entropy and the set of weighting coefficients.

[0065] In some embodiments, the target weighting coefficient corresponding to the target label is negatively correlated with the distance from the target label to the corresponding target cluster center.

[0066] In some embodiments, processing the speech feature representation of a speech sample using a first speech coding model includes: generating multiple segment feature representations corresponding to multiple segments of the speech sample based on the speech feature representations; and processing the multiple segment feature representations using the first speech coding model to generate a set of discrete features corresponding to the multiple segment feature representations, wherein the set of discrete features includes the results output by the intermediate layer of the first speech coding model.

[0067] In some embodiments, process 400 further includes: applying a target mask to the speech feature representation to generate an intermediate feature representation.

[0068] In some embodiments, the speech feature representation includes the spectral features of the speech samples.

[0069] In some embodiments, the first speech coding model includes a speech coding model determined based on a random discrete label pre-training process, and generating label information based on a set of discrete features includes: determining multiple cluster centers by clustering a set of discrete features; and determining a set of labels corresponding to a set of discrete features based on the distances from a set of discrete features to the multiple cluster centers.

[0070] In some embodiments, generating label information based on a set of discrete features includes: determining a set of labels corresponding to a set of discrete features based on the distances from a set of discrete features to multiple preset cluster centers.

[0071] In some embodiments, process 400 further includes: processing the target speech using a trained second speech coding model to generate a target feature representation of the target speech; and processing the target feature representation using a speech decoding model to generate a speech recognition result of the target speech.

[0072] In some embodiments, the speech decoding model includes a language model.

[0073] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 5 A schematic structural block diagram of an apparatus 500 for training a speech coding model according to certain embodiments of the present disclosure is shown. The apparatus 500 may be implemented as or included in an electronic device 110. The various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0074] like Figure 5 As shown, the device 500 includes a discrete feature generation module 510, configured to process the speech feature representation of speech samples using a first speech coding model to generate a set of discrete features. The device 500 also includes a label information generation module 520, configured to generate label information based on the set of discrete features. The label information includes a set of labels indicating the cluster centers corresponding to the respective discrete features. The device 500 further includes a probability information generation module 530, configured to process the intermediate feature representation generated based on the speech feature representation using a second speech coding model to generate probability information corresponding to the label information. The device 500 also includes a training loss determination module 540, configured to determine the training loss based on the label information, probability information, and weight information. The weight information is determined based on the distance from each discrete feature to its corresponding cluster center. The device 500 also includes a model parameter adjustment module 550, configured to adjust the parameters of the second speech coding model based on the training loss.

[0075] In some embodiments, the training loss determination module 540 includes: a first determination module configured to determine the cross-entropy of label information and probability information; a weighting coefficient determination module configured to determine a set of weighting coefficients corresponding to a set of labels based on weight information; and a second determination module configured to determine the training loss based on the cross-entropy and the set of weighting coefficients.

[0076] In some embodiments, the target weighting coefficient corresponding to the target label is negatively correlated with the distance from the target label to the corresponding target cluster center.

[0077] In some embodiments, the discrete feature generation module 510 includes: a segment feature representation generation module configured to generate multiple segment feature representations corresponding to multiple segments of a speech sample based on speech feature representations; and a segment feature representation processing module configured to process the multiple segment feature representations using a first speech coding model to generate a set of discrete features corresponding to the multiple segment feature representations, wherein the set of discrete features includes the results output by the intermediate layer of the first speech coding model.

[0078] In some embodiments, the apparatus 500 further includes a masking module configured to apply a target mask to a speech feature representation to generate an intermediate feature representation.

[0079] In some embodiments, the speech feature representation includes the spectral features of the speech samples.

[0080] In some embodiments, the first speech coding model includes a speech coding model determined based on a random discrete label pre-training process, and the label information generation module 520 includes: a cluster center determination module configured to determine multiple cluster centers by clustering a set of discrete features; and a label determination module configured to determine a set of labels corresponding to a set of discrete features based on the distance from a set of discrete features to the multiple cluster centers.

[0081] In some embodiments, the label information generation module 520 includes a third determining module, configured to determine a set of labels corresponding to a set of discrete features based on the distances from a set of discrete features to multiple preset cluster centers.

[0082] In some embodiments, the apparatus 500 further includes: a target feature representation generation module configured to process the target speech using a trained second speech coding model to generate a target feature representation of the target speech; and a speech recognition result determination module configured to process the target feature representation using a speech decoding model to generate a speech recognition result of the target speech.

[0083] In some embodiments, the speech decoding model includes a language model.

[0084] The units and / or modules included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0085] It should be understood that one or more steps in the above methods can be performed by a suitable electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices may, for example, include the electronic device 110 in FIG1.

[0086] Figure 6 A block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 6 The electronic device 600 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 6 The electronic device 600 shown can be used to implement the electronic device 110 of FIG1 and / or Figure 5 The device 500.

[0087] like Figure 6 As shown, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 600.

[0088] Electronic device 600 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data and accessible within electronic device 600.

[0089] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 6 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0090] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0091] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0092] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0093] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0094] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0095] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0097] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.< / eos> < / bos>

Claims

1. A method for training a speech coding model, comprising: The speech feature representation of the speech sample is processed using the first speech coding model to generate a set of discrete features; Based on the set of discrete features, label information is generated, the label information including a set of labels, the set of labels indicating the cluster centers corresponding to the respective discrete features; The intermediate feature representation generated based on the speech feature representation is processed using the second speech coding model to generate probability information corresponding to the label information; Based on the label information, the probability information, and the weight information, a training loss is determined, wherein the weight information is determined based on the distance from each discrete feature to its corresponding cluster center; and Based on the training loss, the parameters of the second speech coding model are adjusted.

2. The method according to claim 1, wherein determining the training loss based on label information, probability information, and weight information includes: Determine the cross-entropy of label information and probability information; Based on the weight information, a set of weighting coefficients corresponding to the set of labels are determined; as well as The training loss is determined based on the cross-entropy and the set of weighting coefficients.

3. The method according to claim 2, wherein the target weighting coefficient corresponding to the target label is negatively correlated with the distance from the target label to the corresponding target cluster center.

4. The method according to claim 1, wherein processing the speech feature representation of the speech sample using the first speech coding model includes: Based on the speech feature representation, generate multiple segment feature representations corresponding to multiple segments of the speech sample; as well as The first speech coding model is used to process the multiple segment feature representations to generate the set of discrete features corresponding to the multiple segment feature representations, wherein the set of discrete features includes the results output by the intermediate layer of the first speech coding model.

5. The method according to claim 1, further comprising: The target mask is applied to the speech feature representation to generate the intermediate feature representation.

6. The method of claim 5, wherein the speech feature representation includes the spectral features of the speech sample.

7. The method according to claim 1, wherein the first speech coding model comprises a speech coding model determined based on a random discrete label pre-training process, and generating label information based on the set of discrete features comprises: Multiple cluster centers are determined by clustering a set of discrete features; as well as Based on the distance from the set of discrete features to the multiple cluster centers, the set of labels corresponding to the set of discrete features is determined.

8. The method according to claim 1, wherein generating label information based on the set of discrete features comprises: Based on the distances from the set of discrete features to multiple preset cluster centers, the set of labels corresponding to the set of discrete features are determined.

9. The method according to claim 1, further comprising: The target speech is processed using the trained second speech coding model to generate a target feature representation of the target speech; as well as The target feature representation is processed using a speech decoding model to generate a speech recognition result for the target speech.

10. The method of claim 9, wherein the speech decoding model includes a language model.

11. An apparatus for training a speech coding model, comprising: The discrete feature generation module is configured to process the speech feature representation of the speech sample using the first speech coding model to generate a set of discrete features; The label information generation module is configured to generate label information based on the set of discrete features, wherein the label information includes a set of labels, and the set of labels indicates the cluster centers corresponding to the corresponding discrete features; The probability information generation module is configured to process the intermediate feature representation generated based on the speech feature representation using the second speech coding model, so as to generate probability information corresponding to the label information; The training loss determination module is configured to determine the training loss based on the label information, the probability information, and the weight information, wherein the weight information is determined based on the distance from each discrete feature to the corresponding cluster center. as well as The model parameter adjustment module is configured to adjust the parameters of the second speech coding model based on the training loss.

12. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Model training method based on adaptive attention weighted contrast loss function

    CN122116034A