A decoupled voice self-supervised pre-training method

By using a decoupled speech self-supervised pre-training method, independent representations of pitch changes, speaker and content are extracted, which solves the problem of performance imbalance of speech self-supervised learning methods in different downstream tasks, and realizes the flexible adaptation and performance improvement of the model in multi-task scenarios.

CN118841029BActive Publication Date: 2025-11-28TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411011648.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-11-28
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

Existing speech self-supervised learning methods suffer from performance imbalances across different downstream tasks, making it difficult to achieve comprehensive performance improvements across various tasks while maintaining model generalization capabilities, and they also have high resource requirements.

Method used

We employ a decoupled speech self-supervised pre-training method, which progressively extracts independent representations of pitch changes, speaker and content, and utilizes techniques such as residual decoupling and speech masking prediction to construct a flexible speech SSL model, thereby enhancing the model's generalization ability in different downstream tasks.

Benefits of technology

It enables flexible adaptation and fine-grained control of the model in different downstream tasks, improving the performance of tasks such as speech synthesis, recognition, and sentiment analysis, while reducing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118841029B_ABST
    Figure CN118841029B_ABST
Patent Text Reader

Abstract

The application discloses a kind of decoupling voice self-supervision pre-training methods, including pre-training and fine-tuning two stages.Construct with convolution, Transformer, pitch variation processor and speaker information processor as the core self-supervision pre-training model.Input voice, and convolution module is encoded as frame-level feature;Pitch variation processor extracts pitch variation representation, and is excluded from main branch, and is replaced with masking vector and input transformer encoder.In the middle layer of encoder, speaker processor module is added to extract speaker representation, and is excluded from main branch representation.Continue to encode processing, finally map to target speech representation dimension.After first round pre-training, extract middle layer representation, train second K-Means model to generate new pseudo-label target, and carry out second round pre-training.Using weighted summation mechanism obtains task-specific representation, suitable for various downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, and provides a decoupled speech self-supervised pre-training method, a speech feature extractor is trained based on a large amount of unlabeled speech data, and the extracted features can be widely used in various speech signal processing tasks. BACKGROUND

[0002] Speech self-supervised learning (SSL) has become an increasingly important field, which aims to learn general-purpose representations of speech using large amounts of unlabeled speech data, and then provide support for various downstream tasks. Compared with traditional supervised learning methods, SSL does not rely on manually labeled labels, but generates a supervision signal using the speech data itself for pre-training. This method not only reduces the cost of data labeling, but also improves the generalization ability of the model.

[0003] In the SSL method, the generation method and the contrast method are two major ones. The generation method converts the speech signal into feature representation through the construction of an encoder, and trains the encoder by trying to reconstruct the original speech signal from these representations. This method emphasizes the ability to model the details of the speech signal, so it performs well in acoustic tasks (such as speech synthesis, speech enhancement, etc.). However, since it mainly focuses on the reconstruction of the speech signal, it may not perform well in content-related tasks (such as speech recognition, speech emotion analysis, etc.). The contrast method takes a different strategy. They also build an encoder to convert the speech signal into features, but instead of training the encoder by reconstructing the original signal, they optimize the model by comparing the similarity of representations between different inputs or different modules. This method emphasizes the ability to capture macro clustering information of the speech signal, so it performs well in content-related tasks. However, due to the lack of modeling of acoustic details, the contrast method may perform generally in acoustic tasks.

[0004] With the development of multi-modal large language models, the generalizability of speech SSL has received more attention. To evaluate the performance of pre-trained models in multiple domains, researchers have proposed benchmarks such as SUPERB and SUPERB-SG, which cover multiple downstream tasks such as content, speaker, super-language, and acoustic processing. However, although researchers have proposed various SSL strategies for specific tasks, the ability of enhanced models in one task often leads to a decline in their ability in other tasks. This raises an important question: how to achieve comprehensive improvement in task performance while maintaining the generalizability of the model? To solve this problem, researchers have begun to explore the combination of pre-trained models for multiple specific tasks or multi-task pre-training strategies. These methods aim to achieve knowledge transfer and sharing between different tasks by sharing encoders and training processes. However, these methods also face some challenges. First, due to the incompatibility between different tasks, the model may have difficulty finding a unified optimization direction in multi-task learning. Second, these methods may increase resource requirements, including computational resources and storage resources. Finally, although these methods can improve the generalizability of the model to some extent, they do not fundamentally solve the challenge of achieving generalizable speech SSL representation. SUMMARY

[0005] To further explore the generalizability of speech SSL, we need to deeply understand the decoupling of speech information and the task characteristics of different layers in SSL models. Speech information can be gradually decoupled into non-language, super-language, and language information at multiple levels. In practice, speech is often decoupled into speaker, content, and pitch variation components. These information are theoretically independent of each other and can be flexibly combined to meet the needs of different downstream tasks. In addition, the representation of different layers in SSL models also shows a trend of task modularization. The representation of shallow layers usually focuses on acoustic details, while the representation of deep layers contains more context and semantic information. Therefore, by utilizing the independence of speech information and the task characteristics of different layers in SSL models, we can design more flexible and generalizable speech SSL models to better adapt to the needs of different downstream tasks.

[0006] The present invention proposes a decoupled speech self-supervised pre-training method. This method aims to gradually extract independent representations of pitch variation, speaker, and content from speech to enhance the generalizability of the model in different downstream tasks.

[0007] The present invention proposes a decoupled speech self-supervised pre-training method, comprising the following steps:

[0008] Extracting Mel-frequency cepstral coefficients (MFCC) and their first and second order difference features;

[0009] The first K-Means model is trained using the extracted features, and the clustering centers are used as pseudo labels for each frame of speech for subsequent pre-training; meanwhile, the EMA-DINO model is pre-trained, named as the speaker teacher model, to provide guidance for speaker information;

[0010] After inputting the speech data, the frame-level features are first extracted through the convolution module, and the pitch information is extracted from the input speech data in parallel, and the pitch change representation is obtained through the pitch change processor;

[0011] Based on the idea of residual decoupling, the pitch change representation is removed;

[0012] The speech masking prediction method is used for training to ensure that the deep Transformer representation contains more content information;

[0013] In the middle layer of the Transformer, an additional speaker processor is set, and the enhanced speaker information is removed to obtain the speech extraction model;

[0014] The second K-Means model is trained based on the intermediate representation of the speech extraction model to generate new pseudo label targets for the second round of pre-training;

[0015] All intermediate representations of the model are weighted and summed through task-specific weights to obtain specific representations suitable for various downstream speech tasks.

[0016] Further, the pitch information is extracted from the input speech data, and the pitch change representation is obtained through the pitch change processor, which includes:

[0017] The pitch information is actively extracted from the speech data in the self-supervised pre-training model;

[0018] An additional pitch change processor is introduced to extract the pitch change representation from the pitch information.

[0019] Further, the fundamental frequency information F0 of the speech is extracted based on the SWIPE spectral integral method, and the normalized fundamental frequency change rate information P is obtained, as shown in

[0020]

[0021] The normalized fundamental frequency is then processed through a 1D convolution with a three-layer convolution kernel and a 1-layer gated recurrent unit network to obtain the pitch change feature. Since the fundamental frequency change rate contains the intonation information.

[0022] Further, the pitch change representation is removed from the main branch based on the idea of residual decoupling, which includes:

[0023] The output of the main branch convolution module is subtracted from the pitch variation representation in the form of a residual;

[0024] Layer normalization is performed on the representation after the pitch variation information is removed.

[0025] Further, the additional speaker information modeling is set in the middle layer of the Transformer, and the enhanced speaker information is removed from the main branch, including:

[0026] According to the representation output by the nth layer of the Transformer, an additional speaker processor supervised by the speaker teacher network is introduced to extract the speaker information representation;

[0027] The output of the Transformer is subtracted from the speaker information representation in the form of a residual;

[0028] Layer normalization is performed on the representation after the speaker information is removed.

[0029] Further, after removing the pitch variation information and the speaker information, the speech mask prediction method is used for training, including:

[0030] Before inputting the Transformer, the representation of the pitch variation information is randomly and continuously replaced by a mask vector;

[0031] After completing the speaker information removal and the Transformer processing, the similarity between the predicted representation and the representation corresponding to the pseudo label is calculated at the mask frame position, so as to complete the training of the main branch of the model.

[0032] Further, after the first round of pre-training, a second K-Means model is trained according to the intermediate layer representation of the speech extraction model to generate new pseudo label targets for the second round of pre-training, including:

[0033] The representation of the layer with the most content information of the model obtained by the first round of training is extracted for the training of the K-Means model;

[0034] The training steps of the second round of model need to be set to 2 to 4 times of the first round.

[0035] Further, the weighted sum of all intermediate representations of the model is performed by the task-specific weight, including:

[0036] The output of the convolution module and the intermediate output of the Transformer are saved as general representations of the speech, wherein the representation output by the layer selection processor of the pitch variation processor and the speaker information processor is added;

[0037] By using the idea of weighted summation, set the weight distribution of specific tasks, and weighted sum the multi-layer general representation to obtain the speech representation suitable for specific tasks;

[0038] The representation of the specific task is taken as the input of the downstream model, and the downstream model of the specific task is trained based on a small amount of labeled speech.

[0039] Advantages

[0040] From the above specific method design, the present application shows the unique advantages of the Transformer encoder style pre-training model. Through the carefully designed hierarchical task, the method can enhance the model's extraction ability of pitch variation (paralinguistic information), speaker information (non-linguistic information) and content information (linguistic information) at specific layers. At the same time, based on the idea of information decoupling, the method ensures that the model can effectively learn these three kinds of representations and ensure their independence. This independence has important value for many application scenarios, especially for controllable speech synthesis tasks. By independently controlling the pitch, speaker identity and content, more fine-grained and flexible speech generation can be achieved. In addition, after training on a large amount of unlabeled speech data, the representations extracted by different layers of the Transformer can be flexibly combined to adapt to different downstream speech tasks. By adjusting the task-specific weights, the model can be fine-tuned for specific tasks, thereby achieving better results and more controllable performance in different scenarios. This feature makes the method have wide applicability and flexibility, and can meet the needs of various speech processing tasks. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to make the description of the embodiments of the present application clearer, the legends needed when describing the examples of the present application will be introduced below.

[0042] Figure 1 is the implementation schematic diagram of the present application;

[0043] Figure 2 is the implementation schematic diagram when the present application is used for various downstream speech tasks. DETAILED DESCRIPTION

[0044] The pre-training and downstream task fine-tuning embodiments of the present application will be described in detail below in conjunction with the drawings.

[0045] Firstly, we realize that pitch variation, speaker and content information are theoretically independent of each other, so these features can be extracted separately by specific design. When constructing the self-supervised learning (SSL) model, we particularly strengthened the extraction ability of pitch variation and speaker information, and designed a strategy to gradually purify the main branch information.

[0046] Specifically, our decoupled speech self-supervised pre-training method is divided into two stages: pre-training and fine-tuning. In the pre-training stage, we adopt a two-round strategy.

[0047] In the first round, we first extract Mel Frequency Cepstral Coefficients (MFCC) and their first and second order difference features;

[0048] Then, we use these features extracted from speech data to train a K-Means model, taking the cluster centers as pseudo-labels for each frame of speech, which are used in the subsequent pre-training process;

[0049] In addition, we also pre-train an EMA-DINO model on a dataset without speaker labels, named speaker teacher model, to provide guidance for speaker information;

[0050] After obtaining the pseudo-label target and speaker teacher model, we construct a decoupled self-supervised pre-training model with convolution, Transformer, pitch variation processor, and speaker information processor as the core;

[0051] When inputting speech, the model first uses a 7-layer one-dimensional convolution processing module to encode the speech into frame-level abstract features. In this way, the one-dimensional speech signal is converted into frame-level features with a step size of 320 milliseconds;

[0052] In parallel, to extract pitch variation information, we extract the fundamental frequency from the input speech based on the spectral integration method and normalize it to obtain the pitch variation rate information;

[0053] The normalized fundamental frequency is processed by the pitch processor to obtain the pitch variation feature. Since the pitch variation rate contains the information of intonation, this layer of representation is named intonation representation;

[0054] Based on the residual idea, the intonation representation is removed from the speech representation output by the convolution, including: subtracting the pitch variation representation from the output of the main branch convolution module in the form of residual; performing a layer normalization on the representation after removing the pitch variation information;

[0055] In the time frame dimension, the speech representation after removing the intonation representation is randomly and continuously replaced by a trainable masking vector;

[0056] The masked representation is input into the multi-layer Transformer of the encoder structure;

[0057] Among them, after the middle layer Transformer, an additional speaker processor module is added;

[0058] The output of the intermediate layer Transformer will go through a fully connected layer, a frame-level attention statistics layer (FAS), and then an output fully connected layer to obtain the speaker representation;

[0059] The output of the speaker processor module will be supervised by a pre-prepared speaker teacher network. Specifically, the speaker teacher network will obtain a sentence-level speaker representation according to the input complete speech, and then, on the masked frame, by calculating the cosine similarity between the speaker representation extracted by the speaker processor, the speaker processor will be prompted to extract as much speaker information as possible from the speaker representation;

[0060] Based on the idea of residual, the speaker representation is removed;

[0061] The speech representation after removing the speaker information is continuously sent to the encoder for further processing of the features;

[0062] The processed representation is mapped to a final dimension through a linear layer;

[0063] Select the pseudo-label corresponding to each frame of signal;

[0064] According to the pseudo-label, a one-dimensional vector of the final dimension is extracted from the trainable codebook;

[0065] The model is trained by calculating the cosine similarity between the extracted vector in the codebook and the predicted vector in the mask.

[0066] Similar to the speaker loss in speaker decoupling, the content loss Only applied to the masked frame;

[0067] Since the residual decoupling method can only be effective when all modules are trained together. In addition, the speaker processor needs to be supervised by a pre-trained speaker teacher model. The proposed model uses a multi-task loss function for pre-training;

[0068] After the first round of pre-training, the intermediate layer representation of all pre-training data needs to be extracted based on the model obtained in the first round of pre-training;

[0069] Then, based on the same training strategy as before, train the second K-Means to obtain the pseudo-label target for the second round of pre-training;

[0070] Again, use the multi-task loss function in the first round for training to obtain the pre-training representation extraction model;

[0071] Our decoupled self-supervised pre-training method aims to extract multiple representations containing various speech information through a single SSL extractor. These representations can be arbitrarily combined through a weighted summation mechanism, and a small number of weight parameters can obtain task-specific representations for various downstream tasks.

[0072] Specifically, the final representation suitable for various downstream tasks is obtained by weighted summation of multiple Transformer layer features.

[0073] Embodiment: The pre-training scheme proposed by the invention is divided into two rounds.

[0074] After obtaining a large amount of unlabeled speech data, the sampling rate of all speech needs to be converted to 16000.

[0075] The first round needs to extract 13-dimensional MFCC (Mel Frequency Cepstral Coefficients) and its first and second order difference features with a step size of 320 milliseconds.

[0076] Then, 10% of the speech data is extracted from the data, and a 100-class K-Means model is trained according to its corresponding 39-dimensional features. Specifically, MiniBatchKMeans implemented in scikit-learn is used, and the strategy of setting the mini-batch size to 10,000 frames is adopted, and K-Means++ strategy is used for better initialization through 20 random starting points.

[0077] The corresponding cluster centers are used as pseudo-labels for each frame of speech for pre-training.

[0078] An EMA-DINO model is pre-trained on a dataset without speaker labels using the open-source toolkit Wespeaker, named the speaker teacher model. In the case of a thousand-hour pre-training dataset, the parameter configuration of the EMA-DINO teacher network is set to the "ECAPA_TDNN_GLOB_c512" model in Wespeaker. If in the case of a large-scale pre-training task of tens of thousands of hours, the parameter configuration of the EMA-DINO teacher network is set to the "ECAPA_TDNN_GLOB_c1024" model in Wespeaker.

[0079] After training the pseudo-label target and the speaker teacher network, a decoupled self-supervised pre-training model composed of convolution, Transformer, pitch change processor and speaker information processor needs to be built, and the overall pre-training process is as shown in Figure 1 .

[0080] After inputting the speech, first use a 7-layer 1-dimensional convolution processing module to encode the speech into frame-level abstract features, the number of channels of each convolution processing module is 512, the size of the 7-layer convolution kernel is {10, 3, 3, 3, 3, 2, 2} in turn, and the convolution step is {5, 2, 2, 2, 2, 2, 2} in turn, in this way, the one-dimensional speech signal will be converted into frame-level features with a step of 320 milliseconds;

[0081] In parallel, after inputting the speech, the method will also extract the fundamental frequency information F0 of the speech based on the SWIPE spectral integration method, and normalize it to obtain the pitch rate information P, such as

[0082]

[0083] Then the normalized fundamental frequency will pass through a 1-dimensional convolution with a three-layer convolution kernel of 3 and a convolution step of 1 and a 1-layer gated recurrent unit network to obtain the pitch change feature, since the pitch rate contains the information of the intonation;

[0084] Based on the idea of residual, the intonation representation is removed from the speech representation output by convolution;

[0085] In the time frame dimension, according to the possibility of 80% and the masking length of 10 frames, the speech representation after removing the intonation representation is randomly and continuously replaced by a trainable masking vector;

[0086] The masked representation is input into the multi-layer Transformer of the encoder structure, and an additional speaker processor module is added behind the nth Transformer, where the number of layers needs to be set in combination with the total number of layers of the specific model, and the layer containing the most speaker information is selected. In a common configuration, the 4th layer is selected for a 12-layer structure, and the 6th layer is selected for a 24-layer structure;

[0087] The output of the nth Transformer will pass through a fully connected layer, a frame-level attention statistics layer (FAS), and an output fully connected layer to obtain a speaker representation, where the frame-level attention statistics layer calculates the mean and variance of each frame;

[0088] The output of the speaker processor module will be supervised by a pre-prepared speaker teacher network. Specifically, the speaker teacher network will obtain a sentence-level speaker representation from the input complete speech, and then calculate the cosine similarity between the speaker representation extracted by the speaker processor and the speaker representation on the masked frame,

[0089]

[0090] where, is a projection matrix, is the output of the speaker processor, sim(·, ·) denotes the cosine similarity, and σ(·) denotes the sigmoid function. The speaker loss only on the masked frame T m is computed. is used to encourage the speaker processor to extract as much speaker information as possible from the representation;

[0091] The speaker representation is removed from the main branch representation based on the residual idea;

[0092] The speech representation after removing the speaker information is further processed by the encoder;

[0093] The processed representation is mapped to a final dimension through a linear layer;

[0094] The two stages select their corresponding pseudo labels, respectively;

[0095] According to the pseudo label, a one-dimensional vector of the final dimension is extracted from the trainable codebook;

[0096] The model is trained by calculating the cosine similarity between the extracted vector in the codebook and the vector predicted by the mask. Specifically, the pseudo label target is formulated as Z = [z1, z2, …, z T ], where each z ∈ [U] is a U-class variable. After convolution, two feature processors and Transformer processing, the main branch is trained to predict the pseudo label of the masked frame, as follows:

[0097]

[0098] where A c is a projection matrix, is the output of the last Transformer, e u is the embedding vector of the K-Means unit u, and τ is the scaling factor of the logarithm, which is set to 0.1. Similar to the speaker loss in speaker decoupling, the content loss is only applied to the masked frame.

[0099] Since the residual decoupling method can only be effective when all modules are jointly trained. In addition, the speaker processor needs to be jointly supervised by a pre-trained speaker teacher model. The proposed model is pre-trained using the following multi-task loss function:

[0100]

[0101] where, denotes the mean square error of the output representation of the convolution module. and respectively. The hyper-parameters λ f , λ s and λ c are set to 10.0, 1.0 and 1.0 respectively.

[0102] The training of the model uses Adam optimizer with a warm-up learning rate strategy, the learning rate is gradually increased from 0 to 5e -4 in the first 8% steps and then decayed to 0. For data of the order of thousands of hours, the number of training data rounds in the first round is set to 50 to 100 rounds, and for data of the order of tens of thousands of hours, the number of training data rounds is set to 10 to 20.

[0103] After the first round of pre-training is completed, the intermediate layer representation of all pre-training data needs to be extracted based on the model obtained in the first round of pre-training. The number of layers is generally dynamically selected according to the parameter configuration of the model. For common parameter configurations, the 9th layer is selected for a 12-layer Transformer model, and the 18th layer is selected for a 24-layer Transformer. Then, a 500-class K-Means is trained based on the same training strategy as before to obtain the pseudo-label target for the second round of pre-training.

[0104] Then, the multi-task loss function in the first round is used for training again. The number of rounds for the second round of training should be kept twice that of the first round. Finally, the pre-training representation extraction model is obtained.

[0105] The fine-tuning stage is shown in Figure 2 . Given a speech signal, the decoupled speech self-supervised pre-training model uses a multi-layer Transformer encoder structure to extract hierarchical representations O = {O0, O1,..., O n}. Then, the task-specific representation is recombined according to the following formula:

[0106]

[0107] where ω i is the task-specific hierarchical weight. These weights can be learned through task-specific fine-tuning.

[0108] After obtaining the task-specific representation, it can be input into various downstream models and fine-tuned on a small amount of labeled data to achieve excellent performance in corresponding downstream tasks.

[0109] Comparison table of five task performances

[0110]

[0111] The application verifies five downstream tasks, and compared with a contrast model, the speech recognition word error rate is greatly reduced, especially the speaker recognition accuracy and the emotion classification accuracy are very good, and the speech enhancement and timbre conversion are also better than other models.

Claims

1. A decoupled speech self-supervised pre-training method, characterized in that, Includes the following steps: Extract the Mel frequency cepstral coefficients (MFCC) and their first and second order difference features; The extracted features are used to train the first K-Means model, and the cluster centers are used as pseudo-labels for each frame of speech for subsequent pre-training. At the same time, the EMA-DINO model is pre-trained and named the speaker teacher model to provide guidance on speaker information. After inputting speech data, the convolution module first extracts frame-level features, and in parallel, the pitch information is extracted from the input speech data and processed by the pitch change processor to obtain a pitch change representation. Based on the idea of ​​residual decoupling, the pitch change representation is eliminated; Training is performed using speech masking prediction to ensure that the deep Transformer representation contains more content information; An additional speaker processor is set up behind the middle layer of the multi-layer Transformer, and the enhanced speaker information is removed to obtain the speech extraction model; The second K-Means model is trained based on the intermediate layer representation of the speech extraction model to generate new pseudo-label targets, and a second round of pre-training is performed. By weighting all intermediate representations of the model with task-specific weights, a specific representation applicable to various downstream speech tasks can be obtained. The residual decoupling-based approach removes pitch variation representations from the main branch, including: The pitch change is represented by the output of the main branch convolution module minus the pitch change, in the form of residuals. Perform a first-level normalization on the representation after removing pitch variation information; The step of setting up an additional speaker processor after the intermediate layer of the multi-layer Transformer and removing the enhanced speaker information includes: Based on the representation output of the nth layer of the Transformer, an additional speaker processor supervised by the speaker teacher model is introduced to extract speaker information representation. The speaker information representation is subtracted from the Transformer's output in the form of residuals; Perform a first-level normalization on the representation after removing speaker information; After removing pitch variation information and speaker information, training is performed using speech masking prediction, including: Before being input into the Transformer, the representation that removes pitch change information is randomly and continuously replaced with a masking vector; After speaker information removal and Transformer processing, the similarity between the predicted representation and the representation corresponding to the pseudo-label is calculated at the masked frame position to complete the training of the main branch of the model.

2. The decoupled speech self-supervised pre-training method as described in claim 1, characterized in that, The process of extracting pitch information from input speech data and processing it through a pitch change processor to obtain a pitch change representation includes: Actively extract pitch information from speech data in a self-supervised pre-trained model; An additional pitch change processor is introduced to extract pitch change representations from pitch information.

3. The decoupled speech self-supervised pre-training method as described in claim 2, characterized in that, Fundamental frequency information of speech is extracted based on the SWIPE spectral integration method. And normalize it to obtain the fundamental frequency change rate information. ,like Then, the normalized fundamental frequency is passed through three layers of 1D convolution with 3 kernels and 1 stride, and a 1-layer gated recurrent unit network to obtain pitch variation features, since the rate of change of the fundamental frequency contains intonation information.

4. The decoupled speech self-supervised pre-training method as described in claim 1, characterized in that, After the first round of pre-training, a second K-Means model is trained based on the intermediate layer representations of the speech extraction model to generate new pseudo-label targets for the second round of pre-training, including: Using the model obtained from the first round of training, extract the representation of the layer with the most content information for training the K-Means model; The number of training steps for the second round of the model needs to be set to 2 to 4 times that of the first round.

5. The decoupled speech self-supervised pre-training method as described in claim 1, characterized in that, The weighted summation of all intermediate representations of the model using task-specific weights includes: The output of the convolutional module and the intermediate output of the Transformer are saved as a general representation of speech, which incorporates the representation of the layer selection processor output of the pitch variation processor and the speaker information processor. By using the idea of ​​weighted summation, a weight distribution for a specific task is set, and a weighted summation is performed on the multi-layered general representation to obtain a speech representation applicable to a specific task. The downstream model is obtained by using the representation of a specific task as input to the downstream model and training it on a small amount of labeled speech.