Cross-domain voice classification method and device based on feature decoupling and multi-task learning

By building a cross-domain speech classification model, using feature decoupling and multi-task learning methods, the problem of insufficient cross-domain classification performance under multi-source heterogeneous speech data is solved, and high-precision speech classification between different data domains is realized, which is suitable for fields such as multi-source speech recognition, speech health analysis and human-computer interaction.

CN120452429APending Publication Date: 2025-08-08SICHUAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510830708.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the multi-source heterogeneous speech data environment, the cross-domain speech classification performance and generalization capabilities are insufficient, making it difficult to maintain high classification accuracy among different data domains.

Method used

A cross-domain speech classification method based on feature decoupling and multi-task learning is adopted to build a cross-domain speech classification model, including a speech feature encoder module, a data domain classification module, a supervised comparison learning module and a multi-task classification module. The model is trained through a gradient descent algorithm, and a gradient inversion layer and a supervised comparison learning module are introduced to explicitly separate task features and domain features, enhancing discriminant ability and robustness.

Benefits of technology

It significantly improves the model's discrimination and generalization capabilities in cross-domain scenarios, maintains high classification accuracy in complex multi-source speech environments, and is suitable for fields such as multi-source speech recognition, speech health analysis and human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452429A_ABST
    Figure CN120452429A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-domain voice classification method and device based on feature decoupling and multi-task learning. The method comprises the following steps: firstly, acquiring a multi-data-domain voice file, and preprocessing the multi-data-domain voice file to obtain a cross-domain voice classification data set; then, constructing a cross-domain voice classification model which comprises a voice feature encoder module, a data domain classification module, a supervised comparative learning module and a multi-task classification module; then, a joint optimization loss function is constructed based on data field classification loss, supervised contrast learning loss and task classification loss, and a gradient descent algorithm is adopted to train and optimize the cross-domain voice classification model based on the cross-domain voice classification data set and the joint optimization loss function; and finally, inputting to-be-classified voice into the trained cross-domain voice classification model to obtain a voice corresponding category. The discrimination ability and generalization ability of the model in a cross-domain scene are significantly improved, so that the model still maintains high classification precision in a complex multi-source voice environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech classification technology, and in particular to a cross-domain speech classification method and device based on feature decoupling and multi-task learning. Background Art

[0002] With the development of artificial intelligence and speech recognition technology, speech classification, as a core task in intelligent speech processing, has been widely used in fields such as voice command recognition, human-computer interaction, and assisted diagnosis. Traditional speech classification methods rely primarily on manually extracted acoustic features combined with shallow classification models. However, their classification performance and generalization capabilities are severely limited when faced with large-scale, multi-source, and heterogeneous speech data.

[0003] In recent years, the development of deep learning technology has greatly enhanced the capabilities of speech feature representation and classification modeling. Through end-to-end modeling, deep neural networks can automatically extract high-level semantic features from raw speech, achieving remarkable classification results. However, in real-world applications, speech data often comes from different recording devices, speaker environments, speaking styles, or acquisition platforms, resulting in significant distribution shifts between different data domains. This results in a significant performance degradation of models trained on a single data source when applied to new data domains, limiting their application and scalability in multiple scenarios.

[0004] Therefore, in the related technology, there is an urgent need for a method that can improve the cross-domain speech classification performance and generalization ability in a multi-source data environment. Summary of the Invention

[0005] Based on this, it is necessary to provide a cross-domain speech classification method and device based on feature decoupling and multi-task learning, which can improve the cross-domain speech classification performance and generalization ability in a multi-source data environment, in order to address the above technical problems.

[0006] In a first aspect, the present application provides a cross-domain speech classification method based on feature decoupling and multi-task learning. The method comprises: Acquire multi-domain speech files and perform preprocessing to obtain a cross-domain speech classification dataset; Build a cross-domain speech classification model, including a speech feature encoder module, a data domain classification module, a supervised contrastive learning module, and a multi-task classification module; Constructing a joint optimization loss function based on data domain classification loss, supervised contrastive learning loss, and task classification loss, and training and optimizing the cross-domain speech classification model using a gradient descent algorithm based on the cross-domain speech classification dataset and the joint optimization loss function; The speech to be classified is input into the trained cross-domain speech classification model to obtain the corresponding category of the speech.

[0007] Optionally, in one embodiment of the present application, obtaining multi-domain voice files and preprocessing them to obtain a cross-domain voice classification dataset includes: Constructing a cross-data domain data loader, and obtaining a training data set by random sampling based on the data loader, wherein the training data set has a classification label and a data domain label; A set of positive and negative sample pairs is constructed based on the training data set.

[0008] Optionally, in one embodiment of the present application, constructing a cross-domain speech classification model includes: A speech feature encoder module is constructed based on the self-supervised learning speech representation module and the attention statistical pooling module.

[0009] Optionally, in one embodiment of the present application, constructing a cross-domain speech classification model further includes: A data domain classification module is constructed based on the domain feature extraction module and the gradient reversal layer.

[0010] Optionally, in one embodiment of the present application, the supervised contrastive learning module maps the high-dimensional speech embedding representation to the contrastive learning space for supervised comparison, and the multi-task classification module extracts task feature representations directly related to the speech classification task based on the high-dimensional speech embedding.

[0011] Optionally, in one embodiment of the present application, the data domain classification loss and the task classification loss are calculated based on the category and domain category using a normalized exponential function, and the supervised contrastive learning loss is calculated based on the similarity between the anchor sample and the positive and negative samples.

[0012] Optionally, in one embodiment of the present application, constructing a joint optimization loss function based on data domain classification loss, supervised contrastive learning loss, and task classification loss includes: The GradNorm dynamic weight balancing strategy is adopted to determine the loss weight coefficients of the data domain classification loss, supervised contrastive learning loss and task classification loss.

[0013] In a second aspect, the present application also provides a cross-domain speech classification device based on feature decoupling and multi-task learning. The device comprises: The data acquisition and preprocessing module is used to obtain multi-domain speech files and perform preprocessing to obtain a cross-domain speech classification dataset; Cross-domain speech classification model construction module, used to build a cross-domain speech classification model, including a speech feature encoder module, a data domain classification module, a supervised contrastive learning module, and a multi-task classification module; A cross-domain speech classification model training and optimization module is used to construct a joint optimization loss function based on the data domain classification loss, the supervised contrastive learning loss, and the task classification loss, and to train and optimize the cross-domain speech classification model using a gradient descent algorithm based on the cross-domain speech classification dataset and the joint optimization loss function; The cross-domain speech classification module is used to input the speech to be classified into the trained cross-domain speech classification model to obtain the corresponding category of the speech.

[0014] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program and the processor executes the steps of the method described in each of the above embodiments.

[0015] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in each of the above embodiments.

[0016] The cross-domain speech classification method and device based on feature decoupling and multi-task learning first obtains and preprocesses multiple data-domain speech files to obtain a cross-domain speech classification dataset. Next, a cross-domain speech classification model is constructed, comprising a speech feature encoder module, a data-domain classification module, a supervised contrastive learning module, and a multi-task classification module. A joint optimization loss function is then constructed based on the data-domain classification loss, the supervised contrastive learning loss, and the task classification loss. The cross-domain speech classification model is trained and optimized using a gradient descent algorithm based on the cross-domain speech classification dataset and the joint optimization loss function. Finally, the speech to be classified is input into the trained cross-domain speech classification model to obtain the corresponding speech category. In other words, by constructing a task feature branch, a domain adversarial branch, and a contrastive learning branch, multi-angle modeling of speech features is achieved. Task features and domain features are explicitly separated in the network structure, and a gradient reversal layer is introduced to adversarially suppress domain information, significantly improving the model's discriminative and generalization capabilities in cross-domain scenarios. A supervised contrastive learning module is introduced to map task features into a contrastive learning space. Combined with intra-class clustering and inter-class separation mechanisms, this enhances the discriminability and robustness of speech features, enabling the model to maintain high classification accuracy even in complex multi-source speech environments. The proposed cross-domain speech classification model is widely applicable to multiple fields, including multi-source speech recognition, speech health analysis, and human-computer interaction. It exhibits excellent practicality and scalability, making it suitable for high-precision classification tasks involving diverse speech data in complex real-world environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a diagram illustrating an application environment of a cross-domain speech classification method based on feature decoupling and multi-task learning in one embodiment; Figure 21 is a flow chart of a cross-domain speech classification method based on feature decoupling and multi-task learning in one embodiment; Figure 3 Schematic diagram of the structure of a speech feature encoder module in one embodiment; Figure 4 Schematic diagram of the structures of a data domain classification module, a supervised contrastive learning module, and a multi-task classification module in one embodiment; Figure 5 Schematic diagram of a joint optimization loss function in one embodiment; Figure 6 1 is a structural block diagram of a cross-domain speech classification device based on feature decoupling and multi-task learning in one embodiment; Figure 7 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0019] The cross-domain speech classification method based on feature decoupling and multi-task learning provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal communicates with the server through the network. The data storage system can store data that the server needs to process. The data storage system can be integrated on the server or placed on the cloud or other network servers. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented as a standalone server or a server cluster consisting of multiple servers.

[0020] In one embodiment, Figure 2 As shown in the figure, a cross-domain speech classification method based on feature decoupling and multi-task learning is provided. Figure 1 The following steps are used as an example to illustrate the server in the example: S201: Acquire multi-domain speech files and perform preprocessing to obtain a cross-domain speech classification dataset.

[0021] In an embodiment of the present application, first, voice files from multiple data domains are collected, and the collected voice waveform data are preprocessed to construct a cross-domain voice classification dataset.

[0022] Specifically, in one embodiment of the present application, obtaining multi-domain voice files and preprocessing them to obtain a cross-domain voice classification dataset includes: S301: Construct a cross-data domain data loader, and obtain a training data set by random sampling based on the data loader. The training data set has a classification label and a data domain label.

[0023] S303: Constructing a set of positive and negative sample pairs based on the training data set.

[0024] In one embodiment of the present application, a cross-data domain data loader is constructed, and different data domains are represented as :

[0025] in, represents the i-th data field, Represents the data domain The mth sample in each round is obtained from n data domains. The same number of samples p are randomly sampled independently to obtain a set of n samples :

[0026] Then the sampling data of each data domain Merge to get a round of training data sets :

[0027] To achieve joint sampling, ensure that each training batch contains samples from multiple data domains and the samples have classification labels With data field label , used for subsequent module training.

[0028] Based on the current round The loaded samples are used to construct positive and negative sample pairs according to the requirements of supervised contrastive learning. Specifically, the current batch The samples in the same semantic category are combined into a positive sample pair set :

[0029] Combine pairs of samples from different categories or different data domains to form a negative sample pair set :

[0030] in, Representation sample The semantic category of Indicates the data domain to which it belongs.

[0031] S203: Construct a cross-domain speech classification model, including a speech feature encoder module, a data domain classification module, a supervised contrastive learning module, and a multi-task classification module.

[0032] In an embodiment of the present application, after a cross-domain speech classification dataset is constructed, a cross-domain speech classification model (CSCM) is constructed. The model includes a speech feature encoder module (SFEM), a data domain classification module (DDCM), a supervised contrastive learning module (SCLM), and a multi-task classification module (MCM). The speech feature encoder module SFEM obtains high-dimensional speech embedding representations in speech feature vectors based on cross-domain speech data input. The data domain classification module DDCM predicts the data domain to which the data belongs based on domain features and performs adversarial training in combination with a gradient reversal layer to reduce the feature distribution differences between different data domains. The supervised contrastive learning module SCLM optimizes the contrast loss of sample pairs based on supervised contrastive learning, increases the inter-class distance of samples, and reduces the intra-class distance, thereby improving the discriminability of task features. The multi-task classification module MCM extracts task features based on multiple task branches and optimizes them using cross-entropy loss. High-precision cross-domain classification is achieved based on data domain classification and supervised contrastive learning.

[0033] Specifically, in one embodiment of the present application, constructing a cross-domain speech classification model includes: A speech feature encoder module is constructed based on the self-supervised learning speech representation module and the attention statistical pooling module.

[0034] In one embodiment of the present application, Figure 3 As shown in Figure 2, the speech feature encoder module SFEM includes a self-supervised learning speech representation Wav2vec2.0 module and an attention statistics pooling module (ASPM). Specifically, the speech waveform data is obtained by preprocessing. :

[0035] The speech feature vector is extracted by the pre-trained model Wav2vec 2.0 :

[0036] On this basis, in order to meet the expectation of the subsequent task classifier, the input is a fixed-length vector representation (i.e., without the time dimension). ), it is necessary to obtain the original speech feature vector Compression is achieved by performing an attention statistical pooling operation in its temporal dimension. Specifically, a weight set is obtained by assigning weights to each time frame through an attention network consisting of a convolutional layer, a BatchNorm layer, and a nonlinear activation function. :

[0037] According to the weights assigned to Features at each time step According to its corresponding attention weight Perform weighted average and weighted standard deviation calculations to obtain the weighted average eigenvector and weighted standard deviation eigenvector :

[0038]

[0039] The obtained weighted average eigenvector and weighted standard deviation eigenvector In the feature dimension The previous concatenation operation (concat) is performed to obtain a new shared feature vector :

[0040] In one embodiment of the present application, constructing a cross-domain speech classification model further includes: A data domain classification module is constructed based on the domain feature extraction module and the gradient reversal layer.

[0041] In one embodiment of the present application, Figure 4 As shown in the figure, the data domain classification module DDCM mainly includes the domain feature extraction module (DFEM) and the gradient reversal layer. The domain features related to the data domain are obtained by the domain feature extraction module, and the gradient reversal layer is used to perform gradient reversal to suppress domain information. The specific calculation process is as follows:

[0042]

[0043]

[0044] in, represents the domain feature extraction network, 、 Extract network weights and biases for domain features, For the batch normalization (BatchNorm) operation, is the nonlinear activation function ReLU; 、 Represents domain classification weights and biases; represents the gradient reversal layer, represents the predicted probability of the domain label; is the data domain classification loss, is the gradient reversal layer strength coefficient; the gradient reversal layer GRL enables the model to actively ignore domain features and improve cross-domain generalization capabilities.

[0045] In one embodiment of the present application, the supervised contrastive learning module maps the high-dimensional speech embedding representation to the contrastive learning space for supervised comparison, and the multi-task classification module extracts task feature representations directly related to the speech classification task based on the high-dimensional speech embedding.

[0046] In one embodiment of the present application, Figure 4 As shown in Figure 2, the supervised contrastive learning module SCLM further maps the high-dimensional speech embedding representation to the contrastive learning space to enhance the category discriminability of the task features. The calculation process is as follows:

[0047] in, represents the contrast feature mapping function, represents the normalization operation, is the nonlinear activation function ReLU, 、 、 、 represents the domain classification weight, 、 represents the domain classification bias, Represents the normalized feature embedding vector of the i-th sample.

[0048] The similarity of each embedding vector in the high-dimensional feature space is measured by cosine similarity. The similarity matrix is defined as:

[0049] in, 、 denote the normalized feature embeddings of the i-th and j-th samples, respectively, The temperature hyperparameter is used to adjust the smoothness of the similarity distribution, and its value is usually between 0.05 and 0.5. The addition of the temperature parameter τ can effectively control the sensitivity of the supervised contrast loss to difficult and easy-to-distinguish samples.

[0050] The multi-task classification module (MCM) extracts task feature representations directly related to the speech classification task based on high-dimensional speech embeddings to improve the refined representation of the speech classification task. The computational process for feature extraction and classification prediction is as follows:

[0051]

[0052] Where, represents the task feature extraction network, 、 Extract network weights and biases for task features, For the batch normalization (BatchNorm) operation, is the nonlinear activation function ReLU; 、 represents the task classification weight and bias, Represents the predicted probability of the task category.

[0053] It should be noted that if Figure 4 As shown in the figure, the multi-task classification module MCM and the data domain classification module DDCM are composed of a fully connected layer, a BatchNorm layer and a nonlinear activation function respectively; the output task features in the multi-task classification module MCM Used for multi-task classifier; data domain classification module DDCM output domain features For data domain classifier; where, and The feature dimensions represent task characteristics and domain characteristics respectively.

[0054] S205: Constructing a joint optimization loss function based on the data domain classification loss, the supervised contrastive learning loss, and the task classification loss, and training and optimizing the cross-domain speech classification model using a gradient descent algorithm based on the cross-domain speech classification dataset and the joint optimization loss function.

[0055] In the embodiment of this application, Figure 5As shown in the figure, a joint optimization loss function (JOLF) is constructed based on the data domain classification loss, supervised contrastive learning loss, and task classification loss. The joint optimization loss function is used as the final optimization goal. The constructed cross-domain speech classification dataset is input into the initial cross-domain speech classification model. The model is trained using the gradient descent algorithm through multi-branch joint training. After the training is completed, the model with the best training effect is saved.

[0056] In one embodiment of the present application, the data domain classification loss and the task classification loss are calculated based on the category and domain category using a normalized exponential function, and the supervised contrastive learning loss is calculated based on the similarity between the anchor sample and the positive and negative samples.

[0057] In one embodiment of the present application, the task features and domain features are decoupled and represented in the corresponding subspace by minimizing the cross entropy loss function. and domain classification loss They are defined as follows:

[0058]

[0059]

[0060]

[0061] in, is the total number of categories, is the true label, represents the category probability predicted by the model, and Represent weight and bias respectively; is the total number of domain categories, is the domain label, represents the probability of the data domain category predicted by the model, and denote weights and biases respectively.

[0062] The supervised contrast loss is constructed based on the anchor samples and their corresponding positive and negative samples, and features are learned in batches. , first define its corresponding positive sample set and a set containing positive and negative samples The supervised contrast loss is defined as follows:

[0063] in, Represents the index set of all anchor samples in the current batch, Representation sample The set of all positive sample indexes (with the same label); Representation sample The index set of all positive and negative samples except itself; Anchor point samples and its positive sample similarity between Anchor point samples With any sample The similarity between them.

[0064] In one embodiment of the present application, constructing a joint optimization loss function based on data domain classification loss, supervised contrastive learning loss, and task classification loss includes: The GradNorm dynamic weight balancing strategy is adopted to determine the loss weight coefficients of the data domain classification loss, supervised contrastive learning loss and task classification loss.

[0065] In one embodiment of the present application, the joint optimization loss function JOLF adopts the GradNorm dynamic weight balancing strategy. GradNorm dynamically adjusts the loss weight coefficient of each task. , so that the training rate of each task is close to its average level, specifically: For the i-th task, the ratio of its current loss to the initial loss is defined as:

[0066] in, For the The current loss value in the training round, is the initial loss value of training.

[0067] For the i-th task, the current loss is the shared parameter The gradient norm of is:

[0068] The relationship between the target gradient norm and the normalized loss rate is:

[0069] in, is the mean of the gradient norms of all tasks, It is a hyperparameter that adjusts the balance of task training rates.

[0070] GradNorm loss measures the difference between the current gradient norm and the target gradient norm and is defined as:

[0071] By minimizing , realizing task loss weight Dynamic updates.

[0072] To ensure Stability, using constraints ( is a constant, usually set to the number of tasks) to normalize to prevent the task weight from being too large or too small. The final joint loss function expression is:

[0073] in, 、 and They represent the loss weight of each task respectively. Each weight is updated in real time by the GradNorm algorithm to ensure the dynamic balance of training among multiple tasks.

[0074] S207: Input the speech to be classified into the trained cross-domain speech classification model to obtain the corresponding category of the speech.

[0075] In the embodiment of the present application, the speech to be classified is input into the trained cross-domain speech classification model CSCM to obtain the corresponding type of the speech.

[0076] In the above-mentioned cross-domain speech classification method based on feature decoupling and multi-task learning, multiple data-domain speech files are first acquired and preprocessed to obtain a cross-domain speech classification dataset. Next, a cross-domain speech classification model is constructed, comprising a speech feature encoder module, a data-domain classification module, a supervised contrastive learning module, and a multi-task classification module. A joint optimization loss function is then constructed based on the data-domain classification loss, the supervised contrastive learning loss, and the task classification loss. The cross-domain speech classification model is trained and optimized using a gradient descent algorithm based on the cross-domain speech classification dataset and the joint optimization loss function. Finally, the speech to be classified is input into the trained cross-domain speech classification model to obtain the corresponding speech category. In other words, by constructing a task feature branch, a domain adversarial branch, and a contrastive learning branch, multi-angle modeling of speech features is achieved. Task features and domain features are explicitly separated in the network structure, and a gradient reversal layer is introduced to adversarially suppress domain information, significantly improving the model's discriminative and generalization capabilities in cross-domain scenarios. A supervised contrastive learning module is introduced to map task features into a contrastive learning space. Combined with intra-class clustering and inter-class separation mechanisms, this enhances the discriminability and robustness of speech features, enabling the model to maintain high classification accuracy even in complex multi-source speech environments. The proposed cross-domain speech classification model is widely applicable to multiple fields, including multi-source speech recognition, speech health analysis, and human-computer interaction. It exhibits excellent practicality and scalability, making it suitable for high-precision classification tasks involving diverse speech data in complex real-world environments.

[0077] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0078] Based on the same inventive concept, an embodiment of the present application also provides a cross-domain speech classification device based on feature decoupling and multi-task learning for implementing the above-mentioned cross-domain speech classification method based on feature decoupling and multi-task learning. The implementation solution provided by the device is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations in the embodiments of one or more cross-domain speech classification devices based on feature decoupling and multi-task learning provided below can be found in the above-mentioned limitations on the cross-domain speech classification method based on feature decoupling and multi-task learning, and will not be repeated here.

[0079] In one embodiment, Figure 6 As shown, a cross-domain speech classification device 600 based on feature decoupling and multi-task learning is provided, including: a data acquisition and preprocessing module 601, a cross-domain speech classification model construction module 603, a cross-domain speech classification model training optimization module 605 and a cross-domain speech classification module 607, wherein: The data acquisition and preprocessing module 601 is used to obtain multi-domain speech files and perform preprocessing to obtain a cross-domain speech classification dataset.

[0080] The cross-domain speech classification model construction module 603 is used to construct a cross-domain speech classification model, including a speech feature encoder module, a data domain classification module, a supervised contrastive learning module and a multi-task classification module.

[0081] The cross-domain speech classification model training optimization module 605 is used to construct a joint optimization loss function based on the data domain classification loss, the supervised contrastive learning loss and the task classification loss, and to train and optimize the cross-domain speech classification model using a gradient descent algorithm based on the cross-domain speech classification dataset and the joint optimization loss function.

[0082] The cross-domain speech classification module 607 is used to input the speech to be classified into the trained cross-domain speech classification model to obtain the corresponding category of the speech.

[0083] In one embodiment of the present application, the data acquisition and preprocessing module is further configured to: Constructing a cross-data domain data loader, and obtaining a training data set by random sampling based on the data loader, wherein the training data set has a classification label and a data domain label; A set of positive and negative sample pairs is constructed based on the training data set.

[0084] In one embodiment of the present application, the cross-domain speech classification model construction module is further used to: A speech feature encoder module is constructed based on the self-supervised learning speech representation module and the attention statistical pooling module.

[0085] In one embodiment of the present application, the cross-domain speech classification model construction module is further used to: A data domain classification module is constructed based on the domain feature extraction module and the gradient reversal layer.

[0086] In one embodiment of the present application, the supervised contrastive learning module maps the high-dimensional speech embedding representation to the contrastive learning space for supervised comparison, and the multi-task classification module extracts task feature representations directly related to the speech classification task based on the high-dimensional speech embedding.

[0087] In one embodiment of the present application, the data domain classification loss and the task classification loss are calculated based on the category and domain category using a normalized exponential function, and the supervised contrastive learning loss is calculated based on the similarity between the anchor sample and the positive and negative samples.

[0088] In one embodiment of the present application, the cross-domain speech classification model training optimization module is further used to: The GradNorm dynamic weight balancing strategy is adopted to determine the loss weight coefficients of the data domain classification loss, supervised contrastive learning loss and task classification loss.

[0089] Each module in the above-mentioned cross-domain speech classification device based on feature decoupling and multi-task learning can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0090] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, memory, a communication interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal via wired or wireless communication. The wireless communication can be achieved via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a cross-domain speech classification method based on feature decoupling and multi-task learning. The display screen of the computer device can be a liquid crystal display or an electronic ink display. The input device of the computer device can be a touch layer covering the display screen, keys, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse.

[0091] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0092] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0093] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0094] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0095] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0096] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0097] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0098] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A cross-domain speech classification method based on feature decoupling and multi-task learning, characterized in that: The method comprises: Acquire multi-domain speech files and perform preprocessing to obtain a cross-domain speech classification dataset; Build a cross-domain speech classification model, including a speech feature encoder module, a data domain classification module, a supervised contrastive learning module, and a multi-task classification module; Constructing a joint optimization loss function based on data domain classification loss, supervised contrastive learning loss, and task classification loss, and training and optimizing the cross-domain speech classification model using a gradient descent algorithm based on the cross-domain speech classification dataset and the joint optimization loss function; The speech to be classified is input into the trained cross-domain speech classification model to obtain the corresponding category of the speech.

2. The cross-domain speech classification method based on feature decoupling and multi-task learning according to claim 1 is characterized in that The method of obtaining multi-domain voice files and preprocessing to obtain a cross-domain voice classification dataset includes: Constructing a cross-data domain data loader, and obtaining a training data set by random sampling based on the data loader, wherein the training data set has a classification label and a data domain label; A set of positive and negative sample pairs is constructed based on the training data set.

3. The cross-domain speech classification method based on feature decoupling and multi-task learning according to claim 1 is characterized in that The constructing of the cross-domain speech classification model includes: A speech feature encoder module is constructed based on the self-supervised learning speech representation module and the attention statistical pooling module.

4. The cross-domain speech classification method based on feature decoupling and multi-task learning according to claim 1 is characterized in that The constructing of the cross-domain speech classification model further includes: A data domain classification module is constructed based on the domain feature extraction module and the gradient reversal layer.

5. The cross-domain speech classification method based on feature decoupling and multi-task learning according to claim 1 is characterized in that The supervised contrastive learning module maps the high-dimensional speech embedding representation to the contrastive learning space for supervised comparison, and the multi-task classification module extracts task feature representations directly related to the speech classification task based on the high-dimensional speech embedding.

6. The cross-domain speech classification method based on feature decoupling and multi-task learning according to claim 1 is characterized in that The data domain classification loss and task classification loss are calculated based on the category and domain category using a normalized exponential function, and the supervised contrastive learning loss is calculated based on the similarity between the anchor sample and the positive and negative samples.

7. The cross-domain speech classification method based on feature decoupling and multi-task learning according to claim 1 is characterized in that The joint optimization loss function constructed based on data domain classification loss, supervised contrastive learning loss and task classification loss includes: The GradNorm dynamic weight balancing strategy is adopted to determine the loss weight coefficients of the data domain classification loss, supervised contrastive learning loss and task classification loss.

8. A cross-domain speech classification device based on feature decoupling and multi-task learning, characterized in that: The device comprises: The data acquisition and preprocessing module is used to obtain multi-domain speech files and perform preprocessing to obtain a cross-domain speech classification dataset; Cross-domain speech classification model construction module, used to build a cross-domain speech classification model, including a speech feature encoder module, a data domain classification module, a supervised contrastive learning module, and a multi-task classification module; A cross-domain speech classification model training and optimization module is used to construct a joint optimization loss function based on the data domain classification loss, the supervised contrastive learning loss, and the task classification loss, and to train and optimize the cross-domain speech classification model using a gradient descent algorithm based on the cross-domain speech classification dataset and the joint optimization loss function; The cross-domain speech classification module is used to input the speech to be classified into the trained cross-domain speech classification model to obtain the corresponding category of the speech.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • User feedback classification method and electronic equipment

    CN121093097A

  • User feedback classification method and electronic device

    CN121093097B

  • Unmanned aerial vehicle cluster voice command analysis control method and system

    CN121354552A