Pre-training optimization method, device, equipment and medium for emotion prediction model

Through the combination of unsupervised and self-supervised models, the speech emotion recognition performance of the pre-trained model is optimized by using clustering and pseudo-label training technology, solving the problem of poor performance of the pre-trained model on specific tasks, and achieving higher accuracy of emotional information extraction.

CN115457982BActive Publication Date: 2025-05-06PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211082543.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-06
Publication Date
2025-05-06
Estimated Expiration
2042-09-06

AI Technical Summary

Technical Problem

In the prior art, the pre-trained model is not effective when applied to specific speech emotion recognition tasks, especially when extracting emotional information on unlabeled data, and its accuracy is insufficient.

Method used

Speech frame-level features are extracted through the unsupervised sentiment prediction model, clustering obtains the clustering center point and anchor point, and the pseudo-labels are determined using the distance relationship of the anchor point, and the self-supervised sentiment prediction model is trained with these pseudo-labels to optimize the pre-training process.

Benefits of technology

The accuracy of the pre-trained model to extract emotional information from unlabeled data is improved, and the correlation between low-dimensional features and emotional information is enhanced, thereby improving the overall performance of speech emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457982B_ABST
    Figure CN115457982B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of artificial intelligence technology, and in particular to a pre-training optimization method, device, equipment and medium for an emotion prediction model. The method uses an unsupervised emotion prediction model to extract frame-level features of sentence-level speech, uses sentence-level emotion labels of sentence-level speech as emotion categories of frame-level features, clusters frame-level features of all emotion categories to obtain N cluster center points, uses the mean of all frame-level features belonging to the same emotion category as anchor points to obtain M anchor points, calculates the distance between all cluster center points and each anchor point, assigns pseudo labels to frame-level features according to cluster center points and anchor points, trains a self-supervised emotion prediction model based on the pseudo labels of all frame data in the training set, further strengthens the correlation between low-dimensional features and emotion information through clustering, and then uses cluster distribution as emotion pseudo labels and uses this for training, thereby improving the accuracy of the model's prediction of emotion information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application is applicable to the field of artificial intelligence technology, and in particular, relates to a pre-training optimization method, device, equipment and medium for a sentiment prediction model. Background Art

[0002] At present, Speech Emotion Recognition (SER) is an emerging research direction in the field of digital speech signal processing, which has opened up a new path for human-computer interaction and played an important role in many scenarios. Call centers use SER technology to track customer emotions and provide them with better services; in the medical field, diagnostic systems based on SER technology can analyze the degree of depression and pain of patients; there are many other applications that also use efficient SER systems to improve their work efficiency.

[0003] The emotions in human voices are affected by many factors, such as gender, age, speaker, dialect and culture. Therefore, how to better model emotions has always been a key research direction for researchers. Today, methods based on deep learning have become mainstream. Among them, the self-supervised pre-training model provides a high-performance solution. Although pre-training can use large-scale heterogeneous data sets to obtain models with strong performance and good versatility, the effect of applying the pre-training model to specific tasks is not ideal because the pre-training task is not completely consistent with the target task, that is, there is a difference between the pre-training domain and the target domain. In the SER task, a large amount of unlabeled data is usually used for pre-training, and the pre-trained model needs to be able to extract more accurate emotional information from unlabeled data. Therefore, how to optimize the pre-training model to improve the accuracy of the pre-training model in extracting emotional information from unlabeled data has become an urgent problem to be solved. Summary of the invention

[0004] In view of this, the embodiments of the present application provide a pre-training optimization method, device, equipment and medium for a sentiment prediction model to solve the problem of how to pre-train and optimize the pre-training model to improve the accuracy of the pre-training model in extracting sentiment information from unlabeled data.

[0005] In a first aspect, an embodiment of the present application provides a pre-training optimization method for a sentiment prediction model, the pre-training optimization method comprising:

[0006] For any sentence-level speech in the training set, an unsupervised emotion prediction model is used to extract frame-level features corresponding to each frame of data in the sentence-level speech, and the sentence-level emotion label of the sentence-level speech is used as the emotion category of the frame-level feature to obtain the emotion category corresponding to the frame-level features of all frame data in the training set, where all sentence-level speech in the training set are annotated with sentence-level emotion labels;

[0007] Cluster the frame-level features of all emotion categories to obtain N cluster centers, where N is an integer greater than zero;

[0008] The mean of all frame-level features belonging to the same emotion category is taken as the anchor point to obtain M anchor points. The distance between all cluster centers and each anchor point is calculated, where M is an integer greater than zero.

[0009] For any cluster center point, determine the anchor point closest to the cluster center point as the target anchor point, and when the target anchor point meets the preset conditions, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo label of all frame-level features in the cluster center point to obtain the pseudo labels of all frame data in the training set;

[0010] Based on the pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a pre-trained self-supervised emotion prediction model. Both the unsupervised emotion prediction model and the self-supervised emotion prediction model have feature encoders that are aligned with time steps.

[0011] In one embodiment, after calculating the distances between all cluster centers and each anchor point, the method further includes:

[0012] For any cluster center point, determine all anchor points whose distances to the cluster center point are less than a first distance threshold;

[0013] Determining the anchor point closest to the cluster center point as the target anchor point includes:

[0014] An anchor point closest to the cluster center point is determined as a target anchor point from all anchor points whose distances to the cluster center point are less than the first distance threshold.

[0015] In one embodiment, for any cluster center point, determining the anchor point closest to the cluster center point as the target anchor point, and when the target anchor point meets a preset condition, using the sentence-level sentiment label corresponding to the target anchor point as the pseudo label of all frame-level features in the cluster center point includes:

[0016] For any cluster center point, determine the anchor point closest to the cluster center point as the target anchor point;

[0017] Detecting whether the distance between the target anchor point and the cluster center point is less than a second distance threshold;

[0018] If it is detected that the distance between the target anchor point and the cluster center point is less than the second distance threshold, it is determined that the target anchor point meets the preset condition, and the sentence-level sentiment label corresponding to the target anchor point is used as a pseudo label for all frame-level features in the cluster center point.

[0019] In one embodiment, after detecting whether the distance between the target anchor point and the cluster center point is less than a second distance threshold, the method further includes:

[0020] If it is detected that the distance between the target anchor point and the cluster center point is not less than the second distance threshold, it is determined that the target anchor point does not meet the preset condition, and other types of anchor points are created, and the sentence-level sentiment labels defined by the other types of anchor points are used as pseudo labels for all frame-level features in the cluster center point.

[0021] In one embodiment, after the self-supervised emotion prediction model is trained using the training set based on the pseudo labels of all frame data in the training set to obtain a pre-trained self-supervised emotion prediction model, the method further includes:

[0022] For any sentence-level speech in the training set, use the pre-trained self-supervised emotion prediction model to extract the updated frame-level features corresponding to each frame of data in the sentence-level speech, use the sentence-level emotion label of the sentence-level speech as the emotion category of the updated frame-level features, and obtain the emotion categories corresponding to the updated frame-level features of all frame data in the training set;

[0023] Cluster the updated frame-level features of all emotion categories to obtain N updated cluster centers, where N is an integer greater than zero;

[0024] The mean of all updated frame-level features belonging to the same emotion category is taken as the anchor point to obtain M anchor points. The distance between all updated cluster centers and each anchor point is calculated, where M is an integer greater than zero.

[0025] For any updated cluster center point, determine the anchor point closest to the updated cluster center point as the target anchor point, and when the target anchor point meets the preset conditions, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all updated frame-level features in the updated cluster center point to obtain updated pseudo-labels for all frame data in the training set;

[0026] Based on the updated pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a trained self-supervised emotion prediction model.

[0027] In one embodiment, the unsupervised emotion prediction model includes a first feature encoder, a bidirectional LSTM layer and a first fully connected layer, the first feature encoder is composed of a CNN layer, the self-supervised emotion prediction model includes a second feature encoder, a Transformer layer and a second fully connected layer, the second feature encoder is composed of a CNN layer, and the number of CNN layers in the first feature encoder is the same as the number of CNN layers in the second feature encoder; before extracting the frame-level features corresponding to each frame of data in the sentence-level speech for any sentence-level speech in the training set using the unsupervised emotion prediction model, it also includes:

[0028] Using a feature set to train an unsupervised emotion prediction model, the feature set includes sentence-level speech samples and their corresponding emotion labels and emotion category labels;

[0029] Normalizing the output of the first fully connected layer using the softmax function to obtain a predicted value;

[0030] The first cross entropy function is used to measure the loss, and the step of training the unsupervised sentiment prediction model using the feature set is repeated until the result of the loss measurement meets the preset conditions, and the trained unsupervised sentiment prediction model is obtained, wherein the first cross entropy function L g include:

[0031]

[0032] In the formula, Z is the total number of samples, C is the total number of emotion categories, and y i represents the sentiment category label of the corresponding sample i, c j represents the sentiment label, p(c j |X i ) represents the corresponding input feature x i c j The posterior probability prediction value of the class.

[0033] In one embodiment, the second fully connected layer includes two layers of fully connected layers, and the pseudo labels of all frame data in the training set are used as a basis, and the self-supervised emotion prediction model is trained using the training set to obtain a pre-trained self-supervised emotion prediction model, which includes:

[0034] Using all frame data in the training set and their corresponding pseudo labels to train the self-supervised emotion prediction model;

[0035] The loss is measured by the second cross entropy function, and the steps of training the self-supervised emotion prediction model using all the frame data in the training set and their corresponding pseudo labels are repeated until the result of the loss measurement meets the preset conditions, thereby obtaining a pre-trained self-supervised emotion prediction model, wherein the second cross entropy function Lv include:

[0036]

[0037] In the formula, represents the low-dimensional features encoded and outputted by the feature encoder for the frame data, t represents the masked part masked by the second fully connected layer, z t represents the context representation of the masked part extracted using the Transformer layer, Represents the posterior probability prediction value of the context of the masked part.

[0038] In a second aspect, an embodiment of the present application provides a pre-training optimization device for an emotion prediction model, the pre-training optimization device comprising:

[0039] An unsupervised training module is used to extract frame-level features corresponding to each frame of data in any sentence-level speech in a training set using an unsupervised emotion prediction model, and use the sentence-level emotion label of the sentence-level speech as the emotion category of the frame-level feature to obtain the emotion category corresponding to the frame-level features of all frame data in the training set, wherein the sentence-level speech in the training set is annotated with a sentence-level emotion label;

[0040] A feature clustering module is used to cluster the frame-level features of all emotion categories to obtain N cluster center points, where N is an integer greater than zero;

[0041] A feature calculation module is used to take the mean of all frame-level features belonging to the same emotion category as an anchor point, obtain M anchor points, and calculate the distance between all cluster center points and each anchor point, where M is an integer greater than zero;

[0042] A pseudo-label determination module is used to determine, for any cluster center point, the anchor point closest to the cluster center point as the target anchor point, and when the target anchor point meets a preset condition, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all frame-level features in the cluster center point to obtain pseudo-labels for all frame data in the training set;

[0043] A self-supervised training module is used to train a self-supervised emotion prediction model based on the pseudo labels of all frame data in the training set, using the training set to obtain a pre-trained self-supervised emotion prediction model. Both the unsupervised emotion prediction model and the self-supervised emotion prediction model have feature encoders that are aligned with time steps.

[0044] In one embodiment, the pre-training optimization device further comprises:

[0045] An anchor point screening unit, configured to determine, for any cluster center point, all anchor points whose distances to the cluster center point are less than a first distance threshold after calculating the distances between all cluster center points and each anchor point respectively;

[0046] The pseudo-label determination module comprises:

[0047] The first target anchor point determination unit is used to determine, from all anchor points whose distances from the cluster center point are less than the first distance threshold, the anchor point that is closest to the cluster center point as the target anchor point.

[0048] In one embodiment, the pseudo-label determination module includes:

[0049] A second target anchor point determination unit, configured to determine, for any cluster center point, an anchor point that is closest to the cluster center point as a target anchor point;

[0050] A distance detection unit, used to detect whether the distance between the target anchor point and the cluster center point is less than a second distance threshold;

[0051] The first pseudo-label determination unit is used to determine that the target anchor point meets a preset condition if it is detected that the distance between the target anchor point and the cluster center point is less than the second distance threshold, and use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all frame-level features in the cluster center point.

[0052] In one embodiment, the pre-training optimization device further comprises:

[0053] The second pseudo-label determination unit is used to, after detecting whether the distance between the target anchor point and the cluster center point is less than a second distance threshold, determine that the target anchor point does not meet the preset condition if it is detected that the distance between the target anchor point and the cluster center point is not less than the second distance threshold, create other types of anchor points, and use the sentence-level sentiment label defined by the other types of anchor points as the pseudo-label of all frame-level features in the cluster center point.

[0054] In one embodiment, the pre-training optimization device further comprises:

[0055] A fine-tuning module, wherein the fine-tuning module is specifically used for:

[0056] Based on the pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a pre-trained self-supervised emotion prediction model. For any sentence-level speech in the training set, the pre-trained self-supervised emotion prediction model is used to extract updated frame-level features corresponding to each frame data in the sentence-level speech, and the sentence-level emotion label of the sentence-level speech is used as the emotion category of the updated frame-level feature to obtain the emotion category corresponding to the updated frame-level features of all frame data in the training set.

[0057] Cluster the updated frame-level features of all emotion categories to obtain N updated cluster centers;

[0058] The mean of all updated frame-level features belonging to the same emotion category is taken as the anchor point to obtain M anchor points, and the distance between all updated cluster centers and each anchor point is calculated;

[0059] For any updated cluster center point, determine the anchor point closest to the updated cluster center point as the target anchor point, and when the target anchor point meets the preset conditions, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all updated frame-level features in the updated cluster center point to obtain updated pseudo-labels for all frame data in the training set;

[0060] Based on the updated pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a trained self-supervised emotion prediction model.

[0061] In one embodiment, the unsupervised emotion prediction model includes a first feature encoder, a bidirectional LSTM layer and a first fully connected layer, the first feature encoder is composed of a CNN layer, the self-supervised emotion prediction model includes a second feature encoder, a Transformer layer and a second fully connected layer, the second feature encoder is composed of a CNN layer, and the number of CNN layers in the first feature encoder is the same as the number of CNN layers in the second feature encoder; the pre-training optimization device also includes:

[0062] A training module, for training the unsupervised emotion prediction model using a feature set before extracting frame-level features corresponding to each frame of data in the sentence-level speech for any sentence-level speech in the training set using the unsupervised emotion prediction model, wherein the feature set includes sentence-level speech samples and their corresponding emotion labels and their corresponding emotion category labels;

[0063] A normalization module, used to perform softmax function normalization on the output of the first fully connected layer to obtain a predicted value;

[0064] Return to the execution module, which is used to perform loss measurement through the first cross entropy function, and repeat the step of training the unsupervised sentiment prediction model using the feature set until the result of the loss measurement meets the preset conditions, thereby obtaining a trained unsupervised sentiment prediction model, wherein the first cross entropy function L g include:

[0065]

[0066] In the formula, Z is the total number of samples, C is the total number of emotion categories, and y i represents the sentiment category label of the corresponding sample i, c j represents the sentiment label, p(c j |X i ) represents the corresponding input feature x i c j The posterior probability prediction value of the class.

[0067] In one embodiment, the second fully connected layer includes two fully connected layers, and the self-supervised training module is specifically used for:

[0068] Using all frame data in the training set and their corresponding pseudo labels to train the self-supervised emotion prediction model;

[0069] The loss is measured by the second cross entropy function, and the steps of training the self-supervised emotion prediction model using all the frame data in the training set and their corresponding pseudo labels are repeated until the result of the loss measurement meets the preset conditions, thereby obtaining a pre-trained self-supervised emotion prediction model, wherein the second cross entropy function L v include:

[0070]

[0071] In the formula, represents the low-dimensional features encoded and outputted by the feature encoder for the frame data, t represents the masked part masked by the second fully connected layer, z t represents the context representation of the masked part extracted using the Transformer layer, Represents the posterior probability prediction value of the context of the masked part.

[0072] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the pre-training optimization method as described in the first aspect when executing the computer program.

[0073] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the pre-training optimization method as described in the first aspect is implemented.

[0074] Compared with the prior art, the embodiments of the present application have the following beneficial effects: for any sentence-level speech in the training set, the present application uses an unsupervised emotion prediction model to extract frame-level features corresponding to each frame of data in the sentence-level speech, uses the sentence-level emotion label of the sentence-level speech as the emotion category of the frame-level feature, obtains the emotion category corresponding to the frame-level features of all frame data in the training set, clusters the frame-level features of all emotion categories to obtain N cluster center points, uses the mean of all frame-level features belonging to the same emotion category as anchor points to obtain M anchor points, calculates the distance between all cluster center points and each anchor point, and for any cluster center point, determine the anchor point closest to the cluster center point as the target anchor point, when the target anchor point meets the preset conditions, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all frame-level features in the cluster center point, and obtain the pseudo-labels of all frame data in the training set. Based on the pseudo-labels of all frame data in the training set, the self-supervised sentiment prediction model is trained using the training set to obtain a pre-trained self-supervised sentiment prediction model. Clustering is used to further strengthen the correlation between low-dimensional features and sentiment information, and then the cluster distribution is used as the sentiment pseudo-label for training, thereby improving the model's prediction accuracy for sentiment information. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0076] Figure 1 This is a schematic diagram of an application environment of a pre-training optimization method for a sentiment prediction model provided in Example 1 of the present application;

[0077] Figure 2 It is a flowchart of a pre-training optimization method for a sentiment prediction model provided in Example 2 of the present application;

[0078] Figure 3 It is a structural schematic diagram of a pre-training optimization device for a sentiment prediction model provided in Example 3 of the present application;

[0079] Figure 4 It is a structural diagram of a computer device provided in Example 4 of the present application. DETAILED DESCRIPTION

[0080] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0081] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0082] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0083] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0084] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0085] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0086] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0087] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0088] It should be understood that the size of the serial numbers of the steps in the following embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0089] In order to illustrate the technical solution of the present application, a specific embodiment is provided below for illustration.

[0090] The pre-training optimization method of the emotion prediction model provided in the first embodiment of the present application can be applied to Figure 1 In an application environment, a client communicates with a server. The client includes but is not limited to a PDA, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud computer device, a personal digital assistant (PDA), and other computer devices. The server can be implemented by an independent server or a server cluster consisting of multiple servers.

[0091] See also Figure 2 , is a flow chart of a pre-training optimization method for a sentiment prediction model provided in Example 2 of the present application, wherein the pre-training optimization method for the sentiment prediction model is applied to Figure 1 The server in the example, the computer device corresponding to the server connects to the corresponding database to obtain the corresponding training data in the database. The above-mentioned computer device can also be connected to the corresponding client, which is operated by the user, and the user can provide the corresponding training set to the server through the client. Figure 2 As shown, the pre-training optimization method of the emotion prediction model may include the following steps:

[0092] Step S201, for any sentence-level speech in the training set, use the unsupervised emotion prediction model to extract the frame-level features corresponding to each frame data in the sentence-level speech, use the sentence-level emotion label of the sentence-level speech as the emotion category of the frame-level features, and obtain the emotion category corresponding to the frame-level features of all frame data in the training set.

[0093] In the present application, the training set includes at least one sentence-level speech, and each sentence-level speech is annotated with a corresponding sentence-level emotion label. The sentence-level speech can refer to a group of speech data in units of sentences, wherein a sentence can be a group of words, a paragraph, etc. The sentence-level emotion label is the label of a sentence, that is, each frame of data in a group of speech data corresponds to a sentence-level emotion label.

[0094] The unsupervised sentiment prediction model may be a wav2vec model, which includes a feature encoder, a bidirectional LSTM layer, and a fully connected layer, wherein a softmax function may be used for normalization in the fully connected layer.

[0095] The feature encoder is used to align the time step with the self-supervised emotion prediction model. If the self-supervised emotion prediction model adopts the wav2vec2.0 model, the wav2vec2.0 model contains a feature encoder composed of a multi-layer CNN network. Accordingly, the feature encoder in the wav2vec model also needs to have a CNN network with the same number of layers.

[0096] The frame-level feature is the feature of each frame of data, that is, the sentence-level speech is divided into frames of speech, and features are extracted for each frame of speech. A corresponding sentence-level speech corresponds to an emotion label, and the emotion category corresponding to the features of all frames of speech in the sentence-level speech is the emotion label of the sentence-level speech.

[0097] Step S202: cluster the frame-level features of all emotion categories to obtain N cluster center points.

[0098] In the present application, N is an integer greater than zero, and an improved K-means clustering algorithm is used to cluster the frame-level features into N clusters, with the mean of all frame-level features in each cluster corresponding to the cluster center point.

[0099] Specifically, N frame-level features are arbitrarily selected from all frame-level features as initial cluster center points, the distance between each frame-level feature and each initial cluster center is calculated, each frame-level feature is assigned to the cluster center point closest to it, and the centroid of all frame-level features corresponding to each cluster center point is recalculated. If the distance between the newly calculated centroid and the original cluster center point is less than a set threshold, it means that the position of the recalculated centroid does not change much and tends to stabilize and converge. If the distance between the new centroid and the original cluster center point changes greatly, it is necessary to iterate again, and finally obtain N cluster center points, that is, N clusters.

[0100] Step S203, taking the mean of all frame-level features belonging to the same emotion category as an anchor point, obtaining M anchor points, and calculating the distance between all cluster center points and each anchor point.

[0101] In this application, M is an integer greater than zero. Since the extracted frame-level features can already represent certain emotional information, the mean of all frame-level features of the same emotional category is extracted as an anchor point, and M anchor points can be obtained, where M also represents the total number of emotional categories. Using the above clustering results, the distance between the cluster center point of each cluster and each anchor point is calculated.

[0102] Step S204, for any cluster center point, determine the anchor point closest to the cluster center point as the target anchor point. When the target anchor point meets the preset conditions, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all frame-level features in the cluster center point to obtain the pseudo-label of all frame data in the training set.

[0103] Among them, after determining the target anchor point, if the preset condition is defined as the distance between the target anchor point and the corresponding cluster center point is less than a specific value, then satisfying the preset condition means that the distance between the target anchor point and the corresponding cluster center point is less than the specific value.

[0104] The anchor point closest to any cluster center point may be one or more. When there are multiple anchor points, the preset condition can be defined as being selected, that is, an anchor point is randomly selected from multiple anchor points as the selected anchor point, that is, the anchor point that meets the preset condition.

[0105] Optionally, after calculating the distances between all cluster centers and each anchor point, the following steps are also included:

[0106] For any cluster center point, determine all anchor points whose distances to the cluster center point are less than a first distance threshold;

[0107] Determining the anchor point closest to the cluster center as the target anchor point includes:

[0108] From all anchor points whose distances to the cluster center point are less than a first distance threshold, an anchor point whose distance to the cluster center point is closest is determined as a target anchor point.

[0109] Optionally, for any cluster center point, determining the anchor point closest to the cluster center point as the target anchor point, and when the target anchor point meets the preset conditions, using the sentence-level sentiment label corresponding to the target anchor point as the pseudo label of all frame-level features in the cluster center point includes:

[0110] For any cluster center point, determine the anchor point closest to the cluster center point as the target anchor point;

[0111] Detect whether the distance between the target anchor point and the cluster center point is less than a second distance threshold;

[0112] If it is detected that the distance between the target anchor point and the cluster center point is less than the second distance threshold, it is determined that the target anchor point meets the preset condition, and the sentence-level sentiment label corresponding to the target anchor point is used as the pseudo label of all frame-level features in the cluster center point.

[0113] Optionally, after detecting whether the distance between the target anchor point and the cluster center point is less than a second distance threshold, the method further includes:

[0114] If it is detected that the distance between the target anchor point and the cluster center point is not less than the second distance threshold, it is determined that the target anchor point does not meet the preset condition, and other types of anchor points are created, and the sentence-level sentiment labels defined by other types of anchor points are used as pseudo labels for all frame-level features in the cluster center point.

[0115] Among them, if the distance d between the cluster center and the anchor point ij ≤γ, the pseudo-label of the cluster center is mapped to the emotion category corresponding to the anchor point, γ represents the preset threshold, i∈(0,M), j∈(0,N); if the distance is greater than γ, a new other class anchor point is created as a unified category representation of the non-emotion category, and finally there are M+1 pseudo-labels in total.

[0116] Step S205 , based on the pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a pre-trained self-supervised emotion prediction model.

[0117] In this application, both the unsupervised emotion prediction model and the self-supervised emotion prediction model have a time-step aligned feature encoder.

[0118] Optionally, the unsupervised emotion prediction model includes a first feature encoder, a bidirectional LSTM layer and a first fully connected layer, the first feature encoder is composed of a CNN layer, the self-supervised emotion prediction model includes a second feature encoder, a Transformer layer and a second fully connected layer, the second feature encoder is composed of a CNN layer, and the number of CNN layers in the first feature encoder is the same as the number of CNN layers in the second feature encoder; before using the unsupervised emotion prediction model to extract frame-level features corresponding to each frame of data in the sentence-level speech for any sentence-level speech in the training set, it also includes:

[0119] The unsupervised emotion prediction model is trained using a feature set, where the feature set includes sentence-level speech samples and their corresponding emotion labels and emotion category labels;

[0120] Normalize the output of the first fully connected layer using the softmax function to obtain the predicted value;

[0121] The first cross entropy function is used to measure the loss, and the step of training the unsupervised sentiment prediction model using the feature set is repeated until the result of the loss measurement meets the preset conditions, and the trained unsupervised sentiment prediction model is obtained, wherein the first cross entropy function L g include:

[0122]

[0123] In the formula, Z is the total number of samples, C is the total number of emotion categories, and y i represents the sentiment category label of the corresponding sample i, c j represents the sentiment label, p(c j |X i ) represents the corresponding input feature x i c j The posterior probability prediction value of the class.

[0124] Among them, the second feature encoder consists of 7 layers of CNN, and then the context representation is obtained through the Transformer layer, and then the pseudo-label category of the masked part is predicted through a linear multi-head composed of two layers of fully connected layers.

[0125] Optionally, the second fully connected layer includes two fully connected layers, and the self-supervised emotion prediction model is trained using the training set based on the pseudo labels of all frame data in the training set, and the pre-trained self-supervised emotion prediction model includes:

[0126] Use all frame data in the training set and their corresponding pseudo labels to train the self-supervised sentiment prediction model;

[0127] The loss is measured by the second cross entropy function, and the steps of training the self-supervised emotion prediction model using all the frame data in the training set and their corresponding pseudo labels are repeated until the loss measurement result meets the preset conditions, and the pre-trained self-supervised emotion prediction model is obtained, where the second cross entropy function L v include:

[0128]

[0129] In the formula, represents the low-dimensional features encoded and output by the feature encoder, t represents the masked part masked by the second fully connected layer, and z t represents the context representation of the masked part extracted using the Transformer layer, Represents the posterior probability prediction value of the masked part of the context.

[0130] Optionally, after the self-supervised emotion prediction model is trained using the training set based on the pseudo labels of all frame data in the training set to obtain a pre-trained self-supervised emotion prediction model, the method further includes:

[0131] For any sentence-level speech in the training set, use the pre-trained self-supervised emotion prediction model to extract the updated frame-level features corresponding to each frame of the sentence-level speech, use the sentence-level emotion label of the sentence-level speech as the emotion category of the updated frame-level features, and obtain the emotion category corresponding to the updated frame-level features of all the frame data in the training set;

[0132] Cluster the updated frame-level features of all emotion categories to obtain N updated cluster centers, where N is an integer greater than zero;

[0133] The mean of all updated frame-level features belonging to the same emotion category is taken as the anchor point to obtain M anchor points. The distance between all updated cluster centers and each anchor point is calculated, where M is an integer greater than zero.

[0134] For any updated cluster center point, determine the anchor point closest to the updated cluster center point as the target anchor point. When the target anchor point meets the preset conditions, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all updated frame-level features in the updated cluster center point to obtain the updated pseudo-label of all frame data in the training set.

[0135] Based on the updated pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a trained self-supervised emotion prediction model.

[0136] Among them, since the pseudo-label category is related to the emotion class, this method can focus on predicting the emotion information of the masked sequence. After the pre-training is completed, wav2vec2.0 can be directly used to replace wav2vec, and the pre-trained self-supervised emotion prediction model can be fine-tuned according to steps S201 to S205 to obtain a trained self-supervised emotion prediction model.

[0137] In the embodiment of the present application, for any sentence-level speech in the training set, an unsupervised emotion prediction model is used to extract frame-level features corresponding to each frame of data in the sentence-level speech, and the sentence-level emotion label of the sentence-level speech is used as the emotion category of the frame-level feature to obtain the emotion category corresponding to the frame-level features of all frame data in the training set, and the frame-level features of all emotion categories are clustered to obtain N cluster center points, and the mean of all frame-level features belonging to the same emotion category is used as an anchor point to obtain M anchor points, and the distance between all cluster center points and each anchor point is calculated. For any cluster center point, the distance between the cluster center point and the anchor point is determined. The anchor point closest to the target anchor point is taken as the target anchor point. When the target anchor point meets the preset conditions, the sentence-level sentiment label corresponding to the target anchor point is used as the pseudo-label of all frame-level features in the cluster center point to obtain the pseudo-labels of all frame data in the training set. Based on the pseudo-labels of all frame data in the training set, the self-supervised sentiment prediction model is trained using the training set to obtain a pre-trained self-supervised sentiment prediction model. Clustering is used to further strengthen the correlation between low-dimensional features and sentiment information, and the cluster distribution is used as the sentiment pseudo-label for training, thereby improving the model's prediction accuracy for sentiment information.

[0138] Corresponding to the pre-training optimization method of the emotion prediction model in the above embodiment, Figure 3 The structural block diagram of the pre-training optimization device of the emotion prediction model provided in the third embodiment of the present application is shown. The pre-training optimization device is applied to Figure 1 The server in the embodiment of the present invention is connected to a corresponding database by a computer device corresponding to the server, so as to obtain the corresponding training data in the database. The above-mentioned computer device can also be connected to a corresponding client, which is operated by a user, and the user can provide a corresponding training set to the server through the client. For the convenience of explanation, only the part related to the embodiment of the present application is shown.

[0139] See also Figure 3 , the pre-training optimization device comprises:

[0140] The unsupervised training module 31 is used to extract the frame-level features corresponding to each frame of data in the sentence-level speech using the unsupervised emotion prediction model for any sentence-level speech in the training set, and use the sentence-level emotion label of the sentence-level speech as the emotion category of the frame-level features to obtain the emotion categories corresponding to the frame-level features of all the frame data in the training set;

[0141] A feature clustering module 32, used for clustering frame-level features of all emotion categories to obtain N cluster center points, where N is an integer greater than zero;

[0142] A feature calculation module 33 is used to take the mean of all frame-level features belonging to the same emotion category as an anchor point to obtain M anchor points, and calculate the distance between all cluster center points and each anchor point, where M is an integer greater than zero;

[0143] The pseudo-label determination module 34 is used to determine, for any cluster center point, the anchor point closest to the cluster center point as the target anchor point, and when the target anchor point meets the preset conditions, the sentence-level sentiment label corresponding to the target anchor point is used as the pseudo-label of all frame-level features in the cluster center point to obtain the pseudo-label of all frame data in the training set;

[0144] The self-supervised training module 35 is used to train the self-supervised emotion prediction model based on the pseudo labels of all frame data in the training set, and obtain a pre-trained self-supervised emotion prediction model. Both the unsupervised emotion prediction model and the self-supervised emotion prediction model have feature encoders that are aligned with time steps.

[0145] Optionally, the pre-training optimization device further includes:

[0146] An anchor point screening unit, configured to determine, for any cluster center point, all anchor points whose distances to the cluster center point are less than a first distance threshold after calculating the distances between all cluster center points and each anchor point respectively;

[0147] The pseudo-label determination module 34 includes:

[0148] The first target anchor point determination unit is used to determine the anchor point closest to the cluster center point from all anchor points whose distances to the cluster center point are less than a first distance threshold as the target anchor point.

[0149] Optionally, the pseudo-label determination module 34 includes:

[0150] A second target anchor point determination unit is used to determine, for any cluster center point, an anchor point that is closest to the cluster center point as a target anchor point;

[0151] A distance detection unit, used to detect whether the distance between the target anchor point and the cluster center point is less than a second distance threshold;

[0152] The first pseudo-label determination unit is used to determine that the target anchor point meets the preset conditions if it is detected that the distance between the target anchor point and the cluster center point is less than the second distance threshold, and use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all frame-level features in the cluster center point.

[0153] Optionally, the pre-training optimization device further includes:

[0154] The second pseudo-label determination unit is used to determine whether the distance between the target anchor point and the cluster center point is less than a second distance threshold. If it is detected that the distance between the target anchor point and the cluster center point is not less than the second distance threshold, it is determined that the target anchor point does not meet the preset conditions, and other types of anchor points are created, and the sentence-level sentiment labels defined by the other types of anchor points are used as pseudo-labels for all frame-level features in the cluster center point.

[0155] Optionally, the pre-training optimization device further includes:

[0156] Fine-tuning module, which is specifically used for:

[0157] Based on the pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a pre-trained self-supervised emotion prediction model. For any sentence-level speech in the training set, the pre-trained self-supervised emotion prediction model is used to extract the updated frame-level features corresponding to each frame data in the sentence-level speech, and the sentence-level emotion label of the sentence-level speech is used as the emotion category of the updated frame-level features to obtain the emotion category corresponding to the updated frame-level features of all frame data in the training set.

[0158] Cluster the updated frame-level features of all emotion categories to obtain N updated cluster centers;

[0159] The mean of all updated frame-level features belonging to the same emotion category is taken as the anchor point to obtain M anchor points, and the distance between all updated cluster centers and each anchor point is calculated;

[0160] For any updated cluster center point, determine the anchor point closest to the updated cluster center point as the target anchor point. When the target anchor point meets the preset conditions, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all updated frame-level features in the updated cluster center point to obtain the updated pseudo-label of all frame data in the training set.

[0161] Based on the updated pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a trained self-supervised emotion prediction model.

[0162] Optionally, the unsupervised emotion prediction model includes a first feature encoder, a bidirectional LSTM layer and a first fully connected layer, the first feature encoder is composed of a CNN layer, the self-supervised emotion prediction model includes a second feature encoder, a Transformer layer and a second fully connected layer, the second feature encoder is composed of a CNN layer, and the number of CNN layers in the first feature encoder is the same as the number of CNN layers in the second feature encoder; the pre-training optimization device also includes:

[0163] A training module, for training the unsupervised emotion prediction model using a feature set before extracting frame-level features corresponding to each frame of data in the sentence-level speech using the unsupervised emotion prediction model for any sentence-level speech in the training set, wherein the feature set includes sentence-level speech samples and their corresponding emotion labels and emotion category labels;

[0164] The normalization module is used to perform softmax function normalization on the output of the first fully connected layer to obtain the predicted value;

[0165] Return to the execution module, which is used to perform loss measurement through the first cross entropy function, and repeat the step of training the unsupervised sentiment prediction model using the feature set until the result of the loss measurement meets the preset conditions, thereby obtaining a trained unsupervised sentiment prediction model, wherein the first cross entropy function L g include:

[0166]

[0167] In the formula, Z is the total number of samples, C is the total number of emotion categories, and y i represents the sentiment category label of the corresponding sample i, c j represents the sentiment label, p(c j |X i ) represents the corresponding input feature x i c j The posterior probability prediction value of the class.

[0168] Optionally, the second fully connected layer includes two fully connected layers, and the self-supervised training module 35 is specifically used for:

[0169] Use all frame data in the training set and their corresponding pseudo labels to train the self-supervised sentiment prediction model;

[0170] The loss is measured by the second cross entropy function, and the steps of training the self-supervised emotion prediction model using all the frame data in the training set and their corresponding pseudo labels are repeated until the loss measurement result meets the preset conditions, and the pre-trained self-supervised emotion prediction model is obtained, where the second cross entropy function L v include:

[0171]

[0172] In the formula, represents the low-dimensional features encoded and output by the feature encoder, t represents the masked part masked by the second fully connected layer, and z t represents the context representation of the masked part extracted using the Transformer layer, Represents the posterior probability prediction value of the masked part of the context.

[0173] It should be noted that the information interaction, execution process and other contents between the above-mentioned modules are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0174] Figure 4 This is a schematic diagram of the structure of a computer device provided in Example 4 of the present application. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in the above-mentioned pre-training optimization method embodiment of any of the emotion prediction models are implemented.

[0175] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that Figure 4 This is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than those shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.

[0176] The processor may be a CPU, or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0177] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory may be the memory of a computer device, and the internal memory provides an environment for the operation of an operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be a hard disk of a computer device, and in other embodiments may also be an external storage device of a computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on a computer device. Further, the memory may also include both an internal storage unit of a computer device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as program codes of computer programs, etc. The memory may also be used to temporarily store data that has been output or is to be output.

[0178] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the above-mentioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned method embodiment when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0179] The present application implements all or part of the processes in the above-mentioned embodiment method, and may also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiment when executing the computer program product.

[0180] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0181] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0182] In the embodiments provided in the present application, it should be understood that the disclosed devices / computer equipment and methods can be implemented in other ways. For example, the device / computer equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0183] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0184] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A pre-training optimization method for a sentiment prediction model, characterized in that: The pre-training optimization method comprises: For any sentence-level speech in the training set, an unsupervised emotion prediction model is used to extract frame-level features corresponding to each frame of data in the sentence-level speech, and the sentence-level emotion label of the sentence-level speech is used as the emotion category of the frame-level feature to obtain the emotion category corresponding to the frame-level features of all frame data in the training set, where all sentence-level speech in the training set are annotated with sentence-level emotion labels; Cluster the frame-level features of all emotion categories to obtain N cluster centers, where N is an integer greater than zero; The mean of all frame-level features belonging to the same emotion category is taken as the anchor point to obtain M anchor points. The distance between all cluster centers and each anchor point is calculated, where M is an integer greater than zero. For any cluster center point, determine the anchor point closest to the cluster center point as the target anchor point, and when the target anchor point meets the preset conditions, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo label of all frame-level features in the cluster center point to obtain the pseudo labels of all frame data in the training set; Based on the pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a pre-trained self-supervised emotion prediction model. Both the unsupervised emotion prediction model and the self-supervised emotion prediction model have feature encoders that are aligned with time steps.

2. The pre-training optimization method according to claim 1, characterized in that: After calculating the distance between all cluster centers and each anchor point, it also includes: For any cluster center point, determine all anchor points whose distances to the cluster center point are less than a first distance threshold; Determining the anchor point closest to the cluster center point as the target anchor point includes: An anchor point closest to the cluster center point is determined as a target anchor point from all anchor points whose distances to the cluster center point are less than the first distance threshold.

3. The pre-training optimization method according to claim 1, characterized in that: For any cluster center point, determining the anchor point closest to the cluster center point as the target anchor point, and when the target anchor point meets the preset conditions, using the sentence-level sentiment label corresponding to the target anchor point as the pseudo label of all frame-level features in the cluster center point includes: For any cluster center point, determine the anchor point closest to the cluster center point as the target anchor point; Detecting whether the distance between the target anchor point and the cluster center point is less than a second distance threshold; If it is detected that the distance between the target anchor point and the cluster center point is less than the second distance threshold, it is determined that the target anchor point meets the preset condition, and the sentence-level sentiment label corresponding to the target anchor point is used as a pseudo label for all frame-level features in the cluster center point.

4. The pre-training optimization method according to claim 3, characterized in that: After detecting whether the distance between the target anchor point and the cluster center point is less than a second distance threshold, the method further includes: If it is detected that the distance between the target anchor point and the cluster center point is not less than the second distance threshold, it is determined that the target anchor point does not meet the preset condition, and other types of anchor points are created, and the sentence-level sentiment labels defined by the other types of anchor points are used as pseudo labels for all frame-level features in the cluster center point.

5. The pre-training optimization method according to claim 1, characterized in that: After the self-supervised emotion prediction model is trained using the training set based on the pseudo labels of all frame data in the training set to obtain a pre-trained self-supervised emotion prediction model, the method further includes: For any sentence-level speech in the training set, use the pre-trained self-supervised emotion prediction model to extract the updated frame-level features corresponding to each frame of data in the sentence-level speech, use the sentence-level emotion label of the sentence-level speech as the emotion category of the updated frame-level features, and obtain the emotion categories corresponding to the updated frame-level features of all frame data in the training set; Cluster the updated frame-level features of all emotion categories to obtain N updated cluster centers, where N is an integer greater than zero; The mean of all updated frame-level features belonging to the same emotion category is taken as the anchor point to obtain M anchor points. The distance between all updated cluster centers and each anchor point is calculated, where M is an integer greater than zero. For any updated cluster center point, determine the anchor point closest to the updated cluster center point as the target anchor point, and when the target anchor point meets the preset conditions, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all updated frame-level features in the updated cluster center point to obtain updated pseudo-labels for all frame data in the training set; Based on the updated pseudo labels of all frame data in the training set, the self-supervised emotion prediction model is trained using the training set to obtain a trained self-supervised emotion prediction model.

6. The pre-training optimization method according to any one of claims 1 to 5, characterized in that: The unsupervised emotion prediction model includes a first feature encoder, a bidirectional LSTM layer and a first fully connected layer, the first feature encoder is composed of a CNN layer, the self-supervised emotion prediction model includes a second feature encoder, a Transformer layer and a second fully connected layer, the second feature encoder is composed of a CNN layer, and the number of CNN layers in the first feature encoder is the same as the number of CNN layers in the second feature encoder; before extracting the frame-level features corresponding to each frame of data in the sentence-level speech for any sentence-level speech in the training set using the unsupervised emotion prediction model, the method further includes: Using a feature set to train an unsupervised emotion prediction model, the feature set includes sentence-level speech samples and their corresponding emotion labels and emotion category labels; Normalizing the output of the first fully connected layer using the softmax function to obtain a predicted value; The first cross entropy function is used to measure the loss, and the step of training the unsupervised sentiment prediction model using the feature set is repeated until the result of the loss measurement meets the preset conditions, and the trained unsupervised sentiment prediction model is obtained, wherein the first cross entropy function L g include: In the formula, Z is the total number of samples, C is the total number of emotion categories, and y i represents the sentiment category label of the corresponding sample i, c j represents the sentiment label, p(c j |X i ) represents the corresponding input feature x i c j The posterior probability prediction value of the class.

7. The pre-training optimization method according to claim 6, characterized in that: The second fully connected layer includes two layers of fully connected layers, and the pseudo labels of all frame data in the training set are used as a basis, and the self-supervised emotion prediction model is trained using the training set to obtain a pre-trained self-supervised emotion prediction model including: Using all frame data in the training set and their corresponding pseudo labels to train the self-supervised emotion prediction model; The loss is measured by the second cross entropy function, and the steps of training the self-supervised emotion prediction model using all the frame data in the training set and their corresponding pseudo labels are repeated until the result of the loss measurement meets the preset conditions, thereby obtaining a pre-trained self-supervised emotion prediction model, wherein the second cross entropy function L v include: In the formula, represents the low-dimensional features encoded and outputted by the feature encoder for the frame data, t represents the masked part masked by the second fully connected layer, z t represents the context representation of the masked part extracted using the Transformer layer, Represents the posterior probability prediction value of the context of the masked part.

8. A pre-training optimization device for an emotion prediction model, characterized in that: The pre-training optimization device comprises: An unsupervised training module is used to extract frame-level features corresponding to each frame of data in any sentence-level speech in a training set using an unsupervised emotion prediction model, and use the sentence-level emotion label of the sentence-level speech as the emotion category of the frame-level feature to obtain the emotion category corresponding to the frame-level features of all frame data in the training set, wherein the sentence-level speech in the training set is annotated with a sentence-level emotion label; A feature clustering module is used to cluster the frame-level features of all emotion categories to obtain N cluster center points, where N is an integer greater than zero; A feature calculation module is used to take the mean of all frame-level features belonging to the same emotion category as an anchor point, obtain M anchor points, and calculate the distance between all cluster center points and each anchor point, where M is an integer greater than zero; A pseudo-label determination module is used to determine, for any cluster center point, the anchor point closest to the cluster center point as the target anchor point, and when the target anchor point meets a preset condition, use the sentence-level sentiment label corresponding to the target anchor point as the pseudo-label of all frame-level features in the cluster center point to obtain pseudo-labels for all frame data in the training set; A self-supervised training module is used to train a self-supervised emotion prediction model based on the pseudo labels of all frame data in the training set, using the training set to obtain a pre-trained self-supervised emotion prediction model. Both the unsupervised emotion prediction model and the self-supervised emotion prediction model have feature encoders that are aligned with time steps.

9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the pre-training optimization method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the pre-training optimization method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Speech classification network training method and device, computing equipment and storage medium

    CN113593611A

  • Emotion recognition model and training method, device and equipment thereof, and readable storage medium

    CN114757310A