Audio processing method, device and computer program product
The audio processing method integrates frame-level features into audio-level features through a speaker feature extraction network and classifiers to improve the efficiency and accuracy of speaker attribute recognition in audio data, addressing the inefficiencies of separate model approaches.
Patent Information
- Application Number
- CN202210192079.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-02-28
AI Technical Summary
In the prior art, the recognition efficiency and accuracy of the different attribute information of the speaker in audio are low, and multiple independent models are usually used to predict and identify them.
By extracting audio frame features, the feature extraction layer and pooling layer in the speaker feature extraction network are used to convert the frame-level features into audio-level features, and multiple speaker attribute classifiers are input to comprehensively obtain the speaker's classification results under multiple attributes.
It improves the recognition efficiency and accuracy of the speaker's attribute information in the audio, especially the recognition effect of the speaker's gender and age in live broadcast scenarios.
Smart Images

Figure CN114596864B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and in particular, to an audio processing method, a computer device, and a computer program product. Background Art
[0002] With the development of Internet technology, various types of audio data are widely spread on the network, and there is a need to analyze and process the speaker attribute information in the audio to assist in the detection and identification of specific populations. In current technologies, for different attribute information of speakers in audio, multiple independent models are usually used to perform prediction and identification respectively, and the identification efficiency and accuracy are relatively low. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide an audio processing method, a computer device, and a computer program product.
[0004] In a first aspect, the present application provides an audio processing method. The method includes:
[0005] Extracting the features corresponding to each frame of audio in the audio to be processed, and obtaining a plurality of primary first audio frame features;
[0006] Obtaining, through a feature extraction layer in a trained speaker feature extraction network, a plurality of advanced second audio frame features corresponding to the plurality of primary first audio frame features respectively, and converting, through a pooling layer in the speaker feature extraction network, the plurality of advanced second audio frame features into audio features for characterizing the identity characteristics of the speaker in the audio;
[0007] Inputting the audio features into a plurality of trained speaker attribute classifiers, and obtaining a plurality of speaker attribute classification labels respectively output by the plurality of speaker attribute classifiers;
[0008] According to the plurality of speaker attribute classification labels, obtaining classification results of the speaker in the audio to be processed under multiple attributes.
[0009] In a second aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0010] Extract the features corresponding to each frame of the audio to be processed, and obtain multiple primary first audio frame features; obtain multiple advanced second audio frame features corresponding to the multiple primary first audio frame features respectively through the feature extraction layer in the trained speaker feature extraction network, and convert the multiple advanced second audio frame features into audio features used to characterize the identity characteristics of the speaker in the audio through the pooling layer in the speaker feature extraction network; input the audio features into multiple trained speaker attribute classifiers, and obtain multiple speaker attribute classification labels respectively output by the multiple speaker attribute classifiers; according to the multiple speaker attribute classification labels, obtain the classification results of the speaker in the audio to be processed under multiple attributes.
[0011] In a third aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0012] Extract the features corresponding to each frame of the audio to be processed, and obtain multiple primary first audio frame features; obtain multiple advanced second audio frame features corresponding to the multiple primary first audio frame features respectively through the feature extraction layer in the trained speaker feature extraction network, and convert the multiple advanced second audio frame features into audio features used to characterize the identity characteristics of the speaker in the audio through the pooling layer in the speaker feature extraction network; input the audio features into multiple trained speaker attribute classifiers, and obtain multiple speaker attribute classification labels respectively output by the multiple speaker attribute classifiers; according to the multiple speaker attribute classification labels, obtain the classification results of the speaker in the audio to be processed under multiple attributes.
[0013] The above audio processing method, computer device, and computer program product. For each frame in the audio to be processed, the method initially extracts the respective corresponding features to obtain multiple primary first audio frame features. Then, through the feature extraction layer in the trained speaker feature extraction network, multiple second audio frame features corresponding to the multiple first audio frame features are further obtained. Subsequently, through the pooling layer in the speaker feature extraction network, the multiple second audio frame features are uniformly transformed into the audio features of the audio to be processed, thereby uniformly transforming the frame-level features into audio-level features that can represent the identity characteristics of the speaker in the audio. Finally, the audio features of the audio to be processed are simultaneously input into multiple trained speaker attribute classifiers. According to the multiple speaker attribute classification labels respectively output by the multiple speaker attribute classifiers, the classification results of the speaker in the audio under multiple attributes are simultaneously obtained. This solution considers the correlation of various speaker attributes on the audio features representing the speaker identity characteristics, and by means of the speaker feature extraction network, comprehensively obtains the features of the overall audio representing the speaker identity characteristics from the features of each audio frame of the audio, and then inputs the features into each speaker attribute classifier to simultaneously obtain the classification results of the speaker in the audio under multiple attributes, improving the recognition efficiency and accuracy of the speaker attribute information in the audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a schematic flowchart of the audio processing method in one embodiment;
[0015] FIG. 2(a) is a schematic diagram of the speaker feature extraction network and classifier in one embodiment;
[0016] FIG. 2(b) is a schematic diagram of the speaker feature extraction network and classifier in another embodiment;
[0017] Figure 3 is a schematic diagram of the pooling layer processing features in one embodiment;
[0018] Figure 4 is a schematic diagram of the steps of joint training in one embodiment;
[0019] Figure 5 is a schematic diagram of the relationship between pre-training and joint training in one embodiment;
[0020] Figure 6 is a schematic diagram of the curve of the loss function in one embodiment;
[0021] Figure 7 is a schematic diagram of the effect of label distribution smoothing processing in one embodiment;
[0022] Figure 8 is a schematic diagram of the multi-attribute classification recognition result interface in one embodiment;
[0023] Figure 9 The internal structure diagram of a computer device in an embodiment. Specific implementation manners
[0024] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0025] The audio processing method provided by an embodiment of the present application can be executed by computer devices such as terminals and servers. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, and tablet computers; the server can be implemented by an independent server or a server cluster composed of multiple servers. In terms of application scenarios, the audio processing method provided by the present application can be specifically applied to a live broadcast scenario. By analyzing the live real-time audio, the speaker attribute information such as the gender and age of the speaker in the audio is identified or predicted from the audio information dimension, thereby assisting in the detection and identification of specific populations in similar scenarios.
[0026] The audio processing method provided by the present application will be described below in combination with each embodiment and the corresponding accompanying drawings.
[0027] In one embodiment, as Figure 1 shown, an audio processing method is provided, including the following steps:
[0028] Step S101: Extract the features corresponding to each frame of the audio to be processed, and obtain a plurality of primary first audio frame features.
[0029] In this step, after obtaining the audio to be processed, the features of each frame of the audio to be processed (i.e., each frame of the above audio) are preliminarily extracted to obtain the primary features corresponding to each frame of the audio, which are called first audio frame features, so as to obtain a plurality of first audio frame features. The first audio frame features can be, but are not limited to, MFCC, Fbank, raw spectrum features, etc.
[0030] In some embodiments, step S101 specifically includes: performing frame segmentation on the audio to be processed to obtain multiple frames of audio; and obtaining a plurality of first audio frame features according to the frequency domain features corresponding to each frame of audio.
[0031] Specifically, after obtaining the audio to be processed, preprocessing can be performed on the audio to be processed first. The preprocessing can include processing such as encoding format conversion, normalization, and pre-emphasis. Then, windowing, framing, and short-time Fourier transform (STFT) can be performed on the preprocessed audio, so as to divide the preprocessed audio into multiple frames of audio and transform it from the time domain to the frequency domain. Then, feature extraction can be performed, and the frequency domain features corresponding to each frame of audio are used as the first audio frame features, thereby obtaining multiple first audio frame features. This stage is the primary feature extraction stage, and the first audio frame features extracted may include but are not limited to MFCC, Fbank, raw spectrum features, etc.
[0032] Step S102, obtaining multiple advanced second audio frame features corresponding to multiple primary first audio frame features respectively through the feature extraction layer in the trained speaker feature extraction network, and converting the multiple advanced second audio frame features into audio features used to characterize the speaker identity characteristics in the audio through the pooling layer in the speaker feature extraction network.
[0033] This step mainly inputs the multiple first audio frame features obtained in step S101 into a speaker feature extraction network to obtain the audio features of the overall audio to be processed. Specifically, the speaker feature extraction network includes a feature extraction layer and a pooling layer. The main function of the feature extraction layer is to convert the input multiple first audio frame features into multiple second audio frame features. Compared with step S101 which extracts preliminary features, the feature extraction layer extracts high-level features, that is, further obtains high-level second audio frame features based on the primary first audio frame features. The feature extraction layer can adopt a deep neural network, including but not limited to CNN, ResNet, RNN, LSTM, and Transformer, etc. The feature extraction layer transfers the multiple second audio frame features to the pooling layer. The role of the pooling layer is to map the multiple second audio frame features into features with a fixed dimension, that is, through the pooling layer, the frame-level features are converted into audio-level features, realizing the extraction of audio with different durations into audio features with a fixed dimension. And the overall role of the speaker feature extraction network is to convert the input multiple first audio frame features at the frame level into audio-level features representing the speaker identity characteristics in the audio. The speaker identity characteristics refer to who the speaker is. For the construction of the speaker feature extraction network, in specific implementation, for the recognition task of speaker identity characteristics, by training an identification model including the speaker feature extraction network, after the model is trained, the speaker feature extraction network in the model has a strong feature expression ability for speaker identity characteristics. And various speaker attribute information required to be recognized in this application has a strong correlation with the audio features of speaker identity characteristics. Therefore, based on this correlation, this step can obtain the audio features of the overall audio to be processed that are beneficial to subsequent multi-attribute classification by means of the feature extraction layer and the pooling layer in the speaker feature extraction network.
[0034] Regarding the primary features and high-level features of the audio frames in steps S101 and S102 above. Among them, the primary features refer to using traditional features, such as MFCC, pitch, prosody and other features, which are usually extracted manually through methods such as transformation or filtering. The high-level features refer to the feature vectors obtained after passing through a deep neural network (whose network parameters are determined through supervised training with a large amount of labeled data, not manually selected and set), that is, the high-level features are the features extracted by using the non-linear transformation of the deep neural network, and have the characteristics of being more resistant to interference and noise. The intuitive difference between the two mainly lies in whether there is manual intervention. In terms of performance, the high-level abstract features are more robust.
[0035] Step S103: Input the audio features into multiple trained speaker attribute classifiers to obtain multiple speaker attribute classification labels respectively output by the multiple speaker attribute classifiers.
[0036] In this step, the audio features of the entire to-be-processed audio extracted in step S102 can be input into multiple trained speaker attribute classifiers simultaneously. These speaker attribute classifiers can include a speaker age classifier, a speaker gender classifier, etc., which are respectively used to perform different attribute classification tasks. Among them, the multiple speaker attribute classifiers correspondingly output multiple speaker attribute classification labels according to the input audio features.
[0037] Step S104: Obtain the classification results of the speaker in the to-be-processed audio under multiple attributes according to the multiple speaker attribute classification labels.
[0038] In this step, according to the multiple speaker attribute classification labels, the classification results of the speaker in the to-be-processed audio under attributes such as age and gender can be obtained simultaneously, such as obtaining classification results / attribute information such as whether the speaker in the to-be-processed audio is a minor and their gender.
[0039] The above audio processing method considers the correlation of various speaker attributes on the audio features representing the identity characteristics of the speaker, uses a speaker feature extraction network to comprehensively obtain the features of each frame of the audio to obtain the features of the entire audio representing the identity characteristics of the speaker, and then inputs the features into each speaker attribute classifier to simultaneously obtain the classification results of the speaker in the audio under multiple attributes, improving the recognition efficiency and accuracy of the speaker attribute information in the audio.
[0040] For the speaker feature extraction network and the speaker attribute classifiers in steps S102 and S103, there can be two forms of connection. In one embodiment, as shown in Fig. 2(a), the audio features of the to-be-processed audio can be obtained through a speaker feature extraction network as the features shared by each speaker attribute classifier. That is, after preprocessing such as encoding format conversion, normalization, and pre-emphasis on the to-be-processed audio, multiple first audio frame features are extracted, and then the audio features of the to-be-processed audio are obtained through a speaker feature extraction network as the audio features shared by each speaker attribute classifier. The audio features are respectively input into multiple speaker attribute classifiers (speaker attribute classifier 1, speaker attribute classifier 2,..., speaker attribute classifier N), so as to simultaneously obtain the classification results of the speaker in the audio under attributes 1, 2,..., N.
[0041] In another embodiment, as shown in FIG. 2( b ), there may be multiple speaker feature extraction networks, each corresponding to a plurality of speaker attribute classifiers, that is, one speaker feature extraction network corresponds to one speaker attribute classifier, so that each speaker feature extraction network outputs audio features and transmits them to the corresponding speaker attribute classifier for attribute classification, that is, step S103 inputs the audio features into the plurality of speaker attribute classifiers, which may specifically include: inputting the plurality of audio features output by the plurality of speaker feature extraction networks into the speaker attribute classifiers corresponding to the audio features.
[0042] In this embodiment, the audio is preprocessed and the first feature is extracted to obtain a plurality of first audio frame features, which are used as common features and are input into different classification branches respectively. Each classification branch includes a corresponding speaker feature extraction network and its corresponding speaker attribute classifier, such as classification branch 1 includes speaker feature extraction network 1 and speaker attribute classifier 1, speaker feature extraction network 2 and speaker attribute classifier 2, etc. In each classification branch, the corresponding speaker feature extraction network extracts audio features for its branch classification based on the plurality of first audio frame features and transmits them to the corresponding speaker attribute classifier for attribute classification.
[0043] In some embodiments, the pooling layer in the speaker feature extraction network includes an attention random pooling layer; the step S102 of converting the plurality of high-level second audio frame features into audio features for characterizing speaker identity characteristics in the audio through the pooling layer in the speaker feature extraction network specifically includes:
[0044] Multiple high-level second audio frame features are input into the convolutional attention module to obtain the feature weights corresponding to each frame of audio output by the convolutional attention module; multiple second audio frame features and the feature weights corresponding to each frame of audio are input into the attention random pooling layer to obtain the audio features output by the attention random pooling layer.
[0045] In this embodiment, reference Figure 3 In the speaker feature extraction network, the multiple second audio frame features output by the feature extraction layer can generally be reduced by a random pooling layer, that is, the multiple second audio frame features output by the feature extraction layer (corresponding to the audio frame level features) are mapped into audio features with fixed dimensions (corresponding to the audio segment level features). This embodiment replaces the random pooling layer with an attention random pooling layer, which can better extract the overall long-term main features and change features reflecting the audio, extract more robust speaker audio features, and thus improve the accuracy of identifying / predicting the attribute information of the speaker in the audio, such as gender, age, etc. The specific processing includes:
[0046] The second audio frame feature h tInput convolutional attention module. The convolutional attention module obtains the respective feature weights e for each frame of audio according to the following formula t : e t = f(W h t + b)+ k. Then, all the feature weights e t can be normalized through the Softmax function: Finally, the normalized feature weights a t and the second audio frame features h t are input into the attention random pooling layer to calculate the weighted mean and variance, obtaining the audio features. Among them, the weighted mean μ and variance σ in the attention random pooling layer are calculated according to the following formula: Among them, ⊙ represents the Hadamard product; T represents the number of audio frames; t represents the serial number of the audio frame; W and b represent the weights and biases of the neural network in the convolutional attention module, and k represents a random number between 0 and 1.
[0047] Regarding the training steps of the speaker feature extraction network and each speaker attribute classifier in the foregoing embodiments, in one embodiment, in combination with Figure 4 and Figure 5 , it specifically includes:
[0048] Step S401, obtain a pre-trained speaker identification model.
[0049] Among them, the speaker identification model includes a pre-trained feature extraction layer and a pre-trained pooling layer, and the pre-trained feature extraction layer and the pre-trained pooling layer are used to form a speaker feature extraction network. Among them, the pre-trained speaker identification model refers to the speaker identification model trained in the pre-training stage. This speaker identification model can be used to identify the speaker's identity. It is trained with a large number of audio-visual samples without speaker attribute labels such as gender and age. The pre-trained speaker identification model may include a pre-trained speaker feature extraction network and a speaker identity classifier. The speaker feature extraction network specifically includes a pre-trained feature extraction layer and a pre-trained pooling layer. Specifically, in the pre-training stage, the speaker identification model is trained through a large number of audio-visual samples without speaker attribute labels such as gender and age for the first feature extraction. After training is completed, a pre-trained speaker feature extraction network is obtained. The output of the pre-trained speaker feature extraction network can be used as the audio features that can reflect the speaker's identity characteristics. On this basis, the pre-trained speaker feature extraction network is used as the speaker feature extraction network to be trained. The speaker feature extraction network to be trained correspondingly includes a feature extraction layer to be trained and a pooling layer to be trained. Thus, the pre-trained feature extraction layer and the pre-trained pooling layer form a speaker feature extraction network.
[0050] Step S402: Obtain an audio sample and obtain multiple speaker attribute classification labels corresponding to the audio sample.
[0051] Step S403: Extract the primary features corresponding to each frame of audio in the audio sample to obtain multiple first audio frame features of the audio sample.
[0052] Step S404: Jointly train the speaker feature extraction network and multiple speaker attribute classifiers based on the multiple first audio frame features of the audio sample and the multiple speaker attribute classification labels.
[0053] In the above steps S402 to S404, in the speaker attribute classification task of multi-task learning, different from a large number of audio samples, the audio samples used in the attribute classification task need to have multiple corresponding speaker attribute classification labels. The network parameters of the pre-trained speaker feature extraction network are migrated as the initial values of the network parameters of the speaker feature extraction network to be trained in the speaker attribute classification task of multi-task learning. Then, multiple speaker attribute classifiers are configured accordingly, so as to form the speaker feature extraction network to be trained and the multiple speaker attribute classifiers to be trained. Thus, the audio sample and its corresponding multiple speaker attribute classification labels are used for training fine-tuning. That is, for each audio sample and its corresponding multiple speaker attribute classification labels, the audio sample is subjected to first audio frame feature extraction to obtain multiple first audio frame features. Then, based on the multiple first audio frame features of the audio sample and the multiple speaker attribute classification labels, the speaker feature extraction network to be trained and the multiple speaker attribute classifiers to be trained are jointly trained, so as to perform training fine-tuning on the speaker feature extraction network and each speaker attribute classifier for the speaker attribute classification task on the basis of the pre-training for the speaker identity recognition task, so that the trained speaker feature extraction network applicable to multiple speaker attribute classification tasks can extract audio features that match the speaker identity characteristics and attribute characteristics such as gender and age, so as to realize the multi-attribute classification and recognition of the speaker in the audio in the application stage, avoid the overfitting problem in the small data set scenario, and improve the recognition performance in the small data set scenario.
[0054] The solution of this embodiment can solve the problem of model overfitting caused by insufficient training data samples through transfer learning, improve the robustness of the model. By pre-training the speaker identity recognition model, it is beneficial to the rapid convergence of multi-attribute classification recognition training, and at the same time can reduce the probability of overfitting, and also improve the robustness and universality of audio features, and improve the recognition accuracy.
[0055] In one embodiment, the above step S404 specifically includes:
[0056] Input multiple first audio frame features of an audio sample into a speaker feature extraction network; obtain multiple speaker attribute classification label prediction results respectively output by multiple speaker attribute classifiers; based on the multiple speaker attribute classification label prediction results, multiple speaker attribute classification labels, and multiple loss functions respectively corresponding to the multiple speaker attribute classifiers, jointly train the speaker feature extraction network and the multiple speaker attribute classifiers.
[0057] Further, the above-mentioned joint training of the speaker feature extraction network and the multiple speaker attribute classifiers based on the multiple speaker attribute classification label prediction results, multiple speaker attribute classification labels, and multiple loss functions respectively corresponding to the multiple speaker attribute classifiers specifically includes: obtaining multiple classification loss values based on the multiple speaker attribute classification label prediction results, multiple speaker attribute classification labels, and multiple loss functions respectively corresponding to the multiple speaker attribute classifiers; obtaining a comprehensive loss value based on the multiple classification loss values and loss weights respectively corresponding to the multiple speaker attribute classifiers; when the comprehensive loss value does not meet the loss threshold condition, adjusting the parameters of the speaker feature extraction network and the multiple speaker attribute classifiers using the comprehensive loss value until the comprehensive loss value meets the loss threshold condition, and obtaining the trained speaker feature extraction network and the trained multiple speaker attribute classifiers.
[0058] Combine Figure 5, during the training process of the speaker feature extraction network and multiple speaker attribute classifiers, multiple first audio frame features of an audio sample are input into the speaker feature extraction network. The pre-trained feature extraction layer included in the speaker feature extraction network extracts multiple second audio frame features corresponding to the multiple first audio frame features respectively, and the pre-trained pooling layer included in the speaker feature extraction network converts the multiple second audio frame features into the audio feature of the audio sample, and the audio feature of the audio sample is input into multiple speaker attribute classifiers. Then, for each speaker attribute classifier, according to the input audio feature, a prediction result of the speaker attribute classification label is output, so as to obtain multiple prediction results of the speaker attribute classification labels. Each speaker attribute classifier can correspond to a loss function. Therefore, for each speaker attribute classifier, its prediction result of the speaker attribute classification label and the corresponding speaker attribute classification label can be input into its corresponding loss function to obtain the corresponding classification loss value, so as to obtain multiple classification loss values. Each speaker attribute classifier can correspond to a loss weight. Therefore, after obtaining multiple classification loss values, the loss weights are used to weight them to obtain a comprehensive loss value. Then, the speaker feature extraction network and multiple speaker attribute classifiers can be jointly learned and fine-tuned based on the comprehensive loss value. Specifically, when the comprehensive loss value does not meet the loss threshold condition, for example, when the comprehensive loss value is greater than the loss threshold, the parameters of the speaker feature extraction network and multiple speaker attribute classifiers are adjusted using the comprehensive loss value until the comprehensive loss value meets the loss threshold condition, for example, when the comprehensive loss value is less than or equal to the loss threshold, the trained speaker feature extraction network and the trained multiple speaker attribute classifiers are obtained.
[0059] In addition, for the forms of the two speaker feature extraction networks and speaker attribute classifiers shown in FIGS. 2(a) and 2(b), during the training process, the parameters of each speaker feature extraction network and speaker attribute classifier can be adjusted based on the comprehensive loss value until the comprehensive loss value meets the loss threshold condition, and one or more trained speaker feature extraction networks and the trained multiple speaker attribute classifiers are obtained. By using a comprehensive loss value to simultaneously constrain multiple attribute information classification tasks, the recognition of multiple attribute information classification tasks is realized.
[0060] In one embodiment, the following steps are further included: for a target attribute classifier among multiple speaker attribute classifiers, a corresponding loss function is constructed based on the absolute error function and the mean square error loss function.
[0061] Specifically, the target attribute classifier refers to the classifier among multiple speaker attribute classifiers that is used to classify the target attribute. Among them, for the classification of certain attributes such as age, there is a problem of uneven sample distribution in the audio samples, and the attributes with this characteristic are target attributes. In the training process of this embodiment, a loss function corresponding to the target attribute classifier is constructed based on the absolute error function and the mean square error loss function, so that the weight of the samples with more concentrated data samples is reduced, and more attention is paid to the samples with fewer data samples, thereby compensating for the problem caused by the uneven sample distribution. Specifically, the loss function L of the target attribute classifier among multiple speaker attribute classifiers focal can be expressed as:
[0062]
[0063] where f(•) represents the sigmoid function, and α, β, and γ (γ > 0) are adjustable focusing parameters used to adjust the slope and bias of the loss function curve. L ae (y) is the absolute error function, L mse (y) is the mean square error loss function, y represents the classification label of the sample with the target attribute, represents the corresponding prediction result, such as Figure 6 shown, through the loss function L focal , it can make the samples with fewer expected data samples fall in the middle gray area, so that in the training process, the samples with fewer data samples or difficult samples can be more concerned and considered. The focusing parameter γ can reduce the weight of most samples. Among them, when γ = 0, the above loss function L focal is equal to the mean square error loss function, and this loss function L focal can be called the mean square error center loss function.
[0064] In some embodiments, for samples of target attributes such as age, label distribution smoothing can also be performed to overcome the problem of uneven distribution of data samples at the sample level, so as to improve the accuracy of multi-attribute classification recognition. In this embodiment, label distribution smoothing is to smooth the uneven label distribution of target attributes such as age by using the kernel density estimation method of the Gaussian kernel function. It can consider the overlapping information data samples of similar labels to smooth the data distribution of the labels. The following formula calculates the effective label density distribution of the target label (the label of the sample with the target attribute) Here, the Gaussian kernel function is used to calculate the similarity between the current label and the target label, that is, the distance between the two in the target label distribution space.
[0065]
[0066] Among them, Ψ represents the label space, and p(y) represents the number of the label y. In the age label data, for example, the label space can be divided into 100 parts, that is, the minimum age accuracy is 1 year. Refer to Figure 7 , the original label distribution of the audio data sample is shown on the left, which is seriously unbalanced. After the label distribution is smoothed, the distribution shown on the right shows that it can effectively reduce the serious imbalance of the label distribution, thereby reducing the impact of the unbalanced label distribution on the age recognition accuracy.
[0067] Further, in some embodiments, for the loss function of the classifier for classifying and recognizing target attributes such as age, the steps of constructing the corresponding loss function based on the absolute error function and the mean square error loss function in the foregoing embodiments specifically include:
[0068] Construct an initial loss function based on the absolute error function and the mean square error loss function; obtain the loss function weight according to the ratio of the effective label of the target attribute classification to the label of the target attribute classification; obtain the corresponding loss function according to the product of the loss function weight and the initial loss function.
[0069] This embodiment takes into account the influence of the above-mentioned label distribution smoothing and optimizes the loss function corresponding to the classifier of the target attribute during the training process. Specifically, first, an initial loss function is constructed based on the absolute error function and the mean square error loss function, and this initial loss function can correspond to the loss function L in the foregoing embodiments focal =[f(α*L ae (y)-β) γ *L mse (y)], and then multiply by the loss function weight on this basis Among them, the effective label of the target attribute classification is obtained according to the label distribution smoothing process of the target attribute classification label y in the multiple speaker attribute classification labels in the above embodiment. The finally obtained loss function corresponding to the target attribute in this embodiment is This embodiment further optimizes the mean square error center loss function into a weighted mean square error center loss function Thereby, it can better improve the problem of unbalanced data distribution, reduce the decline in recognition performance caused by the unbalanced data sample distribution, reduce the weight of the majority sample data during the training process, and at the same time pay more attention to the sparse samples.
[0070] The audio processing method provided by this application can be applied to accurately and efficiently recognize the classification results of the speaker under multiple attributes in the live real-time audio of the live broadcast scene, etc. For example, for the live real-time audio, the gender and age of the speaker are predicted / recognized simultaneously, and the prediction / recognition results can be displayed on the terminal devices of relevant personnel, such as Figure 8As shown, the specific information presented may include the audio link corresponding to the processed audio, and attribute information such as the gender and age of the speaker in the audio, as well as whether the speaker is a specific population determined based on this attribute information, etc., for further manual review and other processing. This application can achieve intelligent identification of the speaker attribute information in the audio in a vast amount of audio-visual data, greatly reducing the manual review cost and significantly improving the efficiency, accuracy, and timeliness of the review.
[0071] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0072] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structural diagram may be as Figure 9 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as the audio to be processed and the classification results of attribute information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an audio processing method.
[0073] Those skilled in the art can understand that Figure 9 the structure shown in
[0074] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0075] In one embodiment, a computer program product is provided, including a computer program, which when executed by a processor implements the steps in the above method embodiments.
[0076] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0077] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties.
[0078] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0079] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An audio processing method, characterized in that, The method includes: extracting the respective features corresponding to each frame of audio in the audio to be processed to obtain a plurality of primary first audio frame features; obtaining, through a feature extraction layer in a trained speaker feature extraction network, a plurality of advanced second audio frame features respectively corresponding to the plurality of primary first audio frame features, and inputting the plurality of advanced second audio frame features into a convolutional attention module to obtain the feature weight corresponding to each frame of audio output by the convolutional attention module; inputting the plurality of advanced second audio frame features and the feature weight corresponding to each frame of audio into an attention random pooling layer in the speaker feature extraction network to obtain an audio feature output by the attention random pooling layer for characterizing the speaker identity characteristics in the audio; inputting the audio feature into a plurality of trained speaker attribute classifiers to obtain a plurality of speaker attribute classification labels respectively output by the plurality of speaker attribute classifiers; obtaining a classification result of the speaker in the audio to be processed under multiple attributes according to the plurality of speaker attribute classification labels.
2. The method according to claim 1, wherein The number of the speaker feature extraction networks is multiple, and they respectively correspond to the plurality of speaker attribute classifiers; the audio feature includes a plurality of audio features respectively output by the plurality of speaker feature extraction networks; wherein, the inputting the audio feature into a plurality of trained speaker attribute classifiers includes: respectively inputting the plurality of audio features output by the plurality of speaker feature extraction networks into the speaker attribute classifier corresponding to the audio feature.
3. The method according to any one of claims 1 to 2, characterized in that, The method further includes: obtaining a pre-trained speaker identification model; the speaker identification model includes a pre-trained feature extraction layer and a pre-trained pooling layer, and the pre-trained feature extraction layer and the pre-trained pooling layer are used to form the speaker feature extraction network; obtaining an audio sample and obtaining a plurality of speaker attribute classification labels of the audio sample; extracting the respective primary features corresponding to each frame of audio in the audio sample to obtain a plurality of first audio frame features of the audio sample; jointly training the speaker feature extraction network and the plurality of speaker attribute classifiers based on the plurality of first audio frame features of the audio sample and the plurality of speaker attribute classification labels.
4. The method according to claim 3, characterized in that, The jointly training the speaker feature extraction network and the plurality of speaker attribute classifiers based on the plurality of first audio frame features of the audio sample and the plurality of speaker attribute classification labels includes: inputting the plurality of first audio frame features of the audio sample into the speaker feature extraction network, extracting, by the pre-trained feature extraction layer included in the speaker feature extraction network, a plurality of second audio frame features respectively corresponding to the plurality of first audio frame features, and converting, by the pre-trained pooling layer included in the speaker feature extraction network, the plurality of second audio frame features into an audio feature of the audio sample, and inputting the audio feature of the audio sample into the plurality of speaker attribute classifiers; obtaining a plurality of predicted results of speaker attribute classification labels respectively output by the plurality of speaker attribute classifiers; Based on the prediction results of the multiple speaker attribute classification labels and the multiple speaker attribute classification labels, and based on the multiple loss functions respectively corresponding to the multiple speaker attribute classifiers, jointly train the speaker feature extraction network and the multiple speaker attribute classifiers.
5. The method according to claim 4, characterized in that, The jointly training the speaker feature extraction network and the multiple speaker attribute classifiers based on the prediction results of the multiple speaker attribute classification labels and the multiple speaker attribute classification labels, and based on the multiple loss functions respectively corresponding to the multiple speaker attribute classifiers, includes: Based on the prediction results of the multiple speaker attribute classification labels and the multiple speaker attribute classification labels, and based on the multiple loss functions respectively corresponding to the multiple speaker attribute classifiers, obtain multiple classification loss values; Based on the multiple classification loss values and the loss weights respectively corresponding to the multiple speaker attribute classifiers, obtain a comprehensive loss value; When the comprehensive loss value does not meet the loss threshold condition, use the comprehensive loss value to adjust the parameters of the speaker feature extraction network and the multiple speaker attribute classifiers until the comprehensive loss value meets the loss threshold condition, to obtain a trained speaker feature extraction network and trained multiple speaker attribute classifiers.
6. The method according to claim 4, wherein The method further includes: For a target attribute classifier among the multiple speaker attribute classifiers, construct a corresponding loss function based on the absolute error function and the mean square error loss function.
7. The method according to claim 6, wherein The constructing the corresponding loss function based on the absolute error function and the mean square error loss function includes: Construct an initial loss function based on the absolute error function and the mean square error loss function; Based on the ratio of the target attribute classification valid label to the target attribute classification label, obtain a loss function weight; the target attribute classification valid label is obtained by performing label distribution smoothing on the target attribute classification label among the multiple speaker attribute classification labels; Based on the product of the loss function weight and the initial loss function, obtain the corresponding loss function.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Age recognition method, device and equipment and computer readable storage medium
CN111312286A
Voice-based gender and age recognition method and device, equipment and storage medium
CN112002346A
Speech processing method and speech processing device
CN113555010A