Speaker classification method and device based on double-entry data
By extracting the audio features in the double-recorded data and combining with the deep learning model, multi-task loss function training is performed for dialect accents, the problem of insufficient classification accuracy of speakers in the double-recorded data is solved, and higher classification accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510314564.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-17
AI Technical Summary
In the prior art, when classifying the speakers for double-recorded data, the accuracy rate is insufficient, mainly due to inaccurate marking of training data, failure to process dialect accents, and the inability to dynamically adjust the model.
By obtaining historical double-recorded data during the business processing process, extracting audio data and calculating the Mel cepshot coefficient and linear prediction coefficient, feature extraction and fusion are performed. Then, the model is trained by a multitasking loss function based on dialect accent using the SE module and the ResNet18 residual network as the basis. Deploy the trained model to the server and perform migration training until the loss is less than the threshold, real-time speaker classification is achieved.
The accuracy of speaker classification is improved, the model's adaptability to multi-dial accents is enhanced, and real-time data analysis is realized through migration training, which solves the problem of insufficient accuracy of speaker classification in the existing technology.
Smart Images

Figure CN120164487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio data processing, and in particular, to a speaker classification method, a classification device, a computer-readable storage medium, and a service system based on dual-recording data. Background Art
[0002] With the rapid development of the social economy, the demand for customer service quality in various industries is getting higher and higher. In order to ensure customer service quality and improve the user experience, various industries have adopted a variety of technical means to monitor and evaluate customer service quality. Among them, the banking field mostly adopts the method of telephone return visits to users by management personnel. By extracting the business recordings of salespersons and listening to their business processing procedures, service attitudes, and business handling specifications, etc., the customer service quality is monitored and evaluated.
[0003] The traditional method of video return visits for business handling has problems such as large workload, low efficiency, and strong subjectivity of evaluation results for the return visit personnel. Moreover, with the increase in the number of salespersons, the workload of video return visits for business handling has increased sharply. And since video playback needs to be completed manually, the evaluation results of the business handling process of salespersons tend to deviate due to changes in the work intensity, energy, and business volume of the return visit personnel. When using machine learning, it is difficult to generalize manually extracted features to unseen noise types and scenarios, while deep learning usually requires a large amount of labeled data to train the model, and the accuracy of the model depends on the accuracy of the labels, and it is easy to cause large classification deviations of the model due to inaccurate labeling.
[0004] To solve the above problems, a voice interaction method is used to assist the video return visit method, and the return visit work is automatically completed through technical means such as speech recognition technology and speech synthesis technology. Among them, assisting in completing the video return visit work of business handling in the voice interaction method includes: obtaining dual-recording data during the business handling process, extracting audio data from the dual-recording data, performing speaker recognition and annotation based on the audio data, realizing the automatic recognition of the identities of the salesperson and the customer during the business processing process, and the automatic recognition of the business interaction content between the salesperson and the customer, so as to objectively evaluate the standardization of the salesperson's business processing process, customer satisfaction, etc.
[0005] However, in practical applications, speaker recognition and annotation technologies often adopt end-to-end speaker recognition and speaker annotation technologies. This technology uses a speaker recognition model to perform speaker recognition and then completes speaker annotation based on the speaker recognition results. Among them, the training of the speaker recognition model is based on a large amount of speaker identity information and voice data, such as public data sets in the academic field and business data sets within a company, etc. In practical applications, since the speaker identity information in the internal business data sets of each company is that of internal employees of the company, the speaking characteristics of these internal employees, such as accent, speech rate, intonation, etc., all have certain particularities. Therefore, after training the speaker recognition model with the internal business data sets of the company, the recognition effect of the speaker recognition model on the interactive voice data between salespersons and customers is poor.
[0006] In summary, when classifying speakers in dual-record data in the prior art, the accuracy of speaker classification is usually poor due to reasons such as inaccurate training data annotation, failure to specifically process dialect accents, and inability to dynamically adjust the model. Summary of the Invention
[0007] The main objective of this application is to provide a speaker classification method, classification device, computer-readable storage medium, and business system based on dual-record data, so as to at least solve the problem of insufficient accuracy of the method for analyzing dual-record data to achieve speaker classification in the prior art.
[0008] To achieve the above objective, according to one aspect of this application, a speaker classification method based on dual-record data is provided, including: obtaining historical dual-record data during the business handling process to obtain first target data, where the dual-record data includes audio data and video data; extracting the audio data including conversations from the first target data to obtain second target data, and respectively extracting Mel cepstral coefficients and linear prediction coefficients from the second target data to obtain a first coefficient and a second coefficient; respectively performing feature extraction based on the first coefficient and the second coefficient, and performing feature fusion on the extracted features to obtain third target data, and training a speaker classification model based on the third target data to obtain an alternative classification model. Among them, during the training process, the speaker classification model is based on the SE module and the ResNet18 residual network and is trained through a multi-task loss function based on dialect accents; deploying the alternative classification model on a server, and performing transfer training based on the audio data received by the server in real time until the loss of the alternative classification model is less than a first threshold to obtain a target classification model, and performing speaker classification and annotation on the dual-record data received by the server in real time after the model transfer training is completed through the target classification model.
[0009] Optionally, training a speaker classification model based on the third target data to obtain an alternative classification model, including: inputting the third target data into the speaker classification model to obtain a first classification result; determining the dialect type including a dialect accent in the third target data based on the third target data, calculating a loss function value corresponding to each dialect type according to the first classification result to obtain a first loss value; calculating a total loss function value through a joint loss constraint model based on the first loss value to obtain a second loss value: where N is the total number of the dialect types, is the weighting coefficient of the nth type of dialect, l n is the first loss value corresponding to the nth dialect type, L N is the second loss value; according to the second loss value, adjusting the parameters of the speaker classification model through backpropagation, and processing the third target data according to the adjusted speaker classification model until the second loss value is less than a second threshold, and determining the speaker classification model as the alternative classification model.
[0010] Optionally, the speaker classification model includes a ResNet18 residual network, a plurality of convolutional layers and a plurality of SE modules, the convolutional layers and the SE modules are in one-to-one correspondence, inputting the third target data into the speaker classification model to obtain a first classification result, including processing the third target data through the convolutional layers to obtain a first feature, and inputting the first feature into the next convolutional layer until all convolutional layers are processed; inputting the first features processed by each convolutional layer into the corresponding SE modules for processing to obtain a plurality of second features; performing concat fusion on the first feature processed by the last convolutional layer and each second feature to obtain a third feature; processing the third feature through the ResNet18 residual network to obtain a first classification result.
[0011] Optionally, after adjusting the parameters of the speaker classification model through backpropagation, the method further includes: updating the weighting coefficient according to each first loss value:
[0012] Optionally, extracting mel cepstral coefficients based on the second target data respectively to obtain a first coefficient, including: performing pre-emphasis processing on the second target data, and framing the pre-emphasized second target data according to a preset duration to obtain a fourth target data; performing windowing processing on the fourth target data, and performing fast Fourier transform on the windowed fourth target data to obtain a fifth target data; processing the fifth target data through triangular mel filtering, and taking the logarithm of the processing result to obtain a logarithmic mel spectrum; calculating discrete cosine transform on the logarithmic mel spectrum to obtain a first coefficient.
[0013] Optionally, linear prediction coefficients are extracted based on the second target data to obtain second coefficients, including: calculating linear prediction coefficients based on the second target data according to the autocorrelation method.
[0014] Optionally, feature extraction is respectively performed based on the first coefficients and the second coefficients, and the extracted features are fused, including: inputting the first coefficients into a convolutional neural network to obtain fourth features, the convolutional neural network including a convolutional layer and a flattening layer; inputting the second coefficients into the convolutional neural network to obtain fifth features; performing concat fusion on the fourth features and the fifth features to obtain third target data.
[0015] According to another aspect of the present application, there is provided a speaker classification device based on dual-recording data. The device includes: a first acquisition unit for acquiring dual-recording data during the business handling process to obtain first target data, the dual-recording data including recording data and video data; a first processing unit for extracting audio data including conversations from the first target data to obtain second target data, and respectively extracting Mel cepstral coefficients and linear prediction coefficients based on the second target data to obtain first coefficients and second coefficients; a second processing unit for performing feature fusion based on the first coefficients and the second coefficients to obtain third target data, training a speaker classification model based on the third target data to obtain an alternative classification model, wherein during the training process, the speaker classification model is based on the SE module and the ResNet18 residual network and is trained through a multi-task loss function based on dialect accents; a classification unit for deploying the alternative classification model to a server and performing transfer training based on the recording data received by the server in real time until the loss of the alternative classification model is less than a first threshold to obtain a target classification model, and performing speaker classification and annotation on the dual-recording data received by the server in real time after the model transfer training is completed through the target classification model.
[0016] According to still another aspect of the present application, there is provided a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the methods.
[0017] According to yet another aspect of the present application, there is provided a business system, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods.
[0018] Applying the technical solution of the present application in the above-mentioned speaker classification method based on dual-recording data, first, historical dual-recording data in the business handling process is obtained to obtain first target data, and the dual-recording data includes recording data and video data; then, audio data including conversations is extracted from the first target data to obtain second target data, and Mel cepstral coefficients and linear prediction coefficients are respectively extracted based on the second target data to obtain a first coefficient and a second coefficient; after that, feature extraction is respectively performed based on the first coefficient and the second coefficient, and the extracted features are feature-fused to obtain third target data, and a speaker classification model is trained based on the third target data to obtain an alternative classification model. Among them, during the training process, the speaker classification model is based on the SE module and the ResNet18 residual network and is trained through a multi-task loss function based on dialect accents; finally, the alternative classification model is deployed on the server, and transfer training is performed based on the recording data received by the server in real time until the loss of the alternative classification model is less than a first threshold to obtain a target classification model, and the dual-recording data received by the server in real time after the model transfer training is completed is classified and labeled by the target classification model. The present application combines linear prediction coefficients, Mel cepstral coefficients and deep learning models to reduce background noise interference, and involves a multi-task loss function for multi-dialect accent classification to enhance the adaptability to dialect accents, and realizes real-time data analysis through transfer training. This method solves the problem of insufficient accuracy in the method of analyzing dual-recording data to achieve speaker classification in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 FIG. shows a hardware structure block diagram of a mobile terminal for a speaker classification method based on dual-recording data provided in an embodiment of the present application;
[0020] Figure 2 FIG. shows a flowchart of a speaker classification method based on dual-recording data provided in an embodiment of the present application;
[0021] Figure 3 FIG. shows a feature fusion diagram of an alternative classification model provided in an embodiment of the present application;
[0022] Figure 4 FIG. shows a network structure diagram of a target classification model provided in an embodiment of the present application;
[0023] Figure 5 FIG. shows a flowchart of extracting Mel cepstral coefficients provided in an embodiment of the present application;
[0024] Figure 6 FIG. shows an operation flowchart of feature convolution extraction provided in an embodiment of the present application;
[0025] Figure 7 Shows a schematic diagram of flattened layer feature fusion provided according to an embodiment of the present application;
[0026] Figure 8 Shows a structural block diagram of a speaker classification device based on dual-recording data provided according to an embodiment of the present application.
[0027] Among them, the above-mentioned drawings include the following reference numerals:
[0028] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. Detailed implementation manners
[0029] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] For the convenience of description, some nouns or terms related to the embodiments of the present application are described below:
[0033] Dual-recording data: The recording and video recording carried out in the complete sales process of products in the financial industry.
[0034] Dual-recording quality inspection: The manual and / or intelligent detection of the recording and video recording.
[0035] Speaker Diarization (SD): also known as voiceprint segmentation and clustering, is a speech processing technology that aims to solve the problem of "who spoke when", that is, to determine who is speaking at each time point in a speech containing multiple people speaking alternately.
[0036] Squeeze-and-Excitation Networks (SENet): A model can be called SENet after introducing the compression and excitation module (SE Block). SE Block is essentially an attention module that contains global information.
[0037] Multi-task Learning: Learn multiple related tasks together and share some parameters to make the model more generalizable.
[0038] As introduced in the background technology, when dual-recording data is used for speaker classification in the prior art, the accuracy of speaker classification is usually poor due to inaccurate annotation of training data, lack of targeted processing of dialect accents, and inability to dynamically adjust the model. In order to solve the problem of insufficient accuracy of methods for analyzing dual-recording data to achieve speaker classification in the prior art, embodiments of the present application provide a speaker classification method based on dual-recording data, a classification device, a computer-readable storage medium, and a business system.
[0039] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0040] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 1 is a hardware structure block diagram of a mobile terminal of a speaker classification method based on dual-recording data according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.
[0041] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the speaker classification method based on dual-recording data in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0042] In this embodiment, a speaker classification method based on dual-recording data running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0043] Figure 2 is a flowchart of the speaker classification method based on dual-recording data according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0044] Step S201, obtain historical dual-recording data in the business handling process to obtain first target data, where the dual-recording data includes recording data and video data;
[0045] Specifically, collect historical dual-recording data in the business handling process, which includes recording data and video data. For example, assume that a scenario where a customer purchases a financial product in a bank is being processed. During this process, there will be a conversation between the salesperson and the customer, and the conversation will be recorded by recording and video devices. These historical dual-recording data will become the first target data used to train the speaker classification model.
[0046] Step S202: Extract the audio data including conversations from the first target data to obtain the second target data, and respectively extract the Mel cepstral coefficients and linear prediction coefficients from the second target data to obtain the first coefficient and the second coefficient;
[0047] Specifically, extract the audio data including conversations from the first target data to form the second target data. Then, respectively extract the Mel cepstral coefficients (MFCC) and linear prediction coefficients (LPC) from these audio data.
[0048] In a specific implementation, for example, from 10,000 dual-recorded conversations, only focus on the audio data, and extract MFCC and LPC from each conversation. The extraction of MFCC can help understand the spectral characteristics of speech, while LPC can provide useful information about the prediction of speech signals. For example, for one of the dual-recorded conversations, first convert it to the Mel frequency domain, and then calculate its cepstral coefficients to obtain the first coefficient (MFCC). At the same time, use the autocorrelation method to calculate the second coefficient (LPC) to understand the dynamic characteristics of the speech signal.
[0049] Step S203: Respectively perform feature extraction based on the first coefficient and the second coefficient, and perform feature fusion on the extracted features to obtain the third target data. Train the speaker classification model based on the third target data to obtain the alternative classification model, where, during the training process, the speaker classification model is based on the SE module and the ResNet18 residual network, and is trained through a multi-task loss function based on dialect accents;
[0050] Specifically, based on the extracted first coefficient and second coefficient, respectively perform feature extraction through deep learning networks, and then perform feature fusion (using the concat operation) on the outputs of the two networks to form the third target data. This fused feature will be used to train the speaker classification model, and the basic architecture of the model consists of the SE module and the ResNet18 residual network. The SE module can extract important features in the channels, while ResNet18 solves the problem of gradient disappearance that may occur in the training of deep networks through residual connections. The training process uses a multi-task loss function based on dialect accents, which can dynamically adjust the importance of different dialect accent classification tasks to ensure that the model has good classification performance for various dialects.
[0051] In a specific implementation, such as Figure 3As shown, for each piece of audio in the above-mentioned second target data, the MFCC and LPC are subjected to feature extraction through a convolutional neural network, and then the extracted feature vectors are fused through a concat operation to form a more comprehensive feature representation. The fused features will be fed into a ResNet18 network for training. During the training process, a multi-task loss function based on dialect accents is set. Each classification task of each dialect has a corresponding cross-entropy loss, and these losses will be weighted and averaged, and the weights are dynamically adjusted according to the current performance of each dialect task to ensure that the model can achieve a high classification accuracy for all dialects.
[0052] Step S204: Deploy the alternative classification model to the server, and perform transfer training based on the recording data received by the server in real time until the loss of the alternative classification model is less than the first threshold, obtaining the target classification model. Use the target classification model to classify and label the speakers of the dual-recording data received by the server in real time after the model transfer training is completed.
[0053] Specifically, deploy the trained alternative classification model to the server and perform transfer training using the recording data received by the server in real time. During the transfer training process, the model will be fine-tuned according to the new dialect accents, background noise, and speaking styles until the loss of the model is less than the preset first threshold. At this time, the model will be marked as the final target classification model for speaker classification in the real-time quality inspection process.
[0054] Through this embodiment, first, historical dual-recording data in the business handling process is obtained to obtain first target data, where the dual-recording data includes recording data and video data; then, audio data including conversations is extracted from the first target data to obtain second target data, and Mel cepstral coefficients and linear prediction coefficients are respectively extracted from the second target data to obtain a first coefficient and a second coefficient; after that, feature extraction is respectively performed based on the first coefficient and the second coefficient, and the extracted features are fused to obtain third target data, and a speaker classification model is trained based on the third target data to obtain an alternative classification model. Among them, during the training process, the speaker classification model is based on the SE module and the ResNet18 residual network and is trained through a multi-task loss function based on dialect accents; finally, the alternative classification model is deployed on the server, and transfer training is performed based on the recording data received by the server in real time until the loss of the alternative classification model is less than a first threshold to obtain a target classification model, and the target classification model is used to classify and label the dual-recording data received by the server in real time after the model transfer training is completed. This application combines linear prediction coefficients, Mel cepstral coefficients and deep learning models to reduce background noise interference, involves a multi-task loss function for multi-dialect accent classification to enhance the adaptability to dialect accents, and realizes real-time data analysis through transfer training. This method solves the problem of insufficient accuracy in the method of analyzing dual-recording data to achieve speaker classification in the prior art.
[0055] In an alternative embodiment for training the above alternative classification model, step S203 includes:
[0056] Step 2031, input the third target data into the speaker classification model to obtain a first classification result;
[0057] Specifically, the third target data (i.e., the result after MFCC and LPC feature fusion) is fed into the speaker classification model. The model attempts to identify the speaker's identity based on these features and obtains a first classification result.
[0058] Step 2032, determine the dialect type including the dialect accent in the third target data based on the third target data, and calculate the loss function value corresponding to each dialect type according to the first classification result to obtain a first loss value;
[0059] Specifically, based on the third target data, determine the included dialect type. Then, according to the first classification result, calculate the loss function value corresponding to each dialect type, that is, the first loss value. This process involves quantifying the difference between the model prediction result and the actual speaker identity, especially considering the influence of dialect accents.
[0060] Step 2033, calculate the total loss function value through the joint loss constraint model based on the first loss value to obtain a second loss value:
[0061]
[0062] Wherein, N is the total number of dialect types, is the weighted coefficient of the sub-dialect, l n is the first loss value corresponding to the nth dialect type, L N is the second loss value;
[0063] Specifically, through the combined loss function, the loss values of each dialect type are weighted and summed to obtain the total loss value, that is, the second loss value. The weighted coefficient is dynamically adjusted to ensure that the model performs well enough on all dialect types.
[0064] In one embodiment, assume that the first loss value of Mandarin is 0.3, Cantonese is 0.5, and Sichuanese is 0.4. Set the initial weighted coefficient to 1 / N, that is, give the same weight to the loss values of each dialect. In this example, N = 3, so each weighted coefficient is 1 / 3. Calculate the second loss value as: 0.3*(1 / 3)+0.5*(1 / 3)+0.4*(1 / 3) = 0.4. If the second loss value is higher than the preset second threshold (e.g., 0.1), the model parameters need to be adjusted through backpropagation to reduce the total loss value.
[0065] Step 2034, according to the second loss value, adjust the parameters of the speaker classification model through backpropagation, and process the third target data according to the adjusted speaker classification model until the second loss value is less than the second threshold, and determine the speaker classification model as the alternative classification model.
[0066] Specifically, if the classification result of a certain dialect type is not very ideal, that is, the first loss value of this dialect type is relatively high, dynamically adjust the weighted coefficient of this dialect type to increase its influence on the total loss value, which means that the model will pay more attention to the classification performance of this dialect type during training, thereby improving the robustness and generalization ability of the model. The training process will continue until the second loss value is less than the preset second threshold. Once this condition is met, it means that the speaker classification performance of the model on all dialect types has reached the expectation. At this time, the speaker classification model will be determined as the alternative classification model for subsequent real-time dual-recording quality inspection.
[0067] Through the above embodiments, an alternative classification model is trained in this application. This model has high accuracy and robustness in speaker classification of different dialect types, and can effectively handle the dialect accent differences encountered in intelligent dual-recording quality inspection. This not only improves the accuracy of speaker classification, but also enhances the adaptability of the model in practical applications.
[0068] In an alternative embodiment, in order to obtain the above first classification result, as Figure 4 shown, the above step S2031 includes:
[0069] Step S20311: Process the third target data through a convolutional layer to obtain a first feature, and input the first feature into the next convolutional layer until all convolutional layers are processed;
[0070] Specifically, input the third target data (i.e., the fused mel cepstral coefficients and linear prediction coefficients) into the first convolutional layer of the speaker classification model. This convolutional layer performs a convolution operation on the input data through a convolutional kernel to extract local features and obtain a first feature. Then, the first feature is used as input and passed to the next convolutional layer, and so on until all convolutional layers are processed.
[0071] In a specific implementation, assume there are 100 fused audio feature vectors. First, input these vectors into the first convolutional layer of the model. This convolutional layer has multiple different convolutional kernels, and each convolutional kernel performs a convolution operation on the feature vectors to extract local features and generate new feature vectors, that is, the first feature. Then, these first feature vectors will be input into the second convolutional layer of the model, and the above process is repeated until all 5 convolutional layers are processed.
[0072] Step S20312: Input the first feature obtained by processing each convolutional layer into the corresponding SE module for processing to obtain multiple second features;
[0073] Specifically, input each group of first features obtained by processing through the convolutional layer into the corresponding SE module. The SE module can perform "squeeze" and "excitation" operations on the features to recalibrate the features, thereby obtaining multiple enhanced second features.
[0074] In a specific implementation, for the first feature obtained from the first convolutional layer, input it into the first SE module. The SE module compresses the size of the feature vector through global average pooling operation, then learns the importance of these features through a fully connected layer, and reallocates weights to the features (i.e., "excitation"). Finally, multiply the weights by the original feature vector to obtain the enhanced second feature. Perform the same processing on the features obtained from all convolutional layers to obtain a series of second features.
[0075] Step S20313: Concatenate and fuse the first feature processed by the last convolutional layer with each second feature to obtain a third feature;
[0076] Specifically, the features obtained through all convolutional layers are concatenated (i.e., feature splicing) with the second features processed by the corresponding SE module to obtain a third feature containing feature information at all levels.
[0077] In a specific implementation, the feature vectors output by all convolutional layers and the SE module are concatenated, that is, they are linked together on a certain dimension of the feature vectors to form a longer feature vector, namely the third feature. This third feature contains various feature representations from low-level to high-level.
[0078] Step S20314, process the third feature through the ResNet18 residual network to obtain the first classification result.
[0079] Specifically, the third feature will be input into the ResNet18 residual network and further processed through residual blocks to finally obtain the first classification result of the speaker.
[0080] In a specific implementation, the third feature is input into the ResNet18 residual network. The residual connection of the ResNet18 network can avoid the problem of gradient disappearance when training deep networks. The third feature is processed through multiple residual blocks to finally obtain the first classification result of the speaker, that is, to identify who is speaking in the current audio segment.
[0081] Through the above embodiments, the speaker classification model of the present application can make full use of different feature extraction and enhancement technologies, combined with the deep learning ability of the ResNet18 residual network, to achieve accurate classification of speakers. This design not only improves the accuracy of classification, but also enhances the robustness of the model to work under different background noises and dialect accents, and is very suitable for application in the intelligent double-recording quality inspection scenario.
[0082] In order to make the above-mentioned weighting coefficients change with the training effect of the model, after parameter adjustment of the speaker classification model through backpropagation, the above method further includes:
[0083] Step S301, update the weighting coefficients according to each first loss value:
[0084]
[0085] Specifically, represents the weighting coefficient of the Nth task loss function. The initial weighting coefficients are the same, indicating that the importance of the loss of each task to the joint loss is the same. l n represents the loss function of the nth task, Task nDenote the loss function value with the weighted coefficient added. To make the weighted coefficient change with the loss function, the weighted coefficient is updated according to the above formula, and the performance of the model is inversely proportional to the above weighted coefficient.
[0086] It can be understood that this means that the weighted coefficient of the current dialect type is its loss value divided by the sum of the loss values of all dialect types. If the classification error of a certain dialect type is large, its corresponding weighted coefficient will also increase accordingly, so that the model will pay more attention to the classification of this dialect type in subsequent training.
[0087] In the above embodiments, by dynamically updating the weighted coefficient, we can ensure that the speaker classification model pays more attention to those dialect types with poor classification effects during the training process, thereby improving the overall performance of the model on all dialect types.
[0088] To extract the Mel cepstral coefficients, in an alternative embodiment, as Figure 5 shown, the above step S202 includes:
[0089] Step S2021, perform pre-emphasis processing on the second target data, and frame the pre-emphasized second target data according to a preset duration to obtain the fourth target data;
[0090] Specifically, performing pre-emphasis processing on the second target data is to enhance the high-frequency part of the signal because the high-frequency part is of great value in speech recognition. Then, the pre-emphasized audio data is framed according to a preset duration. Usually, the frame length is between 20 - 30 ms, which is to capture the short-term characteristics of the speech signal because the identity characteristics of the speaker can often be accurately captured within a short time window.
[0091] Step S2022, perform windowing processing on the fourth target data, and perform a fast Fourier transform on the windowed fourth target data to obtain the fifth target data;
[0092] Specifically, perform windowing processing on each framed audio data, usually using a Hamming Window or a Hanning Window to reduce the boundary effect between frames. The windowed data is subjected to a fast Fourier transform (FFT) to convert the time-domain signal into a frequency-domain signal to analyze the spectral characteristics of the signal.
[0093] Step S2023, process the fifth target data through triangular Mel filtering, and take the logarithm of the processing result to obtain the logarithmic Mel spectrum;
[0094] Specifically, the spectral data after the fast Fourier transform is processed by triangular Mel filters, which are evenly distributed on the Mel frequency scale and can better simulate the human ear's perception of different frequencies. Taking the logarithm of the result output by the triangular Mel filters gives the log Mel spectrogram.
[0095] Step S2024: Calculate the discrete cosine transform of the log Mel spectrogram to obtain the first coefficients.
[0096] Specifically, perform a discrete cosine transform (DCT) on the log Mel spectrogram to obtain the Mel-frequency cepstral coefficients (MFCCs), which are the first coefficients. These coefficients reflect the dynamic characteristics of the speech signal and are crucial for speaker identification.
[0097] Through the above embodiments, it is possible to extract valuable features for speaker identification - Mel-frequency cepstral coefficients from the audio data. This feature processing flow is based on the best practices of speech signal processing and feature extraction in the prior art, but in terms of the setting of specific parameters (such as the pre-emphasis coefficient, the number of Mel filters, etc.), the present invention may have been optimized and adjusted to better meet the specific requirements of intelligent dual-recording quality inspection.
[0098] In order to extract the autocorrelation coefficients, in an alternative embodiment, the above step S202 further includes:
[0099] Step S2025: Based on the autocorrelation method, calculate the linear prediction coefficients according to the second target data.
[0100] Specifically, use the autocorrelation method to calculate the linear prediction coefficients. The basic idea of the autocorrelation method is to represent the current speech signal sample as a linear combination of several past samples, and estimate the linear prediction coefficients by minimizing the energy of the prediction error. Usually, a prediction order is set when calculating the LPC, and this order determines how many past samples are used to predict the current sample.
[0101] In a specific implementation, it is assumed that the original speech signal has been extracted from a segment of audio data containing a conversation. To calculate the LPC, the prediction order is set to 12, which means that a linear combination of the past 12 samples will be used to predict the current sample. Through autocorrelation method calculation, a set of linear prediction coefficients is obtained. This set of coefficients describes the dynamic characteristics of the speech signal and is very helpful for identifying the speaker's identity. For example, the audio duration is 1 minute and the sampling rate is 16 kHz. First, preprocess this audio, such as noise reduction and pre-emphasis, to enhance the signal quality. Then, use the autocorrelation method to calculate the LPC with the prediction order set to 12. The process of calculating the LPC by the autocorrelation method includes calculating the autocorrelation function, constructing a prediction equation based on the autocorrelation function, and solving the prediction equation to obtain the linear prediction coefficients. That is, first calculate the autocorrelation function, which describes the similarity of the signal at different time delays. Then, construct a prediction equation based on the autocorrelation function. The form of the prediction equation is a linear combination of the autocorrelation function, and the goal is to minimize the energy of the prediction error. Finally, by solving the prediction equation, a set of linear prediction coefficients is obtained, and this set of coefficients is the linear prediction coefficients (the second coefficients) of the second target data.
[0102] Through the above embodiments, the linear prediction coefficients (i.e., the second coefficients) are extracted from the second target data by the autocorrelation method. These coefficients can effectively characterize the dynamic features of the speech signal and are one of the essential features for the speaker classification task.
[0103] To extract the above third target data, in an alternative implementation, as Figure 6 and Figure 7 shown, the above step S203 includes:
[0104] Step S2035, input the first coefficients into a convolutional neural network to obtain the fourth feature. The convolutional neural network includes a convolutional layer and a flattening layer;
[0105] Specifically, the mel-frequency cepstral coefficients (MFCC, the first coefficients) and the linear prediction coefficients (LPC, the second coefficients) are respectively used as inputs and fed into the convolutional neural network. The convolutional neural network includes a convolutional layer and a flattening layer (FlattenLayer). The convolutional layer can learn and extract the local patterns of features, while the flattening layer converts the multi-dimensional features into one-dimensional for subsequent processing.
[0106] Step S2036, input the second coefficients into the convolutional neural network to obtain the fifth feature;
[0107] Specifically, the linear prediction coefficients are also input into the convolutional layer and the flattening layer of the CNN to obtain the fifth feature.
[0108] Step S2037: Concatenate and fuse the fourth feature and the fifth feature to obtain the third target data.
[0109] Specifically, concatenate (feature splicing) and fuse the fourth feature and the fifth feature to obtain a combined feature. The concatenate fusion operation is performed on the one-dimensional vectors of the features. Two feature vectors are concatenated together at a certain dimension to form a longer feature vector. This fusion method can retain the complete information of the original features, introduce new dimensions for model learning, and enhance the model's representation ability.
[0110] Through the above embodiments, key features are extracted from the Mel cepstral coefficients and the linear prediction coefficients, and they are integrated through the concatenate fusion method to form a combined feature (the third target data). This feature fusion method makes full use of the advantages of different acoustic features, provides a more comprehensive and rich input representation for the speaker classification model, and thus improves the classification accuracy and robustness of the model.
[0111] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0112] The embodiment of the present application also provides a speaker classification device based on dual-recording data. It should be noted that the speaker classification device based on dual-recording data in the embodiment of the present application can be used to execute the speaker classification method based on dual-recording data provided by the embodiment of the present application. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0113] The following introduces the speaker classification device based on dual-recording data provided by the embodiment of the present application.
[0114] Figure 8 is a structural block diagram of the speaker classification device based on dual-recording data according to the embodiment of the present application. As Figure 8 shown, the device includes:
[0115] The first acquisition unit 10 is configured to acquire historical dual-recording data during the business handling process to obtain the first target data. The dual-recording data includes recording data and video data;
[0116] Specifically, historical dual-recording data is collected during the business processing, which includes audio data and video data. For example, assume a scenario where a customer is purchasing a wealth management product at a bank. During this process, there will be a conversation between the salesperson and the customer, and the conversation will be recorded by audio and video devices. These historical dual-recording data will become the first target data used to train the speaker classification model.
[0117] The first processing unit 20 is configured to extract the audio data including the conversation from the first target data to obtain the second target data, and respectively extract the Mel cepstral coefficients and the linear prediction coefficients based on the second target data to obtain the first coefficient and the second coefficient;
[0118] Specifically, the audio data including the conversation is extracted from the first target data to form the second target data. Then, the Mel cepstral coefficients (MFCC) and the linear prediction coefficients (LPC) are respectively extracted from these audio data.
[0119] In a specific implementation, for example, from 10,000 dual-recording conversations, only the audio data is concerned, and the MFCC and LPC are extracted from each conversation. The extraction of MFCC can help understand the spectral characteristics of the speech, while the LPC can provide useful information about the prediction of the speech signal. For example, for one of the dual-recording conversations, it is first converted to the Mel frequency domain, and then its cepstral coefficients are calculated to obtain the first coefficient (MFCC). At the same time, the autocorrelation method is used to calculate the second coefficient (LPC) to understand the dynamic characteristics of the speech signal.
[0120] The second processing unit 30 is configured to respectively perform feature extraction based on the first coefficient and the second coefficient, and perform feature fusion on the extracted features to obtain the third target data, and train the speaker classification model based on the third target data to obtain an alternative classification model, where, during the training process, the speaker classification model is based on the SE module and the ResNet18 residual network, and is trained through a multi-task loss function based on dialect accents;
[0121] Specifically, based on the extracted first coefficient and second coefficient, feature extraction is respectively performed through deep learning networks, and then the outputs of the two networks are fused (using the concat operation) to form the third target data. This fused feature will be used to train the speaker classification model, and the basic architecture of the model consists of the SE module and the ResNet18 residual network. The SE module can extract important features in the channels, while ResNet18 solves the problem of gradient disappearance that may occur during the training of deep networks through residual connections. The training process uses a multi-task loss function based on dialect accents, which can dynamically adjust the importance of different dialect accent classification tasks to ensure that the model has good classification performance for various dialects.
[0122] In a specific implementation, such asFigure 3 As shown, for each piece of audio in the above-mentioned second target data, the MFCC and LPC are subjected to feature extraction through a convolutional neural network, and then the extracted feature vectors are fused through a concat operation to form a more comprehensive feature representation. The fused features will be fed into a ResNet18 network for training. During the training process, a multi-task loss function based on dialect accents is set. Each classification task of each dialect has a corresponding cross-entropy loss, and these losses will be weighted and averaged, and the weights are dynamically adjusted according to the current performance of each dialect task to ensure that the model can achieve a high classification accuracy for all dialects.
[0123] A classification unit is used to deploy an alternative classification model on a server and perform transfer training based on the recording data received by the server in real time until the loss of the alternative classification model is less than a first threshold, obtaining a target classification model, and classifying and annotating the speakers of the dual-recording data received by the server in real time after the model transfer training is completed through the target classification model.
[0124] Specifically, the trained alternative classification model is deployed on the server, and transfer training is carried out using the recording data received by the server in real time. During the transfer training process, the model will be fine-tuned according to the new dialect accents, background noises, and speaking styles until the loss of the model is less than a preset first threshold. At this time, the model will be marked as the final target classification model for speaker classification in the real-time quality inspection process.
[0125] Through this embodiment, the first acquisition unit acquires historical dual-recording data during the business handling process to obtain first target data, where the dual-recording data includes audio data and video data; the first processing unit extracts the audio data including conversations from the first target data to obtain second target data, and based on the second target data, extracts Mel cepstral coefficients and linear prediction coefficients respectively to obtain a first coefficient and a second coefficient; the second processing unit performs feature extraction based on the first coefficient and the second coefficient respectively, and performs feature fusion on the extracted features to obtain third target data, and trains a speaker classification model based on the third target data to obtain an alternative classification model. Among them, during the training process, the speaker classification model is based on the SE module and the ResNet18 residual network, and is trained through a multi-task loss function based on dialect accents; the classification unit deploys the alternative classification model on the server, and performs transfer training based on the audio data received by the server in real time until the loss of the alternative classification model is less than the first threshold to obtain a target classification model, and classifies and labels the dual-recording data received by the server in real time after the model transfer training is completed through the target classification model. This application combines linear prediction coefficients, Mel cepstral coefficients and deep learning models to reduce background noise interference, involves a multi-task loss function for multi-dialect accent classification to enhance the adaptability to dialect accents, and realizes real-time data analysis through transfer training. This method solves the problem of insufficient accuracy in the method of analyzing dual-recording data to achieve speaker classification in the prior art.
[0126] In an alternative implementation manner for training the above alternative classification model, the above second processing unit includes:
[0127] A first input module, configured to input the third target data into the speaker classification model to obtain a first classification result;
[0128] Specifically, the third target data (that is, the result after MFCC and LPC feature fusion) is sent into the speaker classification model. The model attempts to identify the speaker's identity based on these features to obtain a first classification result.
[0129] A first calculation module, configured to determine the dialect type including the dialect accent in the third target data based on the third target data, and calculate the loss function value corresponding to each dialect type according to the first classification result to obtain a first loss value;
[0130] Specifically, based on the third target data, determine the included dialect type. Then, according to the first classification result, calculate the loss function value corresponding to each dialect type, that is, the first loss value. This process involves quantifying the difference between the model prediction result and the actual speaker identity, especially considering the influence of dialect accents.
[0131] A second calculation module, configured to calculate a total loss function value through a joint loss constraint model based on the first loss value, and obtain a second loss value:
[0132]
[0133] where N is the total number of dialect types, is the weighting coefficient of the nth type of dialect, and l n is the first loss value corresponding to the nth dialect type, and L N is the second loss value;
[0134] Specifically, through the joint loss function, the loss values of each dialect type are weighted and summed to obtain the total loss value, that is, the second loss value. The weighting coefficient is dynamically adjusted to ensure that the model performs well enough on all dialect types.
[0135] In one embodiment, assume that the first loss value of Mandarin is 0.3, Cantonese is 0.5, and Sichuanese is 0.4. Set the initial weighting coefficient to 1 / N, that is, the same weight is given to the loss value of each dialect. In this example, N = 3, so each weighting coefficient is 1 / 3. Calculate the second loss value as: 0.3*(1 / 3) + 0.5*(1 / 3) + 0.4*(1 / 3) = 0.4. If the second loss value is higher than a preset second threshold (for example, 0.1), the model parameters need to be adjusted through backpropagation to reduce the total loss value.
[0136] A third calculation module, configured to adjust the parameters of the speaker classification model through backpropagation according to the second loss value, and process the third target data according to the adjusted speaker classification model until the second loss value is less than the second threshold, and determine the speaker classification model as an alternative classification model.
[0137] Specifically, if the classification result of a certain dialect type is not very ideal, that is, the first loss value of this dialect type is relatively high, dynamically adjust the weighting coefficient of this dialect type to increase its influence on the total loss value, which means that the model will pay more attention to the classification performance of this dialect type during training, thereby improving the robustness and generalization ability of the model. The training process will continue until the second loss value is less than the preset second threshold. Once this condition is met, it means that the speaker classification performance of the model on all dialect types has reached the expectation. At this time, the speaker classification model will be determined as an alternative classification model for subsequent real-time dual-recording quality inspection.
[0138] Through the above embodiments, an alternative classification model is trained in this application. This model has high accuracy and robustness in classifying speakers of different dialect types and can effectively handle the problem of dialect accent differences encountered in intelligent dual-recording quality inspection. This not only improves the accuracy of speaker classification but also enhances the adaptability of the model in practical applications.
[0139] In order to obtain the above first classification result, in an alternative embodiment, as Figure 4 shown, the above first output module includes:
[0140] A first input sub-module, configured to process the third target data through a convolutional layer to obtain a first feature, and input the first feature into the next convolutional layer until all convolutional layers are processed;
[0141] Specifically, the third target data (i.e., the fused mel cepstral coefficients and linear prediction coefficients) is input into the first convolutional layer of the speaker classification model. This convolutional layer performs a convolutional operation on the input data through a convolutional kernel to extract local features and obtain a first feature. Then, the first feature is used as an input and passed to the next convolutional layer, and so on until all convolutional layers are processed.
[0142] In a specific implementation, assuming there are 100 fused audio feature vectors, these vectors are first input into the first convolutional layer of the model. This convolutional layer has multiple different convolutional kernels, and each convolutional kernel performs a convolutional operation on the feature vectors to extract local features and generate new feature vectors, that is, the first feature. Then, these first feature vectors will be input into the second convolutional layer of the model, and the above process is repeated until all 5 convolutional layers are processed.
[0143] A second input sub-module, configured to input the first features obtained by processing each convolutional layer into the corresponding SE module for processing to obtain multiple second features;
[0144] Specifically, each group of first features obtained by processing through the convolutional layer is input into the corresponding SE module. The SE module can perform "squeeze" and "excitation" operations on the features, that is, learn the importance of the features through global average pooling and a fully connected layer, and recalibrate the features to obtain multiple enhanced second features.
[0145] In a specific implementation, for the first feature obtained from the first convolutional layer, it is input into the first SE module. The SE module compresses the size of the feature vector through global average pooling operation, then learns the importance of these features through a fully connected layer, and re - assigns weights to the features (i.e., "excitation"). Finally, the weights are multiplied by the original feature vector to obtain the enhanced second feature. The same processing is performed on the features obtained from all convolutional layers to obtain a series of second features.
[0146] The first processing sub - module is used to perform concat fusion on the first feature processed by the last convolutional layer and each second feature to obtain a third feature;
[0147] Specifically, perform a concat (i.e., feature splicing) operation on the features obtained from all convolutional layers and the second features processed by the corresponding SE modules to obtain a third feature containing feature information of all levels.
[0148] In a specific implementation, perform a concat operation on the feature vectors output from all convolutional layers and SE modules, that is, link them together on a certain dimension of the feature vectors to form a longer feature vector, which is the third feature. This third feature contains various feature representations from low - level to high - level.
[0149] The second processing sub - module is used to process the third feature through the ResNet18 residual network to obtain a first classification result.
[0150] Specifically, the third feature will be input into the ResNet18 residual network and further processed through residual blocks to finally obtain the first classification result of the speaker.
[0151] In a specific implementation, input the third feature into the ResNet18 residual network. The residual connection of the ResNet18 network can avoid the problem of gradient disappearance when training deep networks. The third feature is processed through multiple residual blocks to finally obtain the first classification result of the speaker, that is, identify who is speaking in the current audio segment.
[0152] Through the above - mentioned embodiments, the speaker classification model of the present application can make full use of different feature extraction and enhancement techniques, combined with the deep - learning ability of the ResNet18 residual network, to achieve accurate classification of speakers. This design not only improves the accuracy of classification but also enhances the robustness of the model in working under different background noises and dialect accents, and is very suitable for application in the intelligent double - recording quality inspection scenario.
[0153] In order to make the above - mentioned weighting coefficients change with the training effect of the model, the above - mentioned device further includes:
[0154] A computing unit, configured to update the weighting coefficients according to each first loss value after adjusting the parameters of the speaker classification model through backpropagation:
[0155]
[0156] Specifically, represents the weighting coefficient of the Nth task loss function. The initial weighting coefficients are the same, indicating that the importance of the loss of each task to the joint loss is the same. l n represents the loss function of the nth task, Task n represents the loss function value with the weighting coefficient added. To make the weighting coefficient change with the loss function, the weighting coefficient is updated according to the above formula. The performance of the model is inversely proportional to the above weighting coefficient.
[0157] It can be understood that this means that the weighting coefficient of the current dialect type is its loss value divided by the sum of the loss values of all dialect types. If the classification error of a certain dialect type is large, its corresponding weighting coefficient will also increase accordingly, so that the model will pay more attention to the classification of this dialect type in subsequent training.
[0158] In the above embodiment, by dynamically updating the weighting coefficients, we can ensure that the speaker classification model pays more attention to those dialect types with poor classification effects during the training process, thereby improving the overall performance of the model on all dialect types.
[0159] In order to extract Mel cepstral coefficients, in an optional implementation manner, as Figure 5 shown, the above first processing unit includes:
[0160] A first processing module, configured to perform pre-emphasis processing on the second target data and frame the pre-emphasized second target data according to a preset duration to obtain fourth target data;
[0161] Specifically, performing pre-emphasis processing on the second target data is to enhance the high-frequency part of the signal because the high-frequency part is of great value in speech recognition. Then, the pre-emphasized audio data is framed according to a preset duration. Usually, the frame length is between 20 - 30 ms. This is to capture the short-term characteristics of the speech signal because the identity characteristics of the speaker can often be accurately captured within a short time window.
[0162] A second processing module, configured to perform windowing processing on the fourth target data and perform a fast Fourier transform on the windowed fourth target data to obtain fifth target data;
[0163] Specifically, windowing processing is performed on each sub-framed audio data, usually using a Hamming Window or a Hanning Window, to reduce the boundary effect between frames. The windowed data is subjected to a Fast Fourier Transform (FFT) to convert the time-domain signal into a frequency-domain signal for analyzing the spectral characteristics of the signal.
[0164] A third processing module is configured to process the fifth target data through triangular Mel filtering and take the logarithm of the processing result to obtain a logarithmic Mel spectrogram.
[0165] Specifically, the spectral data after the Fast Fourier Transform is processed through triangular Mel filters, which are uniformly distributed on the Mel frequency scale and can better simulate the human ear's perception of different frequencies. The logarithm of the result output by the triangular Mel filters is taken to obtain a logarithmic Mel spectrogram.
[0166] A fourth calculation module is configured to calculate a discrete cosine transform of the logarithmic Mel spectrogram to obtain a first coefficient.
[0167] Specifically, a discrete cosine transform (DCT) is performed on the logarithmic Mel spectrogram to obtain Mel Frequency Cepstral Coefficients (MFCCs), i.e., the first coefficient. These coefficients reflect the dynamic characteristics of the speech signal and are crucial for speaker identification.
[0168] Through the above embodiments, it is possible to extract valuable features for speaker identification - Mel Frequency Cepstral Coefficients from audio data. This feature processing flow is based on the best practices of speech signal processing and feature extraction in the prior art, but in terms of the setting of specific parameters (such as the pre-emphasis coefficient, the number of Mel filters, etc.), the present invention may have been optimized and adjusted to better meet the specific requirements of intelligent dual-recording quality inspection.
[0169] In order to extract autocorrelation coefficients, in an optional embodiment, the above first processing unit further includes:
[0170] A fifth calculation module is configured to calculate linear prediction coefficients based on the autocorrelation method according to the second target data.
[0171] Specifically, the autocorrelation method (Autocorrelation Method) is used to calculate the linear prediction coefficients. The basic idea of the autocorrelation method is to represent the current speech signal sample as a linear combination of several past samples, and estimate the linear prediction coefficients by minimizing the energy of the prediction error. Usually, a prediction order is set when calculating the LPC, and this order determines how many past samples are used to predict the current sample.
[0172] In a specific implementation, it is assumed that the original speech signal has been extracted from a segment of audio data containing a conversation. To calculate the LPC, the prediction order is set to 12, which means that a linear combination of the past 12 samples will be used to predict the current sample. Through autocorrelation method calculation, a set of linear prediction coefficients is obtained. This set of coefficients describes the dynamic characteristics of the speech signal and is very helpful for identifying the speaker's identity. For example, the audio duration is 1 minute and the sampling rate is 16 kHz. First, preprocess this audio, such as noise reduction and pre-emphasis, to enhance the signal quality. Then, use the autocorrelation method to calculate the LPC with the prediction order set to 12. The process of calculating the LPC by the autocorrelation method includes calculating the autocorrelation function, constructing a prediction equation based on the autocorrelation function, and solving the prediction equation to obtain the linear prediction coefficients. That is, first calculate the autocorrelation function, which describes the similarity of the signal at different time delays. Then, construct a prediction equation based on the autocorrelation function. The form of the prediction equation is a linear combination of the autocorrelation function, and the goal is to minimize the energy of the prediction error. Finally, by solving the prediction equation, a set of linear prediction coefficients is obtained, and this set of coefficients is the linear prediction coefficients (second coefficients) of the second target data.
[0173] Through the above embodiments, linear prediction coefficients (i.e., second coefficients) are extracted from the second target data by the autocorrelation method. These coefficients can effectively characterize the dynamic features of the speech signal and are one of the essential features for the speaker classification task.
[0174] To extract the above third target data, in an alternative implementation, as Figure 6 and Figure 7 shown, the above second processing unit includes:
[0175] A second input module for inputting the first coefficients into a convolutional neural network to obtain a fourth feature. The convolutional neural network includes a convolutional layer and a flattening layer;
[0176] Specifically, the Mel-frequency cepstral coefficients (MFCC, first coefficients) and the linear prediction coefficients (LPC, second coefficients) are respectively used as inputs and fed into the convolutional neural network. The convolutional neural network includes a convolutional layer and a flattening layer (FlattenLayer). The convolutional layer can learn and extract local patterns of features, while the flattening layer converts multi-dimensional features into one-dimensional for subsequent processing.
[0177] A third input module for inputting the second coefficients into the convolutional neural network to obtain a fifth feature;
[0178] Specifically, the linear prediction coefficients are also input into the convolutional layer and the flattening layer of the CNN to obtain a fifth feature.
[0179] The fourth processing module is used to perform concat fusion on the fourth feature and the fifth feature to obtain the third target data.
[0180] Specifically, the fourth feature and the fifth feature are subjected to concat (feature splicing) fusion to obtain a combined feature. The concat fusion operation is performed on the one-dimensional vector of the feature, and the two feature vectors are concatenated together in a certain dimension to form a longer feature vector. This fusion method can retain the complete information of the original feature, while introducing new dimensions for model learning and enhancing the model's representation ability.
[0181] The above-mentioned speaker classification device based on dual-recording data includes a processor and a memory. The above-mentioned first acquisition unit, first processing unit, second processing unit, classification unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement the corresponding functions. The above-mentioned modules are all located in the same processor; or, the above-mentioned each module is located in different processors in any combination form.
[0182] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and the recognition accuracy of speaker classification for dual-recording data can be improved by adjusting the kernel parameters.
[0183] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0184] The embodiment of the present invention provides a computer-readable storage medium. The above-mentioned computer-readable storage medium includes a stored program, wherein when the above-mentioned program runs, it controls the device where the above-mentioned computer-readable storage medium is located to execute the above-mentioned speaker classification method based on dual-recording data.
[0185] The embodiment of the present invention provides a processor. The above-mentioned processor is used to run a program, wherein when the above-mentioned program runs, it executes the above-mentioned speaker classification method based on dual-recording data.
[0186] The embodiment of the present invention provides a service system. The service system includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the steps of the above-mentioned speaker classification method based on dual-recording data.
[0187] The present application also provides a computer program product, which is suitable for executing a program initialized with at least the steps of the above-mentioned speaker classification method based on dual-recording data when executed on a data processing device.
[0188] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0189] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0190] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0191] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0192] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the functions in the flowFigure 1 one or more processes and / or blocks Figure 1 steps of the functions specified in one or more blocks
[0193] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0194] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0195] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0196] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0197] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0198] 1) The speaker classification method based on dual-recording data of the present application first obtains historical dual-recording data during the business handling process to obtain first target data, where the dual-recording data includes recording data and video data. Then, audio data including conversations is extracted from the first target data to obtain second target data, and Mel cepstral coefficients and linear prediction coefficients are respectively extracted based on the second target data to obtain a first coefficient and a second coefficient. After that, feature extraction is respectively performed based on the first coefficient and the second coefficient, and the extracted features are fused to obtain third target data. A speaker classification model is trained based on the third target data to obtain an alternative classification model. During the training process, the speaker classification model is based on the SE module and the ResNet18 residual network and is trained through a multi-task loss function based on dialect accents. Finally, the alternative classification model is deployed on the server, and transfer training is performed based on the recording data received by the server in real time until the loss of the alternative classification model is less than a first threshold to obtain a target classification model. The speaker in the dual-recording data received by the server in real time after the model transfer training is classified and labeled through the target classification model. The present application combines linear prediction coefficients, Mel cepstral coefficients and deep learning models to reduce background noise interference, involves a multi-task loss function for multi-dialect accent classification to enhance the adaptability to dialect accents, and realizes real-time data analysis through transfer training. This method solves the problem of insufficient accuracy in the method of analyzing dual-recording data to achieve speaker classification in the prior art.
[0199] 2) The speaker classification device based on dual-recording data of the present application, wherein the first acquisition unit acquires historical dual-recording data during the business handling process to obtain first target data, and the dual-recording data includes audio data and video data; the first processing unit extracts the audio data including conversations from the first target data to obtain second target data, and respectively extracts Mel cepstral coefficients and linear prediction coefficients based on the second target data to obtain a first coefficient and a second coefficient; the second processing unit respectively performs feature extraction based on the first coefficient and the second coefficient, and performs feature fusion on the extracted features to obtain third target data, and trains a speaker classification model based on the third target data to obtain an alternative classification model. During the training process, the speaker classification model is based on the SE module and the ResNet18 residual network and is trained through a multi-task loss function based on dialect accents; the classification unit deploys the alternative classification model to the server and performs transfer training based on the audio data received by the server in real time until the loss of the alternative classification model is less than a first threshold to obtain a target classification model, and classifies and labels the dual-recording data received by the server in real time after the model transfer training is completed through the target classification model. The present application combines linear prediction coefficients, Mel cepstral coefficients and deep learning models to reduce background noise interference, and involves a multi-task loss function for multi-dialect accent classification to enhance the adaptability to dialect accents, and realizes real-time data analysis through transfer training. This method solves the problem of insufficient accuracy in the method of analyzing dual-recording data to achieve speaker classification in the prior art.
[0200] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A speaker classification method based on dual-recording data, characterized in that: include: Acquire historical dual-recording data during the business handling process to obtain first target data, wherein the dual-recording data includes audio data and video data; Extracting audio data including dialogue based on the first target data to obtain second target data, and extracting Mel-frequency cepstral coefficients and linear prediction coefficients based on the second target data to obtain first coefficients and second coefficients; Performing feature extraction based on the first coefficient and the second coefficient respectively, and performing feature fusion on the extracted features to obtain third target data, and training a speaker classification model based on the third target data to obtain an alternative classification model, wherein during the training process, the speaker classification model is based on an SE module and a ResNet18 residual network and is trained by a multi-task loss function based on a dialect accent; The alternative classification model is deployed on a server, and migration training is performed based on the recording data received by the server in real time until the loss of the alternative classification model is less than a first threshold, so as to obtain a target classification model. The target classification model is used to classify and label the speakers of the dual-recording data received by the server in real time after the model migration training is completed.
2. The method according to claim 1, characterized in that Training a speaker classification model based on the third target data to obtain an alternative classification model includes: Inputting the third target data into the speaker classification model to obtain a first classification result; Determine the dialect type containing the dialect accent in the third target data based on the third target data, and calculate the loss function value corresponding to each dialect type according to the first classification result to obtain a first loss value; Based on the first loss value, the total loss function value is calculated through the joint loss constraint model to obtain the second loss value: Where N is the total number of dialect types, is the weighting coefficient of the nth dialect, l n is the first loss value corresponding to the nth dialect type, L N is the second loss value; According to the second loss value, the parameters of the speaker classification model are adjusted through back propagation, and the third target data is processed according to the adjusted speaker classification model until the second loss value is less than a second threshold, and the speaker classification model is determined as the alternative classification model.
3. The method according to claim 2, characterized in that The speaker classification model includes the ResNet18 residual network, multiple convolutional layers and multiple SE modules, and the convolutional layers correspond to the SE modules one by one. The third target data is input into the speaker classification model to obtain a first classification result, including Processing the third target data through the convolution layer to obtain a first feature, and inputting the first feature into the next convolution layer until all convolution layers are processed; Inputting the first features obtained by processing each convolutional layer into the corresponding SE module for processing to obtain multiple second features; Concatenate the first feature processed by the last convolutional layer with each of the second features to obtain a third feature; The third feature is processed by the ResNet18 residual network to obtain the first classification result.
4. The method according to claim 2, characterized in that: After adjusting the parameters of the speaker classification model by back propagation, the method further includes: According to each of the first loss values, the weighting coefficient is updated:
5. The method according to claim 1, characterized in that Extracting Mel-frequency cepstral coefficients based on the second target data to obtain first coefficients includes: Performing pre-emphasis processing on the second target data, and dividing the second target data after the pre-emphasis processing into frames according to a preset time length to obtain fourth target data; Performing windowing processing on the fourth target data, and performing fast Fourier transform on the fourth target data after the windowing processing to obtain fifth target data; Processing the fifth target data by using a triangular Mel filter, and taking a logarithm of the processing result to obtain a logarithmic Mel spectrum; A discrete cosine transform is calculated for the logarithmic Mel spectrum to obtain the first coefficient.
6. The method according to claim 1, characterized in that Extracting linear prediction coefficients based on the second target data to obtain second coefficients includes: The linear prediction coefficient is calculated according to the second target data based on an autocorrelation method.
7. The method according to claim 1, characterized in that Extracting features based on the first coefficient and the second coefficient respectively, and fusing the extracted features, including: Inputting the first coefficient into a convolutional neural network to obtain a fourth feature, wherein the convolutional neural network includes a convolution layer and a flattening layer; Inputting the second coefficient into the convolutional neural network to obtain a fifth feature; The fourth feature and the fifth feature are concat-fused to obtain the third target data.
8. A speaker classification device based on dual-recording data, characterized in that: The device comprises: A first acquisition unit is used to acquire dual-recording data during the business handling process to obtain first target data, wherein the dual-recording data includes audio data and video data; A first processing unit is configured to extract audio data including a dialogue based on the first target data to obtain second target data, and to extract Mel-frequency cepstral coefficients and linear prediction coefficients based on the second target data to obtain first coefficients and second coefficients; A second processing unit is used to perform feature fusion based on the first coefficient and the second coefficient to obtain third target data, and train a speaker classification model based on the third target data to obtain an alternative classification model, wherein during the training process, the speaker classification model is based on the SE module and the ResNet18 residual network and is trained by a multi-task loss function based on a dialect accent; A classification unit is used to deploy the alternative classification model on a server, and perform migration training based on the recording data received by the server in real time until the loss of the alternative classification model is less than a first threshold, thereby obtaining a target classification model, and using the target classification model to classify and label the speakers of the dual-recording data received by the server in real time after the model migration training is completed.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
10. A business system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 7.